User interface and technique for changing display manner of object
By detecting interaction conditions in the computer system and adjusting the orientation of display components to indicate or not indicate the user's eye contact, the inefficiency of user interface design in the prior art is solved, enabling faster and more efficient object display and overlay display, and saving energy for battery-powered devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-09-25
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, user interface designs that involve cumbersome and time-consuming operations are inefficient and wasteful of user time and device energy, especially when changing the way objects are displayed.
By detecting the interaction conditions corresponding to the displayed content in the computer system, and adjusting the orientation of the display components to indicate or not indicate the user's eye contact based on different standards met by the interaction conditions, faster and more efficient object display and overlay display can be achieved.
It reduces the cognitive burden on users, improves the efficiency of the user interface, saves power for battery-powered devices, and extends battery charging intervals.
Smart Images

Figure CN122029504A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 541,811, filed September 30, 2023, and U.S. Provisional Patent Application Serial No. 63 / 541,832, filed September 30, 2023, which are incorporated herein by reference in their entirety for all purposes. Background Technology
[0002] Users frequently use computer systems to display objects. These objects include videos, animations, and images. Summary of the Invention
[0003] Existing technologies for using electronic devices to change the display of objects are often cumbersome and inefficient. In some implementations, existing technologies use complex and time-consuming user interfaces that may involve multiple keystrokes or button presses. Some existing technologies require more time than necessary, wasting user time and device power. This latter consideration is particularly important in battery-powered devices.
[0004] Therefore, the present invention provides electronic devices with faster and more efficient methods and interfaces for changing the display of objects and / or for displaying overlays. Such methods and interfaces can optionally supplement or replace other methods for changing the display of objects and / or for displaying overlays. These methods and interfaces reduce the cognitive burden on the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging. Such methods and interfaces can supplement or replace other methods for changing the display of objects.
[0005] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting a first interaction condition corresponding to the content while the content is displayed via the display component; and in response to detecting the first interaction condition corresponding to the content: displaying a representation of a face looking towards the content via the display component based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria; and displaying a representation of the face looking towards a first user detected in a detection field of the one or more input devices based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria.
[0006] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first interaction condition corresponding to content when content is displayed via the display component; and in response to detecting the first interaction condition corresponding to the content: displaying a representation of a face looking towards the content via the display component, based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria; and displaying a representation of the face looking towards a first user detected in the detection field of the one or more input devices, based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria.
[0007] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first interaction condition corresponding to content when content is displayed via the display component; and in response to detecting the first interaction condition corresponding to the content: displaying a representation of a face looking towards the content via the display component, based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria; and displaying a representation of the face looking towards a first user detected in the detection field of the one or more input devices, based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria.
[0008] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first interaction condition corresponding to content when content is displayed via the display component; and in response to detecting the first interaction condition corresponding to the content: displaying a representation of a face looking towards the content via the display component based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria; and displaying a representation of the face looking towards a first user detected in the detection field of the one or more input devices based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria.
[0009] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: when displaying content via the display component, detecting a first interaction condition corresponding to the content; and in response to detecting the first interaction condition corresponding to the content: based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, displaying a representation of a face looking towards the content via the display component; and based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria, displaying a representation of the face looking towards a first user detected in the detection field of the one or more input devices via the display component.
[0010] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first interaction condition corresponding to the content when the content is displayed via the display component; and in response to detecting the first interaction condition corresponding to the content: displaying a representation of a face looking towards the content via the display component based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria; and displaying a representation of the face looking towards a first user detected in the detection field of the one or more input devices based on determining that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from the first set of one or more criteria.
[0011] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: detecting a first input via the one or more input devices while a user interface including a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed with the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed with the first input.
[0012] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via the one or more input devices while a user interface comprising a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed with the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed with the first input.
[0013] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via the one or more input devices while a user interface comprising a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed regarding the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed regarding the first input.
[0014] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting a first input via the one or more input devices while a user interface comprising a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed regarding the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed regarding the first input.
[0015] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting a first input via the one or more input devices while a user interface comprising a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed with the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed with the first input.
[0016] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting a first input via the one or more input devices while a user interface comprising a user interface object representing a portion of a person is displayed via the display component; and in response to detecting the first input: continuing to display the user interface via the display component while altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is expressed with the first input; and continuing to display the user interface via the display component without altering the user interface object in a manner indicating eye contact with the user, based on a determination that consent is not expressed with the first input.
[0017] In some embodiments, a method is described that is performed at a computer system communicating with a display component, a camera, and one or more input devices. In some embodiments, the method includes: upon detecting a first entity in the field of view of a camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed toward the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed toward the portion of the content.
[0018] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices. In some embodiments, the one or more programs include instructions for: upon detecting a first entity in the field of view of a camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
[0019] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices. In some embodiments, the one or more programs include instructions for: upon detecting a first entity in the field of view of a camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
[0020] In some embodiments, a computer system communicating with a display component, a camera, and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: upon detecting a first entity in the field of view of a camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
[0021] In some embodiments, a computer system communicating with a display component, a camera, and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: upon detecting a first entity in the field of view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed toward the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed toward the portion of the content.
[0022] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices. In some embodiments, the one or more programs include instructions for: upon detecting a first entity in the field of view of a camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object is directed toward the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; upon displaying the user interface object indicating the portion of the user, detecting a request to interact with content; and in response to detecting the request to interact with the content: displaying at least a first portion of the content; and displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner to indicate that the user interface object is directed toward the portion of the content.
[0023] In some embodiments, a method is described that is performed at a computer system communicating with a display component and one or more input devices. In some embodiments, the method includes: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on determining that second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on determining that third content different from the first content is to be output and the third content corresponds to the first system digital image.
[0024] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on determining that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on determining that a third content different from the first content is to be output and the third content corresponds to the first system digital image.
[0025] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on the determination that a second content different from the first content should be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on the determination that a third content different from the first content should be output and the third content corresponds to the first system digital image.
[0026] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via the one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on the determination that a second content different from the first content should be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on the determination that a third content different from the first content should be output and the third content corresponds to the first system digital image.
[0027] In some embodiments, a computer system communicating with a display component and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on the determination that a second content different from the first content should be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on the determination that a third content different from the first content should be output and the third content corresponds to the first system digital image.
[0028] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display component and one or more input devices. In some embodiments, the one or more programs include instructions for: displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; outputting first content while displaying the first system digital image; and after outputting the first content and without detecting input via one or more input devices: displaying a second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role, based on determining that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image; and continuing to display the first system digital image via the display component without displaying the second system digital image, based on determining that a third content different from the first content is to be output and the third content corresponds to the first system digital image.
[0029] In some embodiments, a method is described that is performed at a computer system communicating with a display component, an audio generation component, and a motion component. In some embodiments, the method includes: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0030] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component, an audio generation component, and a motion component. In some embodiments, the one or more programs include instructions for: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0031] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display component, an audio generation component, and a motion component. In some embodiments, the one or more programs include instructions for: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0032] In some embodiments, a computer system communicating with a display component, an audio generation component, and a motion component is described. In some embodiments, the computer system includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0033] In some embodiments, a computer system communicating with a display component, an audio generation component, and a motion component is described. In some embodiments, the computer system includes components for performing each of the following steps: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0034] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with a display component, an audio generation component, and a motion component. In some embodiments, the one or more programs include instructions for: outputting audio content via the audio generation component; physically moving a portion of the computer system via the motion component while outputting the audio content; detecting a request to display visual content while the portion of the computer system is physically moved via the motion component; and in response to detecting the request to display the visual content: stopping the physical movement of the portion of the computer system via the motion component; and displaying the visual content via the display component.
[0035] In some embodiments, a method is described that is performed at a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the method includes: detecting, via the one or more input devices, voice input including a description of one or more content attributes corresponding to the content while providing one or more outputs corresponding to a first portion of the content via the one or more output devices; and in response to detecting the voice input including the description of one or more content attributes corresponding to the content, providing one or more outputs via the one or more output devices corresponding to a second portion of the content, without providing output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0036] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting, via the one or more input devices, voice input including a description of one or more content attributes corresponding to the content when providing one or more outputs corresponding to a first portion of content via the one or more output devices; and, in response to detecting the voice input including the description of one or more content attributes corresponding to the content, providing one or more outputs corresponding to a second portion of the content via the one or more output devices, without providing output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0037] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting, via the one or more input devices, voice input including a description of one or more content attributes corresponding to the content when providing one or more outputs corresponding to a first portion of the content via the one or more output devices; and, in response to detecting the voice input including the description of one or more content attributes corresponding to the content, providing one or more outputs corresponding to a second portion of the content via the one or more output devices, without providing output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0038] In some embodiments, a computer system communicating with one or more output devices and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting, via the one or more input devices, voice input including a description of one or more content attributes corresponding to the content when providing one or more outputs corresponding to a first portion of content via the one or more output devices; and in response to detecting the voice input including a description of one or more content attributes corresponding to the content, providing one or more outputs corresponding to a second portion of the content via the one or more output devices, but not providing output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0039] In some embodiments, a computer system communicating with one or more output devices and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting voice input via the one or more input devices a description of one or more content attributes corresponding to the content while providing one or more outputs corresponding to a first portion of the content via the one or more output devices; and in response to detecting voice input including a description of one or more content attributes corresponding to the content, providing one or more outputs via the one or more output devices a second portion of the content, but not providing outputs corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0040] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: detecting, via the one or more input devices, voice input including a description of one or more content attributes corresponding to the content when one or more outputs corresponding to a first portion of the content are provided via the one or more output devices; and in response to detecting the voice input including a description of one or more content attributes corresponding to the content, providing one or more outputs corresponding to a second portion of the content via the one or more output devices, without providing output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
[0041] In some embodiments, a method is described that is performed at a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the method includes: detecting, while providing one or more outputs corresponding to a first portion of content via the one or more output devices, via the one or more input devices that a user's attention is no longer corresponding to the computer system; stopping the provision of one or more outputs corresponding to the first portion of the content in response to detecting that the user's attention is no longer corresponding to the computer system; detecting, while not providing one or more outputs corresponding to the first portion of the content, via the one or more input devices that the user's attention is corresponding to the computer system; and providing, in response to detecting that the user's attention is corresponding to the computer system, one or more outputs corresponding to a second portion of the content via the one or more output devices, the second portion of the content occurring at or before the first portion of the content.
[0042] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: when one or more outputs corresponding to a first portion of content are provided via the one or more output devices, detecting via the one or more input devices that a user's attention is no longer corresponding to the computer system; in response to detecting that the user's attention is no longer corresponding to the computer system, stopping the provision of one or more outputs corresponding to the first portion of the content; when one or more outputs corresponding to the first portion of the content are not provided, detecting via the one or more input devices that the user's attention corresponds to the computer system; and in response to detecting that the user's attention corresponds to the computer system, providing via the one or more output devices one or more outputs corresponding to a second portion of the content, the second portion of the content occurring at or before the first portion of the content.
[0043] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: when one or more outputs corresponding to a first portion of content are provided via the one or more output devices, detecting via the one or more input devices that a user's attention is no longer corresponding to the computer system; in response to detecting that the user's attention is no longer corresponding to the computer system, stopping the provision of one or more outputs corresponding to the first portion of the content; when one or more outputs corresponding to the first portion of the content are not provided, detecting via the one or more input devices that the user's attention corresponds to the computer system; and in response to detecting that the user's attention corresponds to the computer system, providing via the one or more output devices one or more outputs corresponding to a second portion of the content, the second portion of the content occurring at or before the first portion of the content.
[0044] In some embodiments, a computer system communicating with one or more output devices and one or more input devices is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for: detecting, while providing one or more outputs corresponding to a first portion of content via the one or more output devices, via the one or more input devices that a user's attention is no longer corresponding to the computer system; stopping the provision of one or more outputs corresponding to the first portion of the content in response to detecting that the user's attention is no longer corresponding to the computer system; detecting, via the one or more input devices, that a user's attention corresponds to the computer system when no one or more outputs corresponding to the first portion of the content are provided; and providing, via the one or more output devices, one or more outputs corresponding to a second portion of the content, the second portion of the content occurring at or before the first portion of the content, in response to detecting that the user's attention corresponds to the computer system.
[0045] In some embodiments, a computer system communicating with one or more output devices and one or more input devices is described. In some embodiments, the computer system includes components for performing each of the following steps: when providing one or more outputs corresponding to a first portion of content via the one or more output devices, detecting via the one or more input devices that a user's attention is no longer corresponding to the computer system; in response to detecting that the user's attention is no longer corresponding to the computer system, stopping the provision of one or more outputs corresponding to the first portion of the content; when not providing one or more outputs corresponding to the first portion of the content, detecting via the one or more input devices that the user's attention is corresponding to the computer system; and in response to detecting that the user's attention is corresponding to the computer system, providing one or more outputs corresponding to a second portion of the content via the one or more output devices, the second portion of the content being at or before the first portion of the content.
[0046] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more output devices and one or more input devices. In some embodiments, the one or more programs include instructions for: when one or more outputs corresponding to a first portion of content are provided via the one or more output devices, detecting via the one or more input devices that a user's attention is no longer corresponding to the computer system; in response to detecting that the user's attention is no longer corresponding to the computer system, stopping the provision of one or more outputs corresponding to the first portion of the content; when one or more outputs corresponding to the first portion of the content are not provided, detecting via the one or more input devices that the user's attention corresponds to the computer system; and in response to detecting that the user's attention corresponds to the computer system, providing via the one or more output devices one or more outputs corresponding to a second portion of the content, the second portion of the content occurring at or before the first portion of the content.
[0047] Executable instructions for performing these functions may optionally be included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Attached Figure Description
[0048] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals in all the drawings indicate the corresponding parts.
[0049] Figure 1 This is a block diagram illustrating a computer system according to some implementation schemes.
[0050] Figures 2A to 2C These are illustrations of exemplary components and user interfaces of a device 200 according to some implementation schemes.
[0051] Figure 3 This is a block diagram illustrating exemplary components of a device according to some implementation schemes.
[0052] Figure 4 This is a functional diagram of an exemplary actuator device according to some implementation schemes.
[0053] Figure 5 This is a functional diagram of an exemplary intelligent agent system based on some implementation schemes.
[0054] Figures 6A to 6GAn exemplary user interface for changing the display of an object is illustrated according to some implementation schemes.
[0055] Figure 7 This is a flowchart illustrating a method for displaying an object facing a certain direction, based on some implementation schemes.
[0056] Figure 8 This is a flowchart illustrating a method for indicating eye contact with a display object according to some implementation schemes.
[0057] Figure 9 This is a flowchart illustrating a method for removing an emphasized object according to some implementation schemes.
[0058] Figures 10A to 10F Exemplary user interfaces for providing content are illustrated according to some implementation schemes.
[0059] Figure 11 This is a flowchart illustrating a method for displaying a digital image of a system according to some implementation schemes.
[0060] Figure 12 This is a flowchart illustrating a method for selectively moving a portion of a computer system according to some implementation schemes.
[0061] Figure 13 This is a flowchart illustrating a method for navigating content according to some implementation schemes.
[0062] Figure 14 This is a flowchart illustrating a method for pausing content according to some implementation schemes. Detailed Implementation
[0063] The following description illustrates exemplary methods, components, parameters, etc. While specific examples are set forth below, it should be understood that such implementations should not be construed as limiting the scope of this disclosure to the explicit description of the examples set forth herein, but rather as providing illustrative examples.
[0064] Each of the modules and applications identified herein corresponds to a set of executable instructions for performing one or more functions described above and methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) may optionally not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. In some embodiments, a video player module may optionally be combined with a music player module into a single module. In some embodiments, memory may optionally store a subset of the modules and data structures identified above. Furthermore, memory may optionally store additional modules and data structures not described above.
[0065] One or more steps of the method described herein may depend on satisfying one or more conditions. In some embodiments, the method is performed through multiple iterative processes. In some embodiments, the conditional steps may be satisfied in different iterations of the same process and still remain within the scope of the method described herein. In some embodiments, for a given method comprising two steps depending on different conditions, those skilled in the art will understand that the given method should be considered performed even if the process is repeated multiple times until the conditional step is satisfied. In some embodiments, multiple iterations of the process are not required to practice the claims as set forth herein. In some embodiments, the claims of the electronic device, system, or computer-readable medium can be performed without iteratively repeating the process. In some embodiments, the claims of the electronic device, system, or computer-readable medium include instructions for performing one or more steps depending on satisfying one or more conditions. Because such instructions are stored in one or more processors and / or one or more memory locations, the claims of the electronic device, system, or computer-readable medium may include logic for determining whether one or more conditions have been satisfied without requiring the repetition of process steps.
[0066] Although numerical descriptors such as "first" and / or "second" are used below to describe elements, these elements do not correspond to sequential or different representations and should not be limited to the stated numerical terms. In some embodiments, these terms are used only as prefixes to distinguish references to one element from references to another. In some embodiments, "first" device and "second" device may be two separate references to the same device. Conversely, in some embodiments, "first" device and "second" device may be references to two different devices (e.g., not the same device and / or not the same type of device). In some embodiments, the first computer system and the second computer system do not correspond to first and second in time and are merely used to distinguish the two computer systems. Therefore, without departing from the scope of the various described embodiments, the first computer system may be referred to as the second computer system, and the second computer system may be referred to as the first computer system.
[0067] In the description of various elements and examples, the use of certain terms is intended to provide a productive description of the following topics and should not be construed as limiting. As used in describing the various examples herein, the singular forms “a,” “an,” and “the” should not be construed as excluding or precluding the plural forms unless the context clearly indicates otherwise. Similarly, “and / or” is used to cover any and all possible combinations of one or more associated listed items. In some embodiments, “x and / or y” should be interpreted as including “x” or “y” as well as “x and y” as possible permutations. Furthermore, the terms “includes,” “including,” “comprises,” and / or “comprising” used in this specification specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0068] When describing choices and / or logical possibilities, the term "if" may optionally be interpreted, depending on the context, as meaning "when," "in response to determination," "in response to detection," or "according to determination." Similarly, depending on the context, the phrases "if determination..." or "if [the stated condition or event] is optionally interpreted as meaning "in response to determination," "in response to detection," "in response to detection," or "according to determination."
[0069] The processes described below enhance device operability and make user-device interfaces more efficient through various technologies (e.g., by helping users provide correct input and reducing user errors when operating / interacting with the device). These technologies include: providing users with improved feedback (e.g., visual, tactile, auditory, and / or haptic feedback); reducing the number of inputs required to perform an operation; providing additional control options without cluttering the user interface with additional displayed controls; performing an operation without requiring further input (e.g., user input) when a set of conditions are met; and / or other technologies (such as improving the security and / or privacy of the computer system and reducing the aging of one or more parts of the display user interface). These technologies also reduce power consumption and extend device battery life by enabling users to use the device more quickly and efficiently.
[0070] under, Figure 1 , Figures 2A to 2C and Figures 3 to 5 A description of an exemplary device for performing the techniques described herein is provided. Figures 6A to 6G An exemplary user interface for changing the display of an object is illustrated according to some implementation schemes. Figure 7 This is a flowchart illustrating a method for displaying an object facing a certain direction, based on some implementation schemes. Figure 8 This is a flowchart illustrating a method for indicating eye contact with a display object according to some implementation schemes. Figure 9 This is a flowchart illustrating a method for removing an emphasized object according to some implementation schemes. Figures 6A to 6G The user interface in the document is used to illustrate the processes described below, including Figure 7 , Figure 8 and Figure 9 The process in. Figures 10A to 10F Exemplary user interfaces for providing content are illustrated according to some implementation schemes. Figure 11 This is a flowchart illustrating a method for displaying a digital image of a system according to some implementation schemes. Figure 12 This is a flowchart illustrating a method for selectively moving a portion of a computer system according to some implementation schemes. Figure 13 This is a flowchart illustrating a method for navigating content according to some implementation schemes. Figure 14 This is a flowchart illustrating a method for pausing content according to some implementation schemes. Figures 10A to 10F The user interface in the document is used to illustrate the processes described below, including Figure 11 , Figure 12 , Figure 13 and Figure 14 The process in.
[0071] Figure 1A block diagram depicts a computer system 100 (e.g., an electronic device and / or electronic system) comprising a set of electronic components communicating (e.g., connected) with each other (e.g., wired or wireless). It should be understood that computer system 100 is merely one example of a computer system that can be used to perform the functions described below, and one or more other computer systems can be used to perform the functions described below. Furthermore, although... Figure 1 The computer architecture of computer system 100 is described, but other computer architectures of computer systems (e.g., including more components, similar components, and / or fewer components) may be used to perform the functionality described herein.
[0072] In some implementations, computer system 100 may correspond to (e.g., is and / or includes) a system-on-a-chip, a server system, a personal computer system, a smartphone, a smartwatch, a wearable device, a tablet computer, a laptop computer, a fitness tracker, a head-mounted display (HMD) device, a desktop computer, public equipment (e.g., smart speakers, connected thermostats and / or additional home-based computer systems), accessories (e.g., switches, lights, speakers, air conditioners, heaters, window covers, fans, locks, media playback devices, televisions, etc.), controllers, hubs and / or sensors.
[0073] In some embodiments, the sensor includes one or more hardware components capable of detecting (e.g., sensing, generating, and / or processing) information about the physical environment near the sensor. In some embodiments, the sensor may be configured to detect information around the sensor, detect information in one or more directions extending outward from the sensor, and / or detect information based on contact between the sensor and elements of the physical environment. In some embodiments, the hardware components of the sensor include sensing components (e.g., temperature and / or image sensors), transmitting components (e.g., radio and / or laser transmitters), and / or receiving components (e.g., laser and / or radio receivers). In some implementations, the sensors include angle sensors, breakage sensors, flow sensors, force sensors, gas sensors, humidity or moisture sensors, glass breakage sensors, chemical sensors, contact sensors, non-contact sensors, image sensors (e.g., RGB cameras and / or infrared sensors), particle sensors, photoelectric sensors (e.g., ambient light and / or sunlight), positioning sensors (e.g., GPS), precipitation sensors, pressure sensors, proximity sensors, radiation sensors, inertial measurement units, leak sensors, liquid level sensors, metal sensors, microphones, motion sensors, distance or depth sensors (e.g., RADAR, LiDAR), speed sensors, temperature sensors, time-of-flight sensors, torque sensors, ultrasonic sensors, vacancy sensors, presence sensors, voltage and / or current sensors, conductivity sensors, resistivity sensors, capacitance sensors, and / or water sensors. Although in Figure 1 Only a single computer system is depicted, but the functionality described below can be implemented using two or more computer systems operating together. Additionally, in some embodiments, computer system 100 includes one or more sensors as described above, and captures information about the physical environment by combining data from one sensor with data from one or more additional sensors (e.g., which are part of the computer and / or one or more additional computer systems).
[0074] like Figure 1As illustrated, computer system 100 comprises processor subsystem 110, memory 120, and I / O interface 130. Memory 120 corresponds to system memory that communicates with processor subsystem 110. Electronic components constituting computer system 100 are electrically connected via interconnect 150, which allows communication between components of computer system 100. In some embodiments, interconnect 150 may be a system bus, one or more memory locations, and / or additional electrical channels for connecting multiple components of computer system 100. Additionally, I / O interface 130 is connected to I / O device 140 via wired and / or wireless connections. In some embodiments, computer system 100 includes a component comprising I / O interface 130 and I / O device 140, such that the functionality of each component is included within that component. Furthermore, it should be understood that computer system 100 may include one or more I / O interfaces that communicate with one or more I / O devices. In some embodiments, computer system 100 comprises multiple processor subsystems 100s, each processor subsystem being electrically connected via interconnect 150.
[0075] In some embodiments, processor subsystem 110 includes one or more processors or separate processing units capable of executing instructions (e.g., programs, systems, and / or interrupts) to perform the functionality described herein. In some embodiments, operating system-level and / or application-level instructions are executed by processor subsystem 110. In some embodiments, processor subsystem 110 includes one or more components (e.g., implemented as hardware, software, and / or a combination thereof) capable of supporting, interpreting, and / or executing machine learning instructions and / or operations. In some embodiments, computer system 100 may perform operations locally based on a machine learning model. Alternatively or additionally, computer system 100 may communicate with (e.g., perform computations thereto and / or execute corresponding instructions) a remote interactive knowledge base (e.g., processing resources implementing machine learning models, artificial intelligence models, and / or large language models) to perform operations that may otherwise be outside the set of capabilities of computer system 100. In some embodiments, computer system 100 may determine a set of inputs (e.g., instructions, data, and / or parameters) to an interactive knowledge base for performing desired machine learning operations.
[0076] The memory 120, which communicates with the processor subsystem 110, can be implemented using a variety of different physical, non-transitory memory media. In some embodiments, the computer system 100 includes multiple memory components and / or various types of memory components, each of which is directly and / or connected to the processor subsystem 110 via interconnect 150. In some embodiments, the memory 120 can be implemented using removable flash drives, storage arrays, storage area networks (e.g., SANs), flash memory, hard disk storage devices, optical drive storage devices, floppy disk storage devices, removable disk storage devices, random access memory (e.g., SDRAM, DDR SDRAM, RAM-SRAM, EDO RAM, and / or RAMBUS RAM) and / or read-only memory (e.g., PROM and / or EEPROM). Additionally, in some embodiments, the processor subsystem 110 and / or interconnect 150 are connected to a memory controller, which is electrically connected to the memory 120.
[0077] In some embodiments, the instructions may be executed by processor subsystem 110. In this example, memory 120 may include a computer-readable medium (e.g., a non-transitory or transient computer-readable medium) that can be used to store (e.g., configured to store, assigned to store, and / or store) instructions executable by processor subsystem 110. In some embodiments, each instruction stored by memory 120 and executed by processor subsystem 110 corresponds to an operation for performing the functionality described herein. In some embodiments, memory 120 may store program instructions to implement the methods described below (including methods 700, 800, and 900). Figure 7 , Figure 8 and Figure 9 Related functionality.
[0078] As mentioned above, I / O interface 130 may be one or more types of interfaces that enable computer system 100 to communicate with other devices. In some embodiments, I / O interface 130 includes a bridge chip (e.g., a southbridge) connecting the front-side bus to one or more back-side buses. In some embodiments, I / O interface 130 enables communication with one or more I / O devices (exemplified as I / O device 140) via one or more corresponding buses or other interfaces. In some embodiments, I / O devices may include one or more of the following: physical user interface devices (e.g., physical keyboard, mouse, and / or joystick), storage devices (e.g., as described above with respect to memory 120), network interface devices (e.g., to a local area network or wide area network), sensor devices (e.g., as described above with respect to sensors), and / or auditory and / or visual output devices (e.g., screens, speakers, lamps, and / or projectors). In some embodiments, the visual output device is referred to as a display component. In some embodiments, the display component may be configured to provide visual output, such as displaying images on a physical visual medium via an LED display or image projection. As used herein, “display” content includes content that is displayed by sending data (e.g., image data and / or video data) to an integrated or external display component via a wired or wireless connection to visually generate content (e.g., video data rendered and / or decoded by a display controller).
[0079] In some embodiments, computer system 100 includes a component that integrates I / O device 140 with other components (e.g., a component including I / O interface 130 and I / O device 140). In some embodiments, I / O device 140 is separate from other components of computer system 100 (e.g., it is a discrete component). In some embodiments, I / O device 140 includes a network interface device that allows computer system 100 to connect to a network or other computer system (e.g., communicate with it) via wired or wireless means. In some embodiments, the network interface device may include Wi-Fi, Bluetooth, NFC, USB, Thunderbolt, Ethernet, etc. In some embodiments, computer system 100 may utilize NFC connectivity to facilitate banking, credit, financial, token (e.g., fungible or non-fungible tokens) and / or cryptocurrency transactions between computer system 100 and another nearby computer system.
[0080] In some embodiments, I / O device 140 includes components for detecting user (e.g., a person, animal, another computer system different from the computer system, and / or object) and / or input from the detected user (e.g., tap input and / or non-tap input (e.g., verbal input, acoustic request, acoustic command, acoustic statement, swipe input, hold and drag input, gaze input, air gesture, and / or mouse click)). In some embodiments, I / O device 140 enables computer system 100 to identify users associated with and / or not having accounts within the environment. In some embodiments, computer system 100 may detect known users (e.g., users corresponding to accounts) and access information about those users using their accounts. In some embodiments, as part of computer system 100's user detection, computer system 100 detects that a user's account is associated with a group of users (e.g., included in and / or identified relative to that group of users). In some embodiments, computer system 100 may access information associated with a family account in response to detecting a member of a family defined as a group of accounts. In some implementations, the user's account may be connected to additional accounts and / or additional computer systems. In some implementations, computer system 100 may detect such additional computer systems and / or detect such computer systems to detect the user. In some implementations, computer system 100 detects unknown users and enables guest accounts for unknown users to use computer system 100.
[0081] In some embodiments, I / O device 140 includes one or more cameras. In some embodiments, the camera includes an image sensor (e.g., one or more optical sensors and / or one or more depth camera sensors) that provides computer system 100 with the ability to detect user and / or user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)). In some implementations, one or more cameras enable computer system 100 to send image and / or video information to an application. In some implementations, image data captured by the camera enables computer system 100 to complete a video call by sending video data to an application that performs the video call.
[0082] In some embodiments, I / O device 140 includes one or more microphones. In some embodiments, the microphones can be used by 100 to obtain data and / or information from a user without contact input. In some embodiments, the microphones enable computer system 100 to detect verbal and / or voice input from a user. In some embodiments, computer system 100 utilizes voice input to enable personal assistant functionality. In some embodiments, the user makes requests to computer system 100 to perform actions and / or obtain information from the user. In some embodiments, computer system 100 utilizes voice input (e.g., in conjunction with one or more other input and / or output technologies) to request and / or detect information from a user without requiring physical contact between the user and computer system 100.
[0083] In some embodiments, I / O device 140 includes physical input media for a user to interact directly with computer system 100. In some embodiments, the physical input media includes one or more physical buttons (e.g., tactilely pressable buttons and / or touch-sensitive non-pressable components) on and / or connected to computer system 100, mouse and keyboard input methods (e.g., connected to computer system 100 together with and / or separately from one or more I / O interfaces), and / or touch-sensitive display components.
[0084] In some embodiments, I / O device 140 includes one or more components for outputting information (e.g., display components, audio generation components, speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, computer system 100 uses I / O device 140 to transmit information and / or the state of computer system 100. In some embodiments, I / O device 140 includes haptic output components. In some embodiments, the haptic output components may be haptic generation components that enable computer system 100 to convey information to a user who is in contact with computer system 100 (e.g., holding, touching, and / or in the vicinity). In some embodiments, I / O device 140 includes one or more components for outputting visual output (e.g., video, images, animations, 3D rendering, augmented reality overlay, motion graphics, data visualization, digital art, etc.). In some embodiments, content from one or more applications and / or system applications is displayed, and / or desktop applets corresponding to one or more applications are displayed (e.g., controls that display real-time information and / or data).
[0085] In some embodiments, I / O device 140 includes one or more components for outputting audio (e.g., smart speaker, home theater system, soundbar, headphones, earphones, earbuds, speaker, TV speaker, augmented reality headphone speaker, audio jack, optical audio output, Bluetooth audio output, HDMI audio output, audio sensor, etc.). In some embodiments, computer system 100 is capable of outputting audio through one or more speakers. In some embodiments, computer system 100 outputs audio-based content and / or information to a user. In some embodiments, one or more speakers enable spatial audio (e.g., audio output corresponding to the environment (e.g., computer system 100 detects materials and / or objects in the environment and / or computer system 100 changes audio modes, intensities, and / or waveforms to compensate for changing environmental characteristics)).
[0086] Figure 2 to Figure 5Exemplary components and user interfaces of an electronic device 200 according to some embodiments are illustrated. The electronic device 200 (sometimes referred to herein as device 200) may include one or more features of the computer system 100. (Referring to Figures 2 to...) Figure 5 In the described example, device 200 is a laptop computer. In some embodiments, device 200 is not limited to a laptop computer, and those skilled in the art will recognize that device 200 can be one or more other devices (e.g., one or more of the components and / or functions described herein with respect to device 200). In some embodiments, device 200 can be a utility device (such as a smart display, smart speaker, and / or television) and / or a personal device (such as a smartphone, smartwatch, tablet, desktop computer, fitness tracker, and / or head-mounted display). In some embodiments, the utility device is configured to provide functionality to multiple users (e.g., simultaneously and / or at different times). In such embodiments, the utility device can be managed and / or set by a single user. In some embodiments, the personal device is configured to provide functionality to a single user (e.g., once, such as when a single user logs into the personal device).
[0087] Figures 2A to 2C An example is shown of a device 200 located in three different physical locations. For example... Figure 2A As illustrated, device 200 is a laptop computer (also referred to herein as a "laptop"), which includes a base portion 200-2 (e.g., as shown in the image). Figure 2A The device 200 is horizontally placed on a surface such as a table and connected to a base portion 200-2 at a connection 200-3 (e.g., one or more connection points, motor arms, hinges, and / or joints). This connection allows the display portion 200-1 to pivot and / or change orientation relative to the base portion 200-2. In some embodiments, the device 200 may pivot at the connection 200-3 to rotate the display portion 200-1 and / or the device 200 to one or more positions corresponding to a “closed” internal state (e.g., as described below regarding...). Figure 2C(Further description). In some embodiments, the positioning corresponding to the "off" internal state is the positioning of the device 200 in a predetermined pose. In some embodiments, the predetermined pose may include a display portion 200-1 positioned parallel to the base portion 200-2 or forming a predetermined angle (e.g., 60 degrees) with respect to the base portion 200-2. In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., facing downwards, not visible, and / or obscuring the displayed content). In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is not positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., instead positioned in a manner corresponding to the "on" internal state). In some embodiments, when not in a "closed" internal state, device 200 can be positioned within a range of different open positions (e.g., where display portion 200-1 is not parallel to base portion 200-2, and where the area where the content displayed by device 200 is visible and / or unobstructed). It should be recognized that display portion 200-1 being parallel to base portion 200-2 is an example of positioning corresponding to a "closed" internal state of device 200 (e.g., closed positioning). In some embodiments, another configuration may set another orientation of display portion 200-1 relative to base portion 200-2 as a closed positioning of device 200, such as... Figure 2C exemplified.
[0088] Figure 2A The left side illustrates display screen 200-4 (representing the area where device 200 displays content), and the right side illustrates device 200 in the corresponding pose. For example... Figure 2A As illustrated, device 200 is in a first position (e.g., display portion 200-1 is perpendicular to base portion 200-2, forming a 90-degree angle). Figure 2A In this context, display screen 200-4 represents the content currently being displayed (e.g., via a display component) when device 200 is first activated. Figure 2AIn this embodiment, display screen 200-4 illustrates the device 200 in an "on" internal state (e.g., operable, powered, awake, higher power and / or more resource-intensive than the "off" state, and / or activated). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces (e.g., user interface objects, windows, application user interfaces, system user interfaces, controls, and / or other visual content). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces while in an "on" internal state. In some embodiments, in... Figure 2A In this configuration, device 200 is in an "on" internal state, and display screen 200-4 shows a desktop user interface 200-5, including an application window. In some embodiments, the user interface includes (and / or) one or more user interface objects (e.g., windows, icons, and / or other graphical objects). In some embodiments, the user interface (e.g., 200-5) may include one or more graphical objects that are different from and / or the same as the application window.
[0089] Figure 2B Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2B As illustrated, device 200 is in a second position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming an angle of 120 degrees (e.g., more than). Figure 2A (at a larger angle). Figure 2B In the diagram, display screen 200-4 represents the content being displayed when device 200 is in the second position. Display screen 200-4 illustrates the internal state of device 200 being "on" (e.g., with...). Figure 2A (The top diagram shows the same internal state). Figure 2B In the process, device 200 displays (e.g., via display screen 200-4) a desktop user interface 200-5 (e.g., with...). Figure 2A (The same as shown in the image). In some embodiments, device 200 displays a different user interface (e.g., different from desktop user interface 200-5). In some embodiments, although... Figure 2B Example of device 200 in a state of being with Figure 2A Different positioning displays and Figure 2A The same desktop user interface 200-5 exists, but device 200 may display different user interfaces. In some embodiments, device 200 displays a user interface corresponding to (e.g., based on, due to, caused by, involved in, and / or configured to accompany) a physical state (e.g., positioning, location, and / or orientation), including content specific to a particular angle or specific to the current context.
[0090] Figure 2C Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2C As illustrated, device 200 is in a third position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming a 60-degree angle (e.g., compared to...). Figure 2A and Figure 2B (smaller angles)). Figure 2C In the diagram, display screen 200-4 shows the content being displayed when device 200 is in the third position. Figure 2C In the diagram, displays 200-4 illustrate an internal state in which device 200 is "off" (e.g., not operating, not powered, not woken up, not activated, powered off, asleep, hibernating, inactive, and / or disabled). In some embodiments, device 200 does not display (e.g., via displays 200-4) one or more user interfaces (e.g., no visual content is displayed) when it is in the "off" internal state. In some embodiments, device 200 displays (e.g., via displays 200-4) one or more user interfaces (e.g., the same as and / or different from one or more user interfaces displayed when it is in the "on" internal state) (e.g., a user interface specific to the "off" state and / or a way of displaying a user interface not specific to the "off" internal state). Figure 2C In this case, display screen 200-4 is blank because nothing is displayed on the monitor of device 200 (e.g., display screen 200-4 is off and / or does not display the user interface) (e.g., desktop user interface 200-5 is not displayed on display screen 200-4).
[0091] In some embodiments, device 200 includes one or more components (referred herein also as “moving components”) that enable device 200 to perform (e.g., cause and / or control) movement (and / or be moved). In some embodiments, performing movement may include a portion of mobile device 200 (e.g., less or all of its components moving), the entire mobile device 200 (e.g., the entire device (including all its components) moving, such as by changing position), and / or moving one or more other devices and / or components (e.g., devices and / or components communicating with device 200 and / or the moving components of device 200). In some embodiments, device 200 may move automatically (e.g., pivot), cause and / or control movement of display portion 200-1 relative to base portion 200-2, such as moving to... Figures 2A to 2CAny positioning shown. In some embodiments, device 200 performs movement based on its internal state. Performing movement based on internal state enables device 200 to perform new (e.g., otherwise unavailable) interactions. In some embodiments, such new interactions of device 200 can be configured using special features, functions, patterns, and / or procedures that utilize device 200's ability to perform movement. Examples of such interactions include using movement to (e.g., to a user) convey the device's internal state (e.g., on, off, sleep, and / or hibernate) to assist user input (e.g., shorten the distance to the user) and / or enhance the device's interactive behavior (e.g., moving in a specific manner during interaction with the user, conveying information such as importance and / or direction of attention). In some embodiments, the performed movement corresponds to (e.g., caused by, responded to, and / or determined and / or performed based on) one or more of the following: detected input, detected context (e.g., environmental context and / or user context), and / or the device 200's internal state (e.g., internal state and / or a set of multiple internal states). In some implementations, device 200 can move the display portion, causing device 200 to move from a position where... Figure 2A The first positioning shown moves to the position Figure 2B The second positioning is shown. In this example, device 200 can detect that the user has repositioned relative to device 200 (e.g., the user stands up), and in response, device 200 can perform a movement to the second positioning such that the display is at an optimized viewing angle based on the height and / or angle of the user's eye relative to the display of device 200. As another example, device 200 can perform a movement such that device 200 moves from a position where... Figure 2A The illustrated first positioning moves to the position where Figure 2C The illustrated third location. In this example, device 200 may perform a movement to a third location in response to detecting an internal state with reduced activity (e.g., an "off" internal state as described above). In this way, movement of device 200 to one or more locations can indicate the internal state of device 200.
[0092] Figures 2A to 2C An example is illustrated of a device 200 having a display portion capable of moving with one degree of freedom via a connection 200-3 (e.g., a hinge) connecting the display portion 200-1 to a base portion 200-2. In some embodiments, the device 200 includes one or more components having one or more degrees of freedom. In some embodiments, the moving components of the device 200 (e.g., output components that cause and / or allow movement) (e.g., Figure 5Device 200-26C may include multiple degrees of freedom (e.g., six degrees of freedom, including three translational components and three rotational components). In some embodiments, device 200 may be implemented to move the display portion in a telescopic forward or backward motion (e.g., the display portion 200-1 moves forward while the base portion 200-2 remains stationary in space relative to the base portion (e.g., to reduce and / or increase the user's viewing distance)). As yet another example, device 200 may be implemented to move the display portion to rotate about an axis perpendicular to the hinge, such that the display portion can rotate to position the display to follow the user as the user walks around device 200. Although Figures 2A to 2C The example shown illustrates a hinge, but other moving components may be included in device 200, such as actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases. In some embodiments, one or more moving components may enable device 200 to move in different ways, such as rotation (e.g., 0 to 360 degrees), lateral movement (e.g., to the right, left, down, up, and / or any combination thereof), and / or tilting (e.g., 0 to 360 degrees).
[0093] Figure 3 An exemplary block diagram of device 200 is illustrated. In some embodiments, device 200 includes... Figure 1 A, Figure 1 B. Figure 3 and Figure 5 B describes some or all of the components. For example... Figure 3 As illustrated, device 200 has a bus 200-13 that operatively couples I / O segments 200-12 (also referred to as I / O sub-segments and / or I / O interfaces) to processor 200-11 and memory 200-10. For example... Figure 3 As illustrated, I / O section 200-12 is connected to output device 200-16 (also referred to herein as "output component"). In some embodiments, output device 200-16 includes one or more visual output devices (e.g., display components such as monitors, displays, projectors, and / or touch-sensitive displays), one or more tactile output devices (e.g., devices that cause vibration and / or other tactile outputs), one or more audio output devices (e.g., speakers), and / or one or more moving components (e.g., actuators, motors, mechanical linkages, devices that cause and / or allow movement, and / or one or more moving components as described above). Figure 3As illustrated, output device 200-16 includes two exemplary moving components (e.g., a movement controller 200-17 and an actuator 200-18). Actuator 200-18 can be any component that performs (e.g., partial and / or overall) physical movement of a device (e.g., device 200 and / or devices coupled to and / or in contact with that device). Movement controller 200-17 can be any component (e.g., a control device) that controls actuator 200-18 (e.g., provides control signals to it). In some embodiments, movement controller 200-17 can provide control signals that cause actuator 200-18 to actuate (e.g., cause physical movement). In some embodiments, movement controller 200-17 includes one or more logic components (e.g., a processor), one or more feedback components (e.g., sensors), and / or one or more control components (e.g., for applying control signals, such as relays, switches, and / or control lines). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in the same device and / or component (e.g., a dedicated onboard motion controller 200-17 attached to the actuator 200-18). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in different devices and / or components (e.g., one or more processors 200-11 may serve as the motion controller 200-17 for the actuator 200-18). In some embodiments, the motion controller 200-17 and / or the actuator 200-18 are embodied in a device (or one or more devices) other than device 200 (e.g., device 200 is coupled to (e.g., temporarily and / or removably) another device and may instruct the motion controller 200-17 and / or the actuator 200-18 to control the other device). Actuator 200-18 can be used to induce one or more types of mechanical movement (e.g., linear and / or rotary movement) in one or more ways (e.g., using electric, magnetic, hydraulic and / or pneumatic power). Examples of actuator 200-18 may include electromechanical actuators, linear actuators and / or rotary actuators.
[0094] like Figure 3As illustrated, I / O section 200-12 is connected to input device 200-14. In some embodiments, input device 200-14 includes one or more visual input devices (e.g., cameras and / or light sensors), one or more physical input devices (e.g., buttons, sliders, switches, touch-sensitive surfaces, and / or rotatable input mechanisms), one or more audio input devices (e.g., microphones), and / or other input devices (e.g., accelerometers, pressure sensors (e.g., contact strength sensors), distance sensors, temperature sensors, GPS sensors, accelerometers, orientation sensors (e.g., compasses), gyroscopes, motion sensors, and / or biometric sensors). Furthermore, I / O section 200-12 may be connected to communication unit 200-15 for receiving application and operating system data using Wi-Fi, Bluetooth, Near Field Communication (NFC), cellular, and / or other wireless (and / or wired) communication technologies.
[0095] The memory 200-10 of device 200 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 200-11, cause the computer processors to perform the techniques described below, including processes 700, 800, 900, 1100, 1200, 1300, and 1400. Figure 7 , Figure 8 , Figure 9 , Figure 11 , Figure 12 , Figure 13 and Figure 14 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transient computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, and Blu-ray technologies, and persistent solid-state storage such as flash memory and solid-state drives. Electronic device 200 is not limited to Figure 3 The components and configurations may include, but may include, other and / or additional components in a variety of possible configurations, all of which are intended to fall within the scope of this disclosure.
[0096] Figure 4A functional diagram of actuator 200-18B according to some embodiments is illustrated. As described above, actuator 200-18B can be any component that performs physical movement. In some embodiments, actuator 200-18B is operated using inputs including control signal 200-18A and / or energy source 200-18B. In some embodiments, actuator 200-18B can be a rotary actuator that converts electrical energy into rotational movement. This rotational movement can cause the above-described... Figures 2A to 2C The movement of the display portion of the described device 200 (e.g., the counterclockwise rotation of the actuator moves the device 200 to a position with a large angle) Figure 2B The illustrated second positioning), and the clockwise (e.g., counterclockwise) rotational movement of the actuator moves the device 200 to a positioning with a smaller angle (e.g., Figure 2C (The illustrated third positioning). Control signal 200-18A may indicate one or more start and / or stop commands, movement and / or actuation direction, movement and / or actuation speed, movement and / or actuation time, target positioning (e.g., pose and / or position) of movement and / or actuation, and / or one or more other characteristics of movement and / or actuation. In some embodiments, the control signal and the energy source are the same signal and / or input. In some embodiments, one or more additional components (e.g., mechanical and / or electrical) (e.g., removably or permanently) are coupled to actuator 200-18B to influence movement and / or actuation (e.g., mechanical linkages such as lead screws, gears, and / or other components for changing (e.g., switching) the characteristics of movement and / or actuation). In some embodiments, actuator 200-18B includes one or more feedback components (e.g., a positioning sensor, an encoder, an overcurrent sensor, and / or a force sensor) that form part of a feedback loop for modifying and / or stopping movement and / or actuation (e.g., slowing down actuation upon reaching a target position and / or stopping actuation if physical resistance to actuation is detected via a sensor). In some embodiments, one or more feedback components are included (e.g., partially and / or entirely) in a motion controller (e.g., motion controller 200-13) operatively coupled to the actuator.
[0097] Now turn attention to the functionality (e.g., features and / or capabilities) of one or more devices (e.g., computer system 100 and / or electronic device 200). One such functionality is the implementation of an "intelligent agent," which may alternatively be referred to as a software intelligent agent, intelligent intelligent agent, interactive intelligent agent, virtual assistant, intelligent virtual assistant, interactive virtual assistant, personal assistant, intelligent personal assistant, interactive personal assistant, intelligent interactive personal assistant, and / or artificial intelligence (AI) assistant. In some embodiments, an intelligent agent refers to one or more sets of functions implemented in hardware and / or software (e.g., local and / or remote) on an intelligent agent system (e.g., a single device and / or multiple devices). In some embodiments, the intelligent agent performs operations to perceive the environment, acquire knowledge, retrieve knowledge, learn skills, interact with a user, and / or perform tasks. In some embodiments, the intelligent agent may perform these (and / or other) operations in response to user input and / or automatically (e.g., at an appropriate time determined based on the perceived context). An incomplete list of exemplary operations that an intelligent agent can be used with and / or employed therewith includes: tracking a user’s eyes, face, and / or body (e.g., to move with the user and / or identify the user’s intentions and / or activities); detecting, identifying, and / or classifying users in the environment; detecting and / or responding to input (e.g., verbal input, air gestures, and / or physical input, such as touch input and / or force input to physical hardware components (e.g., buttons, knobs, and / or sliders); detecting context (e.g., user context, operational context, and / or environmental context); moving (e.g., changing pose, orientation, orientation, and / or location); performing one or more operations in response to input, context, and / or stimuli (e.g., objects or events that elicit one or more responsive operations on the device (e.g., outside and / or inside the device)); providing intelligent interaction capabilities (e.g., in part due to one or more machine learning (“ML”) models, such as large language models (“LLM”)) to respond to and / or perform operations; and / or (e.g., automatically and / or intelligently) performing tasks (e.g., a set of operations for achieving a specific goal). In some implementations, the agent performs actions in response to contactless input (e.g., air gestures and / or natural language commands). The foregoing list is intended to exemplify actions that can be performed by an agent, but is not intended to be an exhaustive list. Other actions fall within the expected scope of the agent's capabilities. Furthermore, for the purposes of this disclosure, the agent need not include all the functionalities mentioned herein, but may include fewer or more functionalities (e.g., the agent may be implemented on an agent system that does not have mobile functionality but otherwise includes an intelligent personal assistant capable of interacting with a user).
[0098] In some embodiments, a user is one or more of a user, person, object, and / or animal in an environment (e.g., a device) that is perceived (e.g., by the device, including, and / or included in) its environment (e.g., a physical and / or virtual environment). In some embodiments, a user is an entity that is perceived (e.g., by the device, one or more other devices, and / or one or more components thereof). In some embodiments, an entity is something distinguishable from surrounding entities (e.g., components of the environment and / or other users) and / or something considered to be of discrete logic construction via one or more components (e.g., a sensing component and / or other components). In some embodiments, a user is physical and / or virtual. In some embodiments, a physical user may represent a user standing in front of the device and perceived by the device. As another example, a virtual user may represent a digital image in a virtual scene perceived by the device (e.g., a digital image detected in a media stream received by the device and / or captured by the device's camera). Although presented above as an example of “user,” throughout this disclosure, the terms and / or concepts referred to as “person,” “object,” and / or “animal” may be used interchangeably with “user” unless otherwise expressly indicated.
[0099] As an example, and to revisit Figures 2A to 2C An agent, at least partially implemented on device 200, can perform operations that cause the display portion 200-1 of device 200 to move relative to the base portion 200-2. In some embodiments, agent detection (e.g., perceiving and determining that it has occurred) includes the context of the user standing (e.g., based on face detection and tracking); and in response, the agent causes device 200 to open and / or device 200 to open the display portion 200-1 to a greater angle. As another example, the agent can detect verbal input corresponding to (e.g., interpreted as and / or referring to including) a request to move the display (e.g., “Please move my display” or “Please enter sleep mode”); and in response, the agent causes device 200 to move and / or device 200 to move the display portion 200-1.
[0100] Figure 5 A functional diagram of an exemplary intelligent agent system 200-20A is shown. Figure 5 As shown, the agent system 200-20A has a dashed box boundary that surrounds the input component 200-22, the agent component 200-24, and the output component 200-26. In some embodiments, the agent system 200-20A includes more than Figure 5Fewer, more, and / or different components are illustrated. In some embodiments, the agent system 200-20 is implemented on a single device (e.g., computer system 100 and / or device 200). In some embodiments, the agent system 200-20 is implemented on multiple devices. In some embodiments, in Figure 5 One or more components of the agent system 200-20 illustrated and / or described with respect to this figure are external to but operatively coupled to the agent system (e.g., accessories, external devices, external sensors, external actuators, external display components, external speakers, and / or external databases). In some embodiments, one or more components of the agent system 200-20 are local to one or more other components of the agent system 200-20. In some embodiments, one or more components of the agent system 200-20 are remote from one or more other components of the agent system 200-20.
[0101] In some implementations, input components 200-22 include components for performing sensing and / or communication functions of the agent system 200-20. For example... Figure 5 As illustrated, input components 200-22 include one or more sensors 200-22A. The one or more sensors 200-22A may include any components for detecting data corresponding to the physical environment. Examples of the one or more sensors 200-22A may include: cameras, light sensors, microphones, accelerometers, positioning sensors, pressure sensors, temperature sensors, olfactory sensors, and / or contact sensors. This list is not intended to be exhaustive, and the one or more sensors 200-22A may include other sensors not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to detect data corresponding to the physical environment. Figure 5 As illustrated, input component 200-22 includes one or more communication components 200-22B. The one or more communication components 200-22B may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the intelligent agent system 200-20. Communication components 200-22B may be between different devices and / or between components within the same device. Communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, input component 200-22 includes more than Figure 5 The components illustrated herein may be fewer, more, and / or different. In some implementations, input components 200-22 are implemented in hardware and / or software.
[0102] In some implementations, agent components 200-24 include components that manage and / or perform the functions of the agents in agent system 200-20. For example... Figure 5 As illustrated, agent components 200-24 include the following functional components: task flow, coordination and / or orchestration component 200-24A, management component 200-24B, perception component 200-24C, evaluation component 200-24D, interaction component 200-24E, policy and decision-making component 200-24F, knowledge component 200-24G, learning component 200-24H, model component 200-24I, and API component 200-24J. Each of these components is briefly described below. It is important to note that this list of agent components 200-24 is not intended to be exhaustive, and agent components 200-24 may include other functional components not explicitly identified herein, which may be used (e.g., to process, store, and / or transform) any function of the agent, such as those described herein. In some embodiments, agent components 200-24 include more than Figure 5 Fewer, more, and / or different components are illustrated. In some implementations, the agent components 200-24 are implemented in hardware and / or software.
[0103] In some embodiments, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between various components. In some embodiments, operations may include processing data processing task flows to move from perception components 200-24C (e.g., detecting speech input) to model components 200-24I (e.g., processing the detected speech input using a large language model to determine the content and / or intent of the speech input). In some embodiments, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between one or more external components (e.g., resources). In some embodiments, Figure 5 Examples of external components (such as external databases 200-30) are illustrated. In some embodiments, management component 200-24B includes functionality performed by the operating system of the device implementing the intelligent agent system 200-20. In some embodiments, management component 200-24B includes functionality performed by one or more applications of the device implementing the intelligent agent system 200-20.
[0104] In some embodiments, management components 200-24B perform operations that enable the agent system to handle management tasks, such as managing system and / or component updates, managing user accounts, and managing system settings and / or component settings. In some embodiments, management components 200-24B include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, management components 200-24B include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0105] In some embodiments, the sensing components 200-24C perform operations that enable the agent to perceive environmental input. In some embodiments, the operations may include detecting that context and / or environmental conditions have occurred, detecting the presence of a user (e.g., a user, person, object, and / or animal in the environment), detecting input including voice, detecting input including air gestures, detecting facial expressions, detecting user characteristics (e.g., visible and / or invisible), and / or detecting verbal and / or physical cues. In some embodiments, the sensing components 200-24C include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the sensing components 200-24C include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0106] In some embodiments, the evaluation component 200-24D performs operations that enable the agent to process evaluation data (e.g., to determine context, such as user context, environmental context, and / or operational context). In some embodiments, the operations may include evaluating data collected from the perception component 200-24C, the knowledge component 200-24G, the external database 200-30, and / or the teleprocessing resource 200-32. In some embodiments, the evaluation component 200-24D includes functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the evaluation component 200-24D includes functionality performed by one or more applications of the device implementing the agent system 200-20.
[0107] This document refers to environmental context (also referred to herein as "context of the environment" and / or "context corresponding to the environment"). In some embodiments, environmental context is a context based on one or more characteristics of the environment (e.g., user, location, time, weather, and / or lighting). In some embodiments, environmental context may include whether it is raining outside, whether it is daytime, and / or whether the device is currently located in a park. In some embodiments, the device (e.g., using an agent) uses one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device) to determine the environmental context (e.g., currently true, happening, and / or applicable).
[0108] This document refers to user context (also referred to herein as "user context" and / or "context corresponding to the user") (and / or user context). In some embodiments, user context is a context based on one or more characteristics of the user. In some embodiments, user context may include the user's appearance and / or clothing, personality, actions, behaviors, movement, location, and / or pose. In some embodiments, the device (e.g., using an intelligent agent) determines user context (e.g., currently true, happening, and / or applicable) using one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device determines user context based on historical context and / or learned user characteristics, wherein one or more user characteristics are learned and / or stored by the device over a period of time.
[0109] This document refers to an operational context (also referred to herein as "the context of operation" and / or "operational context"). In some embodiments, an operational context is a context based on one or more characteristics of the device's operation (e.g., the device and / or one or more other devices that determine and / or access the operational context). In some embodiments, an operational context may include the internal state of the device (and / or one or more components of the device), the device's internal dialogue (e.g., the device's understanding of the context), the operations performed by the device, and applications and / or processes executed on the device (e.g., running and / or opening). In some embodiments, the device (e.g., using an agent) uses one or more of the following to determine the operational context (e.g., currently true, happening, and / or applicable): detected input (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device (e.g., using an agent) uses one or more internal states (e.g., accessed, retrieved, and / or queried by the device's processes) to determine the operational context (e.g., currently true, happening, and / or applicable).
[0110] In some embodiments, the interaction components 200-24E perform operations that enable the agent to manage and / or perform interactions with a user. In some embodiments, the operations may include determining an appropriate interaction model for a specific context and / or in response to specific inputs. In some embodiments, the interaction components 200-24E include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the interaction components 200-24E include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0111] In some embodiments, the policy and decision components 200-24F perform operations that enable the agent to take actions based on available data. In some embodiments, the operations may include determining which operations to perform and / or which functional components to utilize in response to detected context. In some embodiments, the policy and decision components 200-24F include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the policy and decision components 200-24F include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0112] In some embodiments, knowledge components 200-24G perform operations that enable the agent to access and use stored knowledge. In some embodiments, operations may include indexing, storing, and / or retrieving data from data repositories, databases, and / or other resources. In some embodiments, knowledge components 200-24G include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, knowledge components 200-24G include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0113] In some embodiments, the learning component 200-24H performs operations that enable the agent to learn through experience. In some embodiments, the operations may include observation and / or tracking data, including preferences, routines, user characteristics, and / or environmental characteristics, in a way that such data can be used to inform future actions of the agent and / or its components (e.g., when performing tasks and / or interacting with a user). In some embodiments, the learning component 200-24H includes functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the learning component 200-24H includes functionality performed by one or more applications of the device implementing the agent system 200-20.
[0114] In some embodiments, model components 200-24I perform operations that enable the agent to apply an ML model (e.g., a large language model (LLM)) to process data. In some embodiments, the operations may include storing the ML model, executing the ML model, training and / or retraining the ML model, and / or otherwise managing aspects of implementing the ML model. In some embodiments, model components 200-24I include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, model components 200-24I include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0115] In some implementations, the agent system 200-20 responds to natural language input. For example, the agent system 200-20 responds to natural language input in the form of statements, questions, commands, and / or requests. In some implementations, the agent system 200-20 outputs text and / or speech output provided in natural language or mimicking a natural language style. For example, the agent system 200-20 may use a speech response indicating the current outside temperature at the user's location (e.g., "It's 18 degrees outside") to handle the natural language question "How hot is it outside?". In some implementations, the agent system 200-20 responds to natural language input by providing information (e.g., weather, travel, and / or calendar information) and / or performing tasks (e.g., opening a document, searching a database, and / or opening an application).
[0116] In some implementations, the intelligent agent system 200-20 includes and / or relies on one or more data models to process inputs (e.g., natural language input, gesture input, visual input, and / or other data input) and / or provide outputs (e.g., information output via natural language output, visual output, audio output, and / or text output). Such data models may include user data (e.g., data based on a specific interaction and / or from the user with whom the interaction took place) and / or global data (e.g., general data based on the interaction and / or data from many users) and / or be trained using user data and / or global data. For example, user data (e.g., preferences, prior use of language and / or phrases, calendar entries, contact lists, and / or activity data) can be used to better infer user intent and / or provide responses more likely to resolve user requests. In some implementations, the data models used by the intelligent agent system 200-20 include one or more machine learning components (e.g., hardware and / or software) (e.g., one or more neural networks), are used by one or more machine learning components, and / or are implemented using one or more machine learning components. Such machine learning components can be used to process spoken input to determine words and / or phrases therein, one or more contexts corresponding to the words, user intent corresponding to the words, one or more confidence scores, and / or a set of one or more actions to be taken in response to the spoken input. Similar operations can be performed to process other types of input, such as visual input, data input, and / or text input. Such data models may include machine learning and / or data processing models, including but not limited to natural language processing models, language models, speech recognition models, object recognition models, visual processing models, ontology, task flow models, and / or intent recognition models (e.g., for determining user intent).
[0117] In some implementations, application programming interface (API) components 200-24J perform operations that enable the agent to interface with services, devices, and / or components. In some implementations, operations may include relaying data (e.g., requests, responses, and / or other messages) between data interfaces (e.g., between software programs, between system processes and application processes, between system processes, between application processes, between communication protocols, between clients and servers, between file systems, and / or between components on different sides of a trust boundary). In some implementations, the data interfaces served by API components 200-24J are local (e.g., for a device, such as two application processes exchanging data) and / or remote (e.g., from a device, such as interfaced with a web service via a remote server). In some implementations, API components 200-24J include functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, API components 200-24J include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0118] In some implementations, output components 200-26 include components for performing the output functions of the agent system 200-20. A brief description follows. Figure 5 The exemplary output components are illustrated herein. In some embodiments, output components 200-26 include... Figure 5 The components illustrated may include fewer components, more components, and / or different components. In some implementations, the input components are implemented in hardware and / or software.
[0119] like Figure 5 As illustrated, output components 200-26 include one or more visual output components 200-26A. One or more visual output components 200-26A may include any component used for outputting (e.g., generating, creating, and / or displaying) and / or causing visual output (e.g., visually perceptible output, such as a graphical user interface, playback of visual media content, and / or lighting). Examples of one or more visual output components 200-26A may include: display components, projectors, head-mounted displays (HMDs), light-emitting diodes (“LEDs”), and / or components that create visually perceptible effects (e.g., movement). This list is not intended to be exhaustive, and one or more visual output components 200-26A may include other visual output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output visual output.
[0120] like Figure 5As illustrated, output components 200-26 include one or more audio output components 200-26B. One or more audio output components 200-26B may include any component for outputting (e.g., generating and / or creating) and / or causing audio output (e.g., audibly perceptible output, such as sound, music, speech, and / or audio media content). Examples of one or more audio output components 200-26B may include: speakers, audio amplifiers, tone generators, and / or components that produce audibly perceptible effects (e.g., movement, such as vibration). This list is not intended to be exhaustive, and one or more audio output components 200-26B may include other audio output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) the output audio output.
[0121] like Figure 5 As illustrated, output components 200-26 include one or more motion output components 200-26C (also referred to herein as "motion components"). One or more motion output components 200-26C may include any component for outputting (e.g., generating and / or creating) and / or causing motion output (e.g., output including physical movement of a device and / or another device / component). Examples of one or more motion output components 200-26C may include: motion controllers, actuators, mechanical linkages, electromechanical devices, and / or components that generate physical movement. This list is not intended to be exhaustive, and one or more motion output components 200-26C may include other motion output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output motion output. Figure 5 As illustrated, output components 200-26 include one or more haptic output components 200-26D. One or more haptic output components 200-26D may include any component for outputting (e.g., generating, creating, and / or displaying) and / or causing haptic output (e.g., output using haptically perceptible means, such as vibration, pressure, texture, and / or shape). Examples of one or more haptic output components 200-26D may include: speakers, components that generate vibrations, components that generate texture changes, components that generate pressure changes, and / or components that create perceptible haptic effects. This list is not intended to be exhaustive, and one or more haptic output components 200-26D may include other haptic output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output haptic output.
[0122] like Figure 5As illustrated, output components 200-26 include one or more communication components 200-26E. The one or more communication components 200-26E may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the agent system 200-20. In some embodiments, communication may be between different devices and / or between components within the same device. In some embodiments, communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, the one or more communication components 200-26E include one or more features of one or more communication components 200-22B (e.g., as described above). In some embodiments, the one or more communication components 200-26E are identical to one or more communication components 200-22B (e.g., handling communication inputs and outputs and therefore considered as one or more of input and output components and / or both).
[0123] Throughout this disclosure, reference may be made to moving output (e.g., referred to in various forms such as: movement, device movement, moving output, device motion, motion output, and / or motion output). In some embodiments, output movement (e.g., an output that causes movement) refers to movement of an electronic device (e.g., a portion or component thereof relative to another portion and / or the entire electronic device). In some embodiments, reference is made again to... Figure 2B The movable output can refer to the device 200 actuating the movable component 200-3 to move the display part 200-1 to... Figure 2B The illustrated location (e.g., from) Figure 2A (Positioning within). In some embodiments, the motion output is not (e.g., excluding and / or not only including) tactile output (e.g., tactile motion output). In some embodiments, the motion output is not (e.g., excluding and / or not only including) vibration output. In some embodiments, the motion output is not (e.g., excluding and / or not only including) oscillatory motion (e.g., movement of an actuator that causes vibration solely by repeatedly moving the component along a path within the device). In some embodiments, the motion output includes (e.g., requiring and / or causing) a change in the position and / or pose of at least a portion (and / all) of a component or electronic device. In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device from a first position and / or a first pose to a second position and / or a second pose. In some embodiments, relative to Figures 2A to 2C ,exist Figure 2A , Figure 2B and Figure 2CIn each of these embodiments, display portion 200-1 is shown in a different position (e.g., in space) and pose (e.g., relative to base portion 200-2). In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device to a third position and / or a third pose (e.g., from a first position and / or a first pose and / or from a second position and / or a second pose). In some embodiments, the third position and / or the third pose is the same as the first position and / or the first pose and / or the second position and / or the second pose. In some embodiments, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement is made to return to... Figure 2A The illustrated first positioning. In some implementations, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement continues until it stops at... Figure 2C The illustrated third position.
[0124] Throughout this disclosure, electronic devices can be exemplified (and / or described) as being in different positions and / or poses at different times. In some embodiments, Figure 2A Example of device 200 in the first position, Figure 2B An example is shown of device 200 in the second position, and Figure 2A A device 200 in a third position is illustrated. In some embodiments, the electronic device moves itself between such positions and / or poses (e.g., using a movement output). In some embodiments, the device 200 moves from a first position to a second position under its own power (e.g., using a power supply and one or more actuators to induce movement). Specifically, any examples of electronic devices illustrated and / or described herein in different positions and / or poses (e.g., at different times) should be understood to cover scenarios where the device moves itself between such positions and / or poses (e.g., unless otherwise explicitly stated).
[0125] Throughout this disclosure, reference may be made to “performing output,” “causing output,” and / or “output” (e.g., via one or more output generating devices and / or via one or more output generating components) (and / or similar phrases). In some embodiments, the output (e.g., or variations thereof) includes (and / or) output movement (e.g., moving the output as described above).
[0126] Throughout this disclosure, references may be made to “display,” “cause display,” and / or “output visual content” (e.g., via one or more display components) (and / or similar phrases). In some embodiments, display (e.g., or variations thereof) includes displaying visual content in conjunction with output movement (e.g., moving output as described above).
[0127] Throughout this disclosure, reference may be made to "output audio," "output that causes audio," and / or "provide audio output" (e.g., via one or more audio generation components and / or via one or more audio output devices) (and / or similar phrases). In some embodiments, outputting audio (e.g., or variations thereof) includes outputting audio content in conjunction with output movement (e.g., movement output as described above).
[0128] Throughout this disclosure, reference may be made to the movement (and / or similar phrases) of a digital image (e.g., or other representations of a displayed user, agent, and / or role) (e.g., via one or more display components). In some embodiments, moving a digital image (e.g., or a variant thereof) includes movement in conjunction with output movement (e.g., movement output as described above) to display visual content. In some embodiments, displaying a digital image nodding in agreement may include movement of an electronic device in a manner similar to the movement of the digital image (e.g., simulated nodding). In some embodiments, moving a digital image (e.g., or a variant thereof) includes movement of output movement (e.g., movement output as described above) without displaying visual content. In some embodiments, a device may perform simulated nodding without moving the displayed digital image's movement output (e.g., the digital image does not move relative to the display). Figure 5As illustrated, agent system 200-20 may optionally interface with external components such as external database 200-30, remote processing component 200-32, and / or remote management component 200-34. In some embodiments, external database 200-30 represents one or more functions that provide data storage resources accessible to agent system 200-20. In some embodiments, access to data in external database 200-30 is provided directly to agent system 200-20 (e.g., the agent system manages the database) and / or indirectly to agent system 200-20 (e.g., the database is managed by a different system, but the data stored therein can be provided and / or stored for use by agent system 200-20). In some embodiments, external database 200-30 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., is a database of web services accessible to different agent systems), and / or a combination of dedicated and non-dedicated database resources. In some embodiments, remote processing component 200-32 represents one or more components that serve as data processing resources accessible to agent system 200-20. In some embodiments, access to remote processing component 200-32 is provided directly to agent system 200-20 (e.g., the agent system manages the processing resources) and / or indirectly to agent system 200-20 (e.g., processing resources managed by a different system, but which can provide data processing for the benefit of agent system 200-20). In some embodiments, remote processing component 200-32 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., a processing resource of a web service accessible to a different agent system), and / or a combination of both dedicated and non-dedicated processing resources. Examples of data processing include processing image data (e.g., for feature extraction and / or object detection), processing audio data (e.g., for processing natural language speech input via a large language model), and / or training machine learning algorithms and / or models. In some implementations, remote management component 200-34 represents management functions and / or functions related to management functions. In some implementations, such management functions may include providing component updates (e.g., software and / or firmware updates) to agent system 200-30, managing accounts (e.g., associated permissions, access controls, and / or preferences), synchronizing between different agent systems and / or their components (e.g., enabling agents accessible via multiple devices of a user to provide a consistent user experience across these devices), managing cooperation with other services and / or agent systems, error reporting, managing backup resources to maintain agent system reliability and / or agent availability, and / or other functions required for agent system 200-20 to perform operations, such as those described herein.
[0129] The above text is about Figure 5 The various components of the described intelligent agent system 200-20 represent functional blocks that represent functionality. This functionality may be implemented on the same and / or different hardware (e.g., physical components) and / or by the same and / or different software. In some embodiments, a functional block may be implemented using one or more physical components, devices (e.g., computer system 100 and / or device 200), and / or software programs. In other words, each functional block does not necessarily represent a single, discrete physical component, device, and / or software program, but may be implemented using one or more of these. Furthermore, the intelligent agent system 200-20 may include multiple implementations of the functionality represented by the respective functional blocks. In some embodiments, the intelligent agent system 200-20 may include multiple different model components representing ML models used in different contexts, multiple different API components representing different APIs for different services, and / or multiple different visual output components for outputting different types of visual output.
[0130] Now let’s turn our attention to a discussion of the concepts that may arise regarding the operation of intelligent agents.
[0131] As discussed throughout, the agent may be able to interact with the user. In some implementations, this capability includes the ability to process explicit requests, commands, and / or statements. In some implementations, explicit requests, commands, and / or statements include and / or are interpreted as instructions relating to completing a task (e.g., displaying X, completing task Y, and / or performing operation Z). In some implementations, the agent includes the ability to process implicit requests, commands, and / or statements. In some implementations, implicit requests, commands, and / or statements do not include explicit requests, commands, and / or statements. In some implementations, “I like to go to Europe” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays the itinerary in response to the statement. As another example, “This picture is for my grandmother” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays a suggestion to modify the picture. As another example, “I am tired” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 initiates a meditation session with a sleep meditation application. As another example, "I miss my grandfather" can be interpreted as an implicit request, command, and / or statement, which, upon detection, device 200 can initiate a real-time communication session with the grandfather (e.g., a phone call, video call, and / or text messaging session). In some implementations, implicit requests are more likely to be processed based on one or more current contexts, operational contexts, and / or user contexts, while explicit requests are less likely to be processed based on one or more current contexts, operational contexts, and / or user contexts. In some implementations, the phrase "call my grandfather" can be an explicit request, and in response to detecting this request, device 200 will initiate a real-time communication session with the grandfather, regardless of one or more current contexts, operational contexts, and / or user contexts. However, the phrase "I miss my grandfather" can be an implicit request, and in response to detecting this request, device 200 can display a list of gifts to buy for the grandfather if the user has recently been discussing buying gifts, or can call the grandfather in a different context that does not include the user's recent discussions about buying gifts. In some implementations, the request can include one or more explicit requests and one or more implicit requests. In some implementations, implicit requests are responded to independently of explicit requests; in other implementations, responses to implicit requests depend on explicit requests.
[0132] This document may refer to responses of an intelligent agent output by a device. In some embodiments, the response includes an audio component (e.g., audio output, acoustic output, sound and / or speech) (also referred to herein as a “verbal response,” “audio response,” and / or “acoustic response”) and / or a visual component (e.g., display and / or movement of representations and / or digital images). In some embodiments, the response includes a motion component (e.g., movement of the device). In some embodiments, the response includes a tactile component (e.g., touch and / or vibration).
[0133] This document may refer to internal dialogue, internal context, and / or operational context, which may refer to the dynamic context or dynamic decision-making process of a device, the internal state of device 200, and / or internal data of the device based in part on its decisions. In some embodiments, internal dialogue includes a set of one or more rules, features, detections, and / or observations used by a computer system to generate responses to one or more commands, questions, and / or statements. In some embodiments, the set of one or more rules, features, detections, and / or observations is learned and / or generated via deep learning and / or one or more machine learning algorithms and / or using one or more machine learning and / or system agents. In some embodiments, internal dialogue is generated in real time. In some embodiments, internal dialogue is stored locally and / or via cloud storage. In some embodiments, internal dialogue can be modified, updated, and / or deleted. In some embodiments, internal dialogue is generated based on other internal dialogues.
[0134] This document may refer to (e.g., the personality and / or behavior of an agent, user, and / or role) of a person or entity (or a representation of personality / behavior). In some embodiments, personality and / or behavior refers to one or more characteristics that a device detects, understands, conforms to, applies, and / or tracks. In some embodiments, personality or behavior is used as the basis for performing operations. In some embodiments, an agent may detect a user's personality and respond in a personality-based manner (e.g., outputting different responses in response to different user personalities). As another example, an agent may output responses having characteristics corresponding to one or more characteristics corresponding to personality and / or behavior (e.g., outputting responses in different ways depending on the agent's personality). In some embodiments, such characteristics represent and / or simulate a user's personality, such as how a user acts and / or speaks. In some embodiments, such characteristics approximate a user's personality.
[0135] In some embodiments, the intelligent agent is a system intelligent agent. In some embodiments, the system intelligent agent is an intelligent agent corresponding to an operating system originating from the device (e.g., the device implementing the intelligent agent) and / or a process controlled by the device's operating system. In some embodiments, the intelligent agent is an application intelligent agent. In some embodiments, the application intelligent agent is an intelligent agent corresponding to an application originating from the device (e.g., the device implementing the intelligent agent) (e.g., installed on and / or executed by the device) and / or a process controlled by the device's application.
[0136] This document may refer to representations (e.g., digital avatars and / or digital avatar representations) of agents (e.g., and / or users (e.g., people, objects, and / or animals) and / or user interface objects (e.g., animated characters)). In some embodiments, an agent's representation refers to a set of output characteristics (e.g., visual and / or audio) of the agent (and / or user and / or user interface object). In some embodiments, an agent's representation may include (and / or correspond to) a set of one or more visual characteristics (e.g., facial features of an animated face) and / or one or more audio characteristics (e.g., language and speech characteristics of audio output). In some embodiments, (e.g., the agent's) representation is used to represent the agent's output. In some embodiments, a device implementing an interactive agent outputs audio in the agent's voice and displays an animated face of the agent moving in a manner that simulates the agent speaking the audio output. In this way, the user can feel that they are having a normal conversation with the agent. In some embodiments, the agent's representation includes (or does not include) personality and / or behavioral characteristics (e.g., as described above). In some embodiments, the representation of the intelligent agent may include (and / or correspond to) a set of visual characteristics (e.g., facial features of an animated face) and a set of personality characteristics. In some embodiments, the representation of the intelligent agent includes a set of user characteristics corresponding to the user's visual representation (e.g., representations of the user's appearance, voice, and / or personality used as a digital avatar that appears to move and / or speak). In some embodiments, the representation is a facial representation (e.g., a user interface object outputting features that simulate a human face and / or facial expressions (e.g., for conveying information to a viewer)).
[0137] In some implementations, a role (e.g., the role of an agent and / or digital avatar) refers to a specific set of characteristics represented. In some implementations, a digital avatar may embody the characteristics (e.g., use, application, interaction with, and / or output) of fictional and / or non-fictional characters (e.g., from movies, shows, books, TV series, and / or popular culture).
[0138] In some implementations, (e.g., agent and / or digital avatar) speech refers to one or more characteristics corresponding to sound outputs that are similar to (e.g., representing, imitating, and / or reproducing) spoken language (e.g., attributable to and / or simulated as output by an agent and / or digital avatar). In some implementations, device 200 may output sentences that sound different depending on the speech used. In some implementations, a particular character and / or digital avatar may be configured to use a particular speech (e.g., have a corresponding speech). In some implementations, the particular speech may mimic a user's speech.
[0139] In some implementations, the appearance (e.g., of an agent and / or digital avatar) refers to a set of one or more characteristics corresponding to the visual output representing the digital avatar (and / or agent). In some implementations, device 200 may output a digital avatar having a set of facial features that form an appearance similar to a specific character from a movie.
[0140] In some embodiments, the expression of a digital avatar refers to one or more characteristics corresponding to a specific visual appearance of a user, digital avatar, and / or intelligent agent. In some embodiments, device 200 may output a digital avatar having a set of facial features arranged in a specific manner to give the appearance of a facial expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a frown is an expression of sadness, a smile is an expression of happiness, and / or wide eyes are an expression of surprise). As another example, device 200 may output a digital avatar having a set of body features (e.g., arms and / or legs) arranged in a specific manner to give the appearance of a body expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a gesture is an expression of agreement, covering the eyes is an expression of fear, and shrugging is an expression of lack of knowledge). In some embodiments, expressions include movement of the digital avatar (e.g., a nod is an expression of agreement and / or disagreement). In some embodiments, device 200 may be movable via a motion component to indicate expressions accompanying or not accompanying movement of the digital avatar. In some implementations, the agent performs one or more actions that depend on the user's facial expressions (e.g., detecting whether a person is sad and responding with a kind statement or question). In some implementations, facial expressions (e.g., whether they are used, how they are used, and / or how they are output) depend on personality. In some implementations, the first sex may use a particular facial expression more often than the second sex. As another example, the first sex's expressions (e.g., frowning, smiling, and / or widening eyes) may look different from the second sex's expressions (and / or similar and / or equivalent expressions) (e.g., the first sex smiles with their teeth showing, but the second sex smiles without showing their teeth).
[0141] In some implementations, an agent (e.g., a digital image of the agent and / or an agent system implementing the agent (e.g., hardware and / or software)) mimics the characteristics (e.g., in terms of personality, behavior, facial expressions, and / or voice) of another user, agent, and / or role. In some implementations, mimicry includes mirroring the user (e.g., replicating phrases and / or movements detected from a user interacting with the agent). In some implementations, simulating user characteristics includes attempting to reproduce the user's characteristics (e.g., in exactly the same way and / or in a way that is similar to but not an exact reproduction of the characteristics). In some implementations, an agent simulating voice and / or facial expressions is not required to have the exact same voice and / or facial expressions as the user being simulated (e.g., simply approximating the user's voice and / or facial expressions is sufficient).
[0142] In some implementations, components and / or devices use (e.g., performing actions, making decisions, and / or determining context based on them) learned characteristics (e.g., characteristics of the context, user, and / or environment learned by the device over time (e.g., via detection, prior experience, and / or feedback (e.g., from one or more users))). In some implementations, the characteristics learned over time may include user routines. In such an example, if a particular user requests a summary of any new messages for that user from the agent at the same time each day, the agent may learn to automate actions based on the characteristics of the learned routines (e.g., what data is needed, when data is needed, and / or for which user). In some implementations, the learned characteristics enable the agent (and / or device) to improve its understanding (and / or response to) of the context, user, and / or environment, and / or its understanding of context, user, and / or environment that is otherwise not (and / or will not) understood (e.g., not responded to or responded to incorrectly). In some implementations, the learned characteristics are formed using reinforcement learning (e.g., by and / or for the agent). In some implementations, the learned features correspond to one or more confidence levels, determinisms, and / or rewards (e.g., shaped by one or more reward functions). In some implementations, the learned features (and / or how they are used to influence the output of the agent and / or device) can change over time (e.g., confidence levels, determinisms, and / or rewards change over time). In some implementations, the output of the device before learning a set of learned features can differ from the output of the device after learning a set of learned features. In some implementations, components and / or devices use the learned knowledge. In some implementations, similar to what has been described above regarding learned features, the learned knowledge can refer to information used to update (e.g., enhance, add to, and / or expand) the device's knowledge base (e.g., for use by agents implemented on the device). In some implementations, multiple sets of learned features for a user can be stored and / or used. In some implementations, different sets of learned features for different users can be stored and / or used.
[0143] This document may refer to interactions with an agent (and / or device). In some embodiments, an interaction refers to a set of one or more inputs and / or outputs from a device implementing the agent and one or more users. In some embodiments, an interaction may be a user input (e.g., “Please turn on the light”) and a corresponding output (e.g., turning on the light and / or the device’s response “OK”). In some embodiments, an interaction may include multiple inputs / outputs performed by one or more parties to the interaction (e.g., a device and / or a user). In some embodiments, an interaction may include a first user input (e.g., “Please turn on the light”) and a corresponding first output (e.g., “Which lights?”), and also includes a second user input (e.g., “Kitchen light”) and a second output from the device (e.g., “OK”). In some embodiments, which inputs and / or outputs are considered together as an interaction is based on logical and / or contextual grouping (e.g., interactions within the previous thirty (30) seconds and / or interactions related to turning on the light). As those skilled in the art will understand, interactions may be considered in an implementation-dependent manner (e.g., determining when an interaction is complete may involve determining whether the user is still present (e.g., still talking) and / or whether the user is still talking about the light or has moved on to a different topic). In some implementations, the interaction is the current interaction (e.g., ongoing, currently occurring, and / or active). In some implementations, the interaction is a previous interaction. The examples above describe a device that engages in dialogue with a user. In some implementations, the dialogue is between two or more users (e.g., users in an environment). In some implementations, the device can detect dialogue between users (e.g., users directing voice and responses to each other, rather than to the device).
[0144] In some implementations, the agent (and / or device) determines and / or performs an action based on an intent corresponding to the user. In some implementations, the device detects user input and outputs a response dependent on the intent of the user input. In some implementations, the device detects user input including a pointing gesture detected along with a verbal command to “turn on the light,” and in response, the device turns on a light determined to correspond to the intent of the input (e.g., the light the pointing gesture is pointing to). In some implementations, intent is determined using one or more of the following (e.g., determined by the device detecting the input and / or by one or more other devices): one or more inputs, knowledge (e.g., knowledge about the user learned based on observed behavior, personality, and history of interactions), learned characteristics, and / or context. In some implementations, intent is determined based on one or more types of input (e.g., verbal input, visual input via a camera, and / or contextual input).
[0145] Now turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as computer system 100 and / or electronic device 200).
[0146] Figures 6A to 6G Exemplary user interfaces for modifying the appearance, position, and / or size of user interface objects are illustrated according to some embodiments. The user interfaces in these figures are used to illustrate including... Figure 7 , Figure 8 and Figure 9 The process described below is the process in the middle.
[0147] Figures 6A to 6G A computer system 600 (e.g., a tablet computer) is illustrated. It should be understood that the computer system 600 can be other types of computer systems, such as smartphones, smartwatches, laptops, utilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 600 includes and / or communicates with one or more input devices and / or sensors (e.g., cameras, lidar detectors, motion sensors, infrared sensors, touch-sensitive surfaces, physical input mechanisms (such as buttons or sliders) and / or microphones). Such input devices and / or sensors can be used to detect the presence of a user in the environment, the user's attention, statements from the user, input corresponding to the user, requests from the user, and / or instructions from the user. It should be understood that while some embodiments described herein relate to input as voice input, other types of input can be used with the techniques described herein, such as touch input via touch-sensitive surfaces and air gestures detected via cameras. In some embodiments, computer system 600 includes one or more output devices (e.g., a display screen, projector, touch-sensitive display, speaker, and / or movable components) and / or communicates with such one or more output devices. Such output devices can be used to present information and / or cause different visual changes to computer system 600. In some embodiments, computer system 600 includes one or more movable components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. As described above, such movable components can be used to change the positioning (e.g., location and / or orientation) of computer system 600 and / or a portion of computer system 600 (e.g., including one or more sensors, input components, and / or output components). In some embodiments, computer system 600 includes one or more components and / or features described above with respect to computer system 100 and / or electronic device 200. In some embodiments, computer system 600 includes as described above with respect to… Figure 5The description refers to one or more intelligent agents and / or the functionality of intelligent agents. In some embodiments, computer system 600 is, includes, implements one or more intelligent agent systems and / or communicates with one or more intelligent agent systems, as described above. Figure 5 As described, one or more operations are performed (and / or cause to be performed) by an agent.
[0148] like Figures 6A to 6G As shown, computer system 600 displays a user interface object (e.g., user interface object 604). In some embodiments, user interface object 604 is a representation of an animated face and / or an intelligent agent. An animated face may include eyes, a mouth, a nose, a body, and / or other elements. In some embodiments, user interface object 604 corresponds to an application and / or system process of computer system 600. In such embodiments, it should be appreciated that user interface object 604 may have different appearances. In some embodiments, user interface object 604 may be used with multiple different applications. In other embodiments, different user interface objects may be used with different applications. In some embodiments, as discussed below and... Figures 6A to 6G The user interface object 604 shown can be used for one or more sets of applications, while another user interface object is used for another set of one or more applications.
[0149] Figures 6A to 6D The process is illustrated in which computer system 600 modifies the appearance of user interface object 604 (e.g., modifies a portion of the user interface object, changes its color, changes its shape, changes its size, and / or changes its position) in response to input from a user that computer system 600 agrees to (e.g., detected via one or more of the aforementioned input devices). In some embodiments, computer system 600 may make user interface object 604 appear to: (1) initiate and / or maintain eye contact with the user when computer system agrees to the input, and / or (2) stop and / or avoid eye contact with the user when computer system disagrees with the input. It should be appreciated that other movements of the portion and / or user interface object 604 may be made in response to computer system 600 agreeing to input from the user (e.g., moving a portion of computer system 600 (e.g., including input devices and / or output devices) up and down and / or moving user interface object 604 up and down (e.g., to simulate a nodding motion)) or disagreeing to input from the user (e.g., moving the portion left and right and / or moving user interface object 604 left and right (e.g., to simulate a head shaking motion)).
[0150] In some implementations, computer system 600 accepts input when it includes a correct and / or truthful statement, as determined by computer system 600 and / or another computer system communicating with computer system 600. In some implementations, the input may include an indication that “the sky is blue.” In response to detecting this input, computer system 600 may determine that the sky is blue, and therefore computer system 600 accepts the input. It should be recognized that other criteria may be used to determine whether computer system 600 accepts input, such as whether the input is consistent with one or more user preferences known to computer system 600 (e.g., previously entered by the user and / or detected by computer system 600 in previous interactions with the user). In some implementations, the user's preference may be that the user likes blue based on the user having previously told computer system 600 this. Therefore, when computer system 600 detects a question or statement from the user corresponding to blue (e.g., "Does this blue shirt look good on me?"), computer system 600 initiates eye contact between user interface object 604 and the user (and / or performs some other output indicating agreement, such as outputting an audio "yes" and / or making user interface object 604 nod up and down). When computer system 600 detects a question or statement from the user corresponding to red (e.g., "Does this red shirt look good on me?"), computer system 600 does not cause user interface object 604 to initiate eye contact with the user (and / or makes user interface object 604 perform some other output indicating disagreement, such as outputting an audio "no" and / or making user interface object 604 shake its head left and right). In some embodiments, the user does not need to use the word "blue" in the question or statement; instead, computer system 600 can detect characteristics of the object, such as the color of the shirt, via a camera.
[0151] Figure 6A An example is shown of a computer system 600 displaying a user interface 602. The user interface 602 includes a user interface object 604, as described above. Figure 6AAt this point, computer system 600 detects voice input 606 from the user (e.g., "The sky is currently blue"). Therefore, voice input 606 does not include an explicit instruction to change the appearance of user interface object 604. That is, computer system 600 does not change the appearance of user interface object 604 based on a command from the user (e.g., the user tells user interface object 604 "look at me"). Instead, computer system 600 determines to change the appearance of user interface object 604 based on agreement or disagreement voice input 606. In some embodiments, computer system 600 detects input in the form of a question from the user, rather than a statement. In such embodiments, computer system 600 may change the appearance of user interface object 604 based on the question, such as making user interface object 604 look in the user's direction and / or make eye contact with the user when computer system 600 agrees with the question and / or is answering the question with an affirmative answer (e.g., "yes").
[0152] exist Figure 6B At this point, computer system 600 determines that computer system 600 agrees to voice input 606. For example... Figure 6B As shown, in response to detecting voice input 606 and determining that the computer system 600 agrees to the voice input 606, the computer system 600 changes the appearance of the user interface object 604. In some embodiments, the computer system 600 displays the user interface object 604 as if it is making eye contact with the user (e.g., looking up in the user's direction) by changing the eye positioning of the user interface object 604. In some embodiments, when the computer system 600 changes the appearance of the user interface object 604, the computer system 600 maintains the user interface object 604 at a specific location in the user interface 602 and / or maintains the appearance of one or more other user interface objects (e.g., clocks, icons, backgrounds, and / or information) in the user interface 602.
[0153] In some implementations, the computer system 600 maintains the user interface object 604 in a manner that appears to be making eye contact with the user for a predetermined period of time. That is, after a predetermined period of time has elapsed since the eye contact was made, the computer system 600 stops displaying the user interface object 604 as if it were making eye contact with the user (e.g., as shown in the image). Figure 6C As shown and further described below.
[0154] In some embodiments, computer system 600 stops displaying user interface object 604 as if making eye contact with the user in response to detecting another voice input (e.g., from the user or another user different from the user). That is, in some embodiments, when computer system 600 displays user interface object 604 moving and / or positioning in an eye-contact manner and computer system 600 detects another voice input, computer system 600 stops displaying user interface object 604 as if making eye contact with the user until the other voice input is completed. In some embodiments, when another voice input is detected, computer system 600 stops displaying eye contact with the user regardless of whether the user agrees with the other voice input, thereby indicating to the user that user interface object 604 is restarting, before the computer system acknowledges the other voice input. In some implementations, if the user interface object 604 is making eye contact with the user in response to the computer system 600 detecting the voice input “the sky is blue”, then when the computer system 600 detects a question or statement from the user saying “birds have feathers”, the computer system 600 will stop the eye contact between the user interface object 604 and the user, and then resume eye contact with the user interface object 604 due to agreement to a second question or statement.
[0155] like Figure 6C As shown, the computer system 600 displays the user interface object 604 as looking forward and no longer making eye contact with the user. Figure 6C At this point, the computer system 600 detects voice input 608 from the user (e.g., “The sky is currently green”).
[0156] exist Figure 6D At this point, computer system 600 disagrees with voice input 608 (e.g., based on weather information indicating sky color and / or other information known to computer system 600). Therefore, in response to detecting voice input 608, computer system 600 displays user interface object 604 as if continuing to interact with... Figure 6C The same location shown is directly in front.
[0157] In some implementations, in response to detecting a question or statement from a user and the computer system 600 disagreeing with the question or statement, the computer system 600 does not change the appearance of the user interface object 604 (e.g., the user interface object 604 continues to look forward). In some implementations, in response to detecting a question or statement from a user and the computer system 600 disagreeing with the question or statement, the computer system 600 displays the user interface object 604 moving in a different manner than when the computer system 600 detected its agreement with the question or statement (e.g., looking down, shaking its head, and / or closing its eyes).
[0158] Figures 6D to 6G This illustrates the process of interaction between the user interface object 604 displayed by the computer system 600 and the content displayed by the computer system 600. For example... Figure 6D As shown, computer system 600 displays user interface 602, wherein user interface object 604 is looking directly forward (e.g., in an direction not within user interface 602) and / or in the direction of the user. Figure 6D The user interface object 604 is also illustrated. Figures 6A to 6C The dimensions shown are the same and they occupy most of the user interface 602. Figure 6D At this point, computer system 600 detects voice input 610 from the user (e.g., "How's the weather today?"). Voice input 610 indicates a request for computer system 600 to inform the user of the current weather conditions.
[0159] In some implementations, requests to interact with computer system 600 may be alternative types of input (e.g., air gestures, touch input, and / or gaze input from a user). In some implementations, requests to interact with computer system 600 do not include explicit requests. In some implementations, a user may state "It's hot outside," and in response, computer system 600 displays information informing the user of the outdoor temperature. In some implementations, explicit requests include directly asking computer system 600 a question as a request for information. Explicit requests may also include direct questions expressing a desire for voice output from computer system 600, such as "Please answer this question" or "Can you help me?" like Figure 6E As shown, in response to detecting voice input 610, the computer system 600 removes the user interface object 604 from the... Figure 6D The size shown is reduced to a smaller size in the lower left corner of the user interface 602. In some embodiments, the computer system 600 displays the user interface object 604 as shrunken to the lower left corner of the user interface 602 because the user is standing on the left side of the computer system 600. In some embodiments, if the user is standing on the right side of the computer system 600, the computer system 600 displays the user interface object 604 as shrunken back to the lower right corner of the user interface 602. In some embodiments, the computer system 600 displays the user interface object 604 as shrunken to the lower left corner of the user interface 602 because the computer system 600 will display content on the right side of the user interface 602 (e.g., such as...). Figure 6F (As shown). In some implementations, if the content is to be displayed on the left side of the user interface 602, the computer system 600 displays the user interface object 604 as folded back to the lower right corner of the user interface 602.
[0160] like Figure 6EAs shown, the computer system 600 displays the eye of the user interface object 604 as a way to maintain contact with... Figure 6D The same direction is shown. In some implementations, after (e.g., or simultaneously) the user interface object 604 is collapsed, the computer system displays the user interface object 604 in the direction from which it begins to look at the location where the computer system 600 will display content (e.g., as shown). Figure 6F As shown and discussed further below.
[0161] like Figure 6F As shown, in response to the detection of voice input 610, computer system 600 displays content 612, which in this example is represented as a temperature value of 70 degrees. In some embodiments, computer system 600 displays multiple sets of content (e.g., images, icons, and / or words).
[0162] In some implementations, computer system 600 displays user interface object 604 as if it were looking at the user while the user interacts with computer system 600, and as if it were looking at content output by computer system 600 and / or viewed by the user. For example... Figure 6F As shown, in response to detecting voice input 610 and / or displaying content 612, when the computer system 600 detects that the process of the user receiving and / or looking at each set of content has been completed, the computer system 600 may display the user interface object 604 as if its gaze has moved from one set of content to another. In some embodiments, the computer system 600 displays the user interface object 604 as if its gaze has moved from one set of content to another in order to guide the user's attention from one set of content to another. In some embodiments, in response to detecting voice input 610 and / or displaying content 612, the computer system 600 does not change the positioning or gaze of the user interface object 604 from such a set of content. Figure 6E The positioning has changed as shown.
[0163] It is worth noting that the computer system 600 displays content 612 in a portion of the user interface 602, where the computer system 600 previously displayed, as shown in the following example. Figure 6D The user interface object 604 shown. That is, in Figure 6F In the middle, content 612 occupies most of the area of user interface 602. Similarly, as... Figure 6F As shown, in response to the detection of voice input 610, the computer system 600 changes the direction of the user interface object 604 from looking out of the user interface 602 to looking towards the content 612. In this example, the computer system 600 displays the user interface object 604 facing the direction that is the target direction the user wants to look in. That is, the computer system 600 guides the user's gaze to certain areas of the user interface 602 that are important to the user's request for information.
[0164] In some implementations, when the computer system 600 is not waiting for a response from the user (e.g., input) (e.g., in the process of performing a task or a series of tasks) (e.g., after viewing content 612), the computer system 600 displays the user interface object 604 in an direction other than the user's view. In some implementations, in response to determining that the user is looking at content 612, this other direction is content 612. In some implementations, this other direction is another user within the environment. In some implementations, this other direction is an object within the environment, such as a clock, sofa, and / or television. In some implementations, this other direction is a direction not in the specific direction of something.
[0165] In some implementations, computer system 600 may change the gaze of user interface object 604 from pointing at content 612 to pointing at the user in response to detecting that the user has stopped viewing content 612. In some implementations, computer system 600 displays user interface object 604 as looking at content 612 as an indication to the user that the user should be looking at content 612 (e.g., computer system 600 has displayed content 612 for some time and the user has not yet viewed it) (e.g., computer system 600 is outputting audio such as audio response 614, so the user is not paying attention to and / or viewing content 612).
[0166] As described above, computer system 600 can detect multiple people in the environment. In some embodiments, computer system 600 detects a condition (e.g., a completed task), and in response to this condition, computer system 600 stops displaying user interface object 604 as gaze content 612. After computer system 600 stops displaying user interface object 604 as gaze content 612, computer system 600 displays user interface object 604 as the user initiating voice input 610 in the gaze environment. That is, in some embodiments, computer system 600 directs the gaze of user interface object 604 to the user whose voice computer system 600 has detected, rather than another user who did not initiate voice input 610. In some embodiments, computer system 600 detects interaction from users and, in response, initiates a process to determine which user user user interface object 604 should be displayed as looking at. In some embodiments, a first user may request computer system 600 to display video to a second user, and in response, computer system 600 displays the video and directs the gaze of user interface object 604 to the second user. For example, a first user may provide voice input to computer system 600, and computer system 600 may display user interface object 604 to shift its gaze from the first user providing voice input to another user, indicating that user interface object 604 is waiting for voice input from that other user.
[0167] Similarly, Figure 6F As shown, in response to the detection of voice input 610 and / or the computer system 600 shrinking the user interface object 604 and / or displaying content 612, the computer system 600 outputs an audio response 614 from the user interface object 604 (e.g., "The temperature is 70 degrees"). The audio response 614 tells the user, in text form, the information displayed by the computer system 600 in the form of content 612. That is, the audio response 614 tells the user that the temperature is 70 degrees, which is the temperature displayed in content 612. In some embodiments, the computer system 600 uses the audio response 614 to convey information relevant to the context of content 612. That is, the computer system 600 can use the audio response 614 to provide additional context to content 612 that may help the user understand the meaning of content 612. In some embodiments, the computer system 600 replaces previously displayed content (e.g., icons and / or photos) with content 612. That is, in some embodiments, content 612 replaces the content initially displayed by the computer system 600 with content relevant to the user's request for information.
[0168] like Figure 6G As shown, computer system 600 detects that it has completed its tasks of displaying content 612 and outputting audio response 614. In some embodiments, in response to detecting that computer system 600 has completed its tasks and / or meets a set of criteria (e.g., computer system 600 detects that input from the user is needed, a predetermined time period has elapsed, and / or the user is viewing content 612), computer system 600 displays user interface object 604 with its gaze directed towards the user. In some embodiments, in response to detecting that the user has finished viewing content 612 (e.g., the user has viewed content 612 for a predetermined time period and then looks in another direction), computer system 600 stops displaying content 612 and causes user interface object 604 to be displayed as shown. Figures 6A to 6D The original size shown is then re-displayed. In some implementations, in response to detecting that the user has finished viewing content 612, the computer system 600 shrinks content 612 to a smaller size.
[0169] Figure 7 This is a flowchart illustrating a method for displaying an object oriented in a certain direction using a computer system, according to some implementation schemes. Process 700 is executed at a computer system (e.g., 100, 200, and / or 600). Some operations in process 700 may be combined, some operations may be ordered differently, and some operations may be omitted.
[0170] As described below, process 700 provides an intuitive way to display objects facing a particular direction. This method reduces the cognitive burden on the user when displaying objects facing a particular direction, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to display objects facing a particular direction faster and more efficiently saves power and increases the time interval between battery charging.
[0171] In some embodiments, process 700 is performed at a computer system (e.g., 600) that communicates with display components (e.g., a display screen, projector, and / or touch-sensitive display) and one or more input devices (e.g., a camera, depth sensor, and / or microphone). In some embodiments, the computer system is a watch, phone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device.
[0172] When (and / or subsequently and in response to) content displayed via a display component (e.g., 612) (e.g., sound, media, visual content, audio content, topics, users, discussions, and / or relationships), a computer system (e.g., via one or more input devices) detects (702) a first interaction condition corresponding to the content (e.g., voice, touch, context (e.g., content being discussed, talked about, and / or output) and / or movement of the computer system and / or the user) (e.g., via verbal input (e.g., audible requests, audible commands, and / or audible statements) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)) (e.g., as described above relative to...). Figures 6D to 6F (as described above). In some implementations, the first interaction condition is detected via the one or more input devices.
[0173] In response to (704) detecting a first interaction condition corresponding to the content, and based on determining that the first interaction condition corresponding to the content satisfies a first set of one or more criteria, the computer system displays (706) via a display component a representation of a face (e.g., 604) looking towards the content (e.g., 612) in a direction (e.g., corresponding to the content and / or being at the content) (e.g., in a manner indicating eye contact, in a manner indicating gaze, appearing to be looking, pointing, and / or the eyes, mouth, face, and / or voice appearing to be looking in that direction) (e.g., as described above relative to...). Figure 6F (as described above). In some implementations, the first group of one or more criteria includes criteria that are met when the content is directed to, discussed, commented on, introduced, and / or interacted with.
[0174] In response to (704) the detection of a first interaction condition corresponding to the content, and based on the determination that the first interaction condition corresponding to the content satisfies a second set of one or more criteria different from a first set of one or more criteria (e.g., not satisfying the first set of one or more criteria), the computer system displays (708) via a display component a representation of the face (e.g., 604) looking toward a first user (e.g., user, person, animal, and / or object) detected in the detection field of the one or more input devices (e.g., in a manner indicating eye contact and / or eye gaze) (e.g., corresponding to the first user and / or being at the first user's location) (e.g., in a manner indicating eye contact, in a manner indicating gaze, looking, pointing, and / or eyes, mouth, face, and / or voice appearing to be looking in that direction) (e.g., as described above relative to...). Figure 6G (as described above). In some embodiments, the second set of one or more criteria includes criteria that are met when content is not directed to, discussed, commented on, introduced, and / or interacted with. In some embodiments, the second set of one or more criteria includes criteria that are met when it is determined that a response from a first user is needed, a portion of the content is directed to the first user, consent from the first user is needed, requested, and / or expected, and / or a portion of the content being output relates to the first user. Displaying facial representations in response to interactions enables the computer system to provide the user with confirmation that the interaction has been received, thereby providing improved visual feedback. Displaying facial representations via the computer system in an orientation based on which set of one or more criteria are met enables the computer system to selectively display facial representations in the appropriate orientation based on the type of interaction and allows the computer system to display purpose-specific instructions based on the interaction, thereby providing improved visual feedback, performing actions without further input when a set of conditions has been met, and / or reducing the amount of user input required to perform an action.
[0175] In some implementations, the first set of one or more criteria includes determining that a portion of the content (e.g., 612) has been displayed for at least a predetermined period of time (e.g., as described above relative to...). Figure 6FThe first criterion (e.g., the first criterion is met when the content is initially displayed, when a new portion of the content is initially displayed, and / or when new content is initially displayed) is satisfied (e.g., the first interaction condition is a request to display the content when the content was not previously (e.g., immediately after previously), when a new portion of the content is displayed, and / or when new content is displayed). The display of a facial representation in the direction of the content based on a criterion that a portion of the content is displayed for at least a predetermined amount of time enables the computer system to confirm the content's location on the display component, thereby indicating the content's position to the user and / or reducing distractions on the display component, thus providing improved visual feedback, performing an operation when a set of conditions has been met without further input, and / or reducing the amount of input required to perform the operation.
[0176] In some implementations, the first set of one or more criteria includes a second criterion satisfied when it is determined that a first user is looking in the direction of the content (e.g., 612) (e.g., for at least a threshold amount of time (0.1 seconds to 10 seconds) and / or is looking in that direction and / or is looking in that direction) (e.g., based on the detection of the user (e.g., voice, gaze, and / or eyes) in the detection field of the one or more input devices) (e.g., left, right, up, down, and / or any combination thereof). Figure 6F (e.g., detecting and / or occurring a first interaction condition after and / or during the display of the content). In some embodiments, the interaction is detected and / or occurs before displaying content that satisfies one or more of a first set of criteria or one or more of a second set of criteria. Displaying a facial representation in the direction the user is looking at the content, based on criteria, enables the computer system to selectively display a facial representation in the appropriate direction and / or reduce interference on display components based on detection of the user via one or more input systems, thereby providing improved visual feedback, performing an operation when a set of conditions has been met without further input, and / or reducing the amount of input required to perform the operation.
[0177] In some implementations, the first set of one or more criteria includes a third criterion (e.g., as described above relative to) that is satisfied when it is determined that a first user should view the content (e.g., 612). Figure 6F(The above) (e.g., based on the first user's usage pattern (e.g., based on historical usage and / or based on historical patterns), the usage pattern corresponding to one or more previous interactions of the first user with the content, the content including an indication that the first user is looking at the content being displayed) (e.g., the first user is associated with a usage pattern of looking at new content after a previous response condition (e.g., corresponding to the usage pattern, historically known to be associated with the usage pattern, detected as having previously had the usage pattern), such as when the application is initially launched, when one or more additional inputs are requested about the content being displayed (e.g., to verify that the content is correct, to verify moving to the next content (e.g., new content and / or different parts of the content)), when there is an output problem, when text is requested to be transmitted, and / or when checking text). In some embodiments, determining that the first user should be looking at the content includes detecting that the first user has looked at the content being displayed once or multiple times after the content has been displayed. Displaying facial representations in the direction of content based on criteria that the user should be viewing enables the computer system to confirm that content is being displayed to the user and / or allows the computer system to guide the user's attention to the content, thereby providing improved visual feedback, performing actions without further input when a set of conditions has been met, and / or reducing the amount of input required to perform an action.
[0178] In some implementations, the second set of one or more criteria includes a fourth criterion (e.g., as described above relative to) that is satisfied when it is determined that input has been received (e.g., 610) (e.g., interaction, command, and / or request) (e.g., corresponding to and / or made by the first user) (e.g., verbal input (e.g., audible request, audible command, and / or audible statement) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gesture, and / or mouse click)). Figures 6D to 6E (as described above). In some embodiments, after displaying a representation of the direction the face is looking at the content for a first predefined time period, the computer system displays, via a display component, a representation of the direction the face is looking at the first user detected in the detection field of the one or more input devices. In some embodiments, after detecting, via the one or more input devices, the direction the first user is looking at the content (e.g., for at least a second predefined time period (0.1 seconds to 10 seconds) and / or the direction in which the user is looking at the content and / or the direction in which the user is looking at the content), the computer system displays, via a display component, a representation of the direction the face is looking at the first user detected in the detection field of the one or more input devices. Displaying a representation of the face in the user's direction based on criteria for having received input allows the computer system to confirm that input has been received, thereby providing the user with improved visual feedback and performing an operation without further input once a set of conditions has been met.
[0179] In some implementations, the second set of one or more criteria includes determining when the computer system (e.g., 600) is waiting for a response (e.g., from a first user and / or another user) (e.g., as described above relative to...). Figures 6D to 6G The fifth criterion is met when the computer system needs one or more additional inputs to continue an operation (e.g., waiting for (e.g., and / or waiting for) input) and / or a specific type of input (e.g., tap input, air gesture, voice command, and / or mouse click) (e.g., the computer system needs one or more additional inputs to continue an operation) (e.g., a first user needs to provide a second interaction condition (e.g., confirming a text message, selecting a contact, choosing between different applications, and / or verifying content)). Displaying a facial representation in the user's direction based on the criterion that the computer system is waiting for a response allows the user to infer that new input is needed to continue and / or provides feedback to the user that the computer system needs input (e.g., not displaying additional information that the computer system is waiting for a response), thereby providing the user with improved visual feedback, performing an operation without further input when a set of conditions has been met, and / or enabling the computer system to avoid burn-in of display components.
[0180] In some implementations, in response to detecting a first interaction condition corresponding to the content and based on determining that the first interaction condition corresponding to the content satisfies one or more criteria, including criteria that the computer system (e.g., 600) does not meet while waiting for a response, the computer system displays a representation of the face (e.g., 604) looking in a second direction via a display component, the second direction being different from the direction of the first user detected in the detection field of the one or more input devices (e.g., as described above relative to...). Figure 6B (as described above) (and in some embodiments, the orientation of the content). In some embodiments, the second orientation is within the field of view of the one or more input devices, but not within the detection field of the first user. In some embodiments, the second orientation is in the orientation of the user interface, but not in the orientation of the content. In some embodiments, the second orientation is away from the orientation of the first user. The fact that the computer system does not display a facial representation in a second orientation away from the user, based on the standard of waiting for a response, enables the computer system to provide feedback to the user that the computer system does not require input, without displaying additional information on the user interface, thereby providing the user with improved visual feedback and performing an operation without further input when a set of conditions has been met.
[0181] In some implementations, the second direction corresponds to a second user (e.g., a second user, a person, an animal, and / or an object) detected in the detection field of the one or more input devices (e.g., within the field of view of one or more input systems (e.g., a camera, a depth sensor, a microphone)). In some implementations, the second user is different from the first user (e.g., as described above relative to the first user). Figures 6D to 6G (as described above). In some embodiments, when the representation of the face is displayed via a display component looking in a second direction, the representation of the face is looking at a second user. The fact that the computer system does not display the representation of the face in a second direction different from that of the first user, based on the criterion of waiting for a response, enables the computer system to confirm that a second user has been detected in the detection field of one or more input devices, thereby providing improved visual feedback and performing operations without further input once a set of conditions has been met.
[0182] In some implementations, the second direction corresponds to content (e.g., 612) (e.g., as described above relative to...). Figure 6F (as described above). In some embodiments, when the representation of the face is displayed via a display component looking in a second direction, the representation of the face is looking at the content (and / or at least a portion of the content). The computer system's failure to display the representation of the face in a second direction corresponding to the content, based on the absence of a criterion for waiting for a response, provides feedback to the user that content is being displayed and allows the computer system to draw the user's attention to the content, thereby providing improved visual feedback and performing actions without further input once a set of conditions have been met.
[0183] In some implementations, the second direction corresponds to an object in the detection field of the one or more input devices (e.g., a physical and / or virtual object) (e.g., an object located in and / or detected in the detection field) (e.g., as described above relative to...). Figures 6D to 6G (e.g., in an environment including a first user (e.g., a physical, virtual, or mixed reality environment)). Based on the fact that the computer system does not display a facial representation in a second direction corresponding to an object in the detection field according to a waiting response standard, the computer system is able to confirm that an object has been detected in the detection field of the one or more input devices (without displaying additional information on the user interface to confirm the detection) and / or provide feedback that content has been completed and no response is required (e.g., it is time to exit the application, and / or no new content can be provided), thereby providing improved visual feedback, performing operations without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0184] In some implementations, when (and / or after) the representation of the face (e.g., 604) is looking in the direction of the content (e.g., 612) (and in some implementations, not in the direction of the first user), the computer system (e.g., via one or more input devices) detects a second interaction condition (e.g., 614) corresponding to the content (e.g., and / or corresponding to other content different from the content) (e.g., voice, touch, context (e.g., content being discussed, talked about, and / or output) and / or movement of the computer system and / or the user) (e.g., via verbal input (e.g., verbal input, audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)) (e.g., as described above relative to...). Figure 6F In some embodiments, in response to the detection of a second interaction condition corresponding to the content (e.g., 614), the computer system displays a representation of the face via a display component (e.g., 604) in the direction of the first user detected in the detection field of the one or more input devices (e.g., as described above relative to...). Figure 6G (as described above). In some embodiments, in response to detecting a second interaction condition corresponding to the content, the computer system changes the display of the facial representation from the direction of looking at the content to the direction of looking at the first user detected in the detection field of the one or more input devices. Changing the display of the facial representation from the direction of the content to the direction of the user in response to detecting a new interaction condition allows the computer system to confirm that a second interaction condition has been detected, thereby providing improved visual feedback to the user and performing an operation without further input once a set of conditions has been met.
[0185] In some implementations, when (and / or after) the representation of the face (e.g., 604) is looking in the direction of a user detected in the detection field of the one or more input devices, the computer system detects a third interaction condition (e.g., 614) corresponding to that content (e.g., and / or corresponding to other content) (e.g., as described above relative to other content). Figures 6E to 6F (e.g., via verbal input (e.g., verbal input, audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, in response to detecting a third interaction condition corresponding to the content (e.g., 614), the computer system displays a representation of the face (e.g., 604) looking towards the content (e.g., 612) via a display component (e.g., relative to the content as described above). Figures 6E to 6F(as described above). In some embodiments, the computer system changes the representation of the face from the direction of looking towards a first user detected in the detection field of the one or more input devices to the direction of looking towards the content. Responding to the detection of a new interaction condition by changing the representation of the face from the user's direction to the content's direction enables the computer system to confirm that a new interaction condition has been received and / or to indicate that the user's attention should be directed to the content, thereby providing improved feedback, reducing the amount of input required to perform an operation, and performing an operation without further input once a set of conditions has been met.
[0186] In some embodiments, when (and / or after) displaying content via a display component (e.g., 612) (and in some embodiments, a facial representation (e.g., a direction looking towards the content or the first user and / or another direction not corresponding to the first user or the content)) is made, the computer system detects a fourth interaction condition (e.g., 614) corresponding to the content (e.g., and / or corresponding to other content) (e.g., via verbal input (e.g., audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, in response to detecting a fourth interaction condition (e.g., 614) corresponding to the content, the computer system displays a facial representation (e.g., 604) looking towards a third direction not corresponding to the first user and the content (e.g., as described above relative to the content). Figure 6G (as described above). In some embodiments, the computer system changes the display of the facial representation from a direction looking towards a first user detected in the detection field of the one or more input devices to a third-party direction. In some embodiments, the computer system changes the display of the facial representation from a direction looking towards the first user to a third-party direction. In some embodiments, when the facial representation is displayed looking towards a third-party direction via a display component, the facial representation is not looking at the content and the facial representation is not looking at the first user. In some embodiments, the third-party direction is in the field of view of the one or more input devices, but not in the detection field of the first user or the second user. In some embodiments, the third-party direction is in the direction of the user interface, but not in the direction of the content. Displaying the facial representation in a direction not corresponding to the first user and the content in response to the detection of a fourth interaction enables the computer system to confirm that a new interaction has been received and / or confirm that the new interaction is unrelated to the content or the user (e.g., accidental interaction, unrecognized interaction) and / or detect a new user and / or object, thereby providing improved feedback, reducing the amount of input required to perform an operation, and performing an operation without further input when a set of conditions has been met.
[0187] In some implementations, when (and / or after) a representation of a face is displayed via a display component (e.g., 604) looking toward a third direction, the computer system detects a fifth interaction condition corresponding to that content (e.g., and / or corresponding to other content). Figure 6D (e.g., voice, touch, context (e.g., content being discussed, talked about, and / or output) and / or movement of the computer system and / or the user) (e.g., via verbal input (e.g., verbal input, audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, in response to detecting a fifth interaction condition, the computer system continues to display a representation of the face (e.g., 604) looking in a third direction (e.g., the same direction as when, before, and / or at the time the fifth interaction condition was detected) via a display component (e.g., as described above relative to...). Figures 6D to 6E (as described above). In some embodiments, in response to detecting a fifth interaction condition, the computer system does not change the direction in which the facial representation is looking. In some embodiments, the fifth interaction condition is an interaction condition of a different type from the first interaction condition. Continuing to display the facial representation looking in a fourth direction in response to detecting a fifth interaction corresponding to the content enables the computer system to confirm that the new interaction condition corresponds to the direction the facial representation is already facing (e.g., the direction of the content and / or the user) and / or to confirm that the new interaction condition has been ignored, thereby providing improved feedback, reducing the amount of input required to perform an operation, and performing an operation without further input when a set of conditions has already been met.
[0188] In some implementations, when (and / or after) content is displayed via a display component (e.g., 612), the computer system detects a sixth interaction condition (e.g., voice, touch, context (e.g., content being discussed, talked about, and / or output) and / or movement of the computer system and / or the user) corresponding to that content (e.g., and / or corresponding to other content) (e.g., via verbal input (e.g., audible requests, audible commands, and / or audible statements) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)) (e.g., as described above relative to the above). Figures 6D to 6F In some embodiments, in response to detecting a sixth interaction condition corresponding to the content, and based on determining that the sixth interaction condition corresponding to the content satisfies a second set of one or more criteria, the computer system displays a representation of the face (e.g., 604) via a display component, showing the direction of the face toward the first user detected in the detection field of the one or more input devices (e.g., as described above relative to...). Figures 6F to 6GIn some embodiments, in response to detecting a sixth interaction condition corresponding to the content, and based on determining that the sixth interaction condition corresponding to the content satisfies a fourth group of one or more criteria that differ from a first group of one or more criteria and a second group of one or more criteria (e.g., and / or a third group of one or more criteria), the computer system displays a representation of a face (e.g., 604) via a display component, showing the direction of a third user detected in the detection field of the one or more input devices, wherein the third user is different from the first user (e.g., as described above relative to...). Figures 6F to 6G The computer system can confirm the receipt of a new interaction condition and determine that it corresponds to the first user, and / or confirm that the operation being performed is adapted to a third user detected in the detection field of the one or more input devices, based on a sixth interaction condition including one or more criteria from a second set of standards. This provides improved feedback, reduces the amount of input required to perform the operation, and allows the operation to be performed without further input when a set of conditions has been met. The computer system can also confirm the receipt of a new interaction condition and determine that it corresponds to the third user, and / or confirm that the operation being performed is adapted to a third user detected in the detection field of the one or more input devices. This provides improved feedback, reduces the amount of input required to perform the operation, and allows the operation to be performed without further input when a set of conditions has been met.
[0189] In some embodiments, the computer system (e.g., 600) communicates with moving components (e.g., actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases). In some embodiments, in response to detecting a sixth interaction condition corresponding to the content and based on determining that the sixth interaction condition satisfies a second set of one or more criteria, the computer system moves a portion of the computer system (e.g., 600) (e.g., hardware components (e.g., buttons and / or rotatable input mechanisms), a display, and / or the center and / or another portion of the display) from a first location in the environment to a second location in the environment (e.g., as described above relative to the second location) via the moving components. Figures 6D to 6G (as described).
[0190] In some implementations, when this part of the computer system (e.g., 600) is in the second location, the computer system detects a seventh interaction condition corresponding to that content (e.g., and / or detects that one or more criteria of the second set are no longer met) (e.g., as described above relative to...). Figures 6D to 6GIn some implementations, in response to detecting a seventh interaction condition corresponding to the content and determining that the seventh interaction condition corresponding to the content satisfies a first set of one or more criteria (e.g., as described above relative to...), Figures 6D to 6G As described above), the computer system moves that part of the computer system (e.g., 600) from a second location in the environment to a first location in the environment (e.g., as described above relative to...) via a moving component. Figures 6D to 6G In some embodiments, in response to detecting a seventh interaction condition corresponding to the content and determining that the seventh interaction condition corresponding to the content satisfies a first set of one or more criteria, the computer system displays a representation of a face (e.g., 604) looking toward the displayed content (e.g., 612) via a display component (e.g., relative to the above). Figures 6D to 6G (as described above). In some embodiments, in response to detecting a seventh interaction condition corresponding to the content and determining that the seventh interaction corresponding to the content satisfies one or more criteria of a second set, the computer system does not move that part of the computer system from a second location in the environment to a first location in the environment, nor does it display a representation of the face looking in the direction of the displayed content.
[0191] In some implementations, when this part of the computer system (e.g., 600) is in the second location, the computer system detects an eighth interaction condition corresponding to that content (e.g., 610) (e.g., and / or detects that one or more criteria of the second set are no longer met) (e.g., as described above relative to...). Figures 6D to 6G (as described above). In some implementations, in response to the detection of an eighth interaction condition (e.g., 610) corresponding to the content (e.g., as described above relative to...), Figures 6D to 6G According to the determination that the seventh interaction condition corresponding to the content satisfies one or more criteria of the first group, the computer system moves that part of the computer system (e.g., 600) from a second location in the environment to a first location in the environment via a moving component, while continuing to display a representation of the face (e.g., 604) looking towards the user (e.g., as described above relative to the user's face). Figures 6D to 6G (e.g., the user's orientation is different from the orientation of the content being viewed). In some implementations, in response to detecting an eighth interaction condition corresponding to the content and determining that a seventh interaction corresponding to the content satisfies one or more criteria of a first set, the computer system does not move that part of the computer system from a second location in the environment to a first location in the environment, and / or continues to display the representation of the face looking towards the user via a display component.
[0192] It should be noted that the above text regarding method 700 (for example, Figure 7The details of the process described herein also apply in a similar manner to the methods described below / above. In some embodiments, process 800 may optionally include one or more features of the various methods described above with reference to process 700. In some embodiments, the computer system may use one or more techniques of process 700 to determine consent expressed on a given input, and may use one or more techniques of process 800 to display an object in a manner that indicates eye contact with the user. For the sake of brevity, these details will not be repeated below.
[0193] Figure 8 This is a flowchart illustrating a method for displaying an indication of eye contact with an object using a computer system, according to some embodiments. Process 800 is executed at a computer system (e.g., 100, 200, 600). Some operations in process 800 may be combined, some operations may be ordered differently, and some operations may be omitted.
[0194] As described below, process 800 provides an intuitive way to indicate eye contact with an object. This method reduces the cognitive burden on the user from indicating eye contact with an object, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to display eye contact indications of objects faster and more efficiently saves power and increases the time interval between battery charging.
[0195] In some embodiments, process 800 is performed at a computer system (e.g., 600) that communicates with display components (e.g., a display screen, projector, and / or touch-sensitive display) and one or more input devices (e.g., a camera, depth sensor, and / or microphone). In some embodiments, the computer system is a watch, phone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device.
[0196] When (e.g., after and / or in response to) a user interface (e.g., 602) is displayed via a display component, and the user interface includes user interface objects (e.g., shapes, icons, emojis, and / or digital avatars) representing a part of a person (e.g., 604) (e.g., face, eyes, lips, hands, legs, and / or mouth), the computer system detects (802) a first input (e.g., voice input, tap input, swipe input, and / or air gesture) via one or more input devices (e.g., as described above relative to...). Figure 6A (as described above). In some embodiments, the voice input may be received via a text-to-speech device within a computer system and / or via a second computer system.
[0197] In response to (804) detecting the first input, and based on determining that consent has been expressed on the first input (e.g., 606) (e.g., consent statement (1) is correct and / or (2) is consistent with one or more user preferences) (e.g., consent to an answer to a question means (1) the question is correct and (2) the answer to the question is affirmative (e.g., "yes", "I agree" and / or "I like it")), the computer system continues (806) to display the user interface via a display component (e.g., 602), while altering the user interface object in a manner that indicates eye contact with the user (e.g., person, animal, and / or object) (e.g., pose, manner, and / or visual manner) (e.g., as described above relative to...). Figures 6A to 6C (as described).
[0198] In response to (804) detecting the first input, and based on the determination that no agreement was expressed with respect to the first input (e.g., 606) (e.g., disagreeing that the statement is correct and / or disagreeing that the question is indeed correct / accurate), the computer system continues (808) to display the user interface via a display component (e.g., 602) without altering the user interface object in a manner that indicates eye contact with the user (e.g., as described above relative to...). Figures 6B to 6F When consent is expressed for a first input, displaying a user interface object in a manner that indicates eye contact with the user provides feedback to the user that the computer system has detected and consented to the input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components. When consent is not expressed, displaying a user interface object without changing the user interface object in a manner that indicates eye contact with the user provides feedback to the user that the computer system does not consent to the detected input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0199] In some implementations, continuing to display the user interface (e.g., 602) without altering the user interface object in a manner that indicates eye contact with the user includes: displaying the user interface object via a display component in a manner that does not indicate eye contact with the user (e.g., corresponding to the user and / or at the user) (e.g., as described above relative to...). Figures 6B to 6F(as described above). In some implementations, in response to the detection of a first input, the computer system changes a portion of the user interface object from a direction facing the user to a direction that is not pointing at the user. Displaying the user interface object in a manner that does not indicate eye contact with the user when no consent is expressed provides feedback to the user that the computer system disagrees with the detected input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0200] In some implementations, changing a user interface object in a manner that indicates eye contact with the user (e.g., 604) includes: detecting the location of the user (e.g., a person, animal, and / or object) in a detection field of the computer system (e.g., 600) via the one or more input devices (e.g., relative to the above description). Figures 6A to 6C The detection field is described above; and in some embodiments, the detection field is the field of view of the camera, the detection field of the depth sensor, and / or the detection field of the microphone. A first portion (e.g., eye portion, face portion, nose portion, and / or mouth portion) of the user interface object (e.g., 604) is changed to point (e.g., in its direction, looking at, and / or pointing towards) the user's detected position in the detection field of the computer system (e.g., 600) (e.g., as described above relative to...). Figures 6A to 6G (as described above). When consent is expressed for the first input, the first part of the user interface object is changed to point to the detected location of the user in the detection field of the computer system. Doing so provides the user with feedback that the computer system has detected the input and consented to it, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0201] In some implementations, changing the first portion of the user interface object (e.g., 604) to point to the detected location of the user in the detection field includes: moving the first portion of the user interface object in the detection field via a display component from a location more than a predetermined distance (e.g., 0.1 meters to 10 meters) from the detected location of the user to a location no more than a predetermined distance from the detected location of the user (e.g., as described above relative to the detected location). Figures 6A to 6G(The above describes, for example, tilting the user interface object, moving at least a portion of the user interface object (e.g., eyes, mouth, ears, and / or nose) at an angle to a direction closer to the user's location, changing a portion of the user interface object (e.g., eyes, mouth, ears, and / or nose) to point closer to the user's location, and / or changing the eyes of the user interface object to appear closer to the user's location.) When consent is expressed for a first input, the first portion of the user interface object is moved in the detection field of the computer system from a location more than a predetermined distance from the user's detected location to a location no more than a predetermined distance from the user's detected location. Doing so provides the user with feedback that the computer system has detected and consented to the input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0202] In some implementations, changing a first portion of the user interface object (e.g., 604) to point to the detected location of the user in the detection field includes: displaying via a display component the first portion of the user interface object pointing (e.g., appearing to point to and / or protrude toward) the detected location of the user in the detection field (e.g., in the user's orientation) (e.g., as described above relative to the detected location). Figures 6A to 6G (as described above). When consent is expressed for the first input, the first part of the user interface object is displayed pointing to the user's detected location in the detection field of the computer system. This provides feedback to the user that the computer system has detected and consented to the input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0203] In some implementations, the first input (e.g., 606) does not include an explicit instruction to change the user interface object (e.g., the explicit instruction is not an explicit request to move the user interface object) (e.g., the explicit instruction is not an explicit request to change a part of the user interface object) (e.g., the first input includes a question, statement, and / or any other voice input corresponding to the user interface object, without explicit requests and / or instructions) (e.g., as described above relative to...) Figure 6A (as described).
[0204] In some implementations, displaying the user interface (e.g., 602) when the user interface object is changed in a manner indicating eye contact with the user includes: moving a second portion of the user interface object via a display component (e.g., as described above relative to...). Figures 6E to 6F(e.g., moving a portion of the user interface object in the user's direction when consent is expressed; e.g., moving a portion of the user interface object away from the user when consent is not expressed; e.g., moving a portion of the user interface from a first corresponding position to a second corresponding position different from the first corresponding position). Moving a second portion of the user interface object when consent is expressed for a first input provides the user with feedback that the computer system has detected and consented to the input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0205] In some implementations, when (and / or after) a user interface object is displayed in a manner indicating eye contact with the user (e.g., 604), the computer system detects a second input via the one or more input devices (e.g., 608) (e.g., verbal input (e.g., audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)) (e.g., as described above relative to...). Figures 6A to 6C (as described above). In some embodiments, in response to the detection of a second input (e.g., 608), the computer system continues to display the user interface (e.g., 602) while changing the user interface object via the display component in a first manner (e.g., pose, manner, and / or visual manner) that does not indicate eye contact with the user (e.g., pose, manner, and / or visual manner) (e.g., as described above relative to...). Figures 6A to 6C (e.g., regardless of whether consent is determined) (e.g., before consent is determined) (e.g., immediately after input is detected). Changing the user interface object in a first manner that does not indicate eye contact with the user in response to the detection of a second input allows the computer system to suggest whether to express disagreement and / or consent regarding the second input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components. When the user interface object is changed to indicate eye contact, another voice input is detected; and in response, the user interface object is changed or not changed to indicate eye contact based on whether the user consents.
[0206] In some implementations, when (and / or after) a user interface object is displayed in a manner indicating eye contact with the user (e.g., 604), the computer system detects a third input via the one or more input devices (e.g., 608) (e.g., as described above relative to...). Figures 6A to 6C(e.g., via verbal input (e.g., verbal input, audible requests, audible commands, and / or audible statements) and / or nonverbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, in response to detecting a third input, based on determining that no consent has been expressed regarding the third input (e.g., 608), the computer system continues to display the user interface via a display component (e.g., 602) while changing the user interface object in a second manner that does not indicate eye contact with the user (e.g., 604) (e.g., as described above relative to...). Figures 6A to 6C In some embodiments, expressing consent includes: the computer system verifying and / or confirming that the content of the voice input is true (e.g., true based on facts of publicly available data and / or true based on facts of private data (e.g., user data, data associated with a user account, and / or data associated with one or more groups and / or one or more groups to which the user belongs)), and that the content of the voice input corresponds to the user and / or conforms to one or more characteristics associated with, related to, and / or of the user (e.g., facts, personality traits, likes, dislikes, and / or favorites). In some embodiments, in response to detecting a third input, based on determining that consent has been expressed regarding the third input (e.g., 608), the computer system continues to display the user interface via a display component (e.g., 602) without changing the user interface object in a second manner that does not indicate eye contact with the user (e.g., 604) (e.g., as described above relative to...). Figures 6A to 6C (As stated above). Based on determining whether consent is expressed or not expressed regarding the third input, the user interface object is displayed without changing the user interface object in a second manner that does not indicate eye contact with the user and / or while displaying the user interface object, it is changed in a second manner that does not indicate eye contact with the user. Doing so provides the user with feedback that the computer system has detected the input and allows the computer system to provide feedback on whether the user agrees to the input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0207] In some implementations, when (and / or after) a user interface object is displayed in a manner indicating eye contact with the user (e.g., 604), the computer system detects a fourth input via the one or more input devices (e.g., 608) (e.g., via verbal input (e.g., verbal input, audible request, audible command, and / or audible statement) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gesture, and / or mouse click)) (e.g., as described above relative to...). Figure 6CIn some embodiments, in response to detecting a fourth input (e.g., 608) and based on determining that disagreement was expressed regarding the fourth input (e.g., the user input is incorrect) (e.g., the user gave an incorrect statement) (e.g., the user gave a negative answer and / or statement in response to a positive input (e.g., a response and / or statement is required)), the computer system continues to display the user interface object via a display component in a manner indicating eye contact with the user (e.g., 604) (and / or performs operations indicating the disagreement, such as indicating shaking the head of the user interface object and / or outputting a negative statement (e.g., "no" and / or "I disagree"), as described above relative to... Figure 6C In some embodiments, expressing consent includes: the computer system verifying and / or confirming that the content of the voice input is true (e.g., true based on publicly available data and / or true based on private data (e.g., user data, data associated with a user account, and / or data associated with one or more groups and / or one or more groups to which the user belongs)), and that the content of the voice input corresponds to the user and / or conforms to one or more characteristics associated with, related to, and / or of the user (e.g., facts, personality traits, likes, dislikes, and / or favorites). In some embodiments, in response to detecting a fourth input and determining that no disagreement has been expressed regarding the fourth input, the computer system does not continue to display the user interface object via the display components in a manner indicating eye contact with the user and / or displays the user interface object in a manner not indicating eye contact with the user. Continuing to display the user interface object in a manner indicating eye contact with the user when no consent has been expressed regarding the fourth input provides the user with consistent and / or continuous interaction with the user interface object, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of the display components.
[0208] In some implementations, when (and / or after) a user interface object is displayed in a manner indicating eye contact with the user (e.g., 604), the computer system detects a fifth input via one or more input devices (e.g., as described above relative to...). Figures 6A to 6C In some embodiments, in response to the detection of a fifth input (e.g., 608), based on the determination that disagreement was expressed with respect to the fifth input, the computer system displays a user interface object via a display component in a third manner that does not indicate eye contact with the user (e.g., 604) (e.g., as described above relative to...). Figure 6B(as described above). In some implementations, in response to the detection of a fifth input, and based on the determination that disagreement was expressed regarding the fifth input, the computer system modifies the display of the user interface object from a manner indicating eye contact with the user to a third manner not indicating eye contact with the user. When disagreement was not expressed regarding the fifth input, displaying the user interface object while changing it to the third manner not indicating eye contact with the user provides feedback to the user that the computer system disagrees with the detected input, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0209] In some implementations, in response to the detection of a first input (e.g., 606), the computer system continues to display a portion of the user interface (e.g., 602) excluding the user interface object (e.g., 604) via a display component (e.g., changing a portion of the user interface object while maintaining the currently displayed content) (e.g., changing a portion of the user interface object while maintaining all other currently displayed content on the user interface) (e.g., changing the user interface object from a manner indicating eye contact to a manner not indicating eye contact includes continuing to display that portion of the user interface object), while simultaneously changing the user interface object in a manner indicating eye contact with the user (e.g., as described above relative to...). Figures 6A to 6G (as described above). Continuing to display a portion of the user interface, excluding the user interface object, while the user interface object is changed in a manner that indicates eye contact with the user allows the computer system to maintain the user interface while still providing feedback that the computer system agrees with the detected input, and reducing visual interference from displaying completely different user interfaces, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0210] In some implementations, after changing the user interface object in a manner that indicates eye contact with the user (e.g., 604) (e.g., as described above relative to...), Figures 6A to 6CAs described above, based on determining that no predetermined time period (e.g., 0.1 seconds to 10 seconds) has elapsed since the user interface object (e.g., 604) changed in a manner indicating eye contact with the user, the computer system continues to display the user interface object via the display component in a manner indicating eye contact with the user. In some embodiments, after the user interface object changes in a manner indicating eye contact with the user, based on determining that a predetermined time period has elapsed since the user interface object (e.g., 604) changed in a manner indicating eye contact with the user, the computer system abandons continuing to display the user interface object via the display component in a manner indicating eye contact with the user (e.g., as described above relative to...). Figures 6A to 6C (as described above). In some embodiments, after a predetermined time period, the computer system stops displaying the user interface object in that pose (e.g., returns to a previous pose). In some embodiments, after changing the user interface object in a manner indicating eye contact with the user and based on determining that a predetermined time period has elapsed since the user interface object was changed in a manner indicating eye contact with the user, the computer system changes the user interface object from a manner indicating eye contact with the user to a different manner not indicating eye contact with the user. Continuing to display the user interface object in a manner indicating eye contact with the user based on determining that a predetermined time period has not elapsed since the user interface object was changed in a manner indicating eye contact with the user, and / or abandoning to continue displaying the user interface object in a manner indicating eye contact with the user based on determining that a predetermined time period has elapsed since the user interface object was changed in a manner indicating eye contact with the user, allows the first computer system to automatically change the display of the user interface object based on the recency of the detected input determined by the computer system to be uniform, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation when a set of conditions has been met without further input, and / or allowing the computer system to avoid burn-in of display components.
[0211] In some implementations, the first input (e.g., 606) includes a problem (e.g., as described above relative to...). Figure 6A (as described above). In some implementations, determining that consent has been expressed with respect to the first input (e.g., 606) includes: determining that the question corresponds to an affirmative answer (e.g., as described above relative to...). Figures 6A to 6C (e.g., the answer and / or response is an affirmation, approval, and / or acceptance of the question). In some embodiments, determining that no consent has been expressed with respect to the first input (e.g., 606) includes: determining that the question corresponds to a negative answer (e.g., as described above relative to...). Figures 6A to 6C(e.g., the answer and / or response is disapproval, non-acceptance, and / or non-confirmation of the question). When the first input is determined to be a question with an affirmative answer, displaying the user interface object in a manner that indicates eye contact with the user provides feedback to the user that the computer system has detected the input and determined the question to be correct, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components. When the first input is determined to be a question with a negative answer, displaying the user interface object in a manner that indicates eye contact with the user without changing the user interface object provides feedback to the user that the computer system has detected the input and determined the question to be incorrect, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0212] In some implementations, the first input (e.g., 606) includes a statement (e.g., as described above relative to...). Figure 6A (as described above). In some embodiments, determining that consent has been expressed with respect to the first input (e.g., 606) includes: determining that the statement corresponds to a statement of fact (e.g., a statement that is factually true (e.g., based on data, such as public data, private data, and / or customized data)) (e.g., as described above relative to...). Figures 6A to 6C (as described above). In some embodiments, determining that no consent has been expressed with respect to the first input (e.g., 606) includes: determining that the statement does not correspond to a statement of fact (e.g., as described above relative to...). Figures 6A to 6C (e.g., the answer and / or response is disapproval, non-acceptance, and / or non-confirmation of the question). When the first input is determined to be a statement of fact, displaying the user interface object in a manner that indicates eye contact with the user provides feedback to the user that the computer system has detected the input and determined that the statement conforms to the facts, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components. When the first input is determined not to be a statement of fact, displaying the user interface object in a manner that indicates eye contact with the user without changing the user interface object provides feedback to the user that the computer system has detected the input and determined that the statement does not conform to the facts, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0213] In some implementations, determining that consent has been expressed with respect to the first input (e.g., 606) includes determining that the first input satisfies the first user's user preferences (e.g., user preferences, known preferences, preferences learned over time, historical preferences, and / or preferences based on usage patterns (e.g., how the user interacts with the computer system and / or how the user interacts with one or more applications and / or external devices (e.g., fitness trackers, smart devices (e.g., smart lights, smart locks, and / or smart blinds) and / or speakers))) (e.g., the user likes blue, and eye contact is made with the user when the phrase includes "blue" (e.g., selected and / or chosen)). In some implementations, determining that consent has not been expressed with respect to the first input includes determining that the first input does not satisfy the first user's user preferences (e.g., as described above relative to...). Figures 6A to 6C (For example, if a user prefers blue, the user interface object does not make eye contact with the user when the phrase includes "red" (e.g., selected and / or chosen).
[0214] In some implementations, in response to the detection of a first input via the one or more input devices (e.g., 606) (e.g., as described above relative to...), Figures 6A to 6C As described above, the computer system provides output corresponding to the first input (e.g., 610) via one or more output devices (e.g., speakers, haptic output devices, and / or audio generation components) communicating with the computer system (e.g., 600), such output (e.g., output different from changing the user interface object, audio output, and / or haptic output), wherein: based on determining that consent has been expressed with respect to the first input (e.g., 610), the output includes a first phrase indicating consent to the user's preference (e.g., "yes," "I agree," and / or "you are right") (e.g., and the display component continues to display the user interface while changing the user interface object in a manner indicating eye contact with the user) (e.g., the user likes blue, and when describing a blue wall, the user interface object makes eye contact with the user) (e.g., as described above relative to...). Figures 6A to 6C The output includes, based on the determination that no agreement was expressed with respect to the first input (e.g., 610), a second phrase indicating the user's disagreement preference (e.g., "No," "I disagree," and / or "I don't think that's right") (e.g., as described above relative to...). Figures 6A to 6C(For example, and the display component continues to display the user interface without changing the user interface object in a manner that does not indicate eye contact with the user) (For example, if the user likes blue, the user interface object does not make eye contact with the user when describing a yellow wall). Outputting a first phrase and / or a second phrase when consent is determined upon meeting specified conditions allows the computer system to adapt operations to the user based on the determination of user preferences and provides audible feedback to confirm the determination of the first input (e.g., agreeing with and / or disagreeing with user preferences, and / or whether it is a saved user preference), thereby providing improved feedback, reducing the amount of input required to perform the operation, performing the operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of the display component.
[0215] In some implementations, continuing to display the user interface via a display component (e.g., 602) without changing the user interface object in a manner that indicates eye contact with the user (e.g., 604) does not include outputting a second phrase (e.g., as described above relative to...). Figures 6A to 6C (as described).
[0216] It should be noted that the above text regarding method 800 (for example, Figure 8 The details of the process described herein also apply in a similar manner to the methods described below / above. In some embodiments, process 700 may optionally include one or more features of the various methods described above with reference to process 800. In some embodiments, the computer system may use one or more techniques of process 800 to determine that interaction conditions have been met, and use one or more techniques of process 700 to display content-oriented objects. For the sake of brevity, these details will not be repeated below.
[0217] Figure 9 This is a flowchart illustrating a method for using a computer system to de-emphasize an object, according to some implementation schemes. Process 900 is executed at a computer system (e.g., 100, 200, and / or 600). Some operations in process 900 may be combined, some operations may be changed in order, and some operations may be omitted.
[0218] As described below, process 900 provides an intuitive way to de-emphasize objects. This method reduces the cognitive burden on users when de-emphasing objects, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to de-emphasize objects faster and more efficiently saves power and increases the time interval between battery charging.
[0219] In some embodiments, process 900 is performed at a computer system (e.g., 600) that communicates with display components (e.g., a display screen, projector, and / or touch-sensitive display), cameras (e.g., one or more wide-angle cameras, telephoto cameras, and / or ultra-wide-angle cameras), and one or more input devices (e.g., cameras, depth sensors, and / or microphones). In some embodiments, the computer system is a watch, phone, tablet, fitness tracker, processor, head-mounted display (HMD) device, public utility, media device, speaker, television, and / or personal computing device.
[0220] Upon (and / or subsequently and in response to) the detection of a first entity (e.g., a user, animal, and / or object) in the camera's field of view, the computer system displays (902) a user interface object (e.g., 604) representing a portion of the user (e.g., a person, animal, and / or object) in a first manner to indicate that the user interface object points to the first user in the detection field (e.g., the microphone's detection field, the camera's field of view, and / or the depth sensor's detection field) of the computer system (e.g., 600), wherein the user interface object is displayed at a first size (e.g., relative to the above). Figure 6D (as described).
[0221] When a user interface object that indicates a part of the user is displayed (e.g., 604), the computer system detects (904) a request to interact with the content (e.g., 610) (e.g., voice input including the indication of the content, one or more inputs and / or air gestures performed by a part of the body of an entity) (e.g., via verbal input (e.g., verbal input, audible request, audible command and / or audible statement) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gesture and / or mouse click)).
[0222] In response to (906) detecting a request to interact with the content (e.g., 610), the computer system displays (908) at least a first portion of the content (e.g., 612) (e.g., as described above relative to...). Figure 6F (as described).
[0223] In response to (906) detecting a request to interact with content, the computer system displays (910) a user interface object representing a portion of the user (e.g., 604) in a second size smaller than the first size and in a second manner, to indicate that the user interface object points to a portion of the content (e.g., 612) (e.g., as described above relative to...). Figures 6E to 6F(and not pointing to a first entity in the detection field of the camera). The first entity is detected and a user interface object is displayed, which represents a portion of the user in a first manner pointing to a first user in the detection field of the computer system. This allows the computer system to confirm the detection of the first user, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components. Responding to the detection of a request to interact with the content, displaying at least a first portion of the content enables the computer system to provide interactive content to the user, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components. In response to a detected request to interact with content, the user interface object is displayed in a second size smaller than the first size and in a second manner to indicate that the user interface object points to the content portion. This reduces visual clutter when adding content to the user interface and / or provides the user with an indication of the content to be displayed, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0224] In some implementations, prior to detecting a content (e.g., 610) interaction request, a portion of a user interface object representing the user (e.g., 604) is displayed at a first location (e.g., of a display component and / or user interface). Figure 6D (as described above). In some embodiments, in response to detecting the request (e.g., 610), displaying a portion representing the user (e.g., 604) of a user interface object at a second size smaller than the first size includes stopping the display of the portion of the user interface object at the first position (e.g., as described above relative to...). Figures 6D to 6F In some implementations, in response to detecting a request to interact with the content (e.g., 610), a first portion of the content is displayed at a first location (e.g., as described above relative to...). Figure 6F(as described above). In some embodiments, in response to detecting a request to interact with the content, the computer system replaces the display of that portion of the user interface object at the first location with the display of the first portion of the content at the first location. Responding to the detection of a request to interact with the content and stopping the display of the user interface object representing the user at the first location, and displaying the content portion at the first location, reduces visual clutter when adding content to the user interface and / or avoids displaying a completely different user interface, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation when a set of conditions has already been met without further input, and / or allowing the computer system to avoid burn-in of display components.
[0225] In some implementations, the computer system abandons displaying at least that portion of the content (e.g., as described above relative to) before detecting a request to interact with the content (e.g., 610). Figure 6D (e.g., simultaneously displaying a user interface object representing a portion of the user in a first manner and at a first size). In some embodiments, at least a first portion of the content is not displayed until a request to interact with the content is detected. Not displaying at least a portion of the content before a request to interact with the content is detected reduces visual clutter in the user interface, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0226] In some implementations, the one or more input devices include a microphone. In some implementations, detecting a request to interact with content (e.g., 610) includes receiving, via the microphone, a request corresponding to the content (e.g., as described above). Figure 6D Voice input (e.g., voice commands, voice requests, audible questions, and / or audible statements) (e.g., as described above with respect to process 700 and / or as described above with respect to process 800) (e.g., voice input for initiating a request involving interaction with content). In some embodiments, the voice input may be received via a text-to-speech device within a computer system and / or via a second computer system.
[0227] In some implementations, requests for content interaction do not include explicit requests (e.g., 610) (e.g., as described above relative to...). Figure 6D (e.g., requests to interact with content) (e.g., explicit requests do not include requirements for verification information (text, contacts, etc.)) (e.g., explicit requests include requests to open content that requires interaction with the content, such as "open content" and / or "show content") (e.g., the request does not include commands to interact with the content (e.g., "open content", "move content" and / or "show content")).
[0228] In some implementations, a request to interact with content (e.g., 610) is a first request to interact with a first part of the content (e.g., as described above relative to...). Figure 6D In some embodiments, while (and / or after) displaying at least a first portion of the content and displaying a user interface object representing the user (e.g., 604) at a second size, the computer system detects, via the one or more input devices, a request to interact with a second portion of the content that is different from the first portion (e.g., 612) (e.g., the same content but a different portion and / or new content) (e.g., swipe, tap, press and hold, and / or voice input) (e.g., via verbal input (e.g., verbal input, audible request, audible command, and / or audible statement) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gesture, and / or mouse click) (e.g., as described above relative to...). Figure 6F (as described above). In some embodiments, user input is used based on a computer system to determine that interaction with a second portion of the content is required and / or with a second portion of the content that is not shown in the first portion of the content. In some embodiments, in response to detecting a request to interact with the second portion of the content, the computer system displays at least the second portion of the content via a display component, wherein the second portion of the content replaces the first portion of the content (e.g., as described above relative to...). Figure 6F In some embodiments, in response to detecting a request to interact with a second portion of the content, the computer system stops displaying at least a first portion of the content via a display component. In some embodiments, in response to detecting a request to interact with a second portion of the content, the computer system continues to display, via the display component, a user interface object representing the user portion (e.g., 604) at a second size and in a second manner (e.g., as described above relative to...). Figure 6F (as described above). Displaying at least a second portion of the content, wherein when a request to interact with the second portion of the content is detected, the second portion of the content replaces the first portion of the content. This allows the computer system to switch to new content without displaying a completely different user interface, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0229] In some implementations, after a user interface object representing the user (e.g., 604) is displayed at a second size and in a second manner to indicate that the user interface object points to a portion of content (e.g., 612), and based on determining that one or more criteria of a first set are met (e.g., agreement criteria (e.g., as described above relative to process 800) and / or importance criteria (e.g., whether the user interface object and / or request are important), waiting for user input, and / or completion of the output of second content) (e.g., according to the request), the computer system displays the user interface object representing the user (e.g., the first manner, and / or a non-second manner) via a display component at a second size and in a third manner (e.g., a first manner, and / or a non-second manner) to indicate that the user interface object points to a first user in the detection field of the computer system (e.g., 600) (e.g., as described above relative to process 800) Figures 6F to 6G (as described above). In some embodiments, after displaying a user interface object representing a person in a second size and in a second manner, and based on determining that one or more criteria of a first set are met, the computer system stops displaying in the second manner to indicate that the user interface object points to a portion of the content. In some embodiments, the portion of the content is a first part of the content. Based on meeting one or more criteria of the first set, displaying a user interface object representing a user in a second size and a third manner (e.g., a first manner and / or a non-second manner) to indicate that the user interface object points to a first user in the detection field of the computer system allows the computer system to confirm to the user that one or more criteria (e.g., agreement criteria and / or importance criteria) are met, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0230] In some implementations, after a user interface object representing the user (e.g., 604) is displayed at a second size and in a second manner to indicate that the user interface object points to a portion of the content (e.g., 612), and based on the determination that one or more criteria of the first set are not met (e.g., no consent has been expressed, a request for interaction has been detected, and / or the entity is still interacting with the content), the computer system continues to display the user interface object representing the user at a second size and in a second manner (e.g., second manner and / or non-first manner) via a display component to indicate that the user interface object points to a portion of the content (e.g., as described above relative to the first manner). Figure 6F(as described). Based on the failure to meet one or more criteria in the first set, the user interface object representing the user is displayed in a second size and in a second manner to indicate that the user interface object points to the part of the content. Doing so allows the computer system to confirm to the user that one or more criteria are not met, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0231] In some implementations, displaying a user interface object representing the user portion (e.g., 604) in a second manner and at a second size to indicate that the user interface object points to a content portion (e.g., 612) includes: changing the user interface object representing the user portion to a second size (e.g., as described above relative to the user interface object) before displaying the user interface object representing the user portion in the second manner to indicate that the user interface object points to a content portion. Figures 6E to 6F (as described above). Before displaying the user interface object representing the user's portion in the second manner, the user interface object representing the user's portion is changed to a second size to indicate that the user interface object points to the portion of content. Doing so allows the computer system to smoothly transition to the addition of content without displaying a completely different user interface, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation when a set of conditions has already been met without requiring further input, and / or allowing the computer system to avoid burn-in of display components.
[0232] In some embodiments, displaying a user interface object representing a portion of the user (e.g., 604) via a display component in a first manner to indicate that the user interface object points to a first entity in the detection field includes: displaying via the display component a representation of an eye displayed with a first set of one or more characteristics (e.g., in a first position, in a first shape, and / or in a first orientation and / or distance) (e.g., as described above relative to...). Figures 6D to 6G In some embodiments, displaying a user interface object representing a portion of the user (e.g., 604) via a display component in a second manner to indicate that the user interface object points to a portion of content (e.g., 612) includes: displaying a representation of an eye (e.g., as described above) having a second set of characteristics different from the first set of one or more characteristics (e.g., in a first position, in a first shape, and / or in a first orientation and / or distance). Figure 6F(as described). When a user interface object representing the user's portion is displayed in a first manner to indicate that the user interface object is pointing to a first entity in the detection field, a representation of an eye with one or more characteristics is displayed, and / or when a user interface object representing the user's portion is displayed in a second manner to indicate that the user interface object is pointing to a portion of the content, a representation of an eye with a second set of characteristics is displayed. Doing so allows the computer system to provide an indication of where the user interface object representing the user's portion is being pointed to, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0233] In some embodiments, after at least a first portion of the content is displayed (e.g., 612) and based on a determination that the content should be removed (e.g., a request to close the content, the content can no longer be displayed (e.g., content output completed) and / or the computer system is waiting for input) (and in some embodiments, when displaying a user interface object representing the user's portion at a second size) (e.g., as described above relative to...) Figures 6F to 6G As described above), the computer system stops displaying at least the first part of the content (e.g., 604) (e.g., as described above relative to...). Figures 6D to 6E (as described above). In some embodiments, after at least a first portion of the content has been displayed and based on a determination that the content should be removed, the computer system re-displays the user interface object representing the user portion (e.g., 604) via a display component at a first size (and in some embodiments, in a first manner). Figures 6D to 6E (as described above). When it is determined that content should be removed, at least a first portion of the content is stopped from being displayed, and the user interface object representing the user's portion is re-displayed at a first size. This allows the computer system to correctly use the available space in the user interface, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0234] In some implementations, after at least a first portion of the content is displayed (e.g., 612) and based on the determination that the content should be removed and one or more third-group criteria are met (e.g., agreement criteria (e.g., as described above regarding process 800) and / or importance criteria (e.g., whether the user interface object and / or request is important), waiting for user input and / or completion of the second content output) (e.g., according to the request), the user interface object is displayed in a first manner (e.g., 604) to indicate that the user interface object points to a first entity in the detection field of the computer system (e.g., 600) (e.g., as described above relative to...). Figures 6D to 6G(as described above). When it is determined that content should be removed and one or more criteria of the third group are met, a user interface object is displayed in a first manner to indicate that the user interface object points to a first entity in the detection field of the computer system. Doing so allows the user interface object to confirm that one or more criteria of the third group have been met, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0235] In some implementations, after at least a first portion of the content is displayed (e.g., 612) and based on the determination that the content should be removed and one or more criteria from a fourth group are met (e.g., agreement criteria (e.g., as described above regarding process 800) and / or importance criteria (e.g., whether the user interface object and / or request is important), waiting for user input and / or completion of the second content output) (e.g., according to the request) (e.g., different from one or more criteria from a third group), the user interface object is displayed in a fourth manner (e.g., 604) to indicate that the user interface object does not point to the first entity (e.g., as described above relative to the first entity). Figures 6D to 6G (For example, the fourth approach includes displaying the user interface object to be pointed to in the detection field of one or more input devices, but not in the location where the first user is detected) (For example, the fourth approach includes displaying the user interface object to be pointed to in the direction of the first user interface, but not pointing to content). When it is determined that content should be removed and one or more criteria of the third set are not met, the user interface object is displayed in the fourth approach to indicate that the user interface object is not pointing to the first entity. This allows the user interface object to confirm that one or more criteria of the third set have not been met, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has been met, and / or allowing the computer system to avoid burn-in of display components.
[0236] In some implementations, when a user interface object representing a portion of the user (e.g., 604) is displayed in a second manner, the computer system detects an interaction condition corresponding to the content (e.g., as described above with respect to process 800) (e.g., an entity makes a new request to interact with new content, an entity makes an explicit request to stop viewing the content, and / or an entity makes a request to close the content, and a second entity is detected in the detection field) (e.g., the interaction condition described above with respect to process 800) (e.g., via verbal input (e.g., verbal input, audible requests, audible commands, and / or audible statements) and / or non-verbal input (e.g., swipe input, hold and drag input, gaze input, air gestures, and / or mouse clicks)) (e.g., as described above with respect to... Figure 6FIn some embodiments, in response to detecting the interaction condition, based on determining that the interaction condition meets one or more criteria of a fifth group (e.g., the first entity is interacting with content and / or the first entity makes a new request), the computer system displays a user interface object representing the user's portion (e.g., 604) via a display component in a first manner to indicate that the user interface object is directed toward the first entity in the camera's field of view (e.g., in a direction corresponding to the first entity and / or at the first entity), wherein the first entity initiated the interaction condition (e.g., as described above relative to...). Figures 6D to 6G In some implementations, in response to detecting the interaction condition, based on the determination that the interaction condition corresponding to the content satisfies one or more criteria of a sixth group (e.g., not satisfying one or more criteria of the fifth group) that differ from one or more criteria of the fifth group (e.g., an entity other than the first entity is interacting with the content and / or a second entity makes a new request), the computer system displays the user interface object via a display component in such a manner that the user interface object representing the user's portion (e.g., 604) points to the second entity detected in the field of view of the computer system (e.g., 600) (e.g., in the direction corresponding to the second entity and / or at the second entity), wherein the second entity initiated the interaction condition (e.g., as described above relative to the first entity). Figures 6D to 6G (as described above). When a first entity in the camera's field of view initiates an interaction condition, a user interface object representing the user's portion is displayed in a first manner to indicate that the user interface object points to the first entity, and / or when a second entity in the computer system's field of view initiates an interaction condition, a user interface object representing the user's portion is displayed in a manner to indicate that the user interface object points to the second entity. This allows the computer system to identify who initiated the request to interact with the content, thereby providing improved feedback, reducing the amount of input required to perform an operation, performing an operation without further input when a set of conditions has already been met, and / or allowing the computer system to avoid burn-in of display components.
[0237] It should be noted that the above text regarding method 900 (for example, Figure 9 The details of the process described herein also apply in a similar manner to the methods described below / above. In some embodiments, process 700 may optionally include one or more features of the various methods described above with reference to process 900. In some embodiments, the computer system may use one or more techniques of process 900 to determine that interaction conditions have been met, thereby using one or more techniques of process 700 to display content-oriented objects. For the sake of brevity, these details will not be repeated below.
[0238] Figures 10A to 10FExemplary user interfaces for providing content are illustrated according to some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including... Figures 11 to 14 The process in.
[0239] Figures 10A to 10F A computer system 1000 as a tablet computer is illustrated. It should be understood that the computer system 1000 can be other types of computer systems, such as smartphones, smartwatches, laptops, mobile devices, rotatable devices, smart display devices, utilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 1000 includes one or more sensors (e.g., camera sensors, lidar detectors, motion sensors, infrared sensors, and / or microphones) and / or communicates with one or more sensors. In some embodiments, the computer system 1000 includes one or more output devices (e.g., display screens, projectors, touch-sensitive displays, haptic output devices, and / or speakers) and / or communicates with one or more output devices. In some embodiments, the computer system 1000 includes one or more moving components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with them. In some embodiments, the one or more moving components are configured to move (e.g., one or more physical parts of the computer system 1000). In some implementations, computer system 1000 includes one or more components and / or features described above with respect to computer system 100 and / or electronic device 200.
[0240] like Figures 10A to 10F As shown, computer system 1000 displays user interface 1002. In some embodiments, user interface 1002 is a home screen user interface that includes representations of one or more application controls and / or virtual assistants (e.g., 1004, 1014, 1022, and / or 1026). In some embodiments, user interface 1002 is an educational application user interface. In some embodiments, user interface 1002 may be an application for teaching users how to read. In some embodiments, user interface 1002 is an office assistant user interface. In some embodiments, user interface 1002 may be an application user interface that provides users with updates on the latest field reports. In some embodiments, user interface 1002 is an entertainment application user interface. In some embodiments, user interface 1002 may be an e-book application user interface for reading content aloud to users.
[0241] like Figures 10A to 10FAs shown, the user interface 1002 includes different system digital images at different times, including user interface object 1004 (e.g., in...). Figures 10A to 10B (in the middle), the second system digital image 1014 (for example, in Figure 10C and Figure 10F (in the middle), the third system digital image 1022 (for example, in Figure 10D (in the middle) and the fourth system digital image 1026 (for example, in Figure 10E (in Chinese). Figures 10A to 10F In the examples, the different system digital avatars are anthropomorphic visual representations of artificial intelligence applications and / or virtual assistants that visually change based on content output and / or input, and react to and / or respond to input from a user near an input device of computer system 1000. In some embodiments, a system digital avatar is a digital avatar corresponding to a system process of the computer system (e.g., 1000) (e.g., managed, controlled, output, created, and / or requested by that system process). In some embodiments, a system digital avatar may be displayed while other applications are running in the background. In some embodiments, a system process is a process corresponding to the operating system of the computer system. In some embodiments, a digital avatar corresponds to an application process, which corresponds to a software application installed and / or executed by the computer system. In some embodiments, a system process is different from an application process. In some embodiments, computer system 1000 is able to access and / or output content from one or more system processes and / or one or more application processes, making it appear as if one or more digital avatars (e.g., user interface object 1004, second system digital avatar 1014, third system digital avatar 1022, and / or fourth system digital avatar 1026) are outputting content. In some implementations, the application corresponding to the system's digital avatar can access educational and e-book applications, allowing the system's digital avatar to appear as if it is using e-books for teaching. Similarly, the application corresponding to the system's digital avatar can access news and internet applications, allowing the system's digital avatar to appear as if answering questions posed by users about information in news articles.
[0242] In some embodiments, computer system 1000 displays a system digital avatar having a visual appearance and / or facial expression (e.g., happy, sad, confused, afraid, surprised, and / or calm) corresponding to the content being output by computer system 1000. In some embodiments, if the content being output by computer system 1000 includes a dragon, the system digital avatar may have the appearance of a dragon. Similarly, if the content being output by computer system 1000 is intended to be humorous, the system digital avatar may have a laughing appearance. In some embodiments, computer system 1000 displays different system digital avatars with different appearances (e.g., different colors (e.g., color groups, skin tones, red, orange, yellow, green, blue, and / or purple), textures (e.g., skin, hair, fur, scales, plastic, glass, feathers, and / or wood), accessories (e.g., hats, glasses, monocles, wands, books, collars, bows, wings, halos, and / or crowns), and / or facial types (e.g., human, animal, anthropomorphic object, alien, ordinary face, fantasy creature, and / or a collection of similar faces)). In some embodiments, the computer system 1000 outputs audio corresponding to the voice of the system avatar via one or more output devices (e.g., speakers) included in and / or communicating with the computer system 1000. In some embodiments, if the displayed system avatar has the appearance of a mouse, the computer system 1000 may output a squeaking audio. Similarly, if the displayed system avatar has the appearance of a giant, the computer system 1000 may output a deep, loud voice audio. In some embodiments, the computer system 1000 outputs audio with different sounds (e.g., male voice, female voice, high / low pitch, with / without accent, soft voice, and / or loud voice) corresponding to different system avatars, depending on which system avatar is being displayed and / or what appearance the system avatar has.
[0243] In some embodiments, the computer system 1000 displays a system digital image with one or more movement patterns via a display. In some embodiments, if the system digital image has the appearance of a rabbit, it may be displayed as having a bouncing motion (e.g., the system digital image may bounce up and down when the computer system 1000 outputs content corresponding to it). Similarly, if the content output by the computer system 1000 is intended to create a horror effect, the system digital image may be displayed as trembling. In some embodiments, the computer system 1000 displays different system digital images with different sets of one or more movement patterns (e.g., bouncing, swaying, scaling, and / or dancing) via a display. In some embodiments, the system digital image corresponds to one or more movement patterns of the computer system 1000 (via one or more movement components). In some embodiments, a portion of the computer system 1000 may physically move left and right, creating a visual effect of the system digital image slithering when it has the appearance of a snake. Similarly, a portion of the computer system 1000 may move up and down, creating a visual effect of the system digital image nodding when it responds positively to input from a user. In some implementations, different system digital images correspond to different movement modes (e.g., swaying, bending, rotating, and / or tilting) of a portion of the computer system 1000.
[0244] exist Figures 10A to 10F The example shown illustrates four system digital images. In some implementations, there may be more or fewer than four system digital images. Figures 10A to 10F In some implementations, only one system avatar is displayed at a time. In other implementations, multiple system avatars are displayed simultaneously. In some implementations, if the computer system 1000 is outputting content that includes a dialogue between multiple characters, the computer system 1000 may display multiple system avatars, thereby allowing each system avatar to act as performing a different part of the dialogue and to act as interacting together.
[0245] Figures 10A to 10F The right side includes Figure 1006. Figure 1006 is a visual aid that represents the physical environment including computer system 1000. Figure 1006 includes computer system representation 1008, field of view 1008a, and user representation 1010. Field of view 1008a represents the field of view of one or more forward sensors of computer system 1000. The positioning of user representation 1010 and computer system representation 1008 within Figure 1006 represents the user's real-world positioning relative to computer system 1000.
[0246] like Figure 10AAs shown, computer system 1000 displays user interface 1002 with user interface object 1004. In some embodiments, user interface object 1004 is the default system digital image of computer system 1000. In some embodiments, the default system digital image is a digital image preset (e.g., previously configured) by the publisher of user interface 1002 (e.g., or otherwise selected by default, not based on user input and / or the content being output). In some embodiments, the default system digital image is a digital image preset by the user. In this example, when the content output by computer system 1000 does not correspond to a specific and / or different system digital image, computer system 1000 displays the default system digital image (e.g., user interface object 1004). In some embodiments, when displaying user interface object 1004, the movement of user interface object 1004 is synchronized with the content output by computer system 1000. In some embodiments, computer system 1000 may synchronize the visual movement of the facial features of user interface object 1004 with audio output to create the visual effect that user interface object 1004 is speaking to the user.
[0247] like Figure 10A As shown, computer system 1000 displays user interface object 1004 at the center of user interface 1002, which occupies a large portion of the user interface 1002. In other examples, user interface object 1004 may be displayed in other locations and / or occupy more or less area of user interface 1002. In this example, computer system 1000 displays user interface object 1004 as a representation of a face. In some embodiments, user interface object 1004 is displayed as a representation of a different face. In some embodiments, user interface object 1004 is displayed as a representation of a face that is not human (e.g., an animal, anthropomorphic object, alien, non-descriptive face, fantasy creature, and / or a collection of face-like objects). It should be recognized that other representations of the first system user interface object 1004 may be used with the techniques described herein, and the representation of a face is one example.
[0248] exist Figure 10A As indicated by the user's indication that 1010 is positioned within the field of view 1008a, the user is within the field of view of computer system 1000. Figure 10A At this point, computer system 1000 detects a user. In some implementations, detecting a user causes computer system 1000 to display user interface object 1004 with different facial expressions, such as changing from a default or calm expression to a smiling expression. Figure 10AAt this point, the user moves to a new position within the field of view of the computer system 1000. In some embodiments, the computer system 1000 moves a portion of itself (e.g., a display component and / or other components) to keep the user within the field of view. In some embodiments, this portion of the computer system 1000 may rotate via one or more moving components to keep the user within the field of view as the user physically moves through the space. Figure 10A At this point, computer system 1000 detects audio input 1012. Audio input 1012 corresponds to a voice input request from a user that causes computer system 1000 to output the requested content (e.g., let's read "that dog").
[0249] like Figure 10B As shown, in response to the detection of audio input 1012, computer system 1000 provides output of a first portion of the requested content (e.g., "that dog"). In this example, the requested content includes audio output (e.g., reading "that dog") and visual output (e.g., displaying the text "that dog"). In some embodiments, during the provision of the audio output of the first portion of the requested content, computer system 1000 displays first text 1016 corresponding to the first portion of the requested content. Figure 10B As shown, in response to the detection of audio input 1012, computer system 1000 reduces the size of user interface object 1004 and moves user interface object 1004 to the lower left to make room for first text 1016. In some embodiments, computer system 1000 does not change the size and / or position of user interface object 1004. In some embodiments, computer system 1000 does not display first text 1016. In some embodiments, computer system 1000 does not output audio in the form of reading the requested content aloud. In some embodiments, computer system 1000 outputs audio content that enhances the requested content, but that audio content is not the requested content. In some embodiments, if the requested content corresponds to a ship, computer system 1000 may output the sounds of waves and / or seagulls. In some embodiments, this part of computer system 1000 physically moves in sync with the audio content. In some embodiments, this part of computer system 1000 may move in a manner that mimics the movement of a ship in sync with the sound of waves. Figure 10B For example, the position of 1010 within the field of view 1008a as indicated by the user. Figure 10A The different positions of user 1010 within the field of view 1008a indicate the different locations of the user within the field of view of computer system 1000. Figure 10BEven if the user has moved, the user remains within the field of view of the computer system 1000, enabling the computer system 1000 to detect the user within the field of view. As described below, if the computer system 1000 has already output content, in response to continuing to detect the user within the field of view, the computer system 1000 continues to output content.
[0250] In some implementations, if the computer system 1000 is not facing the user (e.g., the user is not in the field of view), in response to the detection of audio input 1012, the computer system 1000 moves to face the user (e.g., moves that portion until the user is in the field of view). In some implementations, if the computer system 1000 moves to face the user in response to the detection of audio input 1012, the computer system 1000 first outputs a first portion of the requested content, then stops moving (e.g., in response to the detection of the user in the field of view). In some implementations, in response to the detection of audio input 1012, the computer system 1000 may begin outputting "that dog" while rotating to face the user. In some implementations, if the computer system 1000 moves to face the user in response to the detection of audio input 1012, the computer system 1000 outputs a first portion of the requested content after stopping moving (e.g., in response to the detection of the user in the field of view). In some implementations, in response to the detection of audio input 1012, the computer system 1000 may rotate to face the user, stop moving, and then begin outputting "that dog". In some implementations, if the computer system 1000 is moving (e.g., moving to face the user, and / or moving during content output), the computer system 1000 stops moving in response to detecting that the user is within a predefined distance from the computer system 1000. In some implementations, if the computer system 1000 moves in sync with music content, the computer system 1000 stops moving in response to detecting that the user is within a predefined distance (e.g., the user approaches and / or the user reaches out to interact with the computer system 1000).
[0251] like Figure 10C As shown, in response to continued detection of the user within the field of view, after providing the output of the first part of the requested content, the computer system 1000 continues to provide the output of the second part of the requested content immediately following the first part (e.g., if no other input is provided). Figure 10CAs shown, during the output of the second part of the requested content, computer system 1000 displays second text 1018 corresponding to the second part of the requested content. In some embodiments, if computer system 1000 is moving that part of computer system 1000 (e.g., during content output), in response to outputting the second part of the requested content, computer system 1000 stops moving to provide the user with a better view of the displayed content after the change. In some embodiments, after a predetermined time following the cessation of movement in response to outputting the second part of the requested content, computer system 1000 physically moves that part of computer system 1000 during the output of the second part of the requested content. In some embodiments, in response to outputting the second part of the requested content, computer system 1000 may pause movement for a period of time to allow the user time to read the newly displayed text. In some implementations, if the computer system 1000 is moving a portion of itself (e.g., during content output), the computer system 1000 displays an item unrelated to the requested content (e.g., a reminder, notification, and / or call), thereby stopping the movement to allow the user to notice and / or interact with the item. In some implementations, the computer system 1000 may pause movement in response to displaying a calendar reminder, allowing the user to view the calendar reminder.
[0252] exist Figure 10C At that point, computer system 1000 determines that the second part of the requested content corresponds to the second system digital image 1014. For example... Figure 10C As shown, in response to determining that the second part of the requested content corresponds to the second system digital avatar 1014, the computer system 1000 stops displaying the user interface object 1004 and displays the second system digital avatar 1014. In some embodiments, the second system digital avatar 1014 is displayed as having one or more visual characteristics having the same size and / or the same facial expression as the user interface object 1004. In some embodiments, such as Figure 10C As shown, computer system 1000 displays the second system digital image 1014 as having the same characteristics as... Figure 10B The second system digital avatar 1014 has the same facial expression as the user interface object 1004. In other examples, the second system digital avatar 1014 is displayed as having one or more visual characteristics that are different in size and / or facial expression from the user interface object 1004. In some implementations, such as Figure 10C As shown, computer system 1000 displays the second system digital image 1014 as having a higher resolution than... Figure 10B The user interface object 1004 in the image has a larger ear. (Example:) Figure 10CAs shown, the second system digital image 1014 is displayed as a dog to correspond to the second part of the requested content (e.g., the second text 1018). In other examples, the requested content corresponds to a different system digital image, such that the computer system 1000 displays a different system digital image instead of the second system digital image 1014. In some embodiments, the computer system 1000 synchronizes the movement of the second system digital image 1014 (e.g., visually via a display and / or physically via a portion of the computer system 1000) with the second part of the requested content. In some embodiments, since the second part of the requested content corresponds to "the dog was all alone at first," the second system digital image 1014 may appear to be looking around (e.g., visually looking left and right via a display and / or physically rotating left and right via a movement of a portion of the computer system 1000) as part of the content corresponding to the second part of the requested content. In some embodiments, the computer system outputs audio corresponding to the sound of the second system digital image 1014, which is different from the audio corresponding to the sound of the user interface object 1004.
[0253] In some embodiments, the computer system 1000 outputting a second portion of the requested content includes: outputting audio corresponding to the reading of the second text 1018, and then outputting audio that enhances the requested content. In some embodiments, the computer system 1000 may output content that makes the second system digital image 1014 appear to be reading a part of a story and then barking. In some embodiments, in response to outputting audio corresponding to the reading of the second text 1018, the computer system 1000 moves that portion of the computer system 1000 by a first degree of movement. In some embodiments, in response to outputting audio corresponding to the reading of the second text 1018, the computer system 1000 may tilt (e.g., slightly) towards the user. In some embodiments, in response to outputting audio that enhances the story, the computer system 1000 moves that portion of the computer system 1000 by a second degree of movement different from the first degree of movement. In some embodiments, in response to outputting the audio of a dog barking, the computer system may make that portion of the computer system 1000 bounce to mimic the movements of an excited dog. In some embodiments, the second degree of movement is greater than the first degree of movement.
[0254] exist Figure 10CAt this point, computer system 1000 detects audio input 1020. Audio input 1020 corresponds to a voice input request from a user to cause computer system 1000 to switch to outputting a third part of the requested content (e.g., “Go to when the dog is safe”). This voice input request does not include explicit indication of a third part of the requested content (e.g., it does not include the chapter / section name, the chapter / section number associated with the requested content, and / or the chapter / section page number). The voice input includes a description of one or more attributes of the scene. In some embodiments, the description of one or more attributes of the scene includes a description of a portion of the plot of the requested content (e.g., when the dog is safely escaped, when she confronts the witch, and / or when the dragon burns the town). In some embodiments, the user can say “Jump to the part where they find the magic key,” thereby causing computer system 1000 to switch to outputting the portion of the requested content containing the plot where the character finds the magic key. In some embodiments, the description of one or more attributes of the scene includes a description of a portion of the theme of the requested content (e.g., action, calm, violence, love, and / or music). In some implementations, a user may say, "I only want to see the dance segment," causing the computer system 1000 to output only the dance portion of the requested content. In some implementations, the description of one or more attributes of the scene includes a description of a portion of a character's storyline in the requested content (e.g., when a character reforms, when a character betrays a close friend, and / or when a character realizes they have always possessed powers). In some implementations, a user may say, "I want to see the training montage segment," causing the computer system 1000 to output the training montage portion of the requested content. In some implementations, the description of one or more attributes of the scene includes a description of one or more characteristics of the requested content (e.g., everyone is wearing red, a scene in a police station, and / or a section with an elephant). In some implementations, a user may say, "Return to the train scene," causing the computer system 1000 to jump back to outputting the portion of the requested content that includes the train scene.
[0255] In some implementations, the computer system 1000 performs different operations in response to different voice inputs. In some implementations, as described above, in response to a first voice input, the computer system 1000 may change which part of the requested content is being output. In response to a second voice input (e.g., different from the first voice input), the computer system 1000 may modify aspects of the content playback (e.g., volume, speed, language, font size, rhythm, and / or font contrast). In some implementations, while providing the second part of the requested content, if a voice input not corresponding to a request to change to outputting a third part of the requested content is detected, the computer system 1000 continues to provide output corresponding to the second part of the requested content. In some implementations, while the computer system 1000 is providing output corresponding to the second part of the requested content, if the com...
Claims
1. A method, the method comprising: At the computer system communicating with the display component and the one or more input devices: When displaying content via the display component, a first interaction condition corresponding to the content is detected; and In response to detecting the first interaction condition corresponding to the content: Based on the determination that the first interaction condition corresponding to the content satisfies one or more of the first set of criteria, the display component displays a representation of the direction of the face looking at the content. as well as Based on the determination that the first interaction condition corresponding to the content satisfies a second group of one or more criteria that are different from the first group of one or more criteria, the representation of the face is displayed via the display component in the direction of the first user detected in the detection field of the one or more input devices.
2. The method of claim 1, wherein the first set of one or more criteria includes a first criterion satisfied when it is determined that a portion of the content has been displayed for at least a predetermined time period.
3. The method according to any one of claims 1 to 2, wherein the first set of one or more criteria includes a second criterion satisfied when it is determined that the first user is looking in the direction of the content.
4. The method according to any one of claims 1 to 3, wherein the first set of one or more criteria includes a third criterion satisfied when it is determined that the first user should be looking at the content.
5. The method according to any one of claims 1 to 4, wherein the second set of one or more criteria includes a fourth criterion satisfied when it is determined that input has been received.
6. The method according to any one of claims 1 to 5, wherein the second set of one or more criteria includes a fifth criterion satisfied when it is determined that the computer system is waiting for a response.
7. The method according to claim 6, further comprising: In response to detecting a first interaction condition corresponding to the content and determining that the first interaction condition corresponding to the content satisfies one or more criteria, the one or more criteria including criteria satisfied when it is determined that the computer system is not waiting for the response, the representation of the face is displayed via the display component in a second direction, the second direction being different from the direction of the first user detected in the detection field of the one or more input devices.
8. The method of claim 7, wherein the second direction corresponds to a second user detected in the detection field of the one or more input devices, and wherein the second user is different from the first user.
9. The method of claim 7, wherein the second direction corresponds to the content.
10. The method of claim 7, wherein the second direction corresponds to an object in the detection field of the one or more input devices.
11. The method according to any one of claims 1 to 10, the method further comprising: When the face is shown to be looking in the direction of the content, a second interaction condition corresponding to the content is detected; as well as In response to detecting a second interaction condition corresponding to the content, the representation of the face is displayed via the display component, showing the direction in which the first user is detected in the detection field of the one or more input devices.
12. The method according to any one of claims 1 to 11, the method further comprising: When the representation of the face is looking toward the direction of the user detected in the detection field of the one or more input devices, a third interaction condition corresponding to the content is detected; as well as In response to detecting the third interaction condition corresponding to the content, the representation of the face looking towards the content is displayed via the display component.
13. The method according to any one of claims 1 to 12, further comprising: When the content is displayed via the display component, a fourth interaction condition corresponding to the content is detected; as well as In response to detecting the fourth interaction condition corresponding to the content, the representation of the face being looked at a third direction not corresponding to the first user and the content is displayed via the display component.
14. The method according to any one of claims 1 to 13, further comprising: When the representation of the face being viewed in a third direction is displayed via the display component, a fifth interaction condition corresponding to the content is detected; as well as In response to the detection of the fifth interaction condition, the display component continues to display the representation of the face looking toward the third party.
15. The method according to any one of claims 1 to 14, the method further comprising: When the content is displayed via the display component, a sixth interaction condition corresponding to the content is detected; as well as In response to the detection of the sixth interaction condition corresponding to the content: Based on the determination that the sixth interaction condition corresponding to the content satisfies one or more criteria of the second group, the representation of the face is displayed via the display component in the direction of the first user detected in the detection field of the one or more input devices; as well as Based on the determination that the sixth interaction condition corresponding to the content satisfies a fourth group of one or more criteria that are different from the first group of one or more criteria and the second group of one or more criteria, the representation of the face is displayed via the display component in the direction of a third user detected in the detection field of the one or more input devices, wherein the third user is different from the first user.
16. The method of claim 15, wherein the computer system communicates with the mobile component, the method further comprising: In response to detecting the sixth interaction condition corresponding to the content and based on determining that the sixth interaction condition satisfies one or more of the second set of criteria, a portion of the computer system is moved from a first location in the environment to a second location in the environment via the moving component.
17. The method of claim 16, further comprising: When the part of the computer system is in the second location, a seventh interaction condition corresponding to the content is detected; as well as In response to detecting the seventh interaction condition corresponding to the content and determining that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria: Moving the portion of the computer system from the second location in the environment to the first location in the environment via the moving component; and The direction in which the face is looking is displayed, as shown by the display component.
18. The method of claim 16, further comprising: When the part of the computer system is in the second location, an eighth interaction condition corresponding to the content is detected; as well as In response to the detection of the eighth interaction condition corresponding to the content: Based on the determination that the seventh interaction condition corresponding to the content satisfies the first set of one or more criteria, the part of the computer system is moved from the second location in the environment to the first location in the environment via the moving component, while the representation of the face looking towards the user continues to be displayed via the display component.
19. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and the one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 18.
20. A computer system communicating with a display component and the one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 18.
21. A computer system communicating with a display component and the one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 1 to 18.
22. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and the one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 18.
23. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and the one or more input devices, the one or more programs including instructions for: When displaying content via the display component, a first interaction condition corresponding to the content is detected; and In response to detecting the first interaction condition corresponding to the content: Based on the determination that the first interaction condition corresponding to the content satisfies one or more criteria of a first group, the display component displays a representation of the direction the face is looking at the content; and Based on the determination that the first interaction condition corresponding to the content satisfies a second group of one or more criteria that are different from the first group of one or more criteria, the representation of the face is displayed via the display component in the direction of the first user detected in the detection field of the one or more input devices.
24. A computer system communicating with a display component and the one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When displaying content via the display component, a first interaction condition corresponding to the content is detected; and In response to detecting the first interaction condition corresponding to the content: Based on the determination that the first interaction condition corresponding to the content satisfies one or more of the first set of criteria, the display component displays a representation of the direction of the face looking at the content. as well as Based on the determination that the first interaction condition corresponding to the content satisfies a second group of one or more criteria that are different from the first group of one or more criteria, the representation of the face is displayed via the display component in the direction of the first user detected in the detection field of the one or more input devices.
25. A computer system communicating with a display component and the one or more input devices, the computer system comprising: A component for the following operation: when displaying content via the display component, detecting a first interaction condition corresponding to the content; as well as In response to detecting the first interaction condition corresponding to the content: The component is used to: display a facial representation of the direction of looking at the content via the display component, based on the determination that the first interaction condition corresponding to the content satisfies one or more criteria of a first set; and The component is used to: display, via the display component, the representation of the face looking toward the direction of a first user detected in the detection field of the one or more input devices, based on the determination that the first interaction condition corresponding to the content satisfies a second group of one or more criteria that are different from the first group of one or more criteria.
26. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component and the one or more input devices, the one or more programs comprising instructions for: When displaying content via the display component, a first interaction condition corresponding to the content is detected; and In response to detecting the first interaction condition corresponding to the content: Based on the determination that the first interaction condition corresponding to the content satisfies one or more criteria of a first group, the display component displays a representation of the direction the face is looking at the content; and Based on the determination that the first interaction condition corresponding to the content satisfies a second group of one or more criteria that are different from the first group of one or more criteria, the representation of the face is displayed via the display component in the direction of the first user detected in the detection field of the one or more input devices.
27. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: When a user interface including a user interface object representing a portion of a person is displayed via the display component, a first input is detected via the one or more input devices; and In response to the detection of the first input: Based on the determination that consent has been expressed with the first input, the user interface continues to be displayed via the display component, while the user interface object is changed in a manner that indicates eye contact with the user; as well as Based on the determination that no consent was expressed in response to the first input, the user interface continues to be displayed via the display component without altering the user interface object in a manner that indicates eye contact with the user.
28. The method of claim 27, wherein continuing to display the user interface without changing the user interface object in a manner indicative of eye contact with the user comprises: The user interface object is displayed via the display component in a manner that does not indicate eye contact with the user.
29. The method of any one of claims 27 to 28, wherein changing the user interface object in a manner indicating eye contact with the user comprises: The user's location is detected in the detection field of the computer system via the one or more input devices; as well as Change the first part of the user interface object to point to the detected location of the user in the detection field of the computer system.
30. The method of claim 29, wherein changing the first portion of the user interface object to point to the detected location of the user in the detection field comprises: The first portion of the user interface object is moved in the detection field via the display component from a location that is more than a predetermined distance from the user's detected location to a location that is no more than the predetermined distance from the user's detected location.
31. The method according to any one of claims 29 to 30, wherein changing the first portion of the user interface object to point to the detected location of the user in the detection field comprises: The first portion of the user interface object, displayed via the display component, points to the location of the user detected in the detection field.
32. The method according to any one of claims 27 to 31, wherein the first input does not include an explicit instruction to change the user interface object.
33. The method of any one of claims 27 to 32, wherein displaying the user interface when changing the user interface object in a manner indicating eye contact with the user includes moving a second portion of the user interface object via the display component.
34. The method according to any one of claims 27 to 33, the method further comprising: When the user interface object is displayed in a manner indicating eye contact with the user, a second input is detected via the one or more input devices; as well as In response to the detection of the second input, the user interface continues to be displayed while the user interface object is changed via the display component in a first manner that does not indicate eye contact with the user.
35. The method according to any one of claims 27 to 34, the method further comprising: When the user interface object is displayed in a manner that indicates eye contact with the user, a third input is detected via the one or more input devices; as well as In response to the detection of the third input: Based on the determination that no consent was expressed with respect to the third input, the user interface continues to be displayed via the display component, while the user interface object is changed in a second manner that does not indicate eye contact with the user. as well as Based on the determination that consent has been expressed with respect to the third input, the user interface continues to be displayed via the display component without altering the user interface object in the second manner that does not indicate eye contact with the user.
36. The method according to any one of claims 27 to 35, the method further comprising: When the user interface object is displayed in a manner indicating eye contact with the user, a fourth input is detected via the one or more input devices; as well as In response to the detection of the fourth input and based on the determination that disagreement was expressed with respect to the fourth input, the user interface object continues to be displayed via the display component in a manner that indicates eye contact with the user.
37. The method according to any one of claims 27 to 36, the method further comprising: When the user interface object is displayed in a manner indicating eye contact with the user, a fifth input is detected via the one or more input devices; as well as In response to the detection of the fifth input and based on the determination that disagreement was expressed with respect to the fifth input, the user interface object is displayed via the display component in a third manner that does not indicate eye contact with the user.
38. The method according to any one of claims 27 to 37, the method further comprising: In response to detecting the first input, the display component continues to display a portion of the user interface excluding the user interface object, while the user interface object is changed in a manner that indicates eye contact with the user.
39. The method according to any one of claims 27 to 38, further comprising: After changing the user interface object in a manner that indicates eye contact with the user: Based on the determination that no predetermined time period has elapsed since the user interface object changed in the manner indicating eye contact with the user, the user interface object continues to be displayed via the display component in the manner indicating eye contact with the user; as well as Based on the determination that a predetermined time period has elapsed since the user interface object changed in a manner indicating eye contact with the user, the display of the user interface object via the display component in a manner indicating eye contact with the user is discontinued.
40. The method according to any one of claims 27 to 39, wherein: The first input includes a question; The determination that the first input expressed agreement includes: determining that the question corresponds to an affirmative answer; and The determination that no agreement was expressed with respect to the first input includes: determining that the question corresponds to a negative answer.
41. The method according to any one of claims 27 to 40, wherein: The first input includes a statement; The determination that the consent was expressed with respect to the first input includes: determining that the statement corresponds to a statement of fact; and The determination that no consent was expressed with respect to the first input includes: determining that the statement does not correspond to the factual statement.
42. The method of any one of claims 27 to 41, wherein determining that the consent is expressed with respect to the first input comprises determining that the first input satisfies the user preferences of the first user, and wherein determining that the consent is not expressed with respect to the first input comprises determining that the first input does not satisfy the user preferences of the first user.
43. The method according to any one of claims 27 to 42, further comprising: In response to detecting the first input via the one or more input devices, an output corresponding to the first input is provided via one or more output devices communicating with the computer system, wherein: Based on the determination that consent has been expressed with respect to the first input, the output includes a first phrase indicating consent to the user's preference; and If it is determined that no agreement was expressed with respect to the first input, the output includes a second phrase indicating disagreement with the user's preference.
44. The method of claim 43, wherein continuing to display the user interface via the display component without changing the user interface object in a manner indicating eye contact with the user does not include outputting the second phrase.
45. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 27 to 44.
46. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 27 to 44.
47. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 27 to 44.
48. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 27 to 44.
49. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for: When a user interface including a user interface object representing a portion of a person is displayed via the display component, a first input is detected via the one or more input devices; and In response to the detection of the first input: Based on the determination that consent has been expressed regarding the first input, the user interface continues to be displayed via the display component, while the user interface object is changed in a manner indicative of eye contact with the user; and Based on the determination that no consent was expressed in response to the first input, the user interface continues to be displayed via the display component without altering the user interface object in a manner that indicates eye contact with the user.
50. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a user interface including a user interface object representing a portion of a person is displayed via the display component, a first input is detected via the one or more input devices; and In response to the detection of the first input: Based on the determination that consent has been expressed with the first input, the user interface continues to be displayed via the display component, while the user interface object is changed in a manner that indicates eye contact with the user; as well as Based on the determination that no consent was expressed in response to the first input, the user interface continues to be displayed via the display component without altering the user interface object in a manner that indicates eye contact with the user.
51. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for the following operations: detecting a first input via the one or more input devices when a user interface including a user interface object representing a portion of a person is displayed via the display component; and In response to the detection of the first input: The component is used to: continue displaying the user interface via the display component based on the determination that consent has been expressed on the first input, while changing the user interface object in a manner that indicates eye contact with the user; and The component is used to continue displaying the user interface via the display component based on the determination that no consent has been expressed with respect to the first input, without changing the user interface object in a manner that indicates eye contact with the user.
52. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the following operations: When a user interface including a user interface object representing a portion of a person is displayed via the display component, a first input is detected via the one or more input devices; and In response to the detection of the first input: Based on the determination that consent has been expressed regarding the first input, the user interface continues to be displayed via the display component, while the user interface object is changed in a manner indicative of eye contact with the user; and Based on the determination that no consent was expressed in response to the first input, the user interface continues to be displayed via the display component without altering the user interface object in a manner that indicates eye contact with the user.
53. A method, the method comprising: At the computer system that communicates with the display components, camera, and one or more input devices: When a first entity is detected in the field of view of the camera, a user interface object representing a portion of a user is displayed in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; When displaying a user interface object that indicates the user's portion, a request to interact with the content is detected; as well as In response to the detection of the request to interact with the content: Display at least the first part of the content; as well as The user interface object representing the user's portion is displayed at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
54. The method according to claim 53, wherein: Before the request to interact with the content is detected, a portion of the user interface object representing the user's portion is displayed at a first location; In response to detecting the request, displaying the user interface object representing the portion of the user at a second size smaller than the first size includes stopping the display of the portion of the user interface object at the first position; and In response to the detection of the request to interact with the content, the first portion of the content is displayed at the first location.
55. The method according to any one of claims 53 to 54, the method further comprising: Before a request to interact with the content is detected, at least a portion of the content is not displayed.
56. The method of any one of claims 53 to 55, wherein the one or more input devices include a microphone, and wherein detecting the request to interact with the content comprises: Voice input corresponding to the content is received via the microphone.
57. The method of any one of claims 53 to 56, wherein the request for interacting with the content does not include an explicit request.
58. The method of any one of claims 53 to 57, wherein the request to interact with the content is a first request to interact with the first portion of the content, the method further comprising: When displaying at least the first portion of the content and displaying the user interface object representing the user's portion at the second size, a request to interact with a second portion of the content is detected via the one or more input devices, the second portion of the content being different from the first portion of the content; as well as In response to the detection of the request to interact with the second part of the content: At least a second portion of the content is displayed via the display component, wherein the second portion of the content replaces the first portion of the content; as well as The user interface object representing the user's portion continues to be displayed via the display component at the second size and in the second manner.
59. The method according to claims 53 to 58, further comprising: After the user interface object representing the user's portion is displayed at the second size and in the second manner to indicate that the user interface object points to the portion of the content, and based on the determination that one or more first set of criteria are met, the user interface object representing the user's portion is displayed via the display component at the second size and in the third manner to indicate that the user interface object points to the first user in the detection field of the computer system.
60. The method according to claim 59, further comprising: After displaying the user interface object representing the user's portion at the second size and in the second manner to indicate that the user interface object points to the portion of the content, and based on the determination that one or more of the first set of criteria are not met, the user interface object representing the user continues to be displayed via the display component at the second size and in the second manner to indicate that the user interface object points to the portion of the content.
61. The method according to any one of claims 53 to 60, wherein displaying the user interface object representing the portion of the user in the second size and in the second manner to indicate that the user interface object points to the portion of the content comprises: Before displaying the user interface object representing the user's portion in the second manner to indicate that the user interface object points to the portion of the content, the user interface object representing the user's portion is changed to the second size.
62. The method according to any one of claims 53 to 61, wherein: Displaying the user interface object representing the user's portion via the display component in the first manner to indicate that the user interface object points to the first entity in the detection field includes: displaying a representation of an eye shown in a first set of one or more characteristics via the display component; and Displaying the user interface object representing the portion of the user via the display component in the second manner to indicate that the user interface object points to the portion of the content includes: displaying, via the display component, the representation of the eye having a second set of characteristics different from the first set of one or more characteristics.
63. The method according to any one of claims 53 to 63, the method further comprising: After displaying at least the first portion of the content and based on the determination that the content should be removed: Stop displaying at least the first portion of the content; as well as The user interface object representing the user's portion is re-displayed at the first size via the display component.
64. The method according to claim 63, wherein, After displaying at least the first portion of the content and based on the determination that the content should be removed and one or more third set of criteria are met, the user interface object is displayed in the first manner to indicate that the user interface object points to the first entity in the detection field of the computer system.
65. The method according to any one of claims 62 to 63, wherein, After displaying at least the first portion of the content and based on the determination that the content should be removed and one or more criteria of the fourth group, the user interface object is displayed in a fourth manner to indicate that the user interface object does not point to the first entity.
66. The method according to any one of claims 53 to 65, the method further comprising: When displaying the user interface object representing the user's portion in the second manner, the interaction conditions corresponding to the content are detected; as well as In response to the detection of the interaction condition: Based on the determination that the interaction condition satisfies one or more criteria of the fifth group, the user interface object representing the portion of the user is displayed via the display component in the first manner to indicate that the user interface object points to the first entity in the field of view of the camera, wherein the first entity initiated the interaction condition; as well as Based on the determination that the interaction condition corresponding to the content satisfies a sixth group of one or more criteria that are different from the fifth group of one or more criteria, the user interface object representing the user's portion is displayed via the display component in such a way that the user interface object points to a second entity detected in the field of view of the computer system, wherein the second entity initiates the interaction condition.
67. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 53 to 66.
68. A computer system communicating with a display component, a camera, and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 53 to 66.
69. A computer system communicating with a display component, a camera, and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 53 to 66.
70. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 53 to 66.
71. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices, the one or more programs including instructions for performing the following operations: When a first entity is detected in the field of view of the camera, a user interface object representing a portion of a user is displayed in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; When displaying a user interface object that indicates the user's portion, a request to interact with the content is detected; as well as In response to the detection of the request to interact with the content: Display at least the first part of the content; as well as The user interface object representing the user's portion is displayed at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
72. A computer system communicating with a display component, a camera, and one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a first entity is detected in the field of view of the camera, a user interface object representing a portion of a user is displayed in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; When displaying a user interface object that indicates the user's portion, a request to interact with the content is detected; as well as In response to the detection of the request to interact with the content: Display at least the first part of the content; as well as The user interface object representing the user's portion is displayed at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
73. A computer system communicating with a display component, a camera, and one or more input devices, the computer system comprising: Components for the following operation: upon detecting a first entity in the field of view of the camera, displaying a user interface object representing a portion of a user in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; Components for the following operations: when displaying a user interface object that indicates the portion of the user interface, detecting a request to interact with content; as well as In response to the detection of the request to interact with the content: Components for displaying at least a first portion of the content; and A component for the following operation: displaying the user interface object representing the portion of the user at a second size smaller than the first size and in a second manner, to indicate that the user interface object points to the portion of the content.
74. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component, a camera, and one or more input devices, the one or more programs comprising instructions for performing the following operations: When a first entity is detected in the field of view of the camera, a user interface object representing a portion of a user is displayed in a first manner to indicate that the user interface object points to the first user in the detection field of the computer system, wherein the user interface object is displayed at a first size; When displaying a user interface object that indicates the user's portion, a request to interact with the content is detected; as well as In response to the detection of the request to interact with the content: Display at least the first part of the content; as well as The user interface object representing the user's portion is displayed at a second size smaller than the first size and in a second manner to indicate that the user interface object points to the portion of the content.
75. A method, the method comprising: At a computer system that communicates with a display component and one or more input devices: A first system digital image is displayed via the display component, wherein the first system digital image corresponds to a first role; When displaying the digital image of the first system, output the first content; as well as After outputting the first content and without detecting input via the one or more input devices: Based on the determination that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image, the second system digital image is displayed via the display component, wherein the second system digital image corresponds to a second role different from the first role; as well as Based on the determination that a third content different from the first content is to be output, and the third content corresponds to the first system digital image, the first system digital image continues to be displayed via the display component, while the second system digital image is not displayed via the display component.
76. The method according to claim 75, further comprising: After outputting the first content and without detecting input via the one or more input devices: Based on the determination that the second content should be output and that the second content corresponds to the second system digital image, the display of the first system digital image via the display component is stopped.
77. The method according to any one of claims 75 to 76, wherein the first system digital avatar includes a first expression, and wherein the second system digital avatar includes the first expression.
78. The method according to any one of claims 75 to 77, wherein: The computer system communicates with one or more audio output devices; Outputting the first content includes outputting audio corresponding to the digital image of the first system via the one or more audio output devices in a first voice format; and Outputting the second content includes outputting audio corresponding to the digital image of the second system via the one or more audio output devices in a second voice different from the first voice.
79. The method according to any one of claims 75 to 78, wherein the first system digital image has a first appearance, and wherein the second system digital image has a second appearance different from the first appearance.
80. The method according to any one of claims 75 to 79, wherein the first system digital image is a first size, and wherein the second system digital image is the first size.
81. The method according to any one of claims 75 to 79, wherein the first system digital image is a second size, and wherein the second system digital image is a third size different from the second size.
82. The method according to any one of claims 75 to 81, wherein the first content corresponds to the digital image of the first system.
83. The method according to any one of claims 75 to 82, wherein the first content corresponds to a first application, the method further comprising: When displaying the digital image of the first system, output fourth content corresponding to a second application that is different from the first application.
84. The method according to any one of claims 75 to 83, wherein the second content corresponds to a third application, the method further comprising: When displaying the digital image of the second system, output the fifth content corresponding to the fourth application, which is different from the third application.
85. The method according to any one of claims 75 to 84, the method further comprising: After displaying the digital image of the second system: Based on the determination that a sixth content different from the first content is to be output and that the sixth content does not correspond to the second system digital image, the first system digital image is displayed via the display component.
86. The method according to any one of claims 75 to 85, the method further comprising: After outputting the first content and without detecting input via the one or more input devices: Based on the determination that a seventh content different from the first content is to be output, and the seventh content corresponds to a third system digital image different from the first system digital image and the second system digital image, the third system digital image is displayed via the display component.
87. The method according to any one of claims 75 to 86, wherein: The first system digital image includes a first set of mobile modes; and The second system digital avatar includes a second set of action modes that are different from the first set of action modes.
88. The method according to any one of claims 75 to 87, the method further comprising: When displaying the digital image of the first system and outputting the first content, the movement of the digital image of the first system is synchronized with the first content.
89. The method according to any one of claims 75 to 88, the method further comprising: When displaying the digital image of the second system and outputting the second content, the movement of the digital image of the second system is synchronized with the movement of the second content.
90. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 75 to 89.
91. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 75 to 89.
92. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 75 to 89.
93. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 75 to 89.
94. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs including instructions for: A first system digital image is displayed via the display component, wherein the first system digital image corresponds to a first role; When displaying the digital image of the first system, output the first content; as well as After outputting the first content and without detecting input via the one or more input devices: Based on the determination that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image, the second system digital image is displayed via the display component, wherein the second system digital image corresponds to a second role different from the first role; as well as Based on the determination that a third content different from the first content is to be output, and the third content corresponds to the first system digital image, the first system digital image continues to be displayed via the display component, while the second system digital image is not displayed via the display component.
95. A computer system communicating with a display component and one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: A first system digital image is displayed via the display component, wherein the first system digital image corresponds to a first role; When displaying the digital image of the first system, output the first content; as well as After outputting the first content and without detecting input via the one or more input devices: Based on the determination that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image, the second system digital image is displayed via the display component, wherein the second system digital image corresponds to a second role different from the first role; as well as Based on the determination that a third content different from the first content is to be output, and the third content corresponds to the first system digital image, the first system digital image continues to be displayed via the display component, while the second system digital image is not displayed via the display component.
96. A computer system communicating with a display component and one or more input devices, the computer system comprising: Components for displaying a first system digital image via the display component, wherein the first system digital image corresponds to a first role; A component for the following operation: when displaying the digital image of the first system, outputting first content; as well as After outputting the first content and without detecting input via the one or more input devices: The component is used for the following operation: based on determining that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image, the component displays the second system digital image via the display component, wherein the second system digital image corresponds to a second role different from the first role; and The component is used to continue displaying the first system digital image via the display component, without displaying the second system digital image, based on the determination that a third content different from the first content is to be output and the third content corresponds to the first system digital image.
97. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component and one or more input devices, the one or more programs comprising instructions for performing the following operations: A first system digital image is displayed via the display component, wherein the first system digital image corresponds to a first role; When displaying the digital image of the first system, output the first content; as well as After outputting the first content and without detecting input via the one or more input devices: Based on the determination that a second content different from the first content is to be output and the second content corresponds to a second system digital image different from the first system digital image, the second system digital image is displayed via the display component, wherein the second system digital image corresponds to a second role different from the first role; as well as Based on the determination that a third content different from the first content is to be output, and the third content corresponds to the first system digital image, the first system digital image continues to be displayed via the display component, while the second system digital image is not displayed via the display component.
98. A method, the method comprising: At the computer system that communicates with the display component, audio generation component, and motion component: The audio content is output via the audio generation component; When outputting the audio content, a portion of the computer system is physically moved via the moving component; A request to display visual content is detected when the part of the computer system is physically moved via the moving component; as well as In response to the detection of the request to display the visual content: Stop physically moving the portion of the computer system via the moving component; and The visual content is displayed via the display component.
99. The method according to claim 98, wherein: When outputting a portion of the audio content: Based on the determination that a portion of the audio content is being output and that the portion corresponds to a first user, the portion of the computer system is physically moved in a first manner; as well as Based on the determination that a portion of the audio content is being output and that the portion does not correspond to the first user, the portion of the computer system is physically moved in a second manner, different from the first manner.
100. The method according to any one of claims 98 to 99, wherein the audio content is first audio content, the method further comprising: After the physical movement of the part of the computer system via the moving component ceases, a second audio content different from the first audio content is output via the audio generation component; as well as When outputting the second audio content via the audio generation component, the physical movement of the part of the computer system via the moving component is abandoned.
101. The method according to any one of claims 98 to 100, wherein the audio content is third audio content, the method further comprising: After the physical movement of the part of the computer system via the moving component ceases, a fourth audio content different from the third audio content is output via the audio generation component; as well as When the fourth audio content is output via the audio generation component, the portion of the computer system is physically moved via the moving component.
102. The method according to any one of claims 98 to 101, wherein the visual content is first visual content, the method further comprising: When the part of the computer system is physically moved via the moving component, a second visual content different from the first visual content is displayed via the display component.
103. The method of claim 102, wherein the request to display the visual content is a request to change from displaying the second visual content to displaying different visual content.
104. The method according to any one of claims 98 to 103, wherein the portion of the computer system moves according to one or more characteristics of the audio content output by the computer system.
105. The method according to any one of claims 98 to 104, wherein the audio content is fifth audio content, wherein the visual content is third visual content, and the method further comprises: After the physical movement of the part of the computer system via the moving component ceases, while the part of the computer system is being physically moved via the moving component, a sixth audio content different from the fifth audio content is output via the audio generation component. While the sixth audio content is output via the audio generation component and the part of the computer system is physically moved, a request to display fourth visual content that is different from the third visual content is detected. as well as In response to the detection of the request to display the fourth visual content: Based on the determination that the fourth visual content is of the first type of content, the movement of the portion of the computer system via the mobile component is stopped; and Based on the determination that the fourth visual content is a second type of content different from the first type of content, the part of the computer system continues to be physically moved via the moving component.
106. The method according to any one of claims 98 to 105, wherein the visual content is a representation of a digital image.
107. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, an audio generation component, and a mobile component, the one or more programs including instructions for performing the method according to any one of claims 98 to 106.
108. A computer system communicating with a display component, an audio generation component, and a motion component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 98 to 106.
109. A computer system communicating with a display component, an audio generation component, and a motion component, the computer system comprising: Components for performing the method according to any one of claims 98 to 106.
110. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, an audio generation component, and a mobile component, the one or more programs comprising instructions for performing the method according to any one of claims 98 to 106.
111. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display component, an audio generation component, and a motion component, the one or more programs including instructions for: The audio content is output via the audio generation component; When outputting the audio content, a portion of the computer system is physically moved via the moving component; A request to display visual content is detected when the part of the computer system is physically moved via the moving component; as well as In response to the detection of the request to display the visual content: Stop physically moving the portion of the computer system via the moving component; and The visual content is displayed via the display component.
112. A computer system communicating with a display component, an audio generation component, and a motion component, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: The audio content is output via the audio generation component; When outputting the audio content, a portion of the computer system is physically moved via the moving component; A request to display visual content is detected when the part of the computer system is physically moved via the moving component; as well as In response to the detection of the request to display the visual content: Stop physically moving the portion of the computer system via the moving component; and The visual content is displayed via the display component.
113. A computer system communicating with a display component, an audio generation component, and a motion component, the computer system comprising: Components used to output audio content via the audio generation component; A component for physically moving a part of the computer system via the moving component when outputting the audio content; Components for the following operations: detecting a request to display visual content when the portion of the computer system is physically moved via the moving component; and In response to the detection of the request to display the visual content: A component for the following operation: stopping the physical movement of the portion of the computer system via the moving component; and A component used to display the visual content via the display component.
114. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display component, an audio generation component, and a motion component, the one or more programs comprising instructions for performing the following operations: The audio content is output via the audio generation component; When outputting the audio content, a portion of the computer system is physically moved via the moving component; A request to display visual content is detected when the part of the computer system is physically moved via the moving component; as well as In response to the detection of the request to display the visual content: Stop physically moving the portion of the computer system via the moving component; and The visual content is displayed via the display component.
115. A method, the method comprising: At a computer system that communicates with one or more output devices and one or more input devices: When providing one or more outputs corresponding to a first portion of the content via the one or more output devices, voice input including a description of one or more content attributes corresponding to the content is detected via the one or more input devices; as well as In response to detecting the voice input including a description of one or more content attributes corresponding to the content, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, but no output corresponding to a third portion of the content is provided, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
116. The method of claim 115, wherein the one or more output devices include an audio output device, and wherein at least one of the one or more outputs corresponding to the second portion of the content is an audio output provided via the audio output device.
117. The method of any one of claims 115 to 116, wherein the one or more output devices include a first display component, and wherein providing the one or more outputs corresponding to a second portion of the content includes displaying a representation corresponding to the second portion of the content via the first display component.
118. The method of any one of claims 115 to 117, wherein the one or more output devices include a moving component, and wherein providing the one or more outputs includes moving via the moving component.
119. The method according to any one of claims 115 to 118, wherein the first portion of the content is located at a fourth position in the content, and wherein the second portion of the content is located at a fifth position in the content prior to the fourth position.
120. The method according to any one of claims 115 to 119, wherein the first portion of the content is located at a sixth position in the content, and wherein the second portion of the content is located at a seventh position in the content after the sixth position.
121. The method according to any one of claims 115 to 120, wherein the voice input does not include an explicit indication of a second portion of the content.
122. The method according to any one of claims 115 to 121, wherein the first part is in the first subset of the content, and the second part is in the first subset of the content.
123. The method according to any one of claims 115 to 122, wherein the first part is in the second section of the content, and the second part is in a third section of the content, which is different from the second section of the content.
124. The method of any one of claims 115 to 123, wherein the description of one or more content attributes corresponding to the content includes a description of a portion of the plot of the content, and wherein the output corresponding to the second portion of the content relates to the portion of the plot of the content.
125. The method of any one of claims 115 to 124, wherein the description of one or more content attributes corresponding to the content includes a description of a portion of the subject matter of the content, and wherein the output corresponding to a second portion of the content relates to the portion of the subject matter of the content.
126. The method of any one of claims 115 to 125, wherein the description of one or more content attributes corresponding to the content includes a description of a character storyline in the content, and wherein the output corresponding to a second portion of the content relates to the character storyline in the content.
127. The method of any one of claims 115 to 126, wherein the description of one or more content attributes corresponding to the content includes a description of one or more characteristics in the content, and wherein the output corresponding to the second portion of the content relates to the one or more characteristics in the content.
128. The method according to any one of claims 115 to 127, the method further comprising: After outputting one or more outputs corresponding to the second part of the content via the one or more output devices, one or more outputs corresponding to the content portion immediately following the second part of the content are provided via the one or more input devices.
129. The method according to any one of claims 115 to 128, wherein: Based on the description of one or more content attributes corresponding to the content included in the voice input, prior to the one or more content attributes involved in the first part of the content, the second part of the content is positioned in the content prior to the positioning of the first part of the content; and Based on the description of one or more content attributes corresponding to the content included in the voice input, after the one or more content attributes involved in the first part of the content, the second part of the content is located in the content at a position after the position of the first part of the content.
130. The method according to any one of claims 115 to 129, the method further comprising: In response to detecting the voice input including a description of one or more content attributes corresponding to the content: Based on determining that one or more content attributes corresponding to the content include a first event, a first corresponding part of the content is selected as the second part of the content; as well as Based on the determination that the one or more content attributes corresponding to the content include a second event different from the first event, a second corresponding part of the content that is different from the first corresponding part of the content is selected as the second part of the content.
131. The method according to any one of claims 115 to 130, wherein the speech output is a first speech output: While providing one or more outputs corresponding to a first portion of the content via the one or more output devices, a second voice input different from the first voice input is detected via one or more input devices; and In response to the detection of the second voice input, the output corresponding to the first part of the content continues to be provided.
132. The method according to any one of claims 115 to 131, the method further comprising: In response to detecting the voice input including a description of one or more content attributes corresponding to the content, the provision of one or more outputs corresponding to a first portion of the content via the one or more output devices is stopped.
133. The method according to any one of claims 115 to 132, wherein the voice input corresponds to a request to skip a third part of the content.
134. The method according to any one of claims 115 to 133, wherein the voice input corresponds to a request to jump to a second part of the content.
135. The method according to any one of claims 115 to 134, wherein the voice input corresponds to a request to jump from a first portion of the content.
136. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 115 to 135.
137. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 115 to 135.
138. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 115 to 135.
139. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 115 to 135.
140. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs including instructions for: While providing one or more outputs corresponding to a first portion of the content via the one or more output devices, voice input including a description of one or more content attributes corresponding to the content is detected via the one or more input devices; and In response to detecting the voice input including a description of one or more content attributes corresponding to the content, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, but no output corresponding to a third portion of the content is provided, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
141. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When providing one or more outputs corresponding to a first portion of the content via the one or more output devices, voice input including a description of one or more content attributes corresponding to the content is detected via the one or more input devices; as well as In response to detecting the voice input including a description of one or more content attributes corresponding to the content, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, but no output corresponding to a third portion of the content is provided, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
142. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: Components for the following operation: when providing one or more outputs corresponding to a first portion of the content via the one or more output devices, detecting voice input via the one or more input devices including a description of one or more content attributes corresponding to the content; and Components for the following operation: in response to detecting the voice input including a description of one or more content attributes corresponding to the content, providing one or more outputs via the one or more output devices corresponding to a second portion of the content, but not providing an output corresponding to a third portion of the content, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
143. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs comprising instructions for: While providing one or more outputs corresponding to a first portion of the content via the one or more output devices, voice input including a description of one or more content attributes corresponding to the content is detected via the one or more input devices; and In response to detecting the voice input including a description of one or more content attributes corresponding to the content, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, but no output corresponding to a third portion of the content is provided, the third portion of the content being between the first portion of the content and the second portion of the content, wherein the second portion of the content is different from the first portion of the content.
144. A method, the method comprising: At a computer system that communicates with one or more output devices and one or more input devices: When one or more outputs corresponding to the first part of the content are provided via the one or more output devices, the user's attention is detected via the one or more input devices to no longer correspond to the computer system; In response to detecting that the user's attention is no longer corresponding to the computer system, the provision of one or more outputs corresponding to the first portion of the content is stopped; When one or more outputs corresponding to the first part of the content are not provided, the user's attention is detected via the one or more input devices to correspond to the computer system; as well as In response to detecting that the user's attention corresponds to the computer system, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, the second portion of the content being at or before the first portion of the content.
145. The method of claim 144, wherein the one or more output devices include an audio output device, and wherein providing the one or more outputs corresponding to the first portion of the content includes outputting audio via the audio output device.
146. The method of any one of claims 144 to 145, wherein the one or more output devices include a first display component, and wherein providing the one or more outputs corresponding to the first portion of the content includes displaying a representation via the first display component.
147. The method of any one of claims 144 to 146, wherein the one or more output devices include a moving component, and wherein providing the one or more outputs corresponding to the first portion of the content includes moving a first portion of the computer system via the moving component.
148. The method of any one of claims 144 to 147, wherein the second portion of the content is positioned in the content at a location not at the beginning of the first subset of the content.
149. The method according to any one of claims 144 to 148, wherein the first portion of the content is positioned in the content at a location at the beginning of a second subset of the content.
150. The method according to any one of claims 144 to 149, wherein: The first part is located at a first location within a third subset of the content; The second part is located in the content at a second location within the third subset of the content; The first positioning is different from the second positioning; and The first and second locations are not at the endpoint locations of the third subset of the content.
151. The method according to any one of claims 144 to 150, the method further comprising: Before providing the one or more outputs corresponding to the first part of the content, when the user's attention is detected to correspond to the computer system, the one or more outputs corresponding to the second part of the content are provided via the one or more output devices.
152. The method according to any one of claims 144 to 151, the method further comprising: Before detecting that the user's attention no longer corresponds to the computer system, the one or more input devices are used to detect that the user's attention corresponds to the computer system; as well as In response to detecting that the user's attention corresponds to the computer system, after providing the one or more outputs corresponding to the second part of the content, the one or more outputs corresponding to the first part of the content are provided via the one or more output devices.
153. The method according to any one of claims 144 to 152, wherein the computer system communicates with the mobile component, the method further comprising: When outputting audio corresponding to the content, a detection condition is established; as well as In response to the detection of the condition: Based on the determination that the conditions include detecting voice input when outputting audio corresponding to a corresponding scene, and moving via the moving component in a first manner; and Based on the determination that the condition does not include detecting voice input when outputting audio corresponding to the corresponding scene, the moving component moves in a second manner different from the first manner.
154. The method according to any one of claims 144 to 153, wherein the user is a first user, and the method further comprises: When providing one or more outputs corresponding to the content via one or more output devices, it is detected that the attention of a second user no longer corresponds to the computer system while the attention of a third user corresponds to the computer system, wherein the third user is different from the second user; as well as In response to the detection that the attention of the second user no longer corresponds to the computer system while the attention of the third user corresponds to the computer system, one or more outputs corresponding to the content continue to be provided.
155. The method according to any one of claims 144 to 154, wherein the user is a fourth user, and the method further comprises: When providing one or more outputs corresponding to the content via one or more output devices, when it is detected that the attention of the sixth user is no longer corresponding to the computer system, and the sixth user is different from the fifth user; as well as In response to detecting that the attention of the fifth user no longer corresponds to the computer system when the attention of the sixth user corresponds to the computer system: Based on the determination that the fifth user is a user of the first type, continue to provide one or more outputs corresponding to the content; and Based on the determination that the fifth user is a second type of user, different from the first type of user, the provision of one or more outputs corresponding to the content is abandoned.
156. The method of any one of claims 144 to 155, wherein stopping the delivery of the one or more outputs corresponding to the first portion of the content comprises gradually de-emphasizing the one or more outputs corresponding to the first portion of the content via the one or more output devices.
157. The method of any one of claims 144 to 156, wherein stopping the provision of the one or more outputs corresponding to the first portion of the content comprises pausing the one or more outputs corresponding to the first portion of the content via the one or more output devices.
158. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 144 to 157.
159. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 144 to 157.
160. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: Components for performing the method according to any one of claims 144 to 157.
161. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 144 to 157.
162. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs including instructions for: When one or more outputs corresponding to the first part of the content are provided via the one or more output devices, the user's attention is detected via the one or more input devices to no longer correspond to the computer system; In response to detecting that the user's attention is no longer corresponding to the computer system, the provision of one or more outputs corresponding to the first portion of the content is stopped; When one or more outputs corresponding to the first part of the content are not provided, the user's attention is detected via the one or more input devices to correspond to the computer system; as well as In response to detecting that the user's attention corresponds to the computer system, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, the second portion of the content being at or before the first portion of the content.
163. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When one or more outputs corresponding to the first part of the content are provided via the one or more output devices, the user's attention is detected via the one or more input devices to no longer correspond to the computer system; In response to detecting that the user's attention is no longer corresponding to the computer system, the provision of one or more outputs corresponding to the first portion of the content is stopped; When one or more outputs corresponding to the first part of the content are not provided, the user's attention is detected via the one or more input devices to correspond to the computer system; as well as In response to detecting that the user's attention corresponds to the computer system, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, the second portion of the content being at or before the first portion of the content.
164. A computer system that communicates with one or more output devices and one or more input devices, the computer system comprising: Components for the following operation: when providing one or more outputs corresponding to a first portion of the content via the one or more output devices, detecting via the one or more input devices that the user's attention is no longer corresponding to the computer system; A component for the following operation: in response to detecting that the user's attention is no longer corresponding to the computer system, stops providing one or more outputs corresponding to the first part of the content; Components for the following operation: when one or more outputs corresponding to the first part of the content are not provided, detecting via the one or more input devices that the user's attention corresponds to the computer system; and Components for the following operation: in response to detecting that the user's attention corresponds to the computer system, providing one or more outputs via the one or more output devices corresponding to a second portion of the content, the second portion of the content being at or before the first portion of the content.
165. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output devices and one or more input devices, the one or more programs comprising instructions for performing the following operations: When one or more outputs corresponding to the first part of the content are provided via the one or more output devices, the user's attention is detected via the one or more input devices to no longer correspond to the computer system; In response to detecting that the user's attention is no longer corresponding to the computer system, the provision of one or more outputs corresponding to the first portion of the content is stopped; When one or more outputs corresponding to the first part of the content are not provided, the user's attention is detected via the one or more input devices to correspond to the computer system; as well as In response to detecting that the user's attention corresponds to the computer system, one or more outputs corresponding to a second portion of the content are provided via the one or more output devices, the second portion of the content being at or before the first portion of the content.