User interface for wide angle video conferencing

By displaying the communication request interface in the computer system and dynamically adjusting the representation of the camera field of view, the complex and time-consuming problem of real-time video communication session management interface in the prior art is solved, and a faster and more efficient human-machine interface is achieved, saving resources of users and devices.

CN120201156AActive Publication Date: 2025-06-24APPLE INC
View PDF 22 Cites 0 Cited by

Patent Information

Application Number
CN202510503352.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-09-24
Filing Date
2022-01-28
Publication Date
2025-06-24
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

The user interface for managing real-time video communication sessions is complex and time-consuming, resulting in a waste of user time and device energy.

Method used

Provides a faster and more efficient user interface and method that allows the user to opt into a live video communication session by displaying a communication request interface in a computer system and dynamically adjusts the representation of the camera field of view during the session.

Benefits of technology

Reduces the cognitive burden of users and improves the efficiency of the human-machine interface, especially in battery-powered devices, saving power and extending the time interval between battery charging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201156A_ABST
    Figure CN120201156A_ABST
Patent Text Reader

Abstract

The invention relates to a user interface for wide-angle video conferencing. Embodiments are generally directed to a video communication interface for automatically adjusting a display representation of a field of view of a camera in response to detecting a change in a scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of January 28, 2022, the application number of 202280008837.3, and the title of "User Interface for Wide-Angle Video Conferencing". Technical Field

[0002] This disclosure generally relates to computer user interfaces, and more particularly to techniques for managing real-time video communication sessions. Background Art

[0003] A computer system may include hardware and / or software for displaying an interface for a real-time video communication session. Summary of the Invention

[0004] However, some techniques for using an electronic device to manage a real-time video communication session are generally cumbersome and inefficient. For example, some prior art uses complex and time-consuming user interfaces that may include multiple button presses or keystrokes. The prior art takes more time than necessary, which results in wasted user time and device energy. This latter consideration is particularly important in battery-powered devices.

[0005] Accordingly, the present technology provides faster and more efficient methods and interfaces for an electronic device to manage a real-time video communication session. Such methods and interfaces optionally supplement or replace other methods for managing a real-time video communication session. Such methods and interfaces reduce the cognitive burden imposed on the user and result in a more effective human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0006] This document describes exemplary methods. An exemplary method includes, at a computer system in communication with a display generation component, one or more cameras, and one or more input devices: displaying, via the display generation component, a communication request interface that includes: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via the one or more input devices, while the communication request interface is displayed, a set of one or more inputs that includes a selection of the first selectable graphical user interface object; in response to receiving the set of one or more inputs that includes the selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; while the real-time video communication interface is displayed, detecting a change in a scene within a field of view of one or more cameras; and in response to detecting the change in the scene within the field of view of one or more cameras: based on a determination to select to use the first camera mode, adjusting, during the real-time video communication session, a representation of the field of view of one or more cameras based on the detected change in the scene within the field of view of the one or more cameras; and based on a determination to select to use the second camera mode, refraining from adjusting the representation of the field of view of one or more cameras during the real-time video communication session.

[0007] This document describes an example non-transitory computer-readable storage medium. An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a communication request interface that includes: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object when the communication request interface is displayed; in response to receiving the set of one or more inputs including the selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; detecting a change in a scene in the field of view of one or more cameras when the real-time video communication interface is displayed; and in response to detecting a change in a scene in the field of view of one or more cameras: adjusting, based on the detected change in the scene in the field of view of one or more cameras, a representation of the field of view of one or more cameras during the real-time video communication session according to a determination to select to use the first camera mode; and refraining from adjusting a representation of the field of view of one or more cameras during the real-time video communication session according to a determination to select to use the second camera mode.

[0008] This document describes an example transient computer-readable storage medium. An exemplary non-transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a communication request interface that includes: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object when the communication request interface is displayed; in response to receiving the set of one or more inputs including the selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; detecting a change in a scene in the field of view of one or more cameras when the real-time video communication interface is displayed; and in response to detecting a change in a scene in the field of view of one or more cameras: based on a determination to select to use the first camera mode, adjusting a representation of the field of view of one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras; and based on a determination to select to use the second camera mode, forgoing adjusting a representation of the field of view of one or more cameras during the real-time video communication session.

[0009] This document describes an exemplary computer system. An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying, via the display generation component, a communication request interface that includes: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object when the communication request interface is displayed; in response to receiving the set of one or more inputs including the selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; detecting, when the real-time video communication interface is displayed, a change in a scene in the field of view of one or more cameras; and in response to detecting the change in the scene in the field of view of one or more cameras: based on a determination to select to use the first camera mode, adjusting, during the real-time video communication session, the representation of the field of view of one or more cameras based on the detected change in the scene in the field of view of the one or more cameras; and based on a determination to select to use the second camera mode, refraining from adjusting the representation of the field of view of one or more cameras during the real-time video communication session.

[0010] An exemplary computer system, the exemplary computer system comprising: a display generation component; one or more cameras; one or more input devices; means for displaying, via the display generation component, a communication request interface, the communication request interface comprising: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during a real-time video communication session; means for receiving, via the one or more input devices when the communication request interface is displayed, a set of one or more inputs comprising a selection of the first selectable graphical user interface object; means for displaying, in response to receiving the set of one or more inputs comprising the selection of the first selectable graphical user interface object, via the display generation component, a real-time video communication interface for a real-time video communication session; means for detecting, when the real-time video communication interface is displayed, a change in a scene in the field of view of one or more cameras; and means for performing the following in response to detecting the change in the scene in the field of view of one or more cameras: based on a determination to select to use the first camera mode, adjusting, during the real-time video communication session, a representation of the field of view of one or more cameras based on the detected change in the scene in the field of view of the one or more cameras; and based on a determination to select to use the second camera mode, refraining from adjusting the representation of the field of view of one or more cameras during the real-time video communication session.

[0011] An exemplary method includes: at a computer system in communication with a display generation component, one or more cameras, and one or more input devices: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: a representation of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; and detecting, when the real-time video communication interface for the real-time video communication session is displayed, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session; and in response to detecting the set of one or more inputs, initiating the process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.

[0012] An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for the following operations: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; and detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session when displaying the real-time video communication interface for the real-time video communication session; and initiating, in response to detecting the set of one or more inputs, a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.

[0013] An exemplary transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for the following operations: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; and detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session when displaying the real-time video communication interface for the real-time video communication session; and initiating, in response to detecting the set of one or more inputs, a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.

[0014] An exemplary computer system, the exemplary computer system comprising: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; and detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session when displaying the real-time video communication interface for the real-time video communication session; and initiating, in response to detecting the set of one or more inputs, a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.

[0015] An exemplary computer system, the exemplary computer system comprising: a display generation component; one or more cameras; one or more input devices; means for displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; means for detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session when displaying the real-time video communication interface for the real-time video communication session; and means for initiating, in response to detecting the set of one or more inputs, a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.

[0016] An exemplary method includes: at a computer system in communication with a display generation component and one or more cameras: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; when the real-time video communication session is active, capturing, via the one or more cameras, image data of the real-time video communication session; based on determining that a separation amount between a first participant and a second participant satisfies a separation criterion according to the image data of the real-time video communication session captured via the one or more cameras, simultaneously displaying, via the display generation component: a representation of a first portion of the fields of view of the one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the fields of view of the one or more cameras at a second region of the real-time video communication interface different from the first region, wherein the representation of the first portion of the fields of view of the one or more cameras and the representation of the second portion of the fields of view of the one or more cameras are displayed without displaying a representation of a third portion of the fields of view of the one or more cameras between the first portion and the second portion of the fields of view of the one or more cameras; and based on determining that the separation amount between the first participant and the second participant does not satisfy the separation criterion according to the image data of the real-time video communication session captured via the one or more cameras, displaying, via the display generation component, a representation of a fourth portion of the fields of view of the one or more cameras including the first participant and the second participant while maintaining the display of a portion of the fields of view of the one or more cameras between the first participant and the second participant.

[0017] An exemplary non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; capturing, when the real-time video communication session is active, image data of the real-time video communication session via the one or more cameras; determining, based on the image data of the real-time video communication session captured via the one or more cameras, that a separation amount between a first participant and a second participant meets a separation criterion, and simultaneously displaying, via the display generation component: a representation of a first portion of the fields of view of the one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the fields of view of the one or more cameras at a second region of the real-time video communication interface different from the first region, wherein the representation of the first portion of the fields of view of the one or more cameras and the representation of the second portion of the fields of view of the one or more cameras are displayed without displaying a representation of a third portion of the fields of view of the one or more cameras between the first portion of the fields of view of the one or more cameras and the second portion of the fields of view of the one or more cameras; and determining, based on the image data of the real-time video communication session captured via the one or more cameras, that the separation amount between the first participant and the second participant does not meet the separation criterion, and displaying, via the display generation component, a representation of a fourth portion of the fields of view of the one or more cameras including the first participant and the second participant while maintaining the display of a portion of the fields of view of the one or more cameras between the first participant and the second participant.

[0018] An exemplary transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; capturing, via the one or more cameras, image data of the real-time video communication session when the real-time video communication session is active; determining, based on the image data of the real-time video communication session captured via the one or more cameras, that a separation amount between a first participant and a second participant meets a separation criterion, and simultaneously displaying, via the display generation component: a representation of a first portion of the fields of view of the one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the fields of view of the one or more cameras at a second region of the real-time video communication interface different from the first region, wherein the representation of the first portion of the fields of view of the one or more cameras and the representation of the second portion of the fields of view of the one or more cameras are displayed without displaying a representation of a third portion of the fields of view of the one or more cameras between the first portion of the fields of view of the one or more cameras and the second portion of the fields of view of the one or more cameras; and determining, based on the image data of the real-time video communication session captured via the one or more cameras, that the separation amount between the first participant and the second participant does not meet the separation criterion, and displaying, via the display generation component, a representation of a fourth portion of the fields of view of the one or more cameras including the first participant and the second participant while maintaining the display of a portion of the fields of view of the one or more cameras between the first participant and the second participant.

[0019] An exemplary computer system, the exemplary computer system comprising: a display generation component; one or more cameras; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; capturing, via the one or more cameras, image data of the real-time video communication session when the real-time video communication session is active; determining, based on the image data of the real-time video communication session captured via the one or more cameras, that the separation amount between a first participant and a second participant meets a separation criterion, and simultaneously displaying, via the display generation component: a representation of a first portion of the fields of view of the one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the fields of view of the one or more cameras at a second region of the real-time video communication interface different from the first region, wherein the representation of the first portion of the fields of view of the one or more cameras and the representation of the second portion of the fields of view of the one or more cameras are displayed without displaying a representation of a third portion of the fields of view of the one or more cameras between the first portion and the second portion of the fields of view of the one or more cameras; and determining, based on the image data of the real-time video communication session captured via the one or more cameras, that the separation amount between the first participant and the second participant does not meet the separation criterion, and displaying, via the display generation component, a representation of a fourth portion of the fields of view of the one or more cameras including the first participant and the second participant while maintaining the display of a portion of the fields of view of the one or more cameras between the first participant and the second participant.

[0020] An exemplary computer system, the exemplary computer system comprising: a display generation component; one or more cameras; means for displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; means for capturing, when the real-time video communication session is active, image data of the real-time video communication session via the one or more cameras; means for determining, based on the image data of the real-time video communication session captured via the one or more cameras, that a separation amount between a first participant and a second participant meets a separation criterion, and for simultaneously displaying, via the display generation component, the following: a representation of a first portion of the fields of view of the one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the fields of view of the one or more cameras at a second region of the real-time video communication interface different from the first region, wherein the representation of the first portion of the fields of view of the one or more cameras and the representation of the second portion of the fields of view of the one or more cameras are displayed without displaying a representation of a third portion of the fields of view of the one or more cameras between the first portion of the fields of view of the one or more cameras and the second portion of the fields of view of the one or more cameras; and means for determining, based on the image data of the real-time video communication session captured via the one or more cameras, that the separation amount between the first participant and the second participant does not meet the separation criterion, and for displaying, via the display generation component, a representation of a fourth portion of the fields of view of the one or more cameras including the first participant and the second participant while maintaining the display of a portion of the fields of view of the one or more cameras between the first participant and the second participant.

[0021] According to some embodiments, a method is described that is performed at a computer system in communication with one or more output generating components and one or more input devices. The method includes: detecting, via the one or more input devices, a request to display a system interface; in response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including an input pointing to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including an input pointing to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.

[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more output generating components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to display a system interface; in response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including an input pointing to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including an input pointing to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.

[0023] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more output generating components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to display a system interface; in response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including an input pointing to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including an input pointing to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.

[0024] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more input devices; one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: detecting a request to display a system interface via the one or more input devices; in response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including an input pointing to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including an input pointing to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.

[0025] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more input components; means for detecting a request to display a system interface via the one or more input devices; means for, in response to detecting a request to display a system interface, displaying via the one or more output generating components a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; means for, when displaying the system interface having the set of one or more media communication controls, detecting via the one or more input devices a set of one or more inputs including an input pointing to the set of one or more media communication controls; and means for adjusting media communication settings for a corresponding media communication session in response to detecting the set of one or more inputs including an input pointing to the set of one or more media communication controls when the corresponding media communication session has been active for a predetermined amount of time.

[0026] According to some embodiments, a method is described that is performed at a computer system in communication with one or more output generating components, one or more cameras, and one or more input devices. The method includes: displaying via the one or more output generating components a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via the one or more input devices a set of one or more inputs corresponding to a request to change an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras.

[0027] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs executed by one or more processors of a computer system that communicates with one or more output generating components, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the one or more output generating components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and detecting, via the one or more input devices, while displaying the real-time video communication interface for the real-time video communication session, a set of one or more inputs corresponding to a request to change an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras.

[0028] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs executed by one or more processors of a computer system that communicates with one or more output generating components, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the one or more output generating components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and detecting, via the one or more input devices, when displaying the real-time video communication interface for the real-time video communication session, a set of one or more inputs corresponding to a request to change an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras.

[0029] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more cameras; one or more input devices; one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the one or more output generating components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and detecting, via the one or more input devices, when displaying the real-time video communication interface for the real-time video communication session, a set of one or more inputs corresponding to a request to change an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras.

[0030] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more cameras; one or more input devices; and means for displaying, via the one or more output generating components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; means for, when displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and means for changing an appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras in response to detecting the set of one or more inputs.

[0031] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors. Executable instructions for performing these functions are optionally included in a transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0032] Accordingly, a faster and more efficient method and interface for managing a real-time video communication session are provided for a device, thereby enhancing the effectiveness, efficiency, and user satisfaction of such a device. Such methods and interfaces may supplement or replace other methods for managing a real-time video communication session. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals indicate corresponding parts throughout the drawings.

[0034] Figure 1A is a block diagram showing a portable multifunctional device having a touch-sensitive display according to some embodiments.

[0035] Figure 1B is a block diagram showing exemplary components for event handling according to some embodiments.

[0036] Figure 2 Shows a portable multifunctional device with a touch screen according to some embodiments.

[0037] Figure 3 Is a block diagram of an exemplary multifunctional device with a display and a touch-sensitive surface according to some embodiments.

[0038] Figure 4A Shows an exemplary user interface for a menu of an application on a portable multifunctional device according to some embodiments.

[0039] Figure 4B Shows an exemplary user interface for a multifunctional device with a touch-sensitive surface separate from the display according to some embodiments.

[0040] Figure 5A Shows a personal electronic device according to some embodiments.

[0041] Figure 5B Is a block diagram showing a personal electronic device according to some embodiments.

[0042] Figure 5C Shows an exemplary diagram of a communication session between electronic devices according to some embodiments.

[0043] FIG. 6A to FIG. 6Q Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.

[0044] FIG. 7A to FIG. 7B Depicts a flowchart showing a method for managing a real-time video communication session according to some embodiments.

[0045] FIG. 8A to FIG. 8R Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.

[0046] Fig. 9 Is a flowchart showing a method for managing a real-time video communication session according to some embodiments.

[0047] FIG. 10A to FIG. 10J Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.

[0048] Fig.11 Is a flowchart showing a method for managing a real-time video communication session according to some embodiments.

[0049] FIG. 12A to FIG. 12U Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.

[0050] Fig.13is a flowchart showing a method for managing a real-time video communication session according to some embodiments.

[0051] Fig.14 is a flowchart showing a method for managing a real-time video communication session according to some embodiments. Detailed Description

[0052] The following description sets forth exemplary methods, parameters, etc. However, it should be recognized that such description is not intended to limit the scope of the present disclosure, but rather is provided as a description of exemplary embodiments.

[0053] There is a need to provide an electronic device for managing real-time video communication sessions with effective methods and interfaces. Such technologies can reduce the cognitive burden on users participating in video communication sessions, thereby improving productivity. Additionally, such technologies can reduce processor power and battery power otherwise wasted on redundant user input.

[0054] Below, Figure 1A to Figure 1B 、 Figure 2 、 Figure 3 、 FIG. 4A to FIG. 4B and FIG. 5A to FIG. 5C provides a description of exemplary devices for performing techniques for managing real-time video communication sessions. FIG. 6A to FIG. 6Q shows an exemplary user interface for managing a real-time video communication session. FIG. 7A to FIG. 7B depicts a flowchart showing a method for managing a real-time video communication session according to some embodiments. FIG. 6A to FIG. 6Q The user interface in FIG. 7A to FIG. 7B is used to show the processes described below, including the processes in FIG. 8A to FIG. 8R shows an exemplary user interface for managing a real-time video communication session. Fig. 9 is a flowchart showing a method for managing a real-time video communication session according to some embodiments. FIG. 8A to FIG. 8R The user interface in Fig. 9 is used to show the processes described below, which include the processes in FIG. 10A to FIG. 10J shows an exemplary user interface for managing a real-time video communication session. Fig.11 is a flowchart showing a method for managing a real-time video communication session according to some embodiments. Figures 10A-10J The user interface in Fig.11 is used to show the processes described below including the processes in FIG. 12A to FIG. 12U shows an exemplary user interface for managing a real-time video communication session. Fig.13 and Fig.14 are flowcharts showing a method for managing a real-time video communication session according to some embodiments. FIG. 12A to FIG. 12U The user interface in Fig.13 and Fig.14 The process described below of the process in

[0055] In addition, in a method where one or more of the steps described herein depend on one or more conditions being met, it should be understood that the method can be repeated in multiple iterations such that, during the repeated process, all of the conditions that determine the steps in the method are met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and a second step (if the condition is not met), then one of ordinary skill in the art will know to repeat the stated steps until both the condition being met and the condition not being met (in no particular order) occur. Thus, a method described as having one or more steps that depend on one or more conditions being met can be rewritten as a method that repeats until each of the conditions described in the method are met. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium contains instructions for performing conditional operations based on the satisfaction of corresponding one or more conditions and is thus capable of determining whether the possible conditions have been met without explicitly repeating the steps of the method until all of the conditions that determine the steps in the method are met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.

[0056] Although the following description uses terms such as "first", "second", etc. to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch can be named a second touch and similarly a second touch can be named a first touch without departing from the scope of the various described embodiments. In some embodiments, the first touch and the second touch are two separate references to the same touch. In some embodiments, both the first touch and the second touch are touches, but they are not the same touch.

[0057] The terms used in the description of the various embodiments herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms "comprises", "comprising", "includes", and / or "including", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0058] Depending on the context, the term "if" is optionally interpreted to mean "when", "upon", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrases "if determined..." or "if [stated condition or event] is detected" are optionally interpreted to mean "when determining..." or "in response to determining..." or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]".

[0059] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions such as PDA and / or music player functions. Exemplary embodiments of the portable multifunctional device include, but are not limited to, devices from Apple Inc. (Cupertino, California) devices, iPod devices, and Device. Optionally, other portable electronic devices are used, such as a laptop or a tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). In some embodiments, the electronic device is a computer system that communicates (e.g., via wireless communication, via wired communication) with a display generation component. The display generation component is configured to provide a visual output, such as a display via a CRT monitor, a display via an LED monitor, or a display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes displaying content (e.g., video data rendered or decoded by a display controller 156) by transmitting data (e.g., image data or video data) to an integrated or external display generation component via a wired or wireless connection to visually generate the content.

[0060] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick.

[0061] The device generally supports various applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, gaming applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, and / or digital video player applications.

[0062] The various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or within the respective applications. Thus, a common physical architecture of the device, such as a touch-sensitive surface, optionally supports various applications with a user interface that is intuitive and clear to the user.

[0063] Attention is now turned to embodiments of a portable device having a touch-sensitive display. Figure 1AFIG. 0 is a block diagram of a portable multifunctional device 100 having a touch-sensitive display system 112 in accordance with some embodiments. The touch-sensitive display 112 is sometimes called a “touch screen” for convenience and is sometimes known as or called a “touch-sensitive display system”. Device 100 includes memory 102 (which optionally includes one or more computer-readable storage media), memory controller 122, one or more processing units (CPUs) 120, peripheral device interface 118, RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, input / output (I / O) subsystem 106, other input control devices 116, and external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 for detecting the intensity of contacts on device 100 (e.g., a touch-sensitive surface such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 for generating tactile outputs on device 100 (e.g., generating tactile outputs on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate over one or more communication buses or signal lines 103.

[0064] As used in this specification and the claims, the "intensity" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on the touch-sensitive surface, or to a surrogate for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using a variety of methods and a variety of sensors or combinations of sensors. For example, one or more force sensors beneath or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change thereof of the contact area detected on the touch-sensitive surface, the capacitance and / or change thereof of the touch-sensitive surface near the contact, and / or the resistance and / or change thereof of the touch-sensitive surface near the contact are optionally used as surrogates for the force or pressure of a contact on the touch-sensitive surface. In some embodiments, the surrogate measurements of contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measurements). In some embodiments, the surrogate measurements of contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of user input allows a user to access additional device functions that would otherwise be inaccessible on a smaller device with limited footprint, the smaller device being used to (e.g., on a touch-sensitive display) display affordances and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as a knob or button).

[0065] As used in this specification and the claims, the term "haptic output" refers to a physical displacement of the device relative to a previous portion of the device detected by a user using the user's sense of touch, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device. For example, in the case of contact between the device or a component of the device and a surface sensitive to touch by the user (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a sense of touch corresponding to a perceived change in the physical characteristics of the device or the component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or a touchpad) is optionally interpreted by the user as a "press click" or "release click" of a physical actuation button. In some cases, the user will feel a sense of touch, such as a "press click" or "release click", even when the physical actuation button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement does not move. As another example, movement of a touch-sensitive surface will optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface even when there is no change in the smoothness of the touch-sensitive surface. Although such interpretations of touch by the user will be limited by the user's individual sensory perception, many sensory perceptions of touch are common to most users. Thus, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated haptic output corresponds to a physical displacement of the device or a component thereof that would generate the stated sensory perception of a typical (or ordinary) user.

[0066] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0067] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0068] The peripheral device interface 118 can be used to couple the input and output peripheral devices of the device to the CPU 120 and the memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in the memory 102 to perform various functions of the device 100 and process data. In some embodiments, the peripheral device interface 118, the CPU 120, and the memory controller 122 are optionally implemented on a single chip such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0069] The RF (Radio Frequency) circuit 108 receives and transmits RF signals which are also referred to as electromagnetic signals. The RF circuit 108 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with a communication network and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, a memory, and the like. The RF circuit 108 optionally communicates with the network and other devices via wireless communication, and these networks are such as the Internet (also known as the World Wide Web (WWW)), an intranet, and / or a wireless network (such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN)). The RF circuit 108 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as via a short-range communication radio component. The wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), WiMAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols not yet developed as of the date of submission of this document.

[0070] The audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuitry 110 receives audio data from the peripheral interface 118, converts the audio data into an electrical signal, and transmits the electrical signal to the speaker 111. The speaker 111 converts the electrical signal into sound waves audible to humans. The audio circuitry 110 also receives the electrical signal converted from sound waves by the microphone 113. The audio circuitry 110 converts the electrical signal into audio data and transmits the audio data to the peripheral interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuitry 108 by the peripheral interface 118. In some embodiments, the audio circuitry 110 also includes an earphone jack (e.g., Figure 2 212 in

[0071] ). The earphone jack provides an interface between the audio circuitry 110 and a removable audio input / output peripheral device, which is such as an output-only headset or a headset having both an output (e.g., a mono or stereo headset) and an input (e.g., a microphone). Figure 2 The I / O subsystem 106 couples input / output peripheral devices on the device 100, such as the touch screen 112 and other input control devices 116, to the peripheral interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / transmit electrical signals to the other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some embodiments, the input controller 160 is optionally coupled to (or not coupled to) any of the following: a keyboard, an infrared port, a USB port, and a pointing device such as a mouse. One or more buttons (e.g., Figure 2206). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication, via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some embodiments, an air gesture is detected when the user does not touch an input element that is part of the device (or independent of an input element that is part of the device) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the hand of the user), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).

[0072] Quickly pressing the depress button optionally disengages the lock of the touch screen 112 or optionally starts a process of unlocking the device using gestures on the touch screen, as described in U.S. Patent Application No. 11 / 322,549, filed Dec. 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image" (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. Long pressing the depress button (e.g., 206) optionally powers on or powers off the device 100. The functions of the one or more buttons are optionally user-customizable. The touch screen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.

[0073] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from / to the touch screen 112. The touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0074] The touch screen 112 has a touch-sensitive surface, sensor, or group of sensors that accepts input from the user based on haptic and / or tactile contact. The touch screen 112 and the display controller 156 (along with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of that contact) on the touch screen 112 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.

[0075] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, but uses other display technologies in other embodiments. The touch screen 112 and the display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112 to detect contact and any movement or interruption thereof, the variety of touch sensing technologies including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as the technology used in and iPod used in.

[0076] The touch-sensitive display in some embodiments of the touch screen 112 optionally resembles the multi-touch sensitive touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, the touch screen 112 displays visual output from the device 100, while the touch-sensitive touchpad does not provide visual output.

[0077] Touch-sensitive displays in some embodiments of the touch screen 112 are described in the following applications: (1) U.S. Patent Application 11 / 381,313, "Multipoint Touch Surface Controller", filed May 2, 2006; (2) U.S. Patent Application 10 / 840,862, "Multipoint Touchscreen", filed May 6, 2004; (3) U.S. Patent Application 10 / 903,964, "Gestures For Touch Sensitive Input Devices", filed Jul. 30, 2004; (4) U.S. Patent Application 11 / 048,264, "Gestures For Touch Sensitive Input Devices", filed Jan. 31, 2005; (5) U.S. Patent Application 11 / 038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices", filed Jan. 18, 2005; (6) U.S. Patent Application 11 / 228,758, "Virtual Input Device Placement On A Touch Screen User Interface", filed Sep. 16, 2005; (7) U.S. Patent Application 11 / 228,700, "Operation Of A Computer With A Touch Screen Interface", filed Sep. 16, 2005; (8) U.S. Patent Application 11 / 228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard", filed Sep. 16, 2005; and (9) U.S. Patent Application 11 / 367,749, "Multi-Functional Hand-Held Device", filed Mar. 3, 2006. All of these applications are hereby incorporated by reference in their entirety.

[0078] The touch screen 112 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen has a video resolution of approximately 160 dpi. The user optionally uses any suitable object or attachment such as a stylus, finger, etc. to contact the touch screen 112. In some embodiments, the user interface is designed to work primarily through finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts the rough finger-based input into a precise pointer / cursor position or command for performing the action desired by the user.

[0079] In some embodiments, in addition to the touch screen, the device 100 optionally further includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device, which, unlike the touch screen, does not display a visual output. The touchpad is optionally a touch-sensitive surface separate from the touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.

[0080] The device 100 also includes a power system 162 for powering various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.

[0081] The device 100 optionally further includes one or more optical sensors 164. Figure 1AAn optical sensor coupled to the optical sensor controller 158 in the I / O subsystem 106 is shown. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with the imaging module 143 (also referred to as the camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite the touchscreen display 112 on the front of the device, such that the touchscreen display can be used as a viewfinder for still image and / or video image capture. In some embodiments, the optical sensor is located on the front of the device such that an image of the user can optionally be captured for video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that a single optical sensor 164 is used with the touchscreen display for both video conferencing and still image and / or video image capture.

[0082] The device 100 optionally further includes one or more depth camera sensors 175. Figure 1A A depth camera sensor coupled to the depth camera controller 169 in the I / O subsystem 106 is shown. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of an object (e.g., a face) within a scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with the imaging module 143 (also referred to as the camera module), the depth camera sensor 175 is optionally used to determine depth maps of different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located on the front of the device 100 such that an image of the user with depth information can optionally be captured for video conferencing while the user views other video conferencing participants on the touchscreen display, and a selfie with depth map data can be captured. In some embodiments, the depth camera sensor 175 is located on the rear of the device, or on both the rear and front of the device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that the depth camera sensor 175 is used with the touchscreen display for both video conferencing and still image and / or video image capture.

[0083] In some embodiments, a depth map (e.g., a depth map image) contains information (e.g., values) related to the distance of objects in a scene from a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines the position in the Z-axis of the viewpoint where its corresponding two-dimensional pixel lies. In some embodiments, a depth map is composed of pixels, where each pixel is defined by a value (e.g., from 0 to 255). For example, a value of "0" represents a pixel located farthest from the viewpoint (e.g., a camera, an optical sensor, a depth camera sensor) in a "three-dimensional" scene, and a value of "255" represents a pixel located closest to the viewpoint in the "three-dimensional" scene. In other embodiments, the depth map represents the distance between an object in a scene and a plane of the viewpoint. In some embodiments, a depth map includes information about the relative depth of various features of an object of interest in the field of view of a depth camera (e.g., the relative depth of the eyes, nose, mouth, ears of a user's face). In some embodiments, a depth map includes information that enables a device to determine the profile of an object of interest in the z-direction.

[0084] Device 100 optionally further includes one or more contact intensity sensors 165. Figure 1A A contact intensity sensor coupled to an intensity sensor controller 159 in I / O subsystem 106 is shown. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-mechanical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors for measuring the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the rear of device 100, opposite to the touchscreen display 112 located on the front of device 100.

[0085] Device 100 optionally further includes one or more proximity sensors 166. Figure 1AA proximity sensor 166 is shown coupled to the peripheral device interface 118. Alternatively, the proximity sensor 166 is optionally coupled to an input controller 160 in the I / O subsystem 106. The proximity sensor 166 optionally operates as described in the following U.S. patent applications: No. 11 / 241,839, titled "Proximity Detector In Handheld Device"; No. 11 / 240,788, titled "Proximity Detector In Handheld Device"; No. 11 / 620,702, titled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; No. 11 / 586,862, titled "Automated Response To And Sensing Of User Activity In Portable Devices"; and No. 11 / 638,251, titled "Methods And Systems For Automatic Configuration Of Peripherals", which U.S. patent applications are hereby incorporated by reference in their entirety. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touch screen 112.

[0086] Device 100 optionally further includes one or more haptic output generators 167. Figure 1AShows a haptic output generator coupled to the haptic feedback controller 161 in the I / O subsystem 106. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting an electrical signal into a haptic output on the device). The contact intensity sensor 165 receives haptic feedback generation instructions from the haptic feedback module 133 and generates a haptic output on the device 100 that can be felt by a user of the device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., the touch-sensitive display system 112), and optionally generates a haptic output by moving the touch-sensitive surface vertically (e.g., into / out of the surface of the device 100) or laterally (e.g., backward and forward in the same plane as the surface of the device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.

[0087] The device 100 optionally further includes one or more accelerometers 168. Figure 1A Shows an accelerometer 168 coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to the input controller 160 in the I / O subsystem 106. The accelerometer 168 optionally operates as described in the following U.S. Patent Publications: U.S. Patent Publication No. 20050190059, titled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication No. 20060017692, titled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", both of which are hereby incorporated by reference in their entireties. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on an analysis of data received from one or more accelerometers. The device 100 optionally further includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver in addition to the accelerometer 168 for obtaining information about the location and orientation (e.g., portrait or landscape) of the device 100.

[0088] In some embodiments, the software components stored in the memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and application programs (or instruction sets) 136. Additionally, in some embodiments, the memory 102 ( Figure 1A ) or 370 ( Figure 3 ) stores a device / global internal state 157, as shown in Figure 1A and Figure 3 . The device / global internal state 157 includes one or more of the following: an active application state that indicates which applications (if any) are currently active; a display state that indicates what applications, views, or other information occupy the respective regions of the touch screen display 112; a sensor state that includes information obtained from the various sensors and input control devices 116 of the device; and location information related to the location and / or orientation of the device.

[0089] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between the various hardware components and software components.

[0090] The communication module 128 facilitates communication with other devices via one or more external ports 124, and also includes various software components for processing data received by the RF circuit 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices, or indirectly coupled via a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is the same as or similar to and / or compatible with the 30-pin connector used on (a trademark of Apple Inc.) devices, a multi-pin (e.g., 30-pin) connector.

[0091] The contact / motion module 130 optionally detects contact with the touch screen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or a physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of the contact point being represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple finger contact). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.

[0092] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by the user (e.g., determining whether the user has "clicked" an icon). In some embodiments, at least a subset of the intensity thresholds is determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and can be adjusted without changing the physical hardware of the device 100). For example, without changing the touchpad or touch screen display hardware, the mouse "click" threshold of the touchpad or touch screen can be set to any one of a wide range of predefined thresholds. Additionally, in some implementations, software settings are provided to the user of the device for adjusting one or more of the intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by using a system-level click on an "intensity" parameter to adjust multiple intensity thresholds at once).

[0093] The touch / motion module 130 optionally detects gesture inputs made by the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timing, and / or intensities of the detected contacts). Thus, gestures are optionally detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift (lift-off) event.

[0094] The graphics module 132 includes various known software components for presenting and displaying graphics on the touch screen 112 or other display, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual attributes). As used herein, the term "graphics" includes any object that can be displayed to the user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0095] In some embodiments, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes for specifying the graphics to be displayed from an application, etc., and also receives coordinate data and other graphic attribute data as necessary, and then generates screen image data for output to the display controller 156.

[0096] The haptic feedback module 133 includes various software components for generating instructions that are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.

[0097] The text input module 134, which is optionally a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).

[0098] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., provided to the phone 138 for location-based dialing; provided to the camera 143 as picture / video metadata; and provided to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0099] The application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:

[0100] ● A contacts module 137 (sometimes referred to as an address book or contacts list);

[0101] ● A phone module 138;

[0102] ● A video conferencing module 139;

[0103] ● An email client module 140;

[0104] ● An instant messaging (IM) module 141;

[0105] ● A fitness support module 142;

[0106] ● A camera module 143 for still images and / or video images;

[0107] ● An image management module 144;

[0108] ● A video player module;

[0109] ● A music player module;

[0110] ● A browser module 147;

[0111] ● A calendar module 148;

[0112] ● A widget module 149, which optionally includes one or more of the following: a weather widget 149-1, a stock market widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6;

[0113] ● A widget creator module 150 for forming user-created widgets 149-6;

[0114] ● A search module 151;

[0115] ● A video and music player module 152, which combines the video player module and the music player module;

[0116] ● A notes module 153;

[0117] ● A maps module 154; and / or

[0118] ● An online video module 155.

[0119] Examples of other application programs 136 that are optionally stored in the memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice reproduction.

[0120] In conjunction with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the contacts module 137 is optionally used to manage an address book or contact list (e.g., in the application internal state 192 of the contacts module 137 stored in the memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating a phone number, email address, physical address, or other information with a name; associating an image with a name; categorizing and classifying names; providing a phone number or email address to initiate and / or facilitate communication via the phone 138, video conferencing module 139, email 140, or IM 141; and so on.

[0121] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the phone module 138 is optionally used to input a character sequence corresponding to a phone number, access one or more phone numbers in the contacts module 137, modify an entered phone number, dial the corresponding phone number, conduct a session, and disconnect or hang up when the session is complete. As described above, wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies.

[0122] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contacts module 137, and phone module 138, the video conferencing module 139 includes executable instructions to initiate, conduct, and terminate a video conference between the user and one or more other participants according to user instructions.

[0123] In conjunction with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the email client module 140 includes executable instructions to create, send, receive, and manage emails in response to user instructions. In conjunction with the image management module 144, the email client module 140 makes it very easy to create and send emails with static images or video images captured by the camera module 143.

[0124] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134, the instant messaging module 141 includes executable instructions for the following operations: inputting a character sequence corresponding to an instant message, modifying a previously input character, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing the received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0125] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, the GPS module 135, the map module 154, and the music player module, the fitness support module 142 includes executable instructions for creating a fitness (e.g., having time, distance, and / or calorie burn goals); communicating with a fitness sensor (exercise device); receiving fitness sensor data; calibrating the sensors for monitoring fitness; selecting and playing music for the fitness; and displaying, storing, and transmitting fitness data.

[0126] In combination with the touch screen 112, the display controller 156, the optical sensor 164, the optical sensor controller 158, the contact / motion module 130, the graphics module 132, and the image management module 144, the camera module 143 includes executable instructions for the following operations: capturing a still image or video (including a video stream) and storing them in the memory 102, modifying the characteristics of a still image or video, or deleting a still image or video from the memory 102.

[0127] In combination with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, and the camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or video images.

[0128] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, and linking to attachments and other files of web pages.

[0129] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions for creating, displaying, modifying, and storing calendars and data associated with the calendars (e.g., calendar entries, to-do items, etc.) according to user instructions.

[0130] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is a mini-application optionally downloaded and used by the user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or a mini-application created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (HyperText Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (eXtensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).

[0131] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 is optionally used by the user to create widgets (e.g., transforming a user-specified portion of a web page into a widget).

[0132] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, search module 151 includes executable instructions for searching for text, music, sound, images, videos, and / or other files in memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.

[0133] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch screen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0134] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, note module 153 includes executable instructions for creating and managing notes, to-do lists, etc. in accordance with user instructions.

[0135] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data related to stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.

[0136] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, e-mail client module 140, and browser module 147, the online video module 155 includes instructions for performing the following operations: allowing a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on the touch screen or on an external display connected via the external port 124), send an e-mail with a link to a particular online video, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, the instant message module 141 is used instead of the e-mail client module 140 to send a link to a particular online video. Other descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed Jun. 20, 2007, and entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed Dec. 31, 2007, and entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are hereby incorporated by reference in their entirety.

[0137] Each of the above modules and applications corresponds to a set of executable instructions for performing one or more of the above functions and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, the video player module is optionally combined with the music player module into a single module (e.g., Figure 1A the video and music player module 152 in ). In some embodiments, the memory 102 optionally stores a subgroup of the above modules and data structures. In addition, the memory 102 optionally stores additional modules and data structures not described above.

[0138] In some embodiments, device 100 is a device in which a predefined set of functions on the device are performed uniquely via a touchscreen and / or a touchpad. By using the touchscreen and / or the touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on device 100 is optionally reduced.

[0139] The predefined set of functions performed uniquely via the touchscreen and / or the touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 from any user interface displayed on device 100 to a main menu, a home menu, or a root menu. In such embodiments, the touchpad is used to implement a "menu button". In some other embodiments, the menu button is a physical push button or other physical input control device rather than the touchpad.

[0140] Figure 1B is a block diagram showing exemplary components for event handling according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 ) includes an event classifier 170 (e.g., in operating system 126) and corresponding application 136-1 (e.g., any one of the foregoing applications 137 to 151, 155, 380 to 390).

[0141] Event classifier 170 receives event information and determines application 136-1 to which the event information is to be delivered and application view 191 of application 136-1. Event classifier 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is / are currently active, and application internal state 192 is used by event classifier 170 to determine application view 191 to which the event information is to be delivered.

[0142] In some embodiments, application internal state 192 includes additional information such as one or more of the following: recovery information to be used when application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by application 136-1, a state queue for enabling a user to return to a previous state or view of application 136-1, and a repeat / undo queue of previous actions taken by the user.

[0143] The event monitor 171 receives event information from the peripheral interface 118. The event information includes information about sub-events (e.g., a user touch on the touch-sensitive display 112 as part of a multi-touch gesture). The peripheral interface 118 transmits information that it receives from the I / O subsystem 106 or sensors such as the proximity sensor 166, one or more accelerometers 168, and / or the microphone 113 (via the audio circuitry 110). The information that the peripheral interface 118 receives from the I / O subsystem 106 includes information from the touch-sensitive display 112 or a touch-sensitive surface.

[0144] In some embodiments, the event monitor 171 sends requests to the peripheral interface 118 at predetermined intervals. In response, the peripheral interface 118 transmits event information. In other embodiments, the peripheral interface 118 transmits event information only when there is a significant event (e.g., a received input that is above a predetermined noise threshold and / or a received input that exceeds a predetermined duration).

[0145] In some embodiments, the event classifier 170 further includes a hit view determination module 172 and / or an active event recognizer determination module 173.

[0146] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where within one or more of the views a sub-event has occurred. Views are composed of controls and other elements that a user can see on the display.

[0147] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, within which information is displayed and touch-based gestures occur. The application view (of the corresponding application) within which a touch is detected optionally corresponds to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest-level view within which a touch is detected is optionally referred to as the hit view, and the set of events that are recognized as correct inputs is optionally determined at least in part based on the hit view of the initial touch that begins the touch-based gesture.

[0148] The hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest view in the hierarchical structure that should handle the sub-events. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view generally receives all sub-events related to the same touch or input source for which it was identified as the hit view.

[0149] The active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, the active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, the active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views and, thus, determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with a particular view, higher views in the hierarchy will still remain as actively participating views.

[0150] The event dispatcher module 174 distributes event information to event recognizers (e.g., event recognizer 180). In embodiments that include the active event recognizer determination module 173, the event dispatcher module 174 delivers the event information to the event recognizer determined by the active event recognizer determination module 173. In some embodiments, the event dispatcher module 174 stores the event information in an event queue, which is retrieved by the corresponding event receiver 182.

[0151] In some embodiments, the operating system 126 includes the event classifier 170. Alternatively, the application 136-1 includes the event classifier 170. In yet another embodiment, the event classifier 170 is an independent module or part of another module (such as the contact / motion module 130) stored in the memory 102.

[0152] In some embodiments, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the user interface of the application. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other embodiments, one or more of the event recognizers 180 are part of an independent module that is a higher-level object such as a user interface toolkit or from which application 136-1 inherits methods and other properties. In some embodiments, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from event classifier 170. Event handler 190 optionally utilizes or invokes data updater 176, object updater 177, or GUI updater 178 to update the internal state 192 of the application. Alternatively, one or more of the application views 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included within a corresponding application view 191.

[0153] A corresponding event recognizer 180 receives event information (e.g., event data 179) from event classifier 170 and identifies an event based on the event information. Event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event recognizer 180 further includes at least a subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).

[0154] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events such as a touch or a touch movement. Depending on the sub-event, the event information further includes additional information such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally further includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device pose).

[0155] Event comparator 184 compares the event information with predefined event or sub - event definitions and determines an event or sub - event based on the comparison, or determines or updates the status of an event or sub - event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains the definition of an event (e.g., a predefined sequence of sub - events), such as event 1 (187 - 1), event 2 (187 - 2), and others. In some embodiments, the sub - events in an event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi - touch. In one example, the definition of event 1 (187 - 1) is a double - tap on a displayed object. For example, a double - tap includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift - off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on the displayed object, and a second lift - off (touch end) of a predetermined duration. In another example, the definition of event 2 (187 - 2) is a drag on a displayed object. For example, a drag includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch - sensitive display 112, and lift - off of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.

[0156] In some embodiments, event definition 187 includes the definition of an event for a corresponding user interface object. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with the sub - event. For example, in an application view that displays three user interface objects on touch - sensitive display 112, when a touch is detected on touch - sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub - event). If each displayed object is associated with a corresponding event handler 190, event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub - event and the object that triggered the hit test.

[0157] In some embodiments, the definition of the corresponding event (187) also includes a delay action that delays the delivery of the event information until it has been determined that the sub - event sequence does or does not correspond to the event type of the event recognizer.

[0158] When the corresponding event recognizer 180 determines that the sub - event sequence does not match any event in the event definition 186, the corresponding event recognizer 180 enters an event - impossible, event - failed, or event - ended state, after which subsequent sub - events of the touch - based gesture are ignored. In such a case, other event recognizers (if any) that remain active for the hit view continue to track and process sub - events of the ongoing touch - based gesture.

[0159] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists that indicate how the event delivery system should perform sub - event delivery to the actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists that indicate how event recognizers interact with each other or can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists that indicate whether sub - events are delivered to different levels in the view or the programmatic hierarchy.

[0160] In some embodiments, when one or more specific sub - events of an event are recognized, the corresponding event recognizer 180 activates an event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferring sending) sub - events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a token associated with the recognized event, and the event handler 190 associated with that token retrieves the token and executes a predefined process.

[0161] In some embodiments, the event delivery instruction 188 includes a sub - event delivery instruction that delivers event information about the sub - event without activating the event handler. Instead, the sub - event delivery instruction delivers the event information to the event handler associated with the sub - event sequence or to the actively participating view. The event handler associated with the sub - event sequence or with the actively participating view receives the event information and executes a predetermined process.

[0162] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137, or stores video files used in a video player module. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates portions of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on the touch-sensitive display.

[0163] In some embodiments, event handler 190 includes data updater 176, object updater 177, and GUI updater 178, or has access to the data updater, the object updater, and the GUI updater. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.

[0164] It should be understood that the foregoing discussion of event handling for user touches on a touch-sensitive display also applies to other forms of user input for operating multifunctional device 100 using an input device, and not all user input is initiated on a touchscreen. For example, mouse movement and mouse button presses optionally in cooperation with single or multiple keyboard presses or holds; contact movement on a touchpad, such as tapping, dragging, scrolling, etc.; stylus input; movement of the device; verbal instructions; detected eye movement; biometric input; and / or any combination thereof are optionally used as input corresponding to sub-events that define events to be discriminated.

[0165] Figure 2FIG. 0 shows a portable multifunctional device 100 having a touch screen 112 according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this and other embodiments described below, a user is able to select one or more of these graphics by making gestures on the graphics using, for example, one or more fingers 202 (not drawn to scale in the figures) or one or more styli 203 (not drawn to scale in the figures). In some embodiments, selection of one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gestures optionally include one or more taps, one or more swipes (from left to right, right to left, up, and / or down), and / or rolling of a finger that has made contact with the device 100 (from right to left, left to right, up, and / or down). In some implementations or in some cases, inadvertently contacting a graphic does not select the graphic. For example, when the gesture corresponding to selection is a tap, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application.

[0166] Device 100 optionally further includes one or more physical buttons, such as a “home” or menu button 204. As previously described, the menu button 204 is optionally used to navigate to any of a set of applications 136 optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen 112.

[0167] In some embodiments, device 100 includes a touch screen 112, a menu button 204, a depress button 206 for powering on / off the device and for locking the device, one or more volume adjustment buttons 208, a subscriber identity module (SIM) card slot 210, an earphone jack 212, and a docking / charging external port 124. The depress button 206 is optionally used to power on / off the device by depressing the button and holding the button in the depressed state for a predefined time interval; to lock the device by depressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also receives voice input for activating or deactivating certain functions via a microphone 113. Device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of contact on the touch screen 112, and / or one or more tactile output generators 167 for generating tactile output for a user of device 100.

[0168] Figure 3is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, in accordance with some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a gaming system, or a control device (e.g., a home or industrial controller). Device 300 generally includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects system components and controls the communication between them. Device 300 includes an input / output (I / O) interface 330 having a display 340, which is typically a touchscreen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to the haptic output generator 167 described above with reference to Figure 1A ), sensors 359 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors (similar to the contact intensity sensor 165 described above with reference to Figure 1A ). Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU 310. In some embodiments, memory 370 stores programs, modules, and data structures similar to or a subset of the programs, modules, and data structures stored in the memory 102 of the portable multifunctional device 100 ( Figure 1A ). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunctional device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while the memory 102 of the portable multifunctional device 100 ( Figure 1A ) optionally does not store these modules.

[0169] Figure 3Each of the above elements in [the above] is optionally stored in one or more of the memory devices of the previously mentioned memory device. Each of the above modules corresponds to an instruction set for performing the above functions. The above modules or computer programs (e.g., instruction sets or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, the memory 370 optionally stores a subgroup of the above modules and data structures. Additionally, the memory 370 optionally stores additional modules and data structures not described above.

[0170] Attention is now turned to an embodiment of a user interface optionally implemented on, for example, the portable multifunctional device 100.

[0171] Figure 4A An exemplary user interface of an application menu on the portable multifunctional device 100 according to some embodiments is shown. A similar user interface is optionally implemented on the device 300. In some embodiments, the user interface 400 includes the following elements or a subset or superset thereof:

[0172] ● A signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;

[0173] ● Time 404;

[0174] ● A Bluetooth indicator 405;

[0175] ● A battery status indicator 406;

[0176] ● A tray 408 with icons for common applications, such as:

[0177] o An icon 416 labeled "Phone" for the phone module 138, which icon 416 optionally includes an indicator 414 of the number of missed calls or voicemails;

[0178] o An icon 418 labeled "Mail" for the email client module 140, which icon 418 optionally includes an indicator 410 of the number of unread emails;

[0179] o An icon 420 labeled "Browser" for the browser module 147; and

[0180] o An icon 422 labeled "iPod" for the video and music player module 152 (also known as the iPod (trademark of Apple Inc.) module 152); and

[0181] ● Icons for other applications, such as:

[0182] o The icon 424 of the IM module 141 marked as "Message";

[0183] o The icon 426 of the calendar module 148 marked as "Calendar";

[0184] o The icon 428 of the image management module 144 marked as "Photo";

[0185] o The icon 430 of the camera module 143 marked as "Camera";

[0186] o The icon 432 of the online video module 155 marked as "Online Video";

[0187] o The icon 434 of the stock market widget 149-2 marked as "Stock Market";

[0188] o The icon 436 of the map module 154 marked as "Map";

[0189] o The icon 438 of the weather widget 149-1 marked as "Weather";

[0190] o The icon 440 of the alarm clock widget 149-4 marked as "Clock";

[0191] o The icon 442 of the fitness support module 142 marked as "Fitness Support";

[0192] o The icon 444 of the note module 153 marked as "Note"; and

[0193] o The icon 446 of the settings application or module marked as "Settings", which provides access to the settings of the device 100 and its various applications 136.

[0194] It should be noted that Figure 4A The icon labels shown in are merely exemplary. For example, the icon 422 of the video and music player module 152 is marked "Music" or "Music Player". Other labels are optionally used for the various application icons. In some embodiments, the label of the corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.

[0195] Figure 4B A device (e.g., Figure 3 is shown having a touch-sensitive surface 451 (e.g., Figure 3Exemplary user interface on device 300). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting the intensity of contacts on the touch-sensitive surface 451 and / or one or more haptic output generators 357 for generating haptic output for a user of device 300.

[0196] Although some examples below will be given with reference to input on a touch screen display 112 (where a touch-sensitive surface and a display are combined), in some embodiments, the device detects input on a touch-sensitive surface separate from the display, as Figure 4B shown. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451 in ) has a main axis corresponding to the main axis (e.g., Figure 4B 453 in ) on the display (e.g., Figure 4B 452 in ). According to these embodiments, the device detects contact (e.g., Figure 4B in, 460 corresponds to 468 and 462 corresponds to 470) with the touch-sensitive surface 451 at a position corresponding to a corresponding position on the display (e.g., Figure 4B 460 and 462 in ). Thus, when the touch-sensitive surface (e.g., Figure 4B 451 in ) is separate from the display of the multifunctional device (e.g., Figure 4B 450 in ), user input (e.g., contacts 460 and 462 and their movement) detected by the device on this touch-sensitive surface is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.

[0197] Additionally, although the following examples are mainly given with reference to finger input (e.g., finger contact, single-finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of a contact). Another example, a tap gesture is optionally replaced by a mouse click when the cursor is above the position of the tap gesture (e.g., instead of detecting a contact, followed by stopping detection of the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or a mouse and a finger contact are optionally used simultaneously.

[0198] Figure 5AAn exemplary personal electronic device 500 is shown. The device 500 includes a body 502. In some embodiments, the device 500 may include some or all of the features described with respect to devices 100 and 300 (e.g., Figures 1A to 4B ). In some embodiments, the device 500 has a touch-sensitive display screen 504 hereinafter referred to as a touch screen 504. As an alternative or addition to the touch screen 504, the device 500 has a display and a touch-sensitive surface. As in the case of devices 100 and 300, in some embodiments, the touch screen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 504 (or the touch-sensitive surface) may provide output data representative of the intensity of the touch. The user interface of the device 500 may respond to the touch based on the intensity of the touch, meaning that touches of different intensities may invoke different user interface operations on the device 500.

[0199] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.

[0200] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) can be in physical form. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) may allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear device 500.

[0201] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include some or all of the components referred to in Figure 1A , Figure 1B and Figure 3 Device 500 has a bus 512 that operatively couples the I / O section 514 to one or more computer processors 516 and a memory 518. The I / O section 514 may be connected to a display 504 that may have a touch-sensitive component 522 and optionally a strength sensor 524 (e.g., a contact strength sensor). Additionally, the I / O section 514 may be connected to a communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 is optionally a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 508 is optionally a button.

[0202] In some examples, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.

[0203] The memory 518 of personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform, for example, the techniques described below, including processes 700, 900, 1100, 1300, and 1400 ( Fig. 7A , Figure 7B , Fig. 9 , Fig.11 , Fig.13 and Fig.14)。A computer-readable storage medium can be any medium that tangibly contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transient computer-readable storage medium. Non-transient computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. Personal electronic device 500 is not limited to Figure 5B the components and configurations thereof, but may include other components or additional components in a variety of configurations.

[0204] As used herein, the term "indicative representation" refers to a user-interactive graphical user interface object optionally displayed on the display screen of devices 100, 300, and / or 500 ( Figure 1A , Figure 3 and FIG. 5A to FIG. 5C ). For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) each optionally constitute an indicative representation.

[0205] As used herein, the term "focus selector" refers to an input element for indicating the current part of the user interface with which the user is interacting. In some specific implementations including a cursor or other position marker, the cursor acts as the "focus selector" such that when an input (e.g., a press input) is detected on a touch-sensitive surface (e.g., Figure 3 the touchpad 355 in Figure 4B or Figure 1A the touch-sensitive display system 112 in Figure 4AIn some specific implementations of the touch screen 112), the detected contact on the touch screen acts as a "focus selector", such that when an input (e.g., a press input made by the contact) is detected at the position of a specific user interface element (e.g., a button, a window, a slider, or other user interface element) on the touch screen display, the specific user interface element is adjusted according to the detected input. In some specific implementations, the focus moves from one area of the user interface to another area of the user interface without a corresponding movement of the cursor or a movement of the contact on the touch screen display (e.g., moving the focus from one button to another button by using the tab key or arrow keys); in these specific implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is typically a user interface element (or a contact on the touch screen display) that is controlled by the user to deliver the interaction with the user interface that the user anticipates (e.g., by indicating to the device the element of the user interface that the user desires to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or a touch screen), the position of the focus selector (e.g., a cursor, a contact, or a selection box) above the corresponding button will indicate that the user desires to activate the corresponding button (rather than other user interface elements shown on the device display).

[0206] As used in the specification and claims, the term "feature intensity" of a contact refers to a feature of the contact based on one or more intensities of the contact. In some embodiments, the feature intensity is based on a plurality of intensity samples. The feature intensity is optionally based on a predefined number or set of intensity samples collected during a predefined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after detecting the contact, before detecting the lift-off of the contact, before or after detecting the start of movement of the contact, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The feature intensity of the contact is optionally based on one or more of the following: the maximum value of the intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the value at the top 10% of the intensity of the contact, the half maximum value of the intensity of the contact, the 90% maximum value of the intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the feature intensity (e.g., when the feature intensity is the average value of the intensity of the contact over time). In some embodiments, the feature intensity is compared with a set of one or more intensity thresholds to determine whether the user has performed an operation. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact whose feature intensity does not exceed the first threshold results in a first operation, a contact whose feature intensity exceeds the first intensity threshold but does not exceed the second intensity threshold results in a second operation, and a contact whose feature intensity exceeds the second threshold results in a third operation. In some embodiments, the comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to forgo performing the corresponding operation) rather than for determining whether to perform a first operation or a second operation.

[0207] In some embodiments, a portion of a gesture is identified for use in determining the feature intensity. For example, a touch-sensitive surface optionally receives a continuous swiping contact that transitions from a starting position to an ending position at which the contact intensity increases. In this example, the feature intensity of the contact at the ending position is optionally based on only a portion of the continuous swiping contact, rather than the entire swiping contact (e.g., only the portion of the swiping contact at the ending position). In some embodiments, a smoothing algorithm is optionally applied to the intensity of the swiping contact before determining the feature intensity of the contact. For example, the smoothing algorithm optionally includes one or more of the following: an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swiping contact for the purpose of determining the feature intensity.

[0208] Optionally, characterize the contact intensity on the touch-sensitive surface relative to one or more intensity thresholds such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device will perform an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device will perform an operation different from an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, when a contact is detected with a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold, contacts below the nominal contact detection intensity threshold are no longer detected), the device will move a focus selector based on the movement of the contact on the touch-sensitive surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally speaking, unless otherwise stated, these intensity thresholds are consistent between different sets of user interface drawings.

[0209] An increase in contact characteristic intensity from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact characteristic intensity from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact characteristic intensity from an intensity below the contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact characteristic intensity from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a lift-off of the contact from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0210] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or contacts), where the corresponding press input is detected at least in part based on the intensity of the detected contact (or contacts) increasing above a press input intensity threshold. In some embodiments, a corresponding operation is performed in response to detecting the intensity of a corresponding contact increasing above a press input intensity threshold (e.g., the "downstroke" of the corresponding press input). In some embodiments, the press input includes an increase in the intensity of a corresponding contact above a press input intensity threshold and a subsequent decrease in the intensity of the contact below the press input intensity threshold, and a corresponding operation is performed in response to detecting the subsequent decrease in the intensity of the corresponding contact below the press input threshold (e.g., the "upstroke" of the corresponding press input).

[0211] In some embodiments, the device employs strength hysteresis to avoid unexpected inputs sometimes referred to as "jitter", where the device defines or selects a hysteresis strength threshold having a predefined relationship to a press input strength threshold (e.g., the hysteresis strength threshold is X strength units lower than the press input strength threshold, or the hysteresis strength threshold is 75%, 90%, or some reasonable proportion of the press input strength threshold). Thus, in some embodiments, a press input includes the strength of a corresponding contact increasing above the press input strength threshold and the strength of that contact subsequently decreasing below the hysteresis strength threshold corresponding to the press input strength threshold, and a corresponding operation is performed in response to detecting that the strength of the corresponding contact subsequently decreases below the hysteresis strength threshold (e.g., the "upstroke" of the corresponding press input). Similarly, in some embodiments, a press input is detected only when the device detects that the contact strength increases from a strength equal to or lower than the hysteresis strength threshold to a strength equal to or higher than the press input strength threshold and optionally the contact strength subsequently decreases to a strength equal to or lower than the hysteresis strength, and a corresponding operation is performed in response to detecting the press input (e.g., depending on the context, the contact strength increases or the contact strength decreases).

[0212] For ease of explanation, optionally, a description of an operation performed in response to a press input associated with a press input strength threshold or in response to a gesture including a press input is triggered in response to detecting any one of the following various conditions: the contact strength increases above the press input strength threshold, the contact strength increases from a strength lower than the hysteresis strength threshold to a strength higher than the press input strength threshold, the contact strength decreases below the press input strength threshold, and / or the contact strength decreases below the hysteresis strength threshold corresponding to the press input strength threshold. Additionally, in an example where an operation is described as being performed in response to detecting that the strength of a contact decreases below the press input strength threshold, the operation is optionally performed in response to detecting that the strength of the contact decreases below the hysteresis strength threshold corresponding to and less than the press input strength threshold.

[0213] Attention is now turned to embodiments of a user interface ("UI") implemented on an electronic device such as portable multifunctional device 100, device 300, or device 500 and associated processes.

[0214] Figure 5CExemplary diagram depicting a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500, and each device shares one or more data connections 510 (such as an Internet connection, a Wi-Fi connection, a cellular connection, a short-range communication connection, and / or any other such data connection or network) with each other to facilitate real-time communication of audio data and / or video data between the corresponding devices for a period of time. In some embodiments, the exemplary communication session may include a shared data session, whereby data is transferred from one or more of the electronic devices to other electronic devices to enable simultaneous output of the corresponding content at the electronic devices. In some embodiments, the exemplary communication session may include a video conference session, whereby audio data and / or video data is transferred between devices 500A, 500B, and 500C such that users of the corresponding devices can use the electronic devices for real-time communication.

[0215] In Figure 5C this, device 500A represents an electronic device associated with user A. Device 500A communicates (via data connection 510) with devices 500B and 500C, which are associated with user B and user C, respectively. Device 500A includes a camera 501A for capturing video data of the communication session, and a display 504A (e.g., a touch screen) for displaying content associated with the communication session. Device 500A also includes other components, such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.

[0216] Device 500A displays a communication UI 520A via display 504A, which is a user interface for facilitating a communication session (e.g., a video conference session) between device 500B and device 500C. Communication UI 520A includes video feed 525-1A and video feed 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and transferred from device 500B to device 500A and 500C during the communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and transferred from device 500C to device 500A and 500B during the communication session.

[0217] Communication UI 520A includes a camera preview 550A, which is a representation of video data captured via camera 501A at device 500A. Camera preview 550A represents to user A the expected video feeds of user A as displayed at the corresponding devices 500B and 500C.

[0218] The communication UI 520A includes one or more controls 555A for controlling one or more aspects of a communication session. For example, the controls 555A may include controls for muting the audio of the communication session, changing the camera view of the communication session (e.g., changing the camera used to capture the communication session video, adjusting the zoom value), terminating the communication session, applying visual effects to the camera view of the communication session, and activating one or more modes associated with the communication session. In some embodiments, the one or more controls 555A are optionally displayed in the communication UI 520A. In some embodiments, the one or more controls 555A are displayed separately from the camera preview 550A. In some embodiments, the one or more controls 555A are displayed as overlaying at least a portion of the camera preview 550A.

[0219] In Figure 5C this example, device 500B represents an electronic device associated with user B, who communicates with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B for capturing video data of the communication session, and a display 504B (e.g., a touch screen) for displaying content associated with the communication session. Device 500B also includes other components such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.

[0220] Device 500B displays a communication UI 520B similar to the communication UI 520A of device 500A via the touch screen 504B. The communication UI 520B includes video feeds 525-1B and 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during the communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during the communication session. The communication UI 520B also includes: a camera preview 550B, which is a representation of video data captured at device 500B via camera 501B; and one or more controls 555B similar to controls 555A, which are used to control one or more aspects of the communication session. The camera preview 550B represents to user B the expected video feeds of user B as displayed at the respective devices 500A and 500C.

[0221] In Figure 5CIn this case, device 500C represents an electronic device associated with user C, and user C communicates with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C for capturing video data of a communication session, and a display 504C (e.g., a touch screen) for displaying content associated with the communication session. Device 500C also includes other components, such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.

[0222] Device 500C displays a communication UI 520C similar to communication UI 520A of device 500A and communication UI 520B of device 500B via touch screen 504C. Communication UI 520C includes video feed 525-1C and video feed 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during the communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during the communication session. Communication UI 520C further includes: a camera preview 550C, which is a representation of video data captured at device 500C via camera 501C; and one or more controls 555C similar to controls 555A and 555B, which are used to control one or more aspects of the communication session. Camera preview 550C represents to user C the expected video feed of user C displayed at the corresponding devices 500A and 500B.

[0223] Although Figure 5C the diagrams depicted herein represent a communication session among three electronic devices, the communication session can be established between two or more electronic devices, and the number of devices participating in the communication session can change as electronic devices join or leave the communication session. For example, if one of the electronic devices leaves the communication session, the audio data and video data from the device that stops participating in the communication session are no longer represented on the participating devices. For example, if device 500B stops participating in the communication session, there is no data connection 510 between devices 500A and 500C, and there is no data connection 510 between devices 500C and 500B. Additionally, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and video data and audio data are shared among all devices such that each device can output the data transmitted from other devices.

[0224] Figure 5C A diagram depicting an embodiment of a communication session between multiple electronic devices, including FIG. 6A to FIG. 6Q , FIG. 8A to FIG. 8R , FIG. 10A to FIG. 10J and FIG. 12A to FIG. 12U the exemplary communication sessions depicted in FIG. 6A to FIG. 6Q , FIG. 8A to FIG. 8R , FIG. 10A to FIG. 10J and FIG. 12A to FIG. 12U The communication sessions depicted in include two or more electronic devices, even if other electronic devices participating in the communication session are not depicted in the diagram.

[0225] FIG. 6A to FIG. 6Q shows an exemplary user interface for managing a real-time video communication session (e.g., a video conference) according to some embodiments. The user interfaces in these figures are used to illustrate the processes described below that include Fig. 7A and Figure 7B the processes in.

[0226] FIG. 6A to FIG. 6Q shows a device 600 that displays a user interface for managing a real-time video communication session on a display 601 (e.g., a display device or a display generation component). FIG. 6A to FIG. 6Q depicts various embodiments in which the device 600 automatically reconstructs a display portion of the camera's field of view based on conditions detected in a scene within the field of view of the camera when the automatic framing mode is enabled. One or more of the embodiments discussed below with respect to FIG. 6A to FIG. 6Q can be combined with one or more of the embodiments discussed with respect to FIG. 8A to FIG. 8R , FIG. 10A to FIG. 10J and FIG. 12A to FIG. 12U .

[0227] The device 600 includes one or more cameras 602 (e.g., a front camera) that are used to capture image data and optionally depth data of a scene within the field of view of the camera. In some embodiments, the camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens having a relatively short focal length and a wide field of view). In some embodiments, the device 600 includes one or more features of the device 100, 300, or 500.

[0228] In Fig. 6AIn [the situation], device 600 displays a video conferencing request interface 604-1, which depicts an incoming request from "John" to participate in a live video conference. The video conferencing request interface 604-1 includes a camera preview 606, an option menu 608, a framing mode enabling indication 610, and a background blur enabling indication 611. The camera preview 606 is a real-time representation of the video feed from camera 602, which is enabled for output in the video conferencing session (if the incoming request is accepted). In Fig. 6A [the situation], camera 602 is currently enabled, and the camera preview 606 depicts a representation of "Jane", who is currently in front of device 600 and within the field of view of camera 602.

[0229] The option menu 608 includes various selectable options for controlling one or more aspects of the video conference. For example, the mute option 608-1 can be selected to mute the transmission of any audio detected by device 600. The flip option 608-2 can be selected to switch the camera being used for the video conference between camera 602 and one or more different cameras (such as a camera on the opposite side of device 600 from camera 602 (e.g., the rear camera)). The accept option 608-3 can be selected to accept the request to participate in the live video conference. The reject option 608-4 can be selected to reject the request to participate in the live video conference.

[0230] The background blur enabling indication 611 can be selected to enable or disable the background blur mode of the video conferencing session. In Fig. 6A the depicted embodiment, when device 600 receives or initiates a request to participate in a video conference, the background blur mode is disabled by default. Thus, the background blur enabling indication 611 is depicted as having an unselected state, as indicated by the non-bolded enabling indication in Fig. 6A [it]. In some embodiments, when device 600 receives or initiates a request to participate in a video conference, the background blur mode is enabled by default (e.g., and the enabling indication is bolded). Regarding FIG. 12A to FIG. 12U the background blur feature is discussed in more detail.

[0231] The framing mode enabling indication 610 can be selected to enable or disable the auto-framing mode of the video conferencing session. In Fig. 6A the depicted embodiment, when device 600 receives or initiates a request to participate in a video conference, the auto-framing mode is enabled by default. Thus, the framing mode enabling indication 610 is depicted as having a selected state, as indicated by the bolded enabling indication in Fig. 6A [it]. In some embodiments, when device 600 receives or initiates a request to participate in a video conference, the auto-framing mode is disabled by default (e.g., and there is no bolding of the enabling indication).

[0232] When the auto-framing mode is enabled, device 600 detects conditions of a scene within the field of view of an enabled camera (e.g., camera 602), such as the presence and / or location of one or more objects within the field of view of the camera, and adjusts in real time the field of view of the video output for a video conferencing session, as represented in the camera preview (e.g., without moving camera 602 or device 600), based on the detected conditions or changes in the scene within the field of view of the camera (e.g., changes in the location and / or movement of an object during a video conferencing session). Various implementations of the auto-framing mode are discussed throughout this disclosure.

[0233] For example, in the implementation depicted in FIG. 6A to FIG. 6Q , device 600 automatically adjusts (e.g., reframes) the display output video feed field of view to keep the display of one or more objects (e.g., Jane) within the field of view of a camera (e.g., camera 602). Because device 600 automatically adjusts the display portion of the camera's field of view to include the display of Jane, Jane can move around within the scene while participating in the video conference without having to manually adjust the perspective of the outgoing video feed to account for her movement or other changes in the scene. Thus, the participants in the video conference require less interaction with device 600 because device 600 automatically reframes the outgoing video feed such that remote participants receiving the video feed from device 600 can continuously view her as she moves around in her environment. Other benefits of the auto-framing mode will be described in the disclosure below. With respect to FIG. 6A to FIG. 6Q , FIG. 8A to FIG. 8R and FIG. 10A to FIG. 10J , various features of the auto-framing mode are discussed for the implementations depicted. One or more of these features can be combined with other features of the auto-framing mode, as discussed herein.

[0234] Fig. 6A Input 612 (e.g., tap input) on framing mode enablement indication 610 and input 614 on accept option 608-3 are depicted. In response to detecting input 612, device 600 disables the auto-framing mode. In response to detecting input 614 after input 612, device 600 accepts the video conference call and joins the video conference session with the auto-framing mode disabled, as depicted in Figure 6B . If device 600 does not detect input 612, or if input 612 is an input that enables the auto-framing mode (e.g., the framing mode enablement indication 610 is not selected when input 612 is received), then device 600 accepts the video conference call in response to input 614 and joins the video conference session with the auto-framing mode enabled, as depicted in Fig. 6F .

[0235] Figure 6Bdepicts a scene 615, which is a physical environment within the field of view 620 of a camera 602. In Figure 6B the depicted embodiment, Jane 622 is sitting on a couch 621, the device 600 is positioned in front of her (e.g., on a table), and a door 618 is in the background. The field of view 620 represents the field of view of the camera 602 (e.g., the maximum field of view of the camera 602 or the wide-angle field of view of the camera 602), which encompasses the scene 615. A portion 625 indicates the portion of the field of view 620 that is currently being output (or selected for output) for a video conferencing session (e.g., portion 625 represents the display portion of the camera's field of view). Thus, portion 625 indicates the portion of the scene 615 that is currently represented in the displayed video feed depicted in the camera preview 606. The field of view 620 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view, and the portion 625 is sometimes referred to herein as the video feed field of view.

[0236] When an incoming video conferencing request is accepted, the video conferencing request interface 604-1 transforms into the video conferencing interface 604, as Figure 6B shown. The video conferencing interface 604 is similar to the video conferencing request interface 604-1, but is updated to depict, for example, an incoming video feed 623 that includes video data of remote participants in the video conferencing received at the device 600. The video feed 623 includes a representation 623-1 of John as a remote participant in the video conferencing with Jane 622. Compared to Fig. 6A , the camera preview 606 is now reduced in size and shifted towards the upper right corner of the display 601. The camera preview 606 includes a representation 622-1 of Jane and a representation of the environment captured within the portion 625 (e.g., the video feed field of view). The option menu 608 is updated to include a framing mode enablement representation 610. In some embodiments, for example, the framing mode enablement representation 610 is displayed within the camera preview 606, as Fig.6O shown. As Figure 6B shown, the framing mode enablement representation 610 is shown as being in an unselected state, indicating that the auto-framing mode is disabled.

[0237] In some embodiments, when the auto-framing mode is disabled, the device 600 outputs a pre-determined portion of the available field of view of the camera 602 as the video feed field of view. An example of such an embodiment is depicted in FIG. 6B to FIG. 6D , where the portion 625 represents a pre-determined portion located at the center of the available field of view of the camera 602 within the field of view 620. In some embodiments, when the auto-framing mode is disabled, the device 600 outputs the entire field of view 620 as the video feed field of view. An example of such an embodiment is depicted in Fig. 10A .

[0238] In Figure 6B In [FIGURE], Jane is using device 600 to participate in a video conference with John. Similarly, John is using a device that includes one or more features of devices 100, 300, 500, or 600 to participate in a video conference with Jane. For example, John is using a tablet computer similar to device 600 (e.g., FIG. 10H to FIG. 10J and FIG. 12B to FIG. 12N John's tablet computer 600a in [FIGURE]). Thus, John's device displays a video conference interface similar to video conference interface 604, except that the camera preview on John's device shows the video feed captured from John's device (depicted in the video feed 623 in [FIGURE] currently), and the incoming video feed on John's device shows the video feed output from Jane's device 600 (depicted in the camera preview 606 in [FIGURE] currently). Figure 6B In [FIGURE], the device 600 moves relative to the scene 615, and thus the field of view 620 and the portion 625 pivot with the device 600. Since the auto-framing mode is disabled, the device 600 does not automatically adjust the video feed field of view to keep Jane's position fixed within the field of view 620. Instead, the perspective of the video feed moves with the device 600, and Jane 622 is no longer at the center of the video feed field of view, as indicated by the portion 625 and as depicted in the camera preview 606, which shows a background of the scene 615 and a portion of the representation 622-1 of Jane. In the embodiment depicted in [FIGURE], the movement of the device 600 is a pivot; however, the movement of the field of view 620 and the portion 625 can be caused by other movements, such as tilting, rotating, and / or moving the device 600 in a manner (e.g., forward, backward, and / or left and right) such that Jane 622 does not remain within the portion 625. Figure 6B In [FIGURE], the device 600 returns to its original position, and Jane 622 bends down and moves out of the portion 625. Again, since the auto-framing mode is disabled, the device 600 does not automatically adjust the video feed field of view to follow Jane's movement as she moves out of the portion 625. As Jane 622 moves, the video feed field of view remains stationary, and most of the representation 622-1 of Jane is outside the viewfinder in the camera preview 606.

[0239] In Figure 6C In [FIGURE], the device 600 moves relative to the scene 615, and thus the field of view 620 and the portion 625 pivot with the device 600. Since the auto-framing mode is disabled, the device 600 does not automatically adjust the video feed field of view to keep Jane's position fixed within the field of view 620. Instead, the perspective of the video feed moves with the device 600, and Jane 622 is no longer at the center of the video feed field of view, as indicated by the portion 625 and as depicted in the camera preview 606, which shows a background of the scene 615 and a portion of the representation 622-1 of Jane. In the embodiment depicted in [FIGURE], the movement of the device 600 is a pivot; however, the movement of the field of view 620 and the portion 625 can be caused by other movements, such as tilting, rotating, and / or moving the device 600 in a manner (e.g., forward, backward, and / or left and right) such that Jane 622 does not remain within the portion 625. Figure 6C In the embodiment depicted in [FIGURE], the movement of the device 600 is a pivot; however, the movement of the field of view 620 and the portion 625 can be caused by other movements, such as tilting, rotating, and / or moving the device 600 in a manner (e.g., forward, backward, and / or left and right) such that Jane 622 does not remain within the portion 625.

[0240] In Fig.6D In [FIGURE], the device 600 returns to its original position, and Jane 622 bends down and moves out of the portion 625. Again, since the auto-framing mode is disabled, the device 600 does not automatically adjust the video feed field of view to follow Jane's movement as she moves out of the portion 625. As Jane 622 moves, the video feed field of view remains stationary, and most of the representation 622-1 of Jane is outside the viewfinder in the camera preview 606.

[0241] In Fig.6DIn [reference], device 600 detects an input 626 (e.g., a tap input) on the viewfinder mode enable indication 610. In response, device 600 bolds the viewfinder mode enable indication 610 (to indicate its selected / enabled state) and enables the auto - viewfinder mode, as Fig. 6E shown. When the auto - viewfinder mode is enabled, device 600 automatically adjusts the field of view of the displayed video feed based on conditions detected within the scene 615. In the Fig. 6E embodiment depicted, device 600 adjusts the field of view of the displayed video feed to center on Jane's face. Accordingly, device 600 updates the camera preview 606 to include a representation 622 - 1 of Jane centered within the viewfinder frame and a representation 621 - 1 of the couch on which she is sitting in the background. The field of view 620 remains fixed as the position of the camera 602 remains unchanged. However, the position of Jane's face within the field of view 620 does change. Thus, device 600 adjusts (e.g., repositions) the displayed portion of the field of view 620 such that Jane remains positioned within the camera preview 606. This is represented by the repositioning of portion 625 in Fig. 6E such that it is centered on Jane's face. In Fig. 6E , portion 627 corresponds to the previous position of portion 625 and thus represents the portion of the field of view 620 that was previously displayed in the camera preview 606 (before the adjustments resulting from enabling the auto - viewfinder mode).

[0242] Fig. 6F depicts the scene 615 and device 600 when the auto - viewfinder mode is enabled in response to an input 614 accepting an incoming request to join a video conference, while the auto - viewfinder mode is enabled (alternatively in response to an input 626 enabling the auto - viewfinder mode). Accordingly, device 600 displays a representation 622 - 1 of Jane at the center of the camera preview 606.

[0243] In Figure 6G , device 600 moves in a manner similar to that discussed above with respect to Figure 6C . However, because the auto - viewfinder mode is enabled in Figure 6G , device 600 automatically adjusts the video feed field of view (portion 625) relative to the field of view 620 to remain fixed on Jane's face, which has changed its position relative to device 600 and camera 602 in response to the pivot of device 600. Thus, the camera preview 606 continues to display a representation 622 - 1 of Jane at the center of the video feed field of view. The adjustment of the video feed field of view is represented by the change in the position of portion 625 within the field of view 620. For example, when compared to Fig. 6F , the relative position of portion 625 has moved from a centered position within the field of view 620 (represented by portion 627 in Figure 6G ) to Figure 6GThe shifted positions depicted in. In Figure 6G In the embodiment depicted in, the movement of device 600 is a pivot. However, device 600 can automatically adjust the field of view of the displayed video feed in response to other movements, such as tilting, rotating, and / or moving device 600 (e.g., forward, backward, and / or left or right) in a manner such that Jane 622 remains within the field of view 620.

[0244] In Figure 6H Device 600 returns to its original position, and Jane 622 has moved to a standing position near the couch 621. Device 600 detects the updated position of Jane 622 in scene 615 and updates the field of view of the video feed to maintain its position on Jane 622, as shown in camera preview 606. Thus, portion 625 has moved from its previous position represented by portion 627 to an updated position around Jane's face, as Figure 6H Depicted in.

[0245] In some embodiments, the transition from the field of view depicted in camera preview 606 in Figure 6G To the field of view depicted in camera preview 606 in Figure 6H Is performed as a match cut. For example, the transition from camera preview 606 in Figure 6G To camera preview 606 in Figure 6H Is a match cut that is performed when Jane 622 has moved from her position sitting on couch 621 to her standing position near the couch. The result of the match cut is that camera preview 606 appears to transition from the first camera view in Figure 6G To a different camera view in Figure 6H (The different camera view optionally has the same zoom level as the camera view in Figure 6G ). However, the actual field of view of camera 602 (e.g., field of view 620) has not changed. Instead, only the portion of the field of view that is being displayed (portion 625) has changed position within field of view 620.

[0246] Fig.6I Is an embodiment similar to the embodiment depicted in Figure 6H But in which camera preview 606 has a larger, more zoomed-out field of view compared to the camera preview shown in Figure 6H . Specifically, Fig.6I The embodiment depicted in shows a jump cut transition from the camera preview in Figure 6G To the camera preview in Fig.6I . The jump cut transition is depicted by transitioning from camera preview 606 in Figure 6G To camera preview 606 in Fig.6I (Which has a larger (e.g., more zoomed-out) field of view). Thus, Fig.6IThe video feed field of view (represented by portion 625) within Fig.6I is a larger portion of the field of view 620. This is shown by the size difference between portion 625 (corresponding to the camera preview within Fig.6I ) and portion 627 (corresponding to the camera preview within Figure 6G ). Fig.6I The camera preview within Fig.6I and Figure 6G the camera preview within Figure 6G .

[0247] Figure 6H and Fig.6I illustrate specific implementations of transitions between different camera previews. In some implementations, other transitions may be performed, such as continuously moving (e.g., panning and / or zooming) the video feed field of view within the field of view 620 to follow Jane 622 as she moves around the scene 615.

[0248] In Figure 6J , Jane 622 has left the device 600, behind the couch 621. In response to detecting the change in the position of Jane 622, the device 600 performs a transition (e.g., a jump cut transition) where the camera preview 606 depicts a closer view of Jane in the scene 615. In the implementation depicted in Figure 6J , the device 600 zooms in (e.g., reduces the field of view) on Jane as she moves away from the camera (e.g., a threshold distance). In some implementations, the device 600 zooms out (e.g., expands the display field of view) from her as Jane 622 moves towards the camera (e.g., a threshold distance). For example, if Jane 622 were to move from her position in Figure 6J to her previous position in Fig.6I , the camera preview 606 would zoom out to the camera preview depicted in Fig.6I . Figure 6J In the implementation depicted in Figure 6J , the device 600 zooms in on Jane as she moves away from the camera (e.g., a threshold distance). In some implementations, the device 600 zooms out from her as Jane 622 moves towards the camera (e.g., a threshold distance). For example, if Jane 622 were to move from her position in Figure 6J to her previous position in Fig.6I , the camera preview 606 would zoom out to the camera preview depicted in Fig.6I . Figure 6J her position in Figure 6J to her previous position in Fig.6I Fig.6I , then the camera preview 606 would zoom out to the camera preview depicted in Fig.6I Fig.6I .

[0249] In Figure 6K , another object Jack 628 walks into the scene 615. The device 600 continues to display the video conferencing interface 604 with the same camera preview as depicted in Figure 6J . In the implementation depicted in Figure 6K , the device 600 detects Jack 628 within the field of view 620, but maintains the same video feed field of view as Jack 628 moves around the scene or until Jack 628 moves to a specific position within the scene (e.g., closer to the center of the scene). In some implementations, when an additional object is detected within the field of view 620, the device 600 displays a prompt to adjust the camera preview. Examples of this prompt are discussed in more detail below with respect to the implementations depicted in Figure 6P and FIG. 8A to FIG. 8J . Figure 6J the same camera preview as depicted in Figure 6J . In Figure 6K the implementation depicted in Figure 6K , the device 600 detects Jack 628 within the field of view 620, but maintains the same video feed field of view as Jack 628 moves around the scene or until Jack 628 moves to a specific position within the scene (e.g., closer to the center of the scene). In some implementations, when an additional object is detected within the field of view 620, the device 600 displays a prompt to adjust the camera preview. Examples of this prompt are discussed in more detail below with respect to the implementations depicted in Figure 6P and FIG. 8A to FIG. 8J . Figure 6P and FIG. 8A to FIG. 8J Examples of this prompt are discussed in more detail below with respect to the implementations depicted in Figure 6P and FIG. 8A to FIG. 8J .

[0250] In Figure 6LIn [description], device 600 reconstructs camera preview 606 to include a representation 628-1 of Jack 628, who is now standing next to Jane 622 in scene 615. In some embodiments, device 600 automatically adjusts camera preview 606 in response to determining that Jack 628 has stopped moving around scene 615 and / or he is exhibiting behavior indicating a desire to participate in a video conference. Examples of such behavior may include turning attention towards device 600 and / or camera 602, focusing / looking at camera 602 or in the general direction of the camera, remaining stationary (e.g., for at least a particular amount of time), positioning next to a participant in the video conference (e.g., Jane 622), facing device 600, speaking, etc. In some embodiments, when the auto-framing mode is enabled, device 600 automatically adjusts the camera preview in response to detecting a change in the number of objects detected within the field of view 620 (such as when Jack 628 enters scene 615). As Figure 6L depicted in [description], portion 625 represents the adjusted video feed field of view, and portion 627 represents the size of the video feed field of view before adjustment in Figure 6L [description]. Compared to the video feed field of view in Figure 6K [description], the adjusted field of view in Figure 6L [description] is zoomed out and recentered on Jane 622 and Jack 628, as depicted in camera preview 606.

[0251] In Figure 6M [description], Jane 622 begins to move away from Jack 628 and out of the video feed field of view represented by portion 625 and camera preview 606. When Jane's movement is detected, device 600 maintains (e.g., does not adjust) the field of view of camera preview 606. In some embodiments, as Jane moves away from Jack, device 600 readjusts the size of the video feed field of view such that both objects remain within the video feed field of view (camera preview). In some embodiments, after Jane has moved away from Jack, device 600 readjusts the video feed field of view after Jane stops moving such that both objects are within the camera preview.

[0252] In Figure 6N [description], device 600 detects that Jane 622 is no longer within portion 625 of Figure 6M [description] (e.g., portion 627 in Figure 6N [description]), and in response, adjusts the displayed video feed field of view to zoom in on Jack 628. Thus, device 600 displays video conference interface 604, where camera preview 606 has a zoomed-in view of Jack 628. Figure 6N depicts portion 625 having a smaller size than portion 627 to indicate the change in the displayed video feed field of view.

[0253] In some embodiments, when the auto-framing mode is enabled, the device 600 displays one or more prompts to adjust the video feed field of view to include additional participants in response to detecting additional objects within the field of view 620. Examples of such embodiments are described below with respect to FIG. 6O to FIG. 6Q the examples of such embodiments.

[0254] Fig.6O depicts an embodiment similar to the embodiment shown in Figure 6N except that the framing mode enablement indicator 610 is displayed in the camera preview area rather than in the options menu 608. The device 600 displays a framing indicator 630 positioned around the representation 628-1 of Jack to indicate the presence of a face (e.g., Jack's face) detected by the device 600 in the camera preview. In some embodiments, the framing indicator also indicates that the auto-framing mode is enabled and is tracking the detected face when a face is detected within the portion 625, the camera preview 606, and / or the field of view 620. Jack 628 is currently the only participant in the video conference located within the scene 615.

[0255] In Figure 6P the device 600 detects that Jane 622 has entered the scene 615 within the field of view 620. In response, the device 600 updates the video conference interface 604 by displaying an add enablement indicator 632 in the camera preview area. The add indicator 632 can be selected to adjust the displayed video feed field of view to include the additional object detected within the field of view 620.

[0256] In Figure 6P the device 600 detects an input 634 on the add enablement indicator 632 and, in response, adjusts the portion 625 such that the camera preview 606 includes the representation 622-1 of Jane and the representation 628-1 of Jack, as depicted in Figure 6Q In some embodiments, Jane 622 is added as a participant in the video conference. The device 600 also identifies the presence of Jane's face and displays a framing indicator 630 around the representation of Jane's face (in addition to those around the representation of Jack's face) to indicate that the framing mode is enabled and is tracking Jane's face when Jane's face is detected within the portion 625, the camera preview 606, and / or the field of view 620.

[0257] FIG. 7A to FIG. 7BFIG. 0 is a flowchart depicting a method for using an electronic device to manage a real-time video communication session according to some embodiments. Method 700 is performed at a computer system (e.g., a smart phone, a tablet) (e.g., 100, 300, 500, 600) that communicates with a display generation component (e.g., 601) (e.g., a display controller, a touch-sensitive display system, one or more cameras (e.g., 602) (e.g., an infrared camera, a depth camera, a visible light camera) and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0258] As described below, method 700 provides an intuitive way to manage a real-time video communication session. The method reduces the cognitive burden on the user to manage the real-time video communication session, thus creating a more effective human-machine interface. For battery-powered computing devices, enabling the user to manage the real-time video communication session faster and more efficiently saves power and increases the time interval between two battery charges.

[0259] In method 700, the computer system (e.g., 600) displays (702) a communication request interface (e.g., 604-1 in Fig. 6A FIG. 6) (e.g., an interface for an incoming or outgoing real-time video communication session (e.g., a real-time video chat session, a real-time video conference session)) via the display generation component (e.g., 601).

[0260] The computer system (e.g., 600) displays (704) a communication request interface (e.g., 604-1 in Fig. 6A FIG. 6), the communication request interface including a first selectable graphical user interface object (e.g., 608-3) (e.g., an "Accept" affordance) associated with the process for joining a real-time video communication session. In some embodiments, the Accept affordance can be selected to initiate the process for accepting an incoming request to join a real-time video communication session. In some embodiments, the first selectable graphical user interface object can be selected to initiate a "Cancel" affordance for the process of canceling or terminating an outgoing request to join a real-time video communication session.

[0261] The computer system (e.g., 600) (e.g., simultaneously with 704) displays (706) a communication request interface (e.g., 604-1 in Fig. 6AIn 604-1), the second selectable graphical user interface (e.g., "framing mode" enablement indication, "background blur" enablement indication, "dynamic video quality" enablement indication) is associated with a process for selecting between using a first camera mode (e.g., auto-framing mode, background blur mode, dynamic video quality mode) for the one or more cameras and using a second camera mode (e.g., a mode different from the first camera mode (e.g., a mode in which the auto-framing mode is disabled, a mode in which the background blur mode is disabled, and / or a mode in which the dynamic video quality mode is disabled)) for the one or more cameras during a real-time video communication session.

[0262] In some embodiments, the framing mode enablement indication can be selected to enable / disable a mode (e.g., auto-framing mode) that is used to: 1) track the position and / or orientation of one or more objects detected within the field of view of one or more cameras during a real-time video communication session, and 2) automatically adjust the display view of the object based on the tracking of the object during the real-time video communication session.

[0263] In some embodiments, the background blur enablement indication can be selected to enable / disable a mode (e.g., background blur mode) in which a visual effect (e.g., blur, darken, color, mask, desaturate, or otherwise de-emphasize the effect) is applied to the background portion (e.g., 606) of the camera field of view (e.g., camera preview, output video feed of the camera field of view) (e.g., without applying the visual effect to the portion of the camera field of view that includes the representation of the object (e.g., foreground portion) (e.g., 622-1)) during a real-time video communication session.

[0264] In some embodiments, the dynamic video quality indication represents a mode (e.g., a dynamic video quality mode) that can be selected to enable / disable output (e.g., transmit and optionally display) of the camera field of view, where portions have different degrees of compression and / or video quality. For example, a portion of the camera field of view that includes a detected face (e.g., 622-1) is less compressed than a portion of the camera field of view that does not include the detected face. In some embodiments, there is an inverse relationship between the degree of compression and video quality (e.g., greater compression results in lower video quality; less compression results in higher video quality). Thus, a video feed for a real-time video communication session can be transmitted (e.g., via a computer system (e.g., 600)) to a receiving device of a remote participant in the real-time video communication session such that a portion of the camera field of view that includes the detected face can be displayed at a higher video quality at the receiving device than a portion of the camera field of view that does not include the detected face (due to reduced compression of the portion of the camera field of view that includes the detected face and increased compression of the portion of the camera field of view that does not include the detected face). In some embodiments, the computer system changes the amount of compression as the video bandwidth changes (e.g., increases, decreases). For example, the degree of compression of a portion of the camera field of view that does not include the detected face (e.g., the camera feed) changes (e.g., increases or decreases with a corresponding change in bandwidth), while the degree of compression of a portion of the camera field of view that includes the detected face remains constant (or, in some embodiments, changes at a lower rate or in a smaller amount than the portion of the camera field of view that does not include the face).

[0265] When displaying a communication request interface (e.g., Fig. 6A 604-1 in), the computer system (e.g., 600) receives (708) via one or more input devices (e.g., 601) a set of one or more inputs (e.g., 612 and / or 614) including a selection (e.g., 614) of a first selectable graphical user interface object (e.g., 608-3) (e.g., the set of one or more inputs includes a selection of an acceptance indication and optionally includes a selection of a framing mode indication, a background blur indication, and / or a dynamic video quality indication).

[0266] In response to receiving the set of one or more inputs (e.g., 612 and / or 614) including a selection (e.g., 614) of the first selectable graphical user interface object (e.g., 608-3), the computer system (e.g., 600) displays (710) via a display generation component (e.g., 601) a real-time video communication interface (e.g., 604) for the real-time video communication session.

[0267] When displaying a real-time video communication interface (e.g., 604), a computer system (e.g., 600) detects (712) a change (e.g., a change in the position of an object) in a scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602). In some embodiments, the scene includes a representation of an object and optionally includes one or more additional objects within the field of view of the one or more cameras.

[0268] In response to detecting a change (714) in a scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602), the computer system (e.g., 600) performs one or more of steps 716 and 718 in method 700.

[0269] Based on a determination to select to use a first camera mode (e.g., enabled) (e.g., if the first camera mode is default disabled, the set of one or more inputs includes a selection of a second selectable graphical user interface object; if the first camera mode is default enabled, the set of one or more inputs does not include a selection of a second selectable graphical user interface object) (e.g., the viewfinder mode enable indication is in a selected state when the acceptance enable indication is selected), the computer system (e.g., 600) adjusts (716) (e.g., automatically, without user input) the representation of the field of view of the one or more cameras (e.g., 606) (e.g., the display field of view of the one or more cameras) based on the detected change in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras during a real-time video communication session (e.g., automatically adjust the field of view of the one or more cameras (e.g., based on the detected object position) during a real-time video communication session). When selecting to use the first camera mode, adjusting the representation of the field of view of the one or more cameras during a real-time video communication based on the detected change in the scene within the field of view of the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera (e.g., to keep the object / user in view) without additional input from the user. Performing an operation when a set of conditions has been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0270] In some embodiments, adjusting the representation of the field of view of the one or more cameras (e.g., 606) during a real-time video communication session includes: 1) Based on a determination that a first set of criteria is met, including that the scene includes at a first position (within the field of view 620 of the one or more cameras, at Fig. 6F Detected objects (e.g., 622) (e.g., one or more users of a computer system) in (e.g.,), display a representation with a first field of view (e.g., Fig. 6F 606 in) of the real-time video communication interface 604 (e.g., display the real-time video communication interface at a first digital zoom level and a first display portion of the field of view of the one or more cameras) (in some embodiments, the representation of the first field of view includes a representation of the object when the object is located at the first position); and 2) based on determining that a second set of criteria is met, including detecting an object (e.g., Figure 6H 625 in) at a second position different from the first position, display a representation of a second field of view different from the representation of the first field of view (e.g., Figure 6H 606 in) of the real-time video communication interface (e.g., display the real-time video communication interface at a second digital zoom level and / or a second display portion of the field of view of the one or more cameras) (e.g., a representation of the field of view that is zoomed in, zoomed out, and / or panned in a direction relative to the representation of the first field of view) (in some embodiments, the representation of the second field of view includes a representation of the object when the object is located at the second position). In some embodiments, when the first camera mode is selected for use (e.g., enabled), the representation of the field of view automatically changes in response to a detected change in the object's position and / or in response to detecting a second object entering or leaving the field of view of the one or more cameras (e.g., without changing the actual field of view of the one or more cameras). For example, the representation of the field of view changes to track the position of the object and adjust the display position and / or zoom level (e.g., digital zoom level) to more prominently display the object (e.g., changing the digital zoom level to appear to zoom in on the object as it moves away from the camera; changing the digital zoom level to appear to zoom out from the object as it moves toward the camera; changing the display portion of the field of view of the one or more cameras to appear to pan in a particular direction as the object moves in that direction).

[0271] Based on determining that the second camera mode is selected for use (e.g., enabled) (e.g., when the accept gesture representation is selected and the viewfinder mode gesture representation is in an unselected or deselected state), the computer system (e.g., 600) forgoes (718) adjusting the representation of the field of view of the one or more cameras during the real-time video communication session (e.g., as Fig.6D(depicted in, e.g., a detected change in the scene within the field of view of the one or more cameras, e.g., when the first camera mode is disabled, the real-time video communication interface maintains the same (e.g., default) representation of the field of view regardless of whether an object is within the scene in the field of view of the one or more cameras and regardless of where the object is located within the field of view of the one or more cameras). In some embodiments, forgoing adjusting the representation of the field of view of the one or more cameras during a real-time video communication session includes: 1) when (e.g., as determined) an object has a first position within the scene in the field of view of the one or more cameras (e.g., as depicted in Figure 6B (depicted in), displaying a real-time video communication interface having a representation of a first field of view (e.g., Figure 6B 606 in); and 2) when (e.g., as determined) an object has a second position within the scene in the field of view of the one or more cameras (e.g., in Fig.6D ), displaying a real-time video communication interface having a representation of a first field of view (e.g., Fig.6D 606 in). In some embodiments, the representation of the first field of view is a standard or default representation of the field of view that does not change based on changes in the scene (e.g., changes in the position of the object relative to the one or more cameras or a second object entering or leaving the field of view of the one or more cameras).

[0272] In some embodiments, a detected change in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) includes a detected change in a set of attention-based factors of one or more objects (e.g., 622, 628) within the scene (e.g., a first object turns their attention (e.g., focuses on, looks at) the one or more cameras (e.g., based on the gaze position, head position, and / or body position of the first object)). In some embodiments, the computer system (e.g., 600) adjusts the representation of the field of view of the one or more cameras (e.g., 606) during a real-time video communication session based on a detected change in the scene within the field of view of the one or more cameras, including: adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) a detected change in the set of attention-based factors of the one or more objects within the scene (e.g., as in Figure 6L(depicted in). Adjusting the representation of the field of view of one or more cameras during a real-time video communication based on detected changes in the set of attention-based factors of one or more objects in the scene enhances the video communication session experience by automatically adjusting the field of view of the cameras based on the set of attention-based factors of the objects in the scene without additional input from the user. Performing an operation when a set of conditions has been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0273] In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) adjusts (e.g., reframes) the display portion of the field of view of one or more cameras (e.g., 602) based on one or more attention-based factors of objects (e.g., 622, 628) detected within the field of view (e.g., 620) of the one or more cameras (e.g., 606). For example, when a first object turns its attention towards one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include the representation of the first object or focuses on the first object. Conversely, when the attention of the first object shifts away from the one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms in) to exclude the representation of the first object (e.g., if other objects remain within the field of view of the one or more cameras) or focuses on another object.

[0274] In some embodiments, the set of attention-based factors includes a first factor that is based on the detected focal plane of a first object (e.g., 628) among the one or more objects in the scene (e.g., as Figure 6L(depicted in). Adjusting the representation of the field of view of the one or more cameras during a real-time video communication based on the detected focal plane of a first object among the one or more objects in a scene enhances the video communication session experience by automatically adjusting the field of view of the camera when the focal plane of the first object meets a criterion without additional input from a user. Performing an operation without additional user input when a set of conditions have been met enhances the operability of a computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, the attention of the first object is determined based on the focal plane of the first object. For example, if the focal plane of the first object is aligned (e.g., coplanar) with the focal plane of the one or more cameras or the focal plane of another object participating in the real-time video communication session, the first object is considered to be paying attention to the one or more cameras. Thus, the first object is considered an active participant in the real-time video communication session, and the computer system then adjusts the representation of the field of view of the one or more cameras to include the first object in the real-time video communication interface.

[0275] In some embodiments, the set of attention-based factors includes a second factor based on whether a second object (e.g., 628) among the one or more objects in a scene (e.g., 615) (e.g., an object other than the first object) is determined to be looking at the one or more cameras (e.g., 602). Adjusting the representation of the field of view of the one or more cameras during a real-time video communication based on whether the second object in the scene is determined to be looking at the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera when the second object is looking at the camera without additional input from a user. Performing an operation without additional user input when a set of conditions have been met enhances the operability of a computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, the attention of the second object is determined based on whether the second object is looking at the one or more cameras. If so, the second object is considered to be paying attention to the one or more cameras and is thus considered an active participant in the real-time video communication session. Thus, the computer system adjusts the representation of the field of view of the one or more cameras to include the second object in the real-time video communication interface.

[0276] In some embodiments, detected changes in a scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) include detected changes in the quantity (e.g., amount, number) of objects (e.g., 622, 628) in the scene (e.g., detected changes in the quantity of objects detected in the scene that meet a first set of criteria (e.g., the objects are within the field of view of the one or more cameras and are optionally stationary) (e.g., one or more objects enter or leave the scene within the field of view of the one or more cameras)). In some embodiments, adjusting the representation (e.g., 606) of the field of view of the one or more cameras during a real-time video communication session based on detected changes in the scene within the field of view of the one or more cameras includes adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) detected changes in the quantity of objects detected (e.g., meeting a first set of criteria) in the scene (e.g., as depicted in Figure 6L and / or Figure 6N ). Adjusting the representation of the field of view of one or more cameras during a real-time video communication based on detected changes in the quantity of objects detected in the scene enhances the video communication session experience by automatically adjusting the field of view of the camera when the quantity of objects in the scene changes without additional input from the user. Performing an operation without additional user input when a set of conditions has been met enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) adjusts (e.g., reframes) the display portion of the field of view of the one or more cameras based on the quantity of objects detected within the field of view of the one or more cameras. For example, when the quantity of objects detected in the scene increases, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include additional objects (e.g., along with previously detected objects). Similarly, when the quantity of objects detected in the scene decreases, the representation of the field of view of the one or more cameras changes (e.g., zooms in) to capture the objects remaining in the scene.

[0277] In some embodiments, based on a detected change in the number of objects detected in a scene, the representation of the field of view of the one or more cameras (e.g., 606) is adjusted during a real-time video communication session based on a determination of whether an object in the field of view (e.g., 620) is stationary (e.g., relatively stationary; moving no more than a threshold amount of movement in the field of view of the one or more cameras). Adjusting the representation of the field of view of the one or more cameras during real-time video communication based on whether an object in the field of view is stationary enhances the video communication session experience by automatically adjusting the field of view of the camera when the objects in the scene are stationary without additional input from the user. Performing an operation when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0278] In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) considers an object to be a participant in a real-time video communication session when the detected object (e.g., 628) moves no more than a threshold amount of movement. When the object is considered an active participant, the computer system adjusts (e.g., reframes) the display portion of the field of view of the one or more cameras (e.g., 606) to subsequently include a representation of the object (e.g., as depicted in Figure 6L . This prevents the computer system from automatically reframing the display portion of the field of view of the one or more cameras based on irrelevant movement in the scene (e.g., movement caused by an object passing in the background or a child jumping around in the field of view of the one or more cameras), which would otherwise distract the participants / viewers of the real-time video communication session.

[0279] In some embodiments, before the computer system (e.g., 600) detects a change in the scene (e.g., 615) in the field of view of the one or more cameras (e.g., 602), the representation of the field of view of the one or more cameras (e.g., 606) has a first represented field of view (e.g., the computer system is displaying a portion of the field of view of the one or more cameras before detecting the change in the scene). In some embodiments, the change in the scene in the field of view of the one or more cameras includes a third object (e.g., 622) moving from a first portion (e.g., 625 in Figure 6G that corresponds to (e.g., is represented by; is included in) the first represented field of view (e.g., 606 in Figure 6G to a second portion of the field of view of the one or more cameras that does not correspond to (e.g., is not represented by; is not included in) the first represented field of view (e.g., Figure 6H Detected movement of a portion 625) therein (e.g., a third object moves from a portion of the scene corresponding to a portion of the field of view of the one or more cameras shown before the third object moves to a portion of the scene corresponding to a portion of the field of view of the one or more cameras not shown before the third object moves). In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on (in some embodiments, in response to) a detected change in the scene within the field of view of the one or more cameras includes: according to determining that a fourth object (e.g., 628) is not detected in the scene in a first portion of the field of view of the one or more cameras, adjusting the representation of the field of view from a first represented field of view to a second represented field of view corresponding to (e.g., representing; showing; including) a second portion of the field of view of the one or more cameras (e.g., different from the first represented field of view) (e.g., showing a camera preview 606 as depicted in Figure 6H . According to determining that a fourth object (e.g., 628) is detected in the scene in the first portion of the field of view of the one or more cameras (e.g., when Jane 622 leaves the frame, Jack 628 is located in Figure 6M portion 625 therein), abandon adjusting the representation of the field of view from the first represented field of view to the second represented field of view (e.g., continue to show the first represented field of view) (e.g., in Figure 6M , when Jane 622 leaves and Jack 628 remains, device 600 continues to show the camera preview 606 depicting portion 625). After another object (e.g., a third object) leaves the first portion of the field of view, selectively adjusting the representation of the field of view of the one or more cameras from a first represented field of view to a second represented field of view during a real-time video communication session based on whether an object (e.g., a fourth object) is detected in the scene in the first portion of the field of view of the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera based on whether there is an additional object remaining in the first portion of the field of view when another object leaves without additional input from the user. Performing an operation without additional user input when a set of conditions has been met enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating the computer system / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, when the automatic framing mode is enabled, the computer system does not track (e.g., follow; adjust the representation of the field of view in response thereto) the movement of another object that leaves the display field of view when an object remains in the display field of view.

[0280] In some embodiments, prior to detecting a change in a scene (e.g., 615) in the field of view (e.g., 620) of one or more cameras (e.g., 602), a representation (e.g., 606) of the field of view of the one or more cameras has a third represented field of view (e.g., in Fig. 6F )(e.g., the computer system is displaying a portion of the field of view of the one or more cameras prior to detecting a change in the scene). In some embodiments, a change in the scene in the field of view of the one or more cameras includes a fifth object (e.g., 622) moving from a third portion (e.g., Fig. 6F 625 in) of the field of view of the one or more cameras corresponding to (e.g., represented by; included in) the third represented field of view to a fourth portion (e.g., Figure 6H 625 in) of the field of view of the one or more cameras not corresponding to (e.g., not represented by; not included in) the third represented field of view. In some embodiments, adjusting a representation of the field of view of the one or more cameras during a real-time video communication session based on (in some embodiments, in response to) a detected change in the scene in the field of view of the one or more cameras includes: displaying, in a real-time video communication interface (e.g., 604), a representation of the field of view of the one or more cameras having a fourth represented field of view (e.g., Figure 6H 606 in)(e.g., different from the third represented field of view; in some embodiments, including a subset of the third represented field of view), the field of view corresponding to the fourth portion of the field of view of the one or more cameras and including a representation of the fifth object (e.g., 622-1)(e.g., replacing the display of the third represented field of view with the fourth represented field of view that includes a representation of the fifth object (and in some embodiments, includes a subset of the third represented field of view)). Stopping the display of the third represented field of view and displaying the fourth represented field of view corresponding to the fourth portion of the field of view and including a representation of the fifth object enhances the video communication session experience by automatically adjusting the field of view of the camera as the object moves to a different location in the scene to keep the object in view without additional input from the user. Performing an operation without additional user input when a set of conditions has been met enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0281] In some embodiments, adjusting the representation of the field of view of one or more cameras (e.g., 606) during a real-time video communication session based on detected changes in a scene (e.g., 615) within the field of view (e.g., 602) of the one or more cameras further includes stopping displaying, in a real-time video communication interface, the representation of the field of view of the one or more cameras having a third represented field of view (e.g., Fig. 6F 606 in). In some embodiments, when an auto-framing mode is enabled and a computer system (e.g., 600) detects that an object (e.g., 622) moves out of a framing box (e.g., 606) (e.g., out of the represented field of view being displayed), the computer system switches to a different framing box that includes the object (e.g., a representation showing a portion of the camera's field of view). In some embodiments, a cut (e.g., a jump cut) includes a change in zoom level (e.g., a zoom in or a zoom out). For example, a change in the represented field of view is a jump cut that includes a zoom out view that includes the user (e.g., when in a single-person tracking mode). In some embodiments, a cut (e.g., a match cut) includes a change from a first region of the camera's field of view being displayed to a second region of the camera's field of view that includes the user but does not include the first region. In some embodiments, when a second object remains within the framing box after a first object moves out of the framing box, the computer system displays a jump cut to a zoomed-in view of the second object (e.g., when the auto-framing mode is enabled).

[0282] In some embodiments, prior to detecting a change in a scene within the field of view of one or more cameras, the representation of the field of view of the one or more cameras (e.g., 606) has a first zoom value (e.g., a zoom setting (e.g., 1x, 0.5x, 0.7x)) (e.g., as Fig.6I depicted). In some embodiments, a change in a scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) includes a sixth object (e.g., 622) moving from a first position (e.g., Fig.6I 625 in) within the field of view of the one or more cameras to a second position (e.g., Figure 6Jthe movement of 625) therein, the first position corresponding to a representation (e.g., represented by) of the field of view and being a first distance from the one or more cameras, the second position corresponding to a representation (e.g., represented by) of the field of view and being a threshold distance from the one or more cameras (e.g., an object moves (e.g., towards the camera; away from the camera) within the displayed viewfinder to a pre-determined distance from the one or more cameras). In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on a detected change in the scene within the field of view of the one or more cameras includes: displaying, in a real-time video communication interface (e.g., 604), a representation of the field of view of the one or more cameras having a second zoom value different from the first zoom value (e.g., Figure 6J in 606) (e.g., a zoom-in view; a zoom-out view; in some embodiments, including the entire portion of the field of view of the one or more cameras that was previously displayed in a representation of the field of view having the first zoom value but is instead displayed with the second zoom value) (e.g., when an object moves within the originally displayed viewfinder to a pre-determined distance from the one or more cameras, jumping from a first zoom level to a second zoom level). Stopping the display of the representation of the field of view having the first zoom value and displaying the representation of the field of view having the second zoom value enhances the video communication session experience by automatically adjusting the zoom value of the representation of the field of view of the camera as the object moves to different distances from the camera within the scene to keep the object prominently displayed without additional input from the user. Performing an operation without additional user input when a set of conditions has been met enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0283] In some embodiments, adjusting the representation of the field of view of the one or more cameras (e.g., 602) during a real-time video communication session based on a detected change in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras includes stopping in a real-time video communication interface (e.g., Fig.6Ishows a representation of the field of view of one or more cameras with a first zoom value in (e.g., 606). In some embodiments, when an object (e.g., 622) moves toward the camera to a first threshold distance from the camera, the representation of the field of view of one or more cameras is switched to a second zoom value, which is a pulled-back view (e.g., a jump cut to a wide-angle view) of a previously displayed portion of the field of view of the one or more cameras. In some embodiments, when an object moves away from the camera to a second threshold distance from the camera, the representation of the field of view of the one or more cameras is switched to a second zoom value, which is a zoomed-in view of a previously displayed portion of the field of view of the one or more cameras.

[0284] In some embodiments, the computer system (e.g., 600) simultaneously displays a second selectable graphical user interface object (e.g., 610) and a real-time video communication interface (e.g., 604), the real-time video communication interface including one or more other selectable controls (e.g., 608) for controlling the real-time video communication (e.g., an end call button for ending the real-time video communication session, a switch camera button for switching which camera is used for the real-time video communication session, a mute button for muting / unmuting the audio of the user of the device in the real-time video communication session, an effects button for adding / removing visual effects from the real-time video communication session, an add user button for adding a user to the real-time video communication session, and / or a camera on / off button for turning on / off the video of the user in the real-time video communication session). In some embodiments, the second selectable graphical user interface object (e.g., the "framing mode" enabling representation) is continuously displayed during the real-time video communication session.

[0285] In some embodiments, when the real-time video communication interface (e.g., 604) is displayed upon detecting a seventh object (e.g., 628) (e.g., the first participant in the real-time video communication session) in a scene (e.g., 615) in the field of view (e.g., 620) of one or more cameras (e.g., 602), the computer system (e.g., 600) detects an eighth object (e.g., 622) (e.g., the second participant in the real-time video communication session) in the scene in the field of view of the one or more cameras (e.g., detecting an increase in the number of objects in the scene). In response to detecting the eighth object in the scene in the field of view of the one or more cameras, the computer system displays a prompt (e.g., 632) via a display generation component (e.g., 601) (e.g., text for adding the second participant to the real-time video communication session, an enabling representation for adding the second participant, an indication of the second participant (e.g., a framing indication (e.g., 630) in a potential preview (e.g., a blurred area) for showing the identification of the detected additional object), a stacked camera preview window, or as regarding FIG. 8A to FIG. 8Rthe other cues discussed) to adjust the representation of the field of view of the one or more cameras to include a representation of the eighth object (e.g., 622-1) in the real-time video communication interface (e.g., 604) (e.g., as Figure 6Q depicted therein). Responsive to detecting the eighth object in the scene, presenting a cue to adjust the representation of the field of view of the one or more cameras to include a representation of the eighth object in the real-time video communication interface provides feedback to a user of the computer system that an additional object has been detected in the field of view of the one or more cameras and reduces the amount of user input at the computer system by providing an option to automatically adjust the representation of the field of view to include the additional object without the user navigating a settings menu or other additional interface to adjust the represented field of view. Providing improved feedback and reducing the amount of input at the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0286] In some embodiments, prior to detecting a change in a scene (e.g., 615) in the field of view (e.g., 620) of one or more cameras (e.g., 602), a representation (e.g., 606) of the field of view of the one or more cameras has a fifth represented field of view (e.g., in Figure 6K )(e.g., the computer system (e.g., 600) is displaying a portion of the field of view of the one or more cameras prior to detecting the change in the scene). In some embodiments, a change in the scene in the field of view of the one or more cameras includes movement of one or more objects (e.g., 628) detected in the scene. In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on a detected change in the scene in the field of view of the one or more cameras includes: displaying, in the real-time video communication interface (e.g., 604), a sixth represented field of view (e.g., Figure 6La representation of the field of view of the one or more cameras that is different (e.g., different from a fifth represented field of view), such as by adjusting the representation of the field of view of the one or more cameras after the one or more objects have had less than a threshold amount of movement for a predetermined amount of time. Based on determining that the one or more objects have not had less than a threshold amount of movement for at least a threshold amount of time, continue to display, in the real-time video communication interface, a representation of the field of view of the one or more cameras having the fifth represented field of view (e.g., until the one or more objects have had less than a threshold amount of movement for at least a threshold amount of time) (e.g., maintaining the originally displayed viewfinder while one or more of these objects are moving). Selectively adjust, during a real-time video communication session, the representation of the field of view of the one or more cameras from a fifth represented field of view to a sixth represented field of view based on whether one or more objects detected in the scene have had less than a threshold amount of movement for at least a threshold amount of time, such as by automatically adjusting the field of view of the camera when an additional object enters the scene with the intention of participating in the real-time video communication session and not adjusting the field of view when an object enters the scene with the intention of not participating, to enhance the video communication session experience. This also reduces the amount of computation performed by the computer system by eliminating irrelevant adjustments to the represented field of view whenever the number of participants in the scene changes. Performing an operation when a set of conditions have been met without additional user input and reducing the amount of computation performed by the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, when the automatic viewfinder mode is enabled, the computer system maintains the original displayed view until one or more of these objects are stationary.

[0287] In some embodiments, a computer system (e.g., 600, 600a) displays, via a display generation component (e.g., 601, 601a), a representation of a first portion of the field of view of one or more cameras (e.g., 606, 1006, 1056, 1208, 1218) of a respective device of a respective participant in a real-time video communication session (e.g., a portion of the field of view of a camera of a remote participant in the real-time video communication session that includes the detected face of the remote participant (e.g., Figure 12L 1220-1b in; a portion of video feed 1210-1 that includes John's face; a portion of video feed 1023 that includes John's face; a portion of video feed 1053-2 that includes Jane's face)) (e.g., a portion of the field of view of one or more cameras of the computer system that includes the face of a detected object (e.g., Figure 12L1208-2 in; a portion of John's face included in camera preview 1218; a portion of Jane's face included in camera preview 1006; a portion of John's face included in camera preview 1056) and a representation of a second portion of the field of view of one or more cameras of the respective devices of the respective participants (e.g., a portion of the field of view of a remote participant's camera that does not include the detected face of the remote participant (e.g., Figure 12L 1220-1a in; a portion of John's face not included in video feed 1210-1; a portion of John's face not included in video feed 1023; a portion of Jane's face not included in video feed 1053-2) (e.g., a portion of the field of view of one or more cameras of a computer system that does not include the face of a detected object (e.g., Figure 12L 1208-1 in; a portion of John's face not included in camera preview 1218; a portion of Jane's face not included in camera preview 1006; a portion of John's face not included in camera preview 1056)), including determining that a first portion of the field of view of the one or more cameras includes detected features of a respective type (e.g., face; multiple different faces) while not detecting detected features of the respective type in a second portion of the field of view of the one or more cameras of the respective devices of the respective participants (e.g., when a face (or multiple different faces) is detected in the first portion of the field of view but not in the second portion, the second portion is compressed to a greater extent than the first portion (e.g., by a sending device (e.g., a remote participant's device (e.g., 600a, 600); the computer system of an object (e.g., 600, 600a))), such that when a face is detected in the first portion rather than in the second portion, the first portion of the field of view can be displayed at a higher video quality than the second portion (e.g., at a receiving device (e.g., a computer system; a remote participant's device)) (e.g., in Figure 12LIn it, 1220-1b is shown as having a higher video quality than 1220-1a; the portion of video feed 1210-1 that includes John's face is shown as having a higher image quality than the portion of video feed 1210-1 that does not include John's face; the portion of video feed 1053-2 that includes Jane's face is shown as having a higher video quality than the portion of video feed 1053-2 that does not include Jane's face; the portion of video feed 1023 that includes John's face is shown as having a higher video quality than the portion of video feed 1023 that does not include John's face)), presenting the representation of the first portion of the field of view of one or more cameras of the corresponding participant's corresponding device with a reduced compression level (e.g., higher video quality) compared to the representation of the second portion of the field of view of one or more cameras of the corresponding participant's corresponding device. Based on determining that the first portion of the field of view of one or more cameras includes a detected feature of a corresponding type while the corresponding type of detected feature is not detected in the second portion of the field of view of the one or more cameras, presenting the representation of the first portion of the field of view of one or more cameras of the corresponding participant's corresponding device in a real-time video communication session with a reduced compression level compared to the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device. This approach saves computing resources by saving bandwidth and reducing the amount of image data that is processed for display and / or transmission at a high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0288] In some embodiments, a computer system (e.g., 600, 600a) enables a dynamic video quality mode for output (e.g., transmission to a receiving device (e.g., 600a, 600), optionally while displaying at the sending device (e.g., 600, 600a)) portions of a camera field of view (e.g., 606, 1006, 1056, 1208, 1210-1, 1218, 1220-1, 1023, 1053-2) having different video compression levels. In some embodiments, the computer system compresses portions of the camera field of view that do not include one or more faces (e.g., Figure 12L 1220-1a in ; a portion of video feed 1210-1 that does not include John's face; a portion of video feed 1023 that does not include John's face; a portion of video feed 1053-2 that does not include Jane's face) more than portions of the camera field of view that include one or more faces (e.g., Figure 12L1220-1b in; a part of the video feed 1210-1 that includes John's face; a part of the video feed 1023 that includes John's face; a part of the video feed 1053-2 that includes Jane's face) more. In some embodiments, the computer system optionally displays the compressed video feed in the camera preview. In some embodiments, the computer system transmits video feeds with different compression levels during a real-time video communication session, such that the receiving device (e.g., a remote participant) can display a video feed received from the sending device (e.g., the computer system) that has higher video quality parts displayed simultaneously with lower video quality parts, wherein the higher video quality parts of the video feed include faces, and the lower video quality parts of the video feed do not include faces (e.g., 1220-1b is displayed as having a higher video quality than Figure 12L 1220-1a in; the part of the video feed 1210-1 that includes John's face is displayed as having a higher video quality than the part of the video feed 1210-1 that does not include John's face; the part of the video feed 1053-2 that includes Jane's face is displayed as having a higher video quality than the part of the video feed 1053-2 that does not include Jane's face; the part of the video feed 1023 that includes John's face is displayed as having a higher video quality than the part of the video feed 1023 that does not include John's face). Similarly, in some embodiments, the computer system receives compressed video data from a remote device (e.g., the device of a remote participant in a real-time video communication session), and displays a video feed from the remote device with different compression levels, such that the video feed of the remote device can be displayed together with higher video quality parts that include the face of the remote participant and lower video quality parts that do not include the face of the remote participant (displayed simultaneously with the higher quality parts) (e.g., 1220-1b is displayed as having a higher video quality than Figure 12L 1220-1a in; the part of the video feed 1210-1 that includes John's face is displayed as having a higher video quality than the part of the video feed 1210-1 that does not include John's face; the part of the video feed 1053-2 that includes Jane's face is displayed as having a higher video quality than the part of the video feed 1053-2 that does not include Jane's face; the part of the video feed 1023 that includes John's face is displayed as having a higher video quality than the part of the video feed 1023 that does not include John's face). In some embodiments, different levels of compression can be applied to a video feed in which multiple faces are detected. For example, the video feed can have multiple higher quality (less compressed) parts, each corresponding to the location of one of the detected faces.

[0289] In some embodiments, the dynamic video quality mode is independent of the auto-framing mode and the background blur mode, such that the dynamic video quality mode can be enabled and disabled separately from the auto-framing mode and the background blur mode. In some embodiments, the dynamic video quality mode is implemented together with the auto-framing mode, such that when the auto-framing mode is enabled, the dynamic video quality mode is enabled, and when the auto-framing mode is disabled, the dynamic video quality mode is disabled. In some embodiments, the dynamic video quality mode is implemented together with the background blur mode, such that when the background blur mode is enabled, the dynamic video quality mode is enabled, and when the background blur mode is disabled, the dynamic video quality mode is disabled.

[0290] In some embodiments, after a feature of a corresponding type has moved from a first portion of the field of view of one or more cameras of a corresponding participant's corresponding device (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) to a second portion of the field of view of one or more cameras of the corresponding participant's corresponding device (e.g., detecting the movement of the feature of the corresponding type from the first portion of the field of view of the one or more cameras to the second portion of the field of view of the one or more cameras; and, in response to detecting the movement of the feature of the corresponding type from the first portion of the field of view of the one or more cameras to the second portion of the field of view of the one or more cameras), a computer system (e.g., 600, 600a) displays, via a display generation component (e.g., 601, 601a), a representation of the first portion of the field of view of one or more cameras of the corresponding participant's corresponding device and a representation of the second portion of the field of view of one or more cameras of the corresponding participant's corresponding device (e.g., a portion of the field of view including the detected face), including displaying the representation of the first portion of the field of view of the one or more cameras of the corresponding participant's corresponding device at a greater degree of compression (e.g., lower video quality) compared to the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device, based on determining that the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device includes the detected feature of the corresponding type while the detected feature of the corresponding type is not detected in the representation of the first portion of the field of view of the one or more cameras of the corresponding participant's corresponding device (e.g., as the face moves within the field of view of the one or more cameras, the degree of compression of the corresponding portion of the field of view of the one or more cameras changes such that the face (e.g., a portion of the field of view including the face) is output (e.g., transmitted and optionally displayed) at a lower degree of compression than the portion of the field of view not including the face (e.g., as Jane's face moves, video feeds 1053-2 and / or 1220-1 are updated such that her face continues to be displayed at a higher video quality, and the portion of the video feed not including her face (even portions that were previously displayed at a higher quality) is displayed at a lower video quality; as John's face moves, video feeds 1023 and / or 1210-1 are updated such that his face continues to be displayed at a higher video quality, and the portion of the video feed not including his face (even portions that were previously displayed at a higher quality) is displayed at a lower video quality)).After a feature of a corresponding type has moved from a first portion of the field of view of one or more cameras of a corresponding participant's corresponding device to a second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device, in accordance with determining that the second portion of the field of view of the one or more cameras includes a detected feature of the corresponding type while the detected feature of the corresponding type is not detected in the first portion of the field of view of the one or more cameras, displaying a representation of the first portion of the field of view of the one or more cameras of the corresponding participant's corresponding device in a real-time video communication session with an increased degree of compression compared to a representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device. This approach saves computational resources by saving bandwidth and reducing the amount of image data that is processed for display and / or transmission at high image quality as the face moves within the scene. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0291] In some embodiments, the feature of the corresponding type is a face (e.g., a face detected within the field of view of one or more cameras; the face of a remote participant (e.g., the face of Jane in video feeds 1220-1 and / or 1053-2; the face of John in video feeds 1023 and / or 1210-1); the face of an object (e.g., the face of Jane in camera previews 606, 1006, and / or 1208; the face of John in camera previews 1056 and / or 1218)). In some embodiments, displaying a representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) includes displaying the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device with a reduced video quality compared to a representation of the first portion of the field of view of the one or more cameras of the corresponding participant's corresponding device, in accordance with determining that the first portion of the field of view of the one or more cameras includes a detected face while the face is not detected in the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device (e.g., 1220-1a is displayed as having a lower quality than Figure 12Llower video quality in 1220-1b; the portion of video feed 1210-1 that does not include John's face is shown as having lower video quality than the portion of video feed 1210-1 that includes John's face; the portion of video feed 1053-2 that does not include Jane's face is shown as having lower video quality than the portion of video feed 1053-2 that includes Jane's face; the portion of video feed 1023 that does not include John's face is shown as having lower video quality than the portion of video feed 1023 that includes John's face) (e.g., due to reduced compression of the representation of the first portion of the field of view of the one or more cameras). Based on determining that the first portion of the field of view includes the detected face while the face is not detected in the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device, the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device is shown with a video quality that is reduced compared to the representation of the first portion of the field of view, which saves computing resources by saving bandwidth and reducing the amount of image data that is processed for display and / or transmission at a high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0292] In some embodiments, portions of the camera field of view that do not include a detected face (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face) are output (e.g., transmitted and optionally displayed) at a lower image quality than portions of the camera field of view that include a detected face (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1053-2 that includes Jane's face) (due to increased compression of the portions that do not include a detected face). In some embodiments, when no face is detected in the field of view of one or more cameras, the computer system (e.g., 600, 600a) applies a uniform or substantially uniform degree of compression to a first portion and a second portion of the field of view of the one or more cameras, such that a video feed with a uniform or substantially uniform video quality can be output (e.g., both the first portion and the second portion). In some embodiments, when multiple faces are detected in the camera field of view (e.g., multiple participants in a real-time video communication session are detected), the computer system simultaneously applies reduced compression to the portions of the field of view corresponding to the detected faces, such that the faces with a higher image quality can be simultaneously displayed (e.g., at the receiving device). In some embodiments, the computer system applies increased compression to the representation of the second portion of the field of view of the one or more cameras, even if a face is detected in the second portion. For example, the computer system can determine that the face in the second portion is not a participant in the real-time video communication session (e.g., this person is a bystander in the background), and thus does not reduce the degree of compression of the second portion with that face.

[0293] In some embodiments, after a change (e.g., is detected) in the bandwidth for a representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of a corresponding device (e.g., 600, 600a) for a corresponding participant (e.g., in response to detecting), when a corresponding type of feature (e.g., a face) is detected in a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., 622-1; 1220-1b; 1208-2; the portion of 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1053-2 that includes Jane's face), and when the corresponding type of feature is not detected in a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., 621-1; 1220-1a; 1208-1; the portion of 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face), the degree of compression (e.g., the amount of compression) of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed by an amount less than the amount of change in the degree of compression of the representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., when a face is detected in the first portion of the field of view of the one or more cameras and no face is detected in the second portion of the field of view of the one or more cameras, the rate of change of compression (in response to a change in bandwidth (e.g., a decrease in bandwidth)) is less for the first portion of the field of view than for the second portion of the field of view). When a corresponding type of feature is detected in the first portion and not detected in the second portion, changing the degree of compression of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant by an amount less than the amount of change in the degree of compression of the representation of the second portion saves computational resources by saving bandwidth for the first portion of the representation of the field of view of the one or more cameras that includes the corresponding type of feature and reducing the amount of image data that is processed for display and / or transmission at a high image quality. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0294] In some embodiments, after a change (e.g., is detected) in the bandwidth for a representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of a corresponding device (e.g., 600, 600a) of a corresponding participant (e.g., in response to detecting), when a corresponding type of feature is not detected in a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., 621-1; 1220-1a; 1208-1; the portion of 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face), and when a corresponding type of feature is detected in a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., 622-1; 1220-1b; 1208-2; the portion of 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1053-2 that includes Jane's face), the degree of compression (e.g., amount of compression) of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed by an amount greater than the amount by which the degree of compression of the representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed (e.g., when a face is detected in the second portion of the field of view of the one or more cameras and no face is detected in the first portion of the field of view of the one or more cameras, (in response to the change in bandwidth (e.g., bandwidth reduction)) the rate of change of compression of the first portion of the field of view is greater than the rate of change of compression of the second portion of the field of view). Changing the degree of compression of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant by an amount greater than the amount by which the degree of compression of the representation of the second portion of the field of view is changed when a corresponding type of feature is not detected in the first portion and is detected in the second portion saves computational resources by saving bandwidth for the second portion of the representation of the field of view of the one or more cameras that includes the corresponding type of feature and reducing the amount of image data that is processed for display and / or transmission at high image quality. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0295] In some embodiments, in response to a change (e.g., being detected) in the bandwidth for a representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of a corresponding device (e.g., 600, 600a) of a corresponding participant, the quality (e.g., video quality) of a representation of a second portion (e.g., 621-1; 1220-1a; 1208-1; the portion of 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face) of the field of view of one or more cameras of the corresponding device of the corresponding participant is changed by an amount greater than the amount of change in the quality of a representation of a first portion (e.g., 622-1; 1220-1b; 1208-2; the portion of 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1056-2 that includes Jane's face) of the field of view of one or more cameras of the corresponding device of the corresponding participant (in some embodiments, the representation of the first portion does not change in quality or has a nominal amount of change in quality) (e.g., when a face is detected in the first portion of the field of view of one or more cameras and no face is detected in the second portion of the field of view of one or more cameras, in response to a change in bandwidth (e.g., a decrease in bandwidth), the image quality of the second portion changes more than the image quality of the first portion). Changing the quality of the representation of the second portion of the field of view of one or more cameras of the corresponding device of the corresponding participant by an amount greater than the amount of change in the quality of the representation of the first portion saves computing resources by saving bandwidth for the first portion of the representation of the field of view of one or more cameras and reducing the amount of image data that is processed for display and / or transmission at a high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.

[0296] In some embodiments, when a face is detected in a first portion of the field of view of one or more cameras (e.g., portion of 622-1; 1220-1b 1208-2; 1218 and / or 1210-1 that includes John's face; portion of 1006 and / or 1053-2 that includes Jane's face), and not detected in a second portion of the field of view (e.g., 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that does not include John's face; 1006 and / or 1053-2 that does not include Jane's face), the computer system (e.g., 600, 600a) detects a change in available bandwidth (e.g., an increase in bandwidth; a decrease in bandwidth), and in response, adjusts (e.g., increases, decreases) the compression of the second portion of the representation of the field of view of the one or more cameras, without adjusting the compression of the first portion of the representation of the field of view of the one or more cameras. In some embodiments, when a bandwidth change is detected, the computer system adjusts the compression of the first portion at a lower rate than the adjustment of the second portion. In some embodiments, the method includes detecting (e.g., at the respective device of the respective participant) a change in the bandwidth for transmitting a representation of the field of view of one or more cameras of the respective device of the respective participant when a respective type of feature (e.g., a face) is detected in a first portion of the field of view of one or more cameras of the respective device of the respective participant, and when the respective type of feature is not detected in a second portion of the field of view of one or more cameras of the respective device of the respective participant.

[0297] Note that the details of the processes described above with reference to method 700 (e.g., Figures 7A to 7B ) also apply in a similar manner to methods 900, 1100, 1300, and 1400 described below. For example, method 900, method 1100, method 1300, and / or method 1400 optionally include one or more features of the various methods described above with reference to method 700. For the sake of brevity, these details are not repeated below.

[0298] Figures 8A to 8R An exemplary user interface for managing a real-time video communication session (e.g., a video conference) in accordance with some embodiments is shown. The user interfaces in these figures are used to illustrate the processes described herein, including Figure 9 the processes in

[0299] Figures 8A to 8R Device 600 is shown displaying a user interface on display 601 for managing a real-time video communication session, similar to that discussed above with respect to Figures 6A to 6Q Figures 8A to 8R ​Depicts various embodiments in which device 600 prompts a user to adjust the camera field of view in response to detecting another object in the scene when the automatic framing mode is enabled (e.g., the camera preview). The following regarding Figures 8A to 8R One or more of the embodiments discussed in Figures 6A to 6Q , Figures 10A to 10J and Figures 12A to 12U can be combined with one or more of the embodiments discussed in

[0300] Figures 8A to 8J Depicts an exemplary embodiment in which device 600 displays a prompt to adjust the field of view of the video feed to include an additional participant in response to detecting an additional object in scene 615. Figures 8A to 8D Illustrates an embodiment in which the prompt includes displaying stacked camera preview options. Figures 8E to 8G Illustrates an embodiment in which the prompt includes displaying a camera preview with an obscured region and an unobscured region. Figures 8H to 8J Illustrates an embodiment in which the prompt includes displaying an option to switch between single-person framing mode and multi-person framing mode. Other embodiments are also provided (such as the embodiments discussed above regarding Figure 6P ), in which device 600 displays a prompt that includes an affordance (e.g., add affordance 632) that can be selected to adjust the display field of view to include an additional object.

[0301] Figure 8A Depicts an embodiment similar to the embodiment discussed above regarding Figure 6O , except that the video feed 623 now includes a representation 623-2 of Pam instead of a representation 623-1 of John. The automatic framing mode is enabled, as indicated by the bold state of the framing mode affordance 610 and the display of the framing indicator 630.

[0302] In Figure 8A , Jack is using device 600 to participate in a video conference with Pam. Similarly, Pam is using a device that includes one or more features of devices 100, 300, 500, or 600 to participate in a video conference with Jack. For example, Pam is using a tablet similar to device 600. Thus, Pam's device displays a video conference interface similar to video conference interface 604, except that the camera preview on Pam's device displays the video feed captured from Pam's device (currently depicted in the video feed 623 in Figure 8A ), and the incoming video feed on Pam's device displays the video feed output from device 600 (currently depicted in the camera preview 606 in Figure 8A ).

[0303] In Figure 8BIn [the figure], device 600 detects that Jane 622 enters scene 615 within the field of view 620. In response, device 600 updates the video conferencing interface 604 by displaying an auxiliary camera preview 806 that is located behind and offset from the camera preview 606. The stacked appearance of the camera preview 606 and the auxiliary camera preview 806 indicates that multiple video feed fields of view are available for the video conference, and the user can change the displayed field of view. Device 600 indicates that the camera preview 606 is the currently selected or enabled video feed field of view because it is on top of the auxiliary camera preview 806. Thus, the auxiliary camera preview 806 represents an option for adjusting the video feed field of view, which, in this embodiment, is an alternative field of view that includes both Jack 628 and Jane 622 (additional objects that have entered the scene). In some embodiments, different camera previews represent different zoom values, and thus, different camera preview options can also be regarded as different zoom controls / options.

[0304] Device 600 detects an input 804 (e.g., a tap input) on the stacked previews (e.g., on the auxiliary camera preview 606), and in response, updates the video conferencing interface 604 by translating the position of the auxiliary camera preview 806 such that it is no longer behind the camera preview 806, as Figure 8C depicted in [the figure].

[0305] In [the figure], Figure 8C device 600 separately (unstacked) displays both the camera preview 606 and the auxiliary camera preview 806 in the video conferencing interface 604. Device 600 also displays a bold outline 807 to indicate the currently selected video feed field of view, which is the camera preview 606 in Figure 8C [the figure]. The auxiliary camera preview 806 represents an available video feed field of view provided by device 600. Portion 825 represents the portion of the field of view 620 that is displayed in the auxiliary camera preview 806, while portion 625 represents the portion of the field of view 620 that is currently displayed in the camera preview 606. Although the camera preview 806 shows a rendering of portion 825, the camera preview 8060 is not currently selected, and thus, device 600 is not currently outputting a view of portion 825 of the video conference.

[0306] The camera preview 606 includes a portion of the representation 628-1 of Jack and the representation 622-1 of Jane. The auxiliary camera preview 806 is a zoomed-out view (compared to the view in the camera preview 606) that includes the representation 628-1 of Jack and the representation 622-1 of Jane. As previously discussed, the framing indicator 630 is depicted in both the camera preview 606 and the auxiliary camera preview 806 to indicate the detection of the faces of Jane and Jack within the respective video feed fields of view.

[0307] When the camera preview options are inFigure 8C When the unstacked configuration depicted in is displayed, device 600 can maintain available preview options or switch between available preview options in response to user input. For example, if device 600 detects input 811 on camera preview 606, the device continues to use (e.g., output) the video feed field of view represented by camera preview 606, and video conferencing interface 604 returns to Figure 8B the view depicted in. If device 600 detects input 812 on secondary camera preview 806, device 600 switches to (e.g., outputs) the video feed field of view represented by secondary camera preview 806, and the camera preview returns to the stacked configuration where secondary camera preview 806 is on top and camera preview 606 is on the bottom, as Figure 8D depicted in. In some embodiments, when device 600 switches from camera preview 606 to secondary camera preview 806, bold outline 807 moves from camera preview 606 to secondary camera preview 806 to indicate the switch from outputting camera preview 606 to outputting secondary camera preview 806.

[0308] In Figure 8D , portion 625 represents the portion of field of view 620 that is currently being output for the video conference. Since secondary camera preview 806 is selected in response to input 812, portion 625 now corresponds to secondary camera preview 806. Portion 627 represents the previous video feed field of view, which now corresponds to camera preview 606.

[0309] In Figure 8E , device 600 detects that Jane 622 has entered scene 615. Device 600 displays camera preview 606 with an unblurred region 606-1 (represented by boundary 808 and without hatching) and a blurred region 606-2 (represented by boundary 809 and with hatching). The unblurred region 606-1 represents the current video feed field of view, while the blurred region 606-2 represents an additional video feed field of view that is available for the video conference but is not currently being output. Thus, portion 625 corresponds to the field of view of the unblurred region, and portion 825 corresponds to the available field of view of the combined blurred and unblurred regions. The display of the blurred and unblurred regions in camera preview 606 indicates that the video feed field of view can be adjusted. The use of blurring is described as a way to distinguish the current video feed field of view from additional available video feed fields of view. However, these regions can be distinguished by other visual indications and appearances, such as shadows, darkening, highlighting, or other visual obscurations to emphasize or de-emphasize the respective regions. Boundaries 808 and 809 are also used to visually distinguish these regions.

[0310] In Figure 8EIn [description], the unblurred region 606-1 depicts the unoccluded representation 628-1a of Jack (specifically, Jack's face), which is being output for a video conference. The blurred region 606-2 depicts an occluded (e.g., blurred) representation of the available video feed field of view that is included in portion 825 of the field of view 620 but not included in portion 625. For example, in Figure 8E In [description], the blurred region 606-2 depicts an occluded representation 628-1b of Jack's body. In some embodiments, the blurred region and / or the unblurred region includes a framing indicator when the device 600 detects a face in the corresponding region.

[0311] In Figure 8F In [description], Jane 622 has entered portion 825 of the field of view 620, and the device 600 displays an occluded representation 622-1b of Jane in the blurred region 606-2 of the camera preview. The device 600 detects an input 813 on the blurred region 606-2 (or on the framing indicator around Jane's face located in the blurred region 606-2), and in response, adjusts (e.g., expands) the video feed field of view to include the previously blurred region 606-2, as depicted in Figure 8G In some embodiments, the blurred / unblurred regions of the camera preview represent different zoom values of the video feed field of view. Thus, a camera preview 606 that can be selected to switch to a field of view with a different zoom value (e.g., by expanding the unblurred region) can also be considered a zoom control.

[0312] Figure 8G Portion 625 in [description] represents the expanded portion of the field of view 620 that is now being output for the video conference, and portion 627 represents the previously displayed portion of the field of view 620 corresponding to the unblurred portion 606-1 in Figure 8F In [description]. The camera preview 606 now depicts the unoccluded representation 628-1 of Jack and the unoccluded representation 622-1 of Jane.

[0313] In some embodiments, when the conditions for triggering an adjustment are no longer met, the adjustment of the video feed field of view discussed above with respect to Figures 8E to 8G is reversed. For example, if Jane leaves the Figure 8G framing box 625 in [description], the device 600 returns to the state depicted in Figure 8E In [description], where the camera preview 606 includes the unblurred region 606-1 and the blurred region 606-2.

[0314] Figures 8H to 8J Depicts various interfaces of embodiments in which the device 600 switches the auto-framing mode between a single-person framing mode and a multi-person framing mode. In Figure 8HIn [description], when the auto - framing mode is enabled, device 600 detects Jack 628 in scene 615 and displays the video conferencing interface 604, where the framing mode options 830 are depicted in the camera preview 606. The framing mode options 830 include a single - person option 830 - 1 and a multi - person option 830 - 2. When the single - person option 830 - 1 is in the selected state, as Figure 8H depicted in [reference], the single - person framing mode setting is enabled and device 600 keeps the video feed field of view focused on the face of a single user, even when another person is detected in the field of view 620. For example, in Figure 8I [reference], although Jane 622 is now positioned next to Jack 628 in scene 615, device 600 keeps the video feed field of view featuring Jack 628 (represented in the camera preview 606 and section 625), rather than automatically adjusting the video feed field of view to include Jane.

[0315] In Figure 8I [reference], device 600 detects an input 832 on the multi - person option 830 - 2. In response, device 600 switches from the single - person framing mode setting to the multi - person framing mode setting. When the multi - person framing mode setting is enabled, device 600 automatically adjusts the video feed field of view to include additional objects (or a subset thereof) detected in the field of view 620. For example, when device 600 switches to the multi - person framing mode setting, device 600 expands the video feed field of view to include representations 628 - 1 and 622 - 1 of both Jack and Jane, as Figure 8J depicted in the camera preview 606 of [reference]. Thus, Figure 8J section 625 in [reference] represents the expanded video feed field of view resulting from enabling the multi - person framing mode setting, and section 627 represents the previous video feed field of view corresponding to the single - person framing mode setting. In some embodiments, the framing mode options 830 correspond to video feed fields of view with different zoom values, and thus the framing mode options 830 can also be regarded as zoom controls / options.

[0316] In some embodiments, Figures 8H to 8J the transition depicted in [reference] can be combined with a camera preview having blurred and unblurred regions as discussed with respect to Figures 8E to 8G [reference]. For example, device 600 can display a camera preview 606 having blurred and unblurred regions, similar to Figure 8E depicted in [reference], but also including the framing mode options 830 similar to Figure 8H depicted in [reference]. When device 600 is in the single - person framing mode setting, device 600 displays a camera preview 606 having both blurred and unblurred regions, regardless of whether anyone is detected in the blurred portion of the viewfinder (similar to Figure 8E and 8F(depicted in). However, when the device 600 is in the multi-person viewfinder mode setting, the device 600 can convert the camera preview 606 from a blurry and non-blurry appearance to a non-blurry appearance in response to detecting another person in the blurry area (similar to Figure 8F and Figure 8G in the conversion). Similarly, if a person is detected in the blurry area when the device 600 switches from the single-person viewfinder mode to the multi-person viewfinder mode (as shown in Figure 8F ), the device 600 adjusts the video feed field of view to include the previous blurry area, which includes the person previously detected in the blurry area (similar to Figure 8F and Figure 8G in the depicted conversion). In some embodiments, the device 600 can reverse the above conversion. For example, if the device 600 is displaying a camera preview 606 in which both objects are in the field of view (similar to Figure 8J depicted in), and the device 600 detects a selection of the single-person viewfinder option 830-1, the device 600 can adjust the video feed field of view to return to the blurry / unblurry appearance, similar to Figure 8F depicted in.

[0317] Now referring to Figure 8K , the device 600 displays a video conferencing interface 834, which depicts an incoming request to join a live video conference with John and two other remote participants. The video conferencing interface 834 is similar to the video conferencing interface 604, except that multiple participants are active in the video conference session depicted in the video conferencing interface 834. Accordingly, the embodiments described herein with respect to the video conferencing interface 604 can be applied in a similar manner to the video conferencing interface 834. Similarly, the embodiments described herein with respect to the video conferencing interface 834 can be applied in a similar manner to the video conferencing interface 604 and the like.

[0318] In Figure 8K , the device 600 detects an input 835 on the accept option 608-3 while enabling the auto-framing mode (as indicated by the bold appearance of the framing mode enablement 610) and disabling the background blur mode (as indicated by the non-bold appearance of the background blur enablement 611). In response, as Figure 8L depicted in, the device 600 accepts the live video conference call and joins the video conference session with the auto-framing mode enabled and the background blur mode disabled.

[0319] Figure 8LDevice 600 depicts a video conferencing interface 834 showing a camera preview 606 (similar to camera preview 836), as well as incoming video feeds 840-1, 840-2, and 840-3 for each of the corresponding remote participants in a live video conferencing session. Camera preview 836 includes a representation 622-1 of Jane and a framing mode enabling representation 610. In some embodiments, the framing mode enabling representation 610 is selectable in camera preview 836 to enable or disable the auto-framing mode. In some embodiments, the framing mode enabling representation 610 is not selectable until camera preview 836 is displayed in a magnified state, such as Figure 8M depicted in. In some embodiments, device 600 displays the framing mode enabling representation 610 in camera preview 836 when the auto-framing mode is enabled and does not display the enabling representation when the auto-framing mode is disabled. In some embodiments, device 600 persistently displays the framing mode enabling representation 610 and indicates whether the auto-framing mode is enabled by changing the appearance of the framing mode enabling representation (e.g., bolding the enabling representation when in the enabled mode). In some embodiments, the framing mode enabling representation 610 is displayed in option menu 608.

[0320] In Figure 8L , Jane is using device 600 to participate in a video conference with Pam, John, and Jack. Similarly, Pam, John, and Jack each use a corresponding device that includes one or more features of devices 100, 300, 500, or 600 to participate in a video conference with Jane and the other corresponding participants. For example, John, Jack, and Pam each use a tablet similar to device 600. Thus, the devices of the other participants (John, Jack, and Pam) each display a video conferencing interface similar to video conferencing interface 834, except that the camera preview on each corresponding device shows a video feed captured from that user's corresponding device (e.g., Pam's camera preview shows what is currently depicted in video feed 840-3, John's camera preview shows what is currently depicted in video feed 840-1, and Jack's camera preview shows what is currently depicted in video feed 840-2), and the incoming video feeds on the devices of the other participants (John, Jack, and Pam) include the video feed output from device 600 (what is currently depicted in Figure 8L the camera preview 836 in) as well as the video feeds output from the devices of the other participants.

[0321] In Figure 8L , device 600 detects an input 837 on camera preview 836 and, in response, magnifies camera preview 836, as Figure 8M depicted in. InFigure 8M In [reference], device 600 displays a framing mode enabling indication 610 at two locations in the video conferencing interface 834. The framing mode enabling indication 610-1 is displayed in the camera preview 836, and the framing mode enabling indication 610-2 is displayed in the options menu 608. In some embodiments, the framing mode enabling indication 610 is displayed in only one location at any given time (e.g., in the options menu 608 or the camera preview 836).

[0322] In some embodiments, when the auto-framing mode is not available, device 600 displays the framing mode enabling indication 610 with a changed appearance. For example, in Figure 8N the lighting conditions in scene 615 are poor, and in response to detecting the poor lighting conditions, device 600 displays the framing mode enabling indication 610-1 and the framing mode enabling indication 610-2 with a grayed-out appearance to indicate that the auto-framing mode is currently not available. When the lighting conditions improve, device 600 displays the framing mode enabling indication 610-1 and 610-2 with the appearance shown in Figure 8M .

[0323] In Figure 8N device 600 detects an input 839 (e.g., a tap input or a drag gesture) on the options menu 608, and in response, displays the interface depicted in Figure 8O . In some embodiments, device 600 displays the video conferencing interface 834 depicted in Figure 8N in response to detecting an input at other locations of the video conferencing interface 834 depicted in Figure 8L e.g., on the camera preview 836 or at a location in the interface other than the options menu 608).

[0324] In Figure 8O device 600 displays a video conferencing interface 834 with an expanded options menu 845, incoming video feeds 840-2 and 840-3, and a camera preview 836. The expanded options menu 845 includes information and various options for the video conference, including a framing mode option 845-1, which is similar to the framing mode enabling indication 610.

[0325] Now referring to Figure 8P , device 600 displays a video conferencing interface 834 with a magnified camera preview 836 and control options 850. In some embodiments, the interface depicted in Figure 8O is displayed in response to an input on the camera preview 836 in Figure 8PThe interface depicted therein. Control options 850 include a viewfinder mode option 850-1, a 1X zoom option 850-2, and a 0.5X zoom option 850-3. The viewfinder mode option 850-1 is similar to the viewfinder mode enabling indication 610 and is displayed in bold to indicate that the auto viewfinder mode is enabled. Since the auto viewfinder mode is enabled, the camera preview 836 includes a viewfinder indicator 852 (similar to the viewfinder indicator 630) positioned around the face of the representation 622-1 of Jane. The zoom options 850-2 and 850-3 can be selected to manually change the digital zoom level of the video feed field of view. Since the zoom options 850-2 and 850-3 manually adjust the digital zoom of the camera preview 836, selecting a zoom option disables the auto viewfinder mode, as discussed below.

[0326] In Figure 8P , the device 600 detects an input 853 on the zoom option 850-2. In response, the device 600 emphasizes (e.g., bolds and optionally enlarges) the zoom option 850-2 and disables the auto viewfinder mode. In Figure 8P In the depicted embodiment, the 1X zoom level is the zoom setting prior to detecting the input 853. Accordingly, the device 600 continues to display the representation 622-1 at the 1X zoom level. Since the auto viewfinder mode is disabled, the viewfinder mode option 850-1 is de-emphasized (e.g., no longer bold) and the viewfinder indicator 852 no longer appears in Figure 8Q .

[0327] In Figure 8Q , the device 60 detects an input 855 on the zoom option 850-3, and in response, adjusts the digital zoom level, as indicated in Figure 8R . Accordingly, the device 600 emphasizes the zoom option 850-3, de-emphasizes the zoom option 850-2, and displays the camera preview 836 with a 0.5X digital zoom value (zoomed out compared to the camera preview 836 in Figure 8Q ). Portion 625 represents the portion of the field of view 620 that is displayed after the video feed field of view is zoomed out, and portion 627 represents the portion of the field of view 620 that was previously displayed when the zoom option 850-2 was selected.

[0328] Figure 9FIG. 0 is a flowchart showing a method for using an electronic device to manage a real-time video communication session according to some embodiments. Method 900 is executed at a computer system (e.g., a smart phone, a tablet computer) (e.g., 100, 300, 500, 600) that communicates with a display generation component (e.g., a display controller, a touch-sensitive display system), one or more cameras (e.g., 602) (e.g., a visible light camera, an infrared camera, a depth camera), and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0329] As described below, method 900 provides an intuitive way to manage a real-time video communication session. This method reduces the cognitive burden on the user for managing a real-time video communication session, thus creating a more effective human-machine interface. For battery-powered computing devices, it enables the user to manage a real-time video communication session faster and more efficiently, saving power and increasing the time interval between two battery charges.

[0330] In method 900, a computer system (e.g., 600) displays (902) a real-time video communication interface (e.g., 604, 834) for a real-time video communication session (e.g., an interface for a real-time video communication session (e.g., a real-time video chat session, a real-time video conference session, etc.)) via a display generation component (e.g., 601). In some embodiments, the real-time video communication interface includes a real-time preview of the user of the computer system and a real-time representation of one or more participants (e.g., remote users) in the real-time video communication session.

[0331] The computer system (e.g., 600) displays a real-time video communication interface (e.g., 604, 834) that includes (904) representations (e.g., 623, 623-1, 840-1, 840-2, 840-3) of one or more participants (e.g., remote participants in the real-time video communication session) other than the participants (e.g., 622, 628) visible via one or more cameras (e.g., 602) in the real-time video communication session. In some embodiments, the participants visible via the one or more cameras are objects located within the field of view (e.g., 620) of the one or more cameras and represented (e.g., displayed) via the display generation component (e.g., 601) in the real-time video communication session (e.g., 606, ...

Claims

1. A method, comprising: at a computer system in communication with a display generation component, one or more cameras, and one or more input devices: display, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than participants visible via the one or more cameras; and a representation of the field of view of the one or more cameras, the representation of the field of view being visually associated with a visual indication of an option to change the representation of the field of view of the one or more cameras during the real-time video communication session; when displaying the real-time video communication interface for the real-time video communication session, detect, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and in response to detecting the set of one or more inputs, initiate the process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

2. The method according to claim 1, wherein: the visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session includes a set of one or more controls for adjusting a zoom level of the content displayed in the representation of the field of view of the one or more cameras.

3. The method according to claim 2, further comprising: detect, via the one or more input devices, a first input to the set of one or more controls for adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras; in response to detecting the first input to the set of one or more controls, display the set of one or more controls such that a first control option is displayed separately from a second control option; when displaying the set of one or more controls, detect a second input corresponding to a selection of the first control option or the second control option; and in response to detecting the second input, adjust the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on the selection of the first control option or the second control option.

4. The method according to claim 2, wherein The set of one or more controls for adjusting the zoom level of the representation of the field of view of the one or more cameras includes a first zoom control having a first fixed position and a second zoom control having a second fixed position, and the method further comprises: detect, via the one or more input devices, a third input corresponding to a selection of the first zoom control or the second zoom control; and While continuing to display the first zoom control having the first fixed position and the second zoom control having the second fixed position, and in response to detecting the third input, based on the selection of the first zoom control or the second zoom control, adjust the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

5. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Detecting the number of objects within the field of view of the one or more cameras during the real-time video communication session; and Adjusting the representation of the field of view of the one or more cameras based on the number of objects detected within the field of view of the one or more cameras during the real-time video communication session.

6. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on one or more characteristics of the scene within the field of view of the one or more cameras.

7. The method according to claim 1, wherein, The visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session can be selected to enable a first camera mode or a second camera mode, and the method further includes: When displaying the real-time video communication interface for the real-time video communication session, detecting a change in the scene within the field of view of the one or more cameras, including detecting a change in the number of objects in the scene; In response to detecting the change in the scene within the field of view of the one or more cameras: Based on determining that the first camera mode is enabled, adjusting the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene; and Based on determining that the second camera mode is enabled, forgoing adjusting the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene.

8. The method according to claim 1, wherein, The real-time video communication interface includes: A first representation of a first participant in the real-time video communication session, the first representation corresponding to a first portion of the field of view of the one or more cameras; and A second representation of a second participant different from the first participant in the real-time video communication session, the second representation corresponding to a second portion of the field of view of the one or more cameras, wherein the second portion of the field of view of the one or more cameras is different from the first portion of the field of view of the one or more cameras, and wherein the second representation is separated from the first representation.

9. The method according to claim 8, wherein In response to detecting an input on the real-time video communication interface for the real-time video communication session, display the first representation of the first participant and the second representation of the second participant.

10. The method according to claim 1, wherein The representation of the field of view of the one or more cameras includes a graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled.

11. The method according to claim 10, further comprising: When displaying the representation of the field of view of the one or more cameras, detecting an input to the representation of the field of view of the one or more cameras; And In response to detecting the input to the representation of the field of view of the one or more cameras, displaying, via the display generation component, a selectable graphical user interface object that can be selected to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session to a different viewing option of the field of view of the one or more cameras.

12. The method according to claim 10, wherein: When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is available for use, the graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled has a first appearance, and When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is not available for use, the graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled has a second appearance different from the first appearance.

13. The method according to claim 1, wherein The representation of the field of view of the one or more cameras includes: A first display area of the representation of the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and A second display area of the representation of the field of view of the one or more cameras, the second display area corresponding to a second portion of the field of view of the one or more cameras that is different from the first portion of the field of view, wherein the second display area is visually distinct from the first display area.

14. The method according to claim 13, further comprising: Based on determining that the first portion of the field of view of the one or more cameras is selected for the real-time video communication session, displaying the first display area with a visually unobstructed appearance and displaying the second display area of the representation of the field of view with an obstructed appearance; And Based on determining that the second portion of the field of view of the one or more cameras is selected for the real-time video communication session, displaying the second display area with a visually unobstructed appearance.

15. The method according to claim 13, wherein, The representation of the field of view of the one or more cameras includes a graphical element displayed to separate the first display area from the second display area.

16. The method according to claim 13, further comprising: Display one or more indications of one or more faces in the second portion of the field of view of the one or more cameras, wherein the one or more indications are selectable to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

17. The method according to claim 16, wherein: When a face of a first object is detected in the field of view of the one or more cameras and the first portion of the field of view of the one or more cameras is selected for the real-time video communication session, display the one or more indications of the one or more faces in the second portion of the field of view of the one or more cameras, and In response to detecting the one or more faces of an object other than the first object in the second portion of the field of view of the one or more cameras, display the one or more indications of the one or more faces.

18. The method according to claim 16, further comprising: When the first portion of the field of view of the one or more cameras is selected for the real-time video communication session: Detect a selection of the one or more indications of the one or more faces in the second portion of the field of view of the one or more cameras; And In response to detecting the selection of the one or more indications, initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session to include one or more representations of the one or more faces corresponding to the selected indication.

19. The method according to claim 1, wherein: The representation of the field of view of the one or more cameras includes a representation of a first object detected in the field of view of the one or more cameras, and Initiating the process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: automatically adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras to include a representation of a second object detected in the field of view of the one or more cameras that meets a set of criteria, wherein the representation of the second object is displayed simultaneously with the representation of the first object.

20. The method according to claim 1, further comprising: When displaying a representation of the field of view of the one or more cameras having a first display state and detecting a face of an object in a first region of the field of view of the one or more cameras, detect a second object in a second region of the field of view of the one or more cameras; In response to detecting the second object in the second region of the field of view of the one or more cameras, display a second selectable graphical user interface object; Detect a selection of the second selectable graphical user interface object; And In response to detecting the selection of the second selectable graphical user interface object, adjusting a representation of the field of view of the one or more cameras from having the first display state to having a second display state, the second display state being different from the first display state and including a representation of the second object.

21. The method according to claim 1, further comprising: When a mode for automatically adjusting a field of view of content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session based on a change in a position of an object detected in the field of view of the one or more cameras is enabled: Detecting a selection of an option for changing a zoom level of content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session; And In response to detecting the selection of the option for changing the zoom level, disabling the mode for automatically adjusting the field of view of content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session, and adjusting the zoom level of the content displayed in a representation of the field of view of the one or more cameras.

22. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs being configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: Displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: Representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; And A representation of the field of view of the one or more cameras, the representation of the field of view being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; And When displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting a field of view of content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session; And In response to detecting the set of one or more inputs, initiating the process for adjusting the field of view of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session.

23. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, the computer system comprising: One or more processors; And A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: Representations of one or more participants in the real-time video communication session other than the participants visible via the one or more cameras; And A representation of the field of view of the one or more cameras, the representation of the field of view being visually associated with a visual indication of an option to change the representation of the field of view of the one or more cameras during the real-time video communication session; And When displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; And In response to detecting the set of one or more inputs, initiating the process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

24. The non-transitory computer-readable storage medium according to claim 22, wherein: The visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session includes a set of one or more controls for adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras.

25. The non-transitory computer-readable storage medium according to claim 24, the one or more programs further comprising instructions for: Detecting, via the one or more input devices, a first input to the set of one or more controls for adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras; In response to detecting the first input to the set of one or more controls, displaying the set of one or more controls such that a first control option is displayed separately from a second control option; When displaying the set of one or more controls, detecting a second input corresponding to a selection of the first control option or the second control option; And In response to detecting the second input, adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on the selection of the first control option or the second control option.

26. The non-transitory computer-readable storage medium according to claim 24, wherein, The set of one or more controls for adjusting the zoom level of the representation of the field of view of the one or more cameras includes a first zoom control having a first fixed position and a second zoom control having a second fixed position, and the one or more programs further comprising instructions for: Detecting, via the one or more input devices, a third input corresponding to a selection of the first zoom control or the second zoom control; and While continuing to display the first zoom control having the first fixed position and the second zoom control having the second fixed position, and in response to detecting the third input, based on the selection of the first zoom control or the second zoom control, adjust the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

27. The non-transitory computer-readable storage medium according to claim 22, wherein, The visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session can be selected to enable a first camera mode or a second camera mode, and the one or more programs further include instructions for: When displaying the real-time video communication interface for the real-time video communication session, detect a change in the scene in the field of view of the one or more cameras, including detecting a change in the number of objects in the scene; In response to detecting the change in the scene in the field of view of the one or more cameras: Based on determining that the first camera mode is enabled, adjust the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene; And Based on determining that the second camera mode is enabled, forgo adjusting the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene.

28. The non-transitory computer-readable storage medium according to claim 22, wherein, The representation of the field of view of the one or more cameras includes a graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled.

29. The non-transitory computer-readable storage medium according to claim 28, wherein the one or more programs further include instructions for: When displaying the representation of the field of view of the one or more cameras, detect an input to the representation of the field of view of the one or more cameras; and In response to detecting the input to the representation of the field of view of the one or more cameras, display, via the display generation component, a selectable graphical user interface object that can be selected to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session to a different viewfinder option of the field of view of the one or more cameras.

30. The non-transitory computer-readable storage medium according to claim 28, wherein: When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is available for use, the graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled has a first appearance, and When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is not available for use, the graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled has a second appearance different from the first appearance.

31. The non-transitory computer-readable storage medium according to claim 22, wherein: The representation of the field of view of the one or more cameras includes: A first display area of the representation of the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and A second display area of the representation of the field of view of the one or more cameras, the second display area corresponding to a second portion of the field of view of the one or more cameras that is different from the first portion of the field of view, wherein the second display area is visually distinct from the first display area; and The one or more programs further include instructions for: Displaying the first display area with a visually unobstructed appearance and displaying the second display area of the representation of the field of view with an obstructed appearance based on determining that the first portion of the field of view of the one or more cameras is selected for the real-time video communication session; and Displaying the second display area with a visually unobstructed appearance based on determining that the second portion of the field of view of the one or more cameras is selected for the real-time video communication session.

32. The non-transitory computer-readable storage medium according to claim 22, wherein: The representation of the field of view of the one or more cameras includes: A first display area of the representation of the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and A second display area of the representation of the field of view of the one or more cameras, the second display area corresponding to a second portion of the field of view of the one or more cameras that is different from the first portion of the field of view, wherein the second display area is visually distinct from the first display area; and The one or more programs further include instructions for: Displaying one or more indications of one or more faces in the second portion of the field of view of the one or more cameras, wherein the one or more indications can be selected to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

33. The non-transitory computer-readable storage medium according to claim 32, wherein: When a face of a first object is detected in the field of view of the one or more cameras and the first portion of the field of view of the one or more cameras is selected for the real-time video communication session, displaying the one or more indications of the one or more faces in the second portion of the field of view of the one or more cameras, and In response to detecting the one or more faces of an object other than the first object in the second portion of the field of view of the one or more cameras, displaying the one or more indications of the one or more faces.

34. The non-transitory computer-readable storage medium according to claim 32, the one or more programs further include instructions for: When a first portion of the field of view of the one or more cameras is selected for the real-time video communication session: detect a selection of one or more of the indications of one or more faces in a second portion of the field of view of the one or more cameras; and in response to detecting the selection of one or more of the indications, initiate a process for adjusting the field of view of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session to include one or more representations of the one or more faces corresponding to the selected indication.

35. The non-transitory computer-readable storage medium according to claim 22, wherein the one or more programs further comprise instructions for: when displaying a representation of the field of view of the one or more cameras having a first display state and detecting a face of an object in a first region of the field of view of the one or more cameras, detecting a second object in a second region of the field of view of the one or more cameras; in response to detecting the second object in the second region of the field of view of the one or more cameras, displaying a second selectable graphical user interface object; detecting a selection of the second selectable graphical user interface object; and in response to detecting the selection of the second selectable graphical user interface object, adjusting the representation of the field of view of the one or more cameras from having the first display state to having a second display state, the second display state being different from the first display state and including a representation of the second object.

36. The non-transitory computer-readable storage medium according to claim 22, wherein the one or more programs further comprise instructions for: when a mode for automatically adjusting the field of view of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session based on a change in the position of an object detected in the field of view of the one or more cameras is enabled: detecting a selection of an option for changing a zoom level of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session; and in response to detecting the selection of the option for changing the zoom level, disabling the mode for automatically adjusting the field of view of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session and adjusting the zoom level of the content displayed in a representation of the field of view of the one or more cameras.

37. The computer system according to claim 23, wherein: the visual indication of the option for changing the representation of the field of view of the one or more cameras during the real-time video communication session includes one or more controls for adjusting a zoom level of the content displayed in a representation of the field of view of the one or more cameras.

38. The computer system according to claim 37, wherein the one or more programs further comprise instructions for: Detect a first input for the one or more controls for adjusting a zoom level of the content displayed in a representation of the field of view of the one or more cameras via the one or more input devices; In response to detecting the first input for the one or more controls, display the one or more controls such that a first control option is displayed separately from a second control option; While displaying the one or more controls, detect a second input corresponding to a selection of the first control option or the second control option; And In response to detecting the second input, adjust the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on the selection of the first control option or the second control option.

39. The computer system according to claim 37, wherein, The one or more controls for adjusting the zoom level of the representation of the field of view of the one or more cameras include a first zoom control having a first fixed position and a second zoom control having a second fixed position, and the one or more programs further include instructions for: Detecting a third input corresponding to a selection of the first zoom control or the second zoom control via the one or more input devices; and While continuing to display the first zoom control having the first fixed position and the second zoom control having the second fixed position, and in response to detecting the third input, adjust the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on the selection of the first zoom control or the second zoom control.

40. The computer system according to claim 23, wherein, A visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session can be selected to enable a first camera mode or a second camera mode, and the one or more programs further include instructions for: While displaying the real-time video communication interface for the real-time video communication session, detect a change in the scene in the field of view of the one or more cameras, including detecting a change in the number of objects in the scene; In response to detecting the change in the scene in the field of view of the one or more cameras: Based on the determination to enable the first camera mode, adjust the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene; And Based on the determination to enable the second camera mode, forgo adjusting the representation of the field of view of the one or more cameras based on the change in the number of objects in the scene.

41. The computer system according to claim 23, wherein, The representation of the field of view of the one or more cameras includes a graphical indication of whether the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is enabled.

42. The computer system according to claim 41, wherein the one or more programs further include instructions for: When displaying a representation of the field of view of the one or more cameras, detect an input to the representation of the field of view of the one or more cameras; and In response to detecting the input to the representation of the field of view of the one or more cameras, display, via the display generating component, a selectable graphical user interface object that can be selected to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session to a different framing option of the field of view of the one or more cameras.

43. The computer system according to claim 41, wherein: When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is available for use, the graphical indication of whether to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session has a first appearance, and When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is not available for use, the graphical indication of whether to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session has a second appearance different from the first appearance.

44. The computer system according to claim 23, wherein: The representation of the field of view of the one or more cameras includes: A first display area of the representation of the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and A second display area of the representation of the field of view of the one or more cameras, the second display area corresponding to a second portion of the field of view of the one or more cameras that is different from the first portion of the field of view, wherein the second display area is visually distinct from the first display area; and The one or more programs further include instructions for: Displaying the first display area having a visually unobstructed appearance and displaying the second display area of the representation of the field of view having an obstructed appearance based on determining that the first portion of the field of view of the one or more cameras is selected for the real-time video communication session; and Displaying the second display area having a visually unobstructed appearance based on determining that the second portion of the field of view of the one or more cameras is selected for the real-time video communication session.

45. The computer system according to claim 23, wherein: The representation of the field of view of the one or more cameras includes: A first display area of the representation of the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and a second display area of a representation of the field of view of the one or more cameras, the second display area corresponding to a second part of the field of view of the one or more cameras that is different from the first part of the field of view, wherein the second display area is visually distinct from the first display area; and the one or more programs further include instructions for: displaying one or more indications of one or more faces in the second part of the field of view of the one or more cameras, wherein the one or more indications are selectable to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.

46. The computer system of claim 45, wherein: when a face of a first object is detected in the field of view of the one or more cameras and the first part of the field of view of the one or more cameras is selected for the real-time video communication session, displaying the one or more indications of the one or more faces in the second part of the field of view of the one or more cameras, and in response to detecting the one or more faces of an object other than the first object in the second part of the field of view of the one or more cameras, displaying the one or more indications of the one or more faces.

47. The computer system of claim 45, the one or more programs further include instructions for: when the first part of the field of view of the one or more cameras is selected for the real-time video communication session: detecting a selection of the one or more indications of the one or more faces in the second part of the field of view of the one or more cameras; and in response to detecting the selection of the one or more indications, initiating a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session to include one or more representations of the one or more faces corresponding to the selected indication.

48. The computer system of claim 23, the one or more programs further include instructions for: when displaying a representation of the field of view of the one or more cameras having a first display state and detecting a face of an object in a first region of the field of view of the one or more cameras, detecting a second object in a second region of the field of view of the one or more cameras; in response to detecting the second object in the second region of the field of view of the one or more cameras, displaying a second selectable graphical user interface object; detecting a selection of the second selectable graphical user interface object; and in response to detecting the selection of the second selectable graphical user interface object, adjusting the representation of the field of view of the one or more cameras from having the first display state to having a second display state, the second display state being different from the first display state and including a representation of the second object.

49. The computer system according to claim 23, wherein the one or more programs further comprise instructions for: When a mode for automatically adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on a change in the position of an object detected in the field of view of the one or more cameras is enabled: Detecting a selection of an option for changing a zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and In response to detecting the selection of the option for changing the zoom level, disabling the mode for automatically adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session, and adjusting the zoom level of the content displayed in the representation of the field of view of the one or more cameras.

50. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Translating the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on one or more conditions detected in a scene in the field of view of the one or more cameras.

51. The non-transitory computer-readable storage medium according to claim 22, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Translating the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on one or more conditions detected in a scene in the field of view of the one or more cameras.

52. The computer system according to claim 23, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Translating the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on one or more conditions detected in a scene in the field of view of the one or more cameras.

Citation Information

Patent Citations

  • Method and apparatus for integrating manual input

    US20020015024A1

  • Acceleration-based theft detection system for portable electronic devices

    US20050190059A1

  • Methods and apparatuses for operating a portable device based on an accelerometer

    US20060017692A1

  • Gestures for touch sensitive input devices

    US20060026521A1

  • Gestures for touch sensitive input devices

    US20060026536A1