User interface for wide-angle video conferencing
By simplifying the user interface and adjusting camera modes, the complex and inefficient real-time video communication session management problem in existing technologies is solved, improving device efficiency and user satisfaction while saving device energy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-01-28
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies for managing real-time video communication sessions have complex and inefficient user interfaces, resulting in wasted user time and device energy, especially in battery-powered devices.
A simplified user interface is provided, which displays a communication request interface through a display generation component, receives input device selections, detects changes in camera field of view, adjusts the camera mode according to the selection, optimizes the camera field of view representation, and reduces user operation steps and device power consumption.
It enables faster and more efficient real-time video communication session management, reduces the cognitive burden on users, saves device power, and extends battery life.
Smart Images

Figure CN120201156B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on January 28, 2022, with application number 202280008837.3 and entitled "User Interface for Wide-Angle Video Conferencing". Technical Field
[0002] This disclosure relates generally to computer user interfaces, and more specifically to techniques for managing real-time video communication sessions. Background Technology
[0003] The computer system may include hardware and / or software for displaying an interface for real-time video communication sessions. Summary of the Invention
[0004] However, some technologies used to manage real-time video communication sessions using electronic devices are often cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may involve multiple keystrokes or button presses. These existing technologies require more time than necessary, resulting in wasted user time and device power. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, this technology provides electronic devices with a faster and more efficient method and interface for managing real-time video communication sessions. Such methods and interfaces can optionally complement or replace other methods for managing real-time video communication sessions. These methods and interfaces reduce the cognitive burden on users and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging.
[0006] This article describes an example method. An exemplary method includes, at a computer system communicating with a display generation component, one or more cameras, and one or more input devices: displaying a communication request interface via the display generation component, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode and using a second camera mode for one or more cameras during the real-time video communication session; receiving, while displaying the communication request interface, a set of one or more inputs including a selection of the first selectable graphical user interface object via one or more input devices; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying a real-time video communication interface for the real-time video communication session via the display generation component; while displaying the real-time video communication interface, detecting changes in the scene in the field of view of one or more cameras; and in response to detecting changes in the scene in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on the detected changes in the scene in the field of view of one or more cameras, depending on determining that the first camera mode is selected; and abandoning the adjustment of the representation of the field of view of one or more cameras during the real-time video communication session, depending on determining that the second camera mode is selected.
[0007] This document describes an example non-transitory computer-readable storage medium. An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with display generating components, one or more cameras, and one or more input devices. The one or more procedures include instructions for: displaying a communication request interface via a display generation component, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object, while displaying the communication request interface; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying a real-time video communication interface for the real-time video communication session via the display generation component; detecting changes in the scene in the field of view of one or more cameras while displaying the real-time video communication interface; and in response to detecting changes in the scene in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on the detected changes in the scene in the field of view of one or more cameras, depending on determining that the first camera mode is selected; and abandoning the adjustment of the representation of the field of view of one or more cameras during the real-time video communication session, depending on determining that the second camera mode is selected.
[0008] This document describes an example transient computer-readable storage medium. An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with display generating components, one or more cameras, and one or more input devices. The one or more procedures include instructions for: displaying a communication request interface via a display generation component, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object, while displaying the communication request interface; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying a real-time video communication interface for the real-time video communication session via the display generation component; detecting changes in the scene in the field of view of one or more cameras while displaying the real-time video communication interface; and in response to detecting changes in the scene in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on the detected changes in the scene in the field of view of one or more cameras, depending on determining that the first camera mode is selected; and abandoning the adjustment of the representation of the field of view of one or more cameras during the real-time video communication session, depending on determining that the second camera mode is selected.
[0009] This document describes an exemplary computer system. An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors. The one or more procedures include instructions for: displaying a communication request interface via a display generation component, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode and using a second camera mode for one or more cameras during a real-time video communication session; receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object, while displaying the communication request interface; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying a real-time video communication interface for the real-time video communication session via the display generation component; detecting changes in the scene in the field of view of one or more cameras while displaying the real-time video communication interface; and in response to detecting changes in the scene in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on the detected changes in the scene in the field of view of one or more cameras, depending on determining that the first camera mode is selected; and abandoning the adjustment of the representation of the field of view of one or more cameras during the real-time video communication session, depending on determining that the second camera mode is selected.
[0010] An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; means for displaying a communication request interface via the display generation component, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode and using a second camera mode for one or more cameras during the real-time video communication session; means for receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object when displaying the communication request interface; and means for using... A means for displaying a real-time video communication interface for a real-time video communication session via a display generation component in response to receiving one or more inputs, including a selection of a first selectable graphical user interface object; means for detecting changes in the scene in the field of view of one or more cameras while displaying the real-time video communication interface; and means for performing the following in response to detecting changes in the scene in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on the detected changes in the scene in the field of view of one or more cameras, according to determining that a first camera mode is selected; and abandoning the adjustment of the representation of the field of view of one or more cameras during the real-time video communication session, according to determining that a second camera mode is selected.
[0011] An exemplary method includes, at a computer system communicating with a display generation component, one or more cameras, and one or more input devices: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session, in addition to participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations being visually associated with visual indications of options for changing the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via the one or more input devices a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0012] An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including the simultaneous display of: representations of one or more participants in the real-time video communication session, in addition to participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations being visually associated with visual indications of options for changing the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via the one or more input devices a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0013] An exemplary transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including the simultaneous display of: representations of one or more participants in the real-time video communication session, in addition to participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations being visually associated with visual indications of options for changing the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via the one or more input devices a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0014] An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session, other than those visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations being visually associated with visual indications of options for changing the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via the one or more input devices a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0015] An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; means for displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session other than participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations being visually associated with visual indications of options for changing the representations of the fields of view of the one or more cameras during the real-time video communication session; means for detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session when displaying the real-time video communication interface for the real-time video communication session; and means for initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session in response to detecting the set of one or more inputs.
[0016] An exemplary method includes: at a computer system communicating with a display generation component and one or more cameras: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the fields of view of one or more cameras; capturing image data of the real-time video communication session via the one or more cameras while the real-time video communication session is active; determining, based on the image data of the real-time video communication session captured via the one or more cameras, that a separation criterion is met between a first participant and a second participant, and simultaneously displaying via the display generation component: a representation of a first portion of the field of view of one or more cameras at a first region of the real-time video communication interface; and a representation of a second portion of the field of view of the one or more cameras at a second region of the real-time video communication interface, distinct from the first region. A representation of a second portion of the field of view of one or more cameras in the region, wherein a representation of a first portion of the field of view of one or more cameras and a representation of a second portion of the field of view of one or more cameras are displayed, but a representation of a third portion of the field of view of one or more cameras located between the first portion of the field of view of one or more cameras and the second portion of the field of view of one or more cameras is not displayed; and based on image data from a real-time video communication session captured via one or more cameras, it is determined that the separation amount between a first participant and a second participant does not meet the separation criterion, and a representation of a fourth portion of the field of view of one or more cameras including the first participant and the second participant is displayed via a display generation component, while maintaining the display of a portion of the field of view of one or more cameras located between the first participant and the second participant.
[0017] An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the fields of view of one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and simultaneously displaying, via the display generation component, a representation of a first portion of the field of view of the one or more cameras at a first region of the real-time video communication interface, and a representation of one or more of the fields of view at a second region of the real-time video communication interface, different from the first region, based on a determination that a separation criterion is met between a first participant and a second participant using the image data of the real-time video communication session captured via the one or more cameras. A representation of the second portion of the field of view of one or more cameras, wherein a representation of the first portion of the field of view of one or more cameras and a representation of the second portion of the field of view of one or more cameras are displayed, but a representation of the third portion of the field of view of one or more cameras located between the first portion of the field of view of one or more cameras and the second portion of the field of view of one or more cameras is not displayed; and based on image data of a real-time video communication session captured via one or more cameras, it is determined that the separation amount between the first participant and the second participant does not meet the separation criterion, and a representation of the fourth portion of the field of view of one or more cameras including the first participant and the second participant is displayed via a display generation component, while maintaining the display of a portion of the field of view of one or more cameras located between the first participant and the second participant.
[0018] An exemplary transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the fields of view of one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and simultaneously displaying, via the display generation component, a representation of a first portion of the field of view of the one or more cameras at a first region of the real-time video communication interface, and a representation of one or more of the fields of view at a second region of the real-time video communication interface, different from the first region, based on a determination that a separation criterion is met between a first participant and a second participant according to the image data of the real-time video communication session captured via the one or more cameras. A representation of the second portion of the field of view of one or more cameras, wherein a representation of the first portion of the field of view of one or more cameras and a representation of the second portion of the field of view of one or more cameras are displayed, but a representation of the third portion of the field of view of one or more cameras located between the first portion of the field of view of one or more cameras and the second portion of the field of view of one or more cameras is not displayed; and based on image data of a real-time video communication session captured via one or more cameras, it is determined that the separation amount between the first participant and the second participant does not meet the separation criterion, and a representation of the fourth portion of the field of view of one or more cameras including the first participant and the second participant is displayed via a display generation component, while maintaining the display of a portion of the field of view of one or more cameras located between the first participant and the second participant.
[0019] An exemplary computer system includes: a display generation component; one or more cameras; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and simultaneously displaying, via the display generation component, a representation of a first portion of the field of view of the one or more cameras at a first region of the real-time video communication interface, and a representation of one or more of the fields of view at a second region of the real-time video communication interface, different from the first region. A representation of the second portion of the field of view of one or more cameras, wherein a representation of the first portion of the field of view of one or more cameras and a representation of the second portion of the field of view of one or more cameras are displayed, but a representation of the third portion of the field of view of one or more cameras located between the first portion of the field of view of one or more cameras and the second portion of the field of view of one or more cameras is not displayed; and based on image data of a real-time video communication session captured via one or more cameras, it is determined that the separation amount between the first participant and the second participant does not meet the separation criterion, and a representation of the fourth portion of the field of view of one or more cameras including the first participant and the second participant is displayed via a display generation component, while maintaining the display of a portion of the field of view of one or more cameras located between the first participant and the second participant.
[0020] An exemplary computer system includes: a display generation component; one or more cameras; means for displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the fields of view of the one or more cameras; means for capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and means for simultaneously displaying, via the display generation component, the following: a representation of a first portion of the field of view of the one or more cameras at a first region of the real-time video communication interface, and a representation of the field of view of the one or more cameras at a different region from the first region of the real-time video communication session, based on determining that a separation criterion is met between a first participant and a second participant. A representation of a second portion of the field of view of one or more cameras in a second region of a communication interface, wherein a representation of a first portion of the field of view of one or more cameras and a representation of a second portion of the field of view of one or more cameras are displayed, but a representation of a third portion of the field of view of one or more cameras located between the first portion of the field of view of one or more cameras and the second portion of the field of view of one or more cameras is not displayed; and means for determining, based on image data of a real-time video communication session captured via one or more cameras, that the separation amount between a first participant and a second participant does not meet a separation criterion, to display, via a display generation component, a representation of a fourth portion of the field of view of one or more cameras including the first participant and the second participant, while maintaining a display of a portion of the field of view of one or more cameras located between the first participant and the second participant.
[0021] According to some embodiments, a method is described that is executed at a computer system communicating with one or more output generation components and one or more input devices. The method includes: detecting a request to display a system interface via the one or more input devices; in response to detecting the request to display the system interface, displaying a system interface via the one or more output generation components, including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, comprising: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein these media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for the predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting a set of one or more inputs via the one or more input devices including inputs pointing to the set of one or more media communication controls; and when a corresponding media communication session has been active for a predetermined amount of time, adjusting media communication settings for the corresponding media communication session in response to detecting the set of one or more inputs including inputs pointing to the set of one or more media communication controls.
[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output generation components and one or more input devices. The one or more programs include instructions for: detecting a request to display a system interface via the one or more input devices; and, in response to detecting the request to display the system interface, displaying a system interface via the one or more output generation components, including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: determining that a media communication session has been active for a predetermined amount of time, wherein the plurality of simultaneously displayed controls include a set of one or more media communication controls, wherein these media communication controls provide access to media communication... Access to the settings, which determine how the computer system processes media during a media communication session; and, based on the determination that the media communication session is not active for a predetermined amount of time, displaying multiple simultaneously displayed controls instead of displaying the group of one or more media communication controls; when displaying a system interface having the group of one or more media communication controls, detecting a group of one or more inputs via one or more input devices, including inputs pointing to the group of one or more media communication controls; and when the corresponding media communication session is active for a predetermined amount of time, adjusting the media communication settings for the corresponding media communication session in response to detecting the group of one or more inputs including inputs pointing to the group of one or more media communication controls.
[0023] According to some embodiments, a transient computer-readable storage medium is described. This transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more output generation components and one or more input devices. The one or more programs include instructions for: detecting a request to display a system interface via the one or more input devices; and, in response to detecting the request to display the system interface, displaying a system interface via the one or more output generation components, including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: determining that a media communication session has been active for a predetermined amount of time, wherein the plurality of simultaneously displayed controls include a set of one or more media communication controls, wherein these media communication controls provide access to media communication... Access to the settings, which determine how the computer system processes media during a media communication session; and, based on the determination that the media communication session is not active for a predetermined amount of time, displaying multiple simultaneously displayed controls instead of displaying the group of one or more media communication controls; when displaying a system interface having the group of one or more media communication controls, detecting a group of one or more inputs via one or more input devices, including inputs pointing to the group of one or more media communication controls; and when the corresponding media communication session is active for a predetermined amount of time, adjusting the media communication settings for the corresponding media communication session in response to detecting the group of one or more inputs including inputs pointing to the group of one or more media communication controls.
[0024] According to some embodiments, a computer system is described. The computer system includes: one or more output generation components; one or more input devices; one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: detecting a request to display a system interface via the one or more input devices; and, in response to detecting the request to display the system interface, displaying a system interface via the one or more output generation components, including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, including: determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein these media communication controls provide access to... Access to media communication settings that determine how the computer system processes media during a media communication session; and, based on the determination that the media communication session is not active for a predetermined amount of time, displaying multiple simultaneously displayed controls instead of displaying the group of one or more media communication controls; when displaying a system interface having the group of one or more media communication controls, detecting a group of one or more inputs, including inputs pointing to the group of one or more media communication controls, via one or more input devices; and, when the corresponding media communication session is active for a predetermined amount of time, adjusting the media communication settings for the corresponding media communication session in response to detecting the group of one or more inputs, including inputs pointing to the group of one or more media communication controls.
[0025] A computer system is described based on some implementation schemes. The computer system includes: one or more output generation components; one or more input components; means for detecting a request to display a system interface via the one or more input devices; means for displaying a system interface including multiple simultaneously displayed controls for controlling different system functions of the computer system via the one or more output generation components in response to detecting the request to display the system interface, including: determining that a media communication session has been active for a predetermined amount of time, the multiple simultaneously displayed controls including a set of one or more media communication controls, wherein these media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and determining that the media communication session has not been active for a predetermined amount of time, displaying the multiple simultaneously displayed controls without displaying the set of one or more media communication controls; means for detecting a set of one or more inputs including inputs pointing to the set of one or more media communication controls via the one or more input devices when displaying a system interface having the set of one or more media communication controls; and means for adjusting media communication settings for a corresponding media communication session in response to detecting the set of one or more inputs including inputs pointing to the set of one or more media communication controls when the corresponding media communication session has been active for a predetermined amount of time.
[0026] According to some embodiments, a method is described that is performed at a computer system communicating with one or more output generation components, one or more cameras, and one or more input devices. The method includes: displaying a real-time video communication interface for a real-time video communication session via the one or more output generation components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with indications of options for initiating a process for changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the representation of the field of view of the one or more cameras, during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and, while displaying the real-time video communication interface for the real-time video communication session, detecting via one or more input devices a set of one or more inputs corresponding to a request to change the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the representation of the field of view of the one or more cameras; and, in response to detecting the set of one or more inputs, changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the representation of the field of view of the one or more cameras.
[0027] According to some implementation schemes, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs executed by one or more processors of a computer system communicating with one or more output generation components, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the one or more output generation components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with indications of options for initiating a process for changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras, during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the field of view representation of the one or more cameras of the computer system; and, when displaying the real-time video communication interface for the real-time video communication session, detecting via one or more input devices a set of one or more inputs corresponding to a request to change the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras; and, in response to detecting the set of one or more inputs, changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras.
[0028] According to some implementation schemes, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs executed by one or more processors of a computer system communicating with one or more output generating components, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the one or more output generating components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with indications of options for initiating a process for changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras, during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, which differs from the representation of the field of view of the one or more cameras of the computer system; and, when displaying the real-time video communication interface for the real-time video communication session, detecting via one or more input devices a set of one or more inputs corresponding to a request to change the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras; and, in response to detecting the set of one or more inputs, changing the appearance of a portion of the field of view representation of one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras.
[0029] According to some embodiments, a computer system is described. The computer system includes: one or more output generation components; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for displaying a real-time video communication interface for a real-time video communication session via the one or more output generation components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with indications of options for initiating a process for changing, in addition to the field of view displayed on the one or more cameras, during the real-time video communication session. The appearance of a portion of the field of view of one or more cameras other than objects displayed in the representation; and the representation of one or more participants in a real-time video communication session, which differs from the representation of the field of view of one or more cameras of a computer system; and, when displaying a real-time video communication interface for a real-time video communication session, detecting and changing a request corresponding to a request via one or more input devices to change the appearance of a portion of the field of view of one or more cameras other than objects displayed in the representation of the field of view of one or more cameras; and, in response to detecting the one or more inputs, changing the appearance of a portion of the field of view of one or more cameras other than objects displayed in the representation of the field of view of one or more cameras.
[0030] According to some embodiments, a computer system is described. The computer system includes: one or more output generation components; one or more cameras; one or more input devices; and means for displaying a real-time video communication interface for a real-time video communication session via the one or more output generation components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of the field of view of the one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with indications of options for initiating a process for changing the appearance of a portion of the field of view representation of the one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras, during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; means for detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change the appearance of a portion of the field of view representation of the one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras, while displaying the real-time video communication interface for the real-time video communication session; and means for changing the appearance of a portion of the field of view representation of the one or more cameras, excluding objects displayed in the field of view representation of the one or more cameras, in response to detecting the set of one or more inputs.
[0031] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0032] Therefore, faster and more efficient methods and interfaces are provided for managing real-time video communication sessions, thereby improving the effectiveness, efficiency, and user satisfaction of such devices. These methods and interfaces can complement or replace other methods used for managing real-time video communication sessions. Attached Figure Description
[0033] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.
[0034] Figure 1A This is a block diagram illustrating a portable multi-functional device with a touch-sensitive display according to some embodiments.
[0035] Figure 1B This is a block diagram illustrating exemplary components for event handling according to some implementation schemes.
[0036] Figure 2 A portable multi-functional device with a touchscreen is shown according to some embodiments.
[0037] Figure 3 This is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface according to some implementation schemes.
[0038] Figure 4A An exemplary user interface for a menu on an application on a portable multifunction device, according to some implementation schemes, is shown.
[0039] Figure 4B An exemplary user interface for a multifunctional device having a touch-sensitive surface separate from the display is shown according to some embodiments.
[0040] Figure 5A A personal electronic device according to some implementation schemes is shown.
[0041] Figure 5B This is a block diagram illustrating a personal electronic device according to some implementation schemes.
[0042] Figure 5C An exemplary diagram of a communication session between electronic devices according to some implementation schemes is shown.
[0043] Figures 6A to 6Q An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0044] Figures 7A to 7B A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is depicted.
[0045] Figures 8A to 8R An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0046] Figure 9 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes.
[0047] Figures 10A to 10J An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0048] Figure 11 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes.
[0049] Figures 12A to 12U An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0050] Figure 13This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes.
[0051] Figure 14 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Detailed Implementation
[0052] The following description illustrates exemplary methods, parameters, etc. However, it should be understood that such description is not intended to limit the scope of this disclosure, but is provided as a description of exemplary embodiments.
[0053] There is a need for electronic devices that provide efficient methods and interfaces for managing real-time video communication sessions. Such technologies can reduce the cognitive burden on users participating in video communication sessions, thereby increasing productivity. Furthermore, these technologies can reduce processor and battery power that would otherwise be wasted on redundant user input.
[0054] under, Figures 1A to 1B , Figure 2 , Figure 3 , Figures 4A to 4B and Figures 5A to 5C A description of an exemplary device for performing techniques for managing real-time video communication sessions is provided. Figures 6A to 6Q An exemplary user interface for managing real-time video communication sessions is shown. Figures 7A to 7B A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is depicted. Figures 6A to 6Q The user interface in the document is used to illustrate the process described below, including Figures 7A to 7B The process in. Figures 8A to 8R An exemplary user interface for managing real-time video communication sessions is shown. Figure 9 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Figures 8A to 8R The user interface is used to illustrate the processes described below, including... Figure 9 The process in. Figures 10A to 10J An exemplary user interface for managing real-time video communication sessions is shown. Figure 11 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Figures 10A-10J The user interface in the document is used to display including Figure 11 The process described below is the process in the middle. Figures 12A to 12U An exemplary user interface for managing real-time video communication sessions is shown. Figure 13 and Figure 14 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Figures 12A to 12U The user interface in the document is used to display including Figure 13 and Figure 14 The process described below is the process in the middle.
[0055] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if a condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0056] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch may be named a second touch and similarly, a second touch may be named a first touch, without departing from the scope of the various described embodiments. In some embodiments, a first touch and a second touch are two separate references to the same touch. In some embodiments, both a first touch and a second touch are touches, but they are not the same touch.
[0057] The terminology used in the description of the various embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “includes”, “including”, “comprises”, and / or “comprising”, when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0058] Depending on the context, the term "if" may optionally be interpreted as meaning "when," "at," or "in response to determination" or "in response to detection." Similarly, depending on the context, the phrases "if determination..." or "if detection [the stated condition or event]" may optionally be interpreted as meaning "in response to determination..." or "in response to detection [the stated condition or event]."
[0059] This document describes implementations of electronic devices, user interfaces for such devices, and associated processes for using such devices. In some implementations, the device is a portable communication device, such as a mobile phone, that also includes other functionalities such as PDA and / or music player functionality. Exemplary implementations of portable multi-functional devices include, but are not limited to, those from Apple Inc. (Cupertino, California). Devices, iPod Equipment, and Device. Optionally, other portable electronic devices may be used, such as laptops or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). In some embodiments, the electronic device is a computer system that communicates with a display generating component (e.g., via wireless or wired communication). The display generating component is configured to provide visual output, such as a display via a CRT monitor, a display via an LED monitor, or a display via image projection. In some embodiments, the display generating component is integrated with the computer system. In some embodiments, the display generating component is separate from the computer system. As used herein, “display” content includes displaying content (e.g., video data rendered or decoded by display controller 156) by transmitting data (e.g., image data or video data) to an integrated or external display generating component via a wired or wireless connection to visually generate content.
[0060] In the following discussion, an electronic device comprising a display and a touch-sensitive surface is described. However, it should be understood that the electronic device may optionally include one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick.
[0061] The device typically supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, phone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camcorder applications, web browsing applications, digital music player applications, and / or digital video player applications.
[0062] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or adjusted and / or varied within the respective applications. In this way, the device's common physical architecture (such as the touch-sensitive surface) optionally utilizes a user interface that is intuitive and clear to the user to support various applications.
[0063] Now let’s turn our attention to implementation schemes for portable devices with touch-sensitive displays. Figure 1AThis is a block diagram illustrating a portable multi-functional device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 is sometimes referred to as a “touchscreen” for convenience, and is sometimes referred to as or called a “touch-sensitive display system.” Device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more haptic output generators 167 for generating haptic output on device 100 (e.g., generating haptic output on a touch-sensitive surface such as the touch-sensitive display system 112 of device 100 or the touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0064] As used in this specification and claims, the term "intensity" of contact on a tactile surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on a tactile surface, or to a substitute (alternative) for the force or pressure of a contact on a tactile surface. The intensity of contact has a range of values that includes at least four different values and more typically hundreds of different values (e.g., at least 256). The intensity of contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the tactile surface are optionally used to measure the force at different points on the tactile surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the tactile surface. Alternatively, the size and / or variation of the contact area detected on the touch-sensitive surface, the capacitance and / or variation of the touch-sensitive surface near the contact, and / or the resistance and / or variation of the touch-sensitive surface near the contact may optionally be used as substitutes for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether an intensity threshold (e.g., the intensity threshold is described in units corresponding to the substitute measurement) has been exceeded. In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure) has been exceeded. Using the intensity of the contact as an attribute of user input allows the user to access additional device functions that would otherwise be inaccessible to the user on a smaller device with limited physical space, which is used (e.g., on a touch-sensitive display) to display an indication and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).
[0065] As used in this specification and claims, the term "haptic output" refers to the physical displacement of a device relative to a previous part of the device, the physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or the displacement of a component relative to the center of mass of the device, which is detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with a touch-sensitive surface that has been physically pressed (e.g., displaced) by the user's movement does not move. For example, even when the smoothness of the tactile surface remains unchanged, the movement of the tactile surface can optionally be interpreted or sensed by the user as the "roughness" of the tactile surface. While such interpretations of touch by users will be limited by the individualized sensory perceptions of the user, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of a user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception of a typical (or ordinary) user.
[0066] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both, including one or more signal processing and / or application-specific integrated circuits.
[0067] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls other components of device 100 to access memory 102.
[0068] Peripheral interface 118 can be used to couple the device's input and output peripherals to CPU 120 and memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0069] RF (Radio Frequency) circuit 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 108 converts electrical signals into electromagnetic signals and vice versa, and communicates with communication networks and other communication devices via these electromagnetic signals. RF circuit 108 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 108 optionally communicates wirelessly with networks and other devices, such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). RF circuit 108 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via near-field communication radio components. Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution, Pure Data (EV-DO), HSPA, HSPA+, Dual-Unit HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE...). 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence with Extended Utility (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol that has not been developed as of the date of this document submission.
[0070] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and device 100. Audio circuitry 110 receives audio data from peripheral interface 118, converts the audio data into electrical signals, and transmits the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuitry 110 converts the electrical signals into audio data and transmits the audio data to peripheral interface 118 for processing. Audio data is optionally retrieved by peripheral interface 118 from and / or transmitted to memory 102 and / or RF circuitry 108. In some embodiments, audio circuitry 110 also includes a headset jack (e.g., ...). Figure 2 (212 in the text). The headset jack provides an interface between the audio circuitry 110 and a removable audio input / output peripheral device, such as an output-only headphone or a headset with both output (e.g., a single-ear or dual-ear headphone) and input (e.g., a microphone).
[0071] I / O subsystem 106 couples input / output peripherals on device 100, such as touchscreen 112 and other input control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, depth camera controller 169, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / send electrical signals to the other input control device 116. The other input control device 116 optionally includes physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click dials, etc. In some embodiments, input controller 160 is optionally coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, and pointing device such as mouse. One or more buttons (e.g., Figure 2 Optionally, 208) includes an increase / decrease button for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include a push-button (e.g., Figure 2(Ref. 206 in the original text). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication or via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some implementations, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes a hand moving a predetermined amount and / or speed in a predetermined posture, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0072] A quick press of the push button optionally disengages the touchscreen 112 from its lock or optionally initiates a process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application 11 / 322,549 (i.e., U.S. Patent No. 7,657,849), filed December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. A long press of the push button (e.g., 206) optionally powers the device 100 on or off. The function of one or more buttons is optionally user-customizable. The touchscreen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0073] The touch-sensitive display 112 provides input and output interfaces between the device and the user. The display controller 156 receives electrical signals from and / or sends electrical signals to the touchscreen 112. The touchscreen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0074] Touchscreen 112 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or haptic contact. Touchscreen 112 and display controller 156 (along with any associated modules and / or instruction set in memory 102) detect contact on touchscreen 112 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touchscreen 112. In an exemplary embodiment, the contact point between touchscreen 112 and the user corresponds to the user's finger.
[0075] Touchscreen 112 optionally employs LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies are used in other embodiments. Touchscreen 112 and display controller 156 optionally employ any of a variety of touch sensing technologies now known or to be developed hereafter, along with other proximity sensor arrays or other elements for determining one or more points of contact with touchscreen 112, to detect contact and any movement or interruption thereof. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that from Apple Inc. (Cupertino, California). and iPod The technology used.
[0076] In some embodiments of the touchscreen 112, the touch-sensitive display optionally resembles a multi-touch touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is incorporated herein by reference in its entirety. However, the touchscreen 112 displays visual output from the device 100, while the touch-sensitive touchpad does not provide visual output.
[0077] The touch-sensitive display in some embodiments of the touchscreen 112 is described in the following applications: (1) U.S. Patent Application 11 / 381,313, filed May 2, 2006, “Multipoint Touch Surface Controller”; (2) U.S. Patent Application 10 / 840,862, filed May 6, 2004, “Multipoint Touchscreen”; (3) U.S. Patent Application 10 / 903,964, filed July 30, 2004, “Gestures For Touch Sensitive Input Devices”; (4) U.S. Patent Application 11 / 048,264, filed January 31, 2005, “Gestures For Touch Sensitive Input Devices”; and (5) U.S. Patent Application 11 / 038,590, filed January 18, 2005, “Mode-Based Graphical User Interfaces For Touch Sensitive Input”. (6) U.S. Patent Application No. 11 / 228,758, filed September 16, 2005, “Virtual Input Device Placement On A Touch Screen User Interface”; (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”; and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, “Multi-Functional Hand-Held Device”. The full text of all these applications is incorporated herein by reference.
[0078] Touchscreen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. Users optionally use any suitable object or accessory such as a stylus, finger, etc., to interact with touchscreen 112. In some embodiments, the user interface is designed to operate primarily through finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor locations or commands to perform the user-desired actions.
[0079] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. Optionally, the touchpad is a touch-sensitive surface separate from the touchscreen 112, or an extension of the touch-sensitive surface formed by the touchscreen.
[0080] The device 100 also includes a power system 162 for supplying power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.
[0081] The device 100 optionally also includes one or more optical sensors 164. Figure 1AAn optical sensor 164 is shown coupled to an optical sensor controller 158 in the I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with an imaging module 143 (also called a camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite to a touchscreen display 112 on the front of the device, allowing the touchscreen display to be used as a viewfinder for still image and / or video image acquisition. In some embodiments, the optical sensor is located on the front of the device, allowing images of the user to be optionally acquired for video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the location of the optical sensor 164 can be changed by the user (e.g., by rotating the lenses and sensor within the device housing), allowing a single optical sensor 164 to be used with the touchscreen display for both video conferencing and still image and / or video image acquisition.
[0082] The device 100 optionally also includes one or more depth camera sensors 175. Figure 1A A depth camera sensor is shown coupled to a depth camera controller 169 in I / O subsystem 106. Depth camera sensor 175 receives data from the environment to create a 3D model of an object (e.g., a face) within the scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with imaging module 143 (also referred to as a camera module), depth camera sensor 175 may optionally be used to determine depth maps of different portions of an image captured by imaging module 143. In some embodiments, the depth camera sensor is located at the front of device 100, such that user images with depth information are optionally acquired for video conferencing while the user views other video conferencing participants on a touchscreen display, and selfies with depth map data are captured. In some embodiments, depth camera sensor 175 is located at the rear of the device, or both the rear and front of device 100. In some embodiments, the position of depth camera sensor 175 may be changed by the user (e.g., by rotating a lens and sensor within the device housing), such that depth camera sensor 175 is used in conjunction with a touchscreen display for both video conferencing and still image and / or video image acquisition.
[0083] In some implementations, the depth map (e.g., a depth map image) contains information (e.g., values) relating to the distance of objects in the scene from the viewpoint (e.g., a camera, optical sensor, depth camera sensor). In one implementation of the depth map, each depth pixel defines the location of its corresponding two-dimensional pixel on the Z-axis of the viewpoint. In some implementations, the depth map consists of pixels, where each pixel is defined by a value (e.g., 0 to 255). For example, a "0" value represents the pixel furthest from the viewpoint (e.g., a camera, optical sensor, depth camera sensor) in the "3D" scene, and a "255" value represents the pixel closest to the viewpoint in the "3D" scene. In other implementations, the depth map represents the distance between objects in the scene and the plane of the viewpoint. In some implementations, the depth map includes information about the relative depth of various features of the object of interest within the field of view of the depth camera (e.g., the relative depth of the eyes, nose, mouth, and ears of a user's face). In some implementations, the depth map includes information that enables the device to determine the contour of the object of interest in the z-direction.
[0084] The device 100 may optionally also include one or more contact strength sensors 165. Figure 1A A contact strength sensor is shown coupled to a strength sensor controller 159 in I / O subsystem 106. The contact strength sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact strength sensor 165 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact strength sensor is located on the rear of device 100, opposite to the touchscreen display 112 located on the front of device 100.
[0085] The device 100 optionally also includes one or more proximity sensors 166. Figure 1AA proximity sensor 166 coupled to a peripheral device interface 118 is shown. Alternatively, the proximity sensor 166 may optionally be coupled to an input controller 160 in an I / O subsystem 106. The proximity sensor 166 may optionally perform as described in the following U.S. patent applications: No. 11 / 241,839, entitled "Proximity Detector In Handheld Device"; No. 11 / 240,788, entitled "Proximity Detector In Handheld Device"; No. 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; No. 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and No. 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals", the entire contents of which are incorporated herein by reference. In some implementations, the proximity sensor is turned off and the touchscreen 112 is disabled when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).
[0086] The device 100 may optionally also include one or more tactile output generators 167. Figure 1AA haptic output generator coupled to a haptic feedback controller 161 in I / O subsystem 106 is shown. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting electrical signals into haptic outputs on the device). A contact intensity sensor 165 receives haptic feedback generation instructions from a haptic feedback module 133 and generates a haptic output on device 100 that can be felt by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a haptic surface (e.g., haptic display system 112) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 100) or laterally (e.g., backward and forward in the same plane as the surface of device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.
[0087] The device 100 may optionally also include one or more accelerometers 168. Figure 1A An accelerometer 168 coupled to a peripheral device interface 118 is shown. Alternatively, the accelerometer 168 may be coupled to an input controller 160 in an I / O subsystem 106. The accelerometer 168 may optionally perform as described in the following U.S. patent publications: U.S. Patent Publication No. 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication No. 20060017692, entitled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from one or more accelerometers. Device 100 may optionally include, in addition to the accelerometer 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for acquiring information about the location and orientation (e.g., portrait or landscape) of device 100.
[0088] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application program (or instruction set) 136. Furthermore, in some embodiments, memory 102 ( Figure 1A ) or 370 ( Figure 3 Storage device / global internal state 157, such as Figure 1A and Figure 3 As shown in the figure. Device / global internal state 157 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, indicating what applications, views or other information occupy various areas of the touch screen display 112; sensor state, including information obtained from the device's various sensors and input control devices 116; and position information relating to the device's position and / or orientation.
[0089] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.
[0090] The communication module 128 facilitates communication with other devices via one or more external ports 124 and includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is connected to… (Trademark of Apple Inc.) The same or similar and / or compatible multi-pin (e.g., 30-pin) connectors used in Apple Inc. devices.
[0091] The contact / motion module 130 optionally detects contact with the touchscreen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., touchpads or physical click-based rotary dials). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether a contact has occurred (e.g., detecting a finger press event), determining the contact intensity (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or a contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple-finger contact). In some implementations, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.
[0092] In some implementations, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by a user (e.g., determining whether the user has “clicked” an icon). In some implementations, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a specific physical actuator and can be adjusted without changing the physical hardware of device 100). For example, the mouse “click” threshold of a touchpad or touchscreen can be set to any threshold in a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some specific implementations, the user of the device is provided with software settings for adjusting one or more intensity thresholds in a set (e.g., by adjusting the individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using system-level clicks on the “intensity” parameter).
[0093] The touch / motion module 130 optionally detects gesture input performed by the user. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event.
[0094] The graphics module 132 includes various known software components for rendering and displaying graphics on the touchscreen 112 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0095] In some implementations, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes from an application or the like to specify the graphic to be displayed, and, if necessary, also receives coordinate data and other graphic attribute data, and then generates screen image data for output to the display controller 156.
[0096] The haptic feedback module 133 includes various software components for generating instructions which are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.
[0097] Optionally, the text input module 134, a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).
[0098] GPS module 135 determines the location of the device and provides that information for use in various applications (e.g., to phone 138 for use in location-based dialing; to camera 143 as image / video metadata; and to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).
[0099] Application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:
[0100] ●Contacts module 137 (sometimes called address book or contact list);
[0101] ● Telephone module 138;
[0102] ●Video conferencing module 139;
[0103] ●Email client module 140;
[0104] ●Instant Messaging (IM) module 141;
[0105] ● Fitness support module 142;
[0106] ● Camera module 143 for still images and / or video images;
[0107] ●Image Management Module 144;
[0108] ●Video player module;
[0109] ●Music player module;
[0110] ● Browser module 147;
[0111] ● Calendar module 148;
[0112] ● Widget module 149, which optionally includes one or more of the following: weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets acquired by the user, and user-created widgets 149-6.
[0113] ● Wrapper module 150 for creating user-created widgets 149-6;
[0114] ●Search module 151;
[0115] ● Video and music player module 152, which combines a video player module and a music player module;
[0116] ●Notes module 153;
[0117] ●Map module 154; and / or
[0118] ● Online video module 155.
[0119] Examples of other applications 136 that may be optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital rights management, speech recognition, and speech duplication.
[0120] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the contact module 137 is optionally used to manage an address book or contact list (e.g., in the application internal state 192 of the contact module 137 stored in memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0121] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, telephone module 138 is optionally used to input character sequences corresponding to telephone numbers, access one or more telephone numbers in contact module 137, modify entered telephone numbers, dial corresponding telephone numbers, initiate conversations, and disconnect or hang up when a conversation is completed. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0122] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contact module 137, and telephone module 138, video conferencing module 139 includes executable instructions to initiate, conduct, and terminate video conferences between the user and one or more other participants based on user instructions.
[0123] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands. Combined with image management module 144, email client module 140 makes it very easy to create and send emails containing still images or video images captured by camera module 143.
[0124] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for: inputting a character sequence corresponding to an instant message, modifying previously input characters, transmitting a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving an instant message, and viewing a received instant message. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photographs, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Services (EMS). As used herein, "instant message" refers to both telephone-based messages (e.g., messages sent using SMS or MMS) and internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0125] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, fitness support module 142 includes executable instructions for creating fitness activities (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (executive devices); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness activities; and displaying, storing, and transmitting fitness data.
[0126] In conjunction with the touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, the camera module 143 includes executable instructions for: capturing still images or videos (including video streams) and storing them in memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from memory 102.
[0127] Incorporating the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0128] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as links to attachments and other files on web pages.
[0129] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions to create, display, modify, and store calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.
[0130] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is optionally a microapplication downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or a user-created microapplication (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).
[0131] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 can optionally be used by the user to create widgets (e.g., to convert user-specified portions of a webpage into widgets).
[0132] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.
[0133] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions allowing users to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0134] Incorporating the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the note-taking module 153 includes executable instructions for creating and managing notes, to-do items, etc., according to user instructions.
[0135] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and map-related data (e.g., driving directions, data related to shops and other points of interest at or near a specific location, and other location-based data) according to user instructions.
[0136] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, the online video module 155 includes instructions for performing the following actions: allowing users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen or on an external display connected via external port 124), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 141 is used instead of the email client module 140 to send links to specific online videos. Further descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are incorporated herein by reference in their entirety.
[0137] Each of the above modules and applications corresponds to an executable set of instructions for performing one or more of the functions described above and the methods described in this patent application (e.g., computer-implemented methods and other information processing methods as described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module (e.g., Figure 1A (e.g., video and music player module 152). In some embodiments, memory 102 optionally stores subgroups of the aforementioned modules and data structures. Additionally, memory 102 optionally stores other modules and data structures not described above.
[0138] In some implementations, device 100 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push-buttons, dials, etc.) on device 100 can be optionally reduced.
[0139] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 from any user interface displayed on device 100 to the main menu, home menu, or root menu. In such embodiments, a "menu button" is implemented using a touchpad. In some other embodiments, the menu button is a physical push-button or other physical input control device, rather than a touchpad.
[0140] Figure 1B This is a block diagram illustrating exemplary components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 This includes an event classifier 170 (e.g., in operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 137 to 151, 155, 380 to 390).
[0141] Event classifier 170 receives event information and determines the application 136-1 and application view 191 of application 136-1 to which the event information should be delivered. Event classifier 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information should be delivered.
[0142] In some implementations, the application internal state 192 includes additional information such as one or more of the following: recovery information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 136-1, a state queue for enabling the user to return to the previous state or view of the application 136-1, and a repeat / undo queue for the user's previous actions.
[0143] Event monitor 171 receives event information from peripheral device interface 118. The event information includes information about sub-events (e.g., user touches on touch-sensitive display 112 as part of a multi-touch gesture). Peripheral device interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via audio circuitry 110). The information received by peripheral device interface 118 from I / O subsystem 106 includes information from touch-sensitive display 112 or touch-sensitive surfaces.
[0144] In some implementations, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other implementations, peripheral device interface 118 transmits event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or receiving input for a predetermined duration).
[0145] In some implementations, the event classifier 170 also includes a hit view determination module 172 and / or an activity event recognizer determination module 173.
[0146] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.
[0147] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a procedural level within the application's procedural or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events identified as correct input is optionally determined at least in part based on the hit view of the initial touch that initiates a touch-based gesture.
[0148] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.
[0149] The activity event recognizer determination module 173 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 173 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 173 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, higher views in the hierarchy will still remain actively participating views.
[0150] Event assigner module 174 assigns event information to event identifiers (e.g., event identifier 180). In embodiments that include active event identifier determination module 173, event assigner module 174 delivers event information to the event identifier determined by active event identifier determination module 173. In some embodiments, event assigner module 174 stores event information in an event queue, which is retrieved by the corresponding event receiver 182.
[0151] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or part of another module (such as contact / motion module 130) stored in memory 102.
[0152] In some implementations, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other implementations, one or more of the event recognizers 180 are part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some implementations, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handlers 190 optionally utilize or invoke the data updater 176, the object updater 177, or the GUI updater 178 to update the application's internal state 192. Alternatively, one or more application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0153] The corresponding event recognizer 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies the event based on the event information. The event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event recognizer 180 also includes at least one subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0154] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events, such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a longitudinal orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device orientation).
[0155] Event comparator 184 compares event information with predefined event or sub-event definitions and determines the event or sub-event based on the comparison, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (187-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off of a predetermined duration (touch end). In another example, event 2 (187-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) on the displayed object for a predetermined duration, movement of the touch on the touch-sensitive display 112, and lifting off the touch (end of touch). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0156] In some implementations, event definition 187 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub-event and the object that triggered the hit test.
[0157] In some implementations, the definition of the corresponding event (187) also includes a delay action that delays the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.
[0158] When the corresponding event recognizer 180 determines that the sub-event sequence does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.
[0159] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing how or how event recognizers can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0160] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some implementations, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and delaying) the sub-event to the corresponding hit view. In some implementations, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with that flag retrieves the flag and executes a predefined process.
[0161] In some implementations, event delivery instruction 188 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and executes a predetermined process.
[0162] In some implementations, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video player module. In some implementations, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates parts of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on touch-sensitive display.
[0163] In some implementations, event handler 190 includes, or has access to, a data updater 176, an object updater 177, and a GUI updater 178. In some implementations, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other implementations, they are included in two or more software modules.
[0164] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 100 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.
[0165] Figure 2A portable multifunction device 100 with a touchscreen 112 is shown according to some embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 100. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.
[0166] Device 100 optionally also includes one or more physical buttons, such as a "home" or menu button 204. As previously described, menu button 204 is optionally used to navigate to any application 136 of a set of applications optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touchscreen 112.
[0167] In some embodiments, device 100 includes a touchscreen 112, a menu button 204, a push-button 206 for powering on / off and locking the device, one or more volume control buttons 208, a SIM card slot 210, a headset jack 212, and a docking / charging external port 124. The push-button 206 is optionally used to power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In another embodiment, device 100 also accepts voice input via microphone 113 for activating or deactivating certain functions. Device 100 also optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on the touchscreen 112, and / or one or more haptic output generators 167 for generating haptic outputs for a user of device 100.
[0168] Figure 3This is a block diagram of an exemplary multi-functional device with a display and a touch-sensitive surface according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 300 includes an input / output (I / O) interface 330 with a display 340, which is typically a touchscreen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, and a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to the reference above). Figure 1A The tactile output generator 167 and sensor 359 (e.g., optical sensor, accelerometer, proximity sensor, touch sensor and / or contact intensity sensor (similar to the one mentioned above)) are described. Figure 1A The contact strength sensor 165 is described above. The memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. The memory 370 optionally includes one or more storage devices located remotely from the CPU 310. In some embodiments, the memory 370 stores information related to the portable multifunction device 100. Figure 1A The memory 370 stores programs, modules, and data structures similar to those in the memory 102 of the portable multifunction device 100, or subsets thereof. Additionally, the memory 370 optionally stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunction device 100. For example, the memory 370 of the device 300 optionally stores a drawing module 380, a rendering module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while the portable multifunction device 100 (… Figure 1A The memory 102 may optionally not store these modules.
[0169] Figure 3Each of the elements described above is optionally stored in one or more memory devices of the previously mentioned memory devices. Each of the modules described above corresponds to a set of instructions for performing the functions described above. The modules or computer programs described above (e.g., instruction sets or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subgroup of the modules and data structures described above. In addition, memory 370 optionally stores additional modules and data structures not described above.
[0170] Now let’s turn our attention to the implementation of the user interface, which is optionally implemented on, for example, a portable multifunction device 100.
[0171] Figure 4A An exemplary user interface for an application menu on a portable multifunction device 100 according to some embodiments is shown. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements or a subset or superset thereof:
[0172] ●Signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0173] ●Time 404;
[0174] ●Bluetooth indicator 405;
[0175] ● Battery status indicator 406;
[0176] ● Tray 408 features icons for commonly used applications, such as:
[0177] The telephone module 138 has an icon 416 labeled "telephone", which optionally includes an indicator 414 indicating the number of missed calls or voicemail messages.
[0178] The email client module 140 has an icon 418 labeled "Mail" which optionally includes an indicator 410 for the number of unread emails.
[0179] The browser module 147 has an icon 420 labeled "Browser"; and
[0180] o Video and music player module 152 (also known as iPod module 152, a trademark of Apple Inc.) with icon 422 labeled "iPod"; and
[0181] ● Icons of other applications, such as:
[0182] o IM module 141's icon 424 labeled "Message";
[0183] The calendar module 148 has an icon 426 labeled "Calendar";
[0184] The image management module 144 has an icon 428 labeled "Photo".
[0185] The icon 430 of the camera module 143 is labeled "camera";
[0186] o The icon 432 of the online video module 155, which is labeled "Online Video";
[0187] The icon 434 labeled "Stock Market" in the stock market widget 149-2;
[0188] The icon 436 labeled "map" in the map module 154;
[0189] The weather widget 149-1 has icon 438 labeled "weather";
[0190] The alarm clock widget 149-4 has an icon 440 labeled "clock";
[0191] The icon 442 of the fitness support module 142 is labeled "fitness support";
[0192] The icon 444 labeled "Notes" in the o-notes module 153; and
[0193] o A settings icon 446 labeled "Settings" is used to set up applications or modules, providing access to settings for device 100 and its various applications 136.
[0194] It should be pointed out that, Figure 4A The icon labels shown are merely exemplary. For example, icon 422 of video and music player module 152 is labeled "Music" or "Music Player". Other labels may be optionally used for various application icons. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to that particular application icon.
[0195] Figure 4B A touch-sensitive surface 451 (e.g., separate from the display 450 (e.g., touchscreen display 112)) is shown. Figure 3 Devices such as tablets or touchpads (e.g., 355) Figure 3An exemplary user interface on the device 300. The device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of the sensors 359) for detecting the intensity of contact on the tactile surface 451 and / or one or more tactile output generators 357 for generating tactile outputs for the user of the device 300.
[0196] While some examples of input on a reference touchscreen display 112 (which combines a touch-sensitive surface and a display) are given below, in some implementations the device detects input on a touch-sensitive surface separate from the display, such as... Figure 4B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a spindle (e.g., on the display (e.g., 450)). Figure 4B The spindle corresponding to 453 in the middle (e.g., Figure 4B (452 in the example). According to these embodiments, the device detects the position corresponding to the corresponding position on the display (e.g., in the example). Figure 4B In the middle, 460 corresponds to 468 and 462 corresponds to 470) is in contact with the touch-sensitive surface 451 (e.g., Figure 4B (460 and 462 in the text). Thus, when the touch-sensitive surface (e.g., ...) Figure 4B 451) and the display of a multi-functional device (e.g., Figure 4B When 450 is separated from the touch-sensitive surface, user input detected by the device on that touch-sensitive surface (e.g., touches 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may optionally be used for other user interfaces described herein.
[0197] Additionally, while the examples below are primarily given with reference to finger input (e.g., finger touch, single-finger tap, finger swipe), it should be understood that in some implementations, one or more of these finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the swipe path (e.g., instead of movement of the touch). Similarly, a tap gesture may optionally be replaced by a mouse click while the cursor is over the location of the tap gesture (e.g., instead of detection of touch, followed by cessation of touch detection). Likewise, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.
[0198] Figure 5AAn exemplary personal electronic device 500 is illustrated. Device 500 includes a body 502. In some embodiments, device 500 may include components relative to devices 100 and 300 (e.g., Figures 1A to 4B Some or all of the features described herein. In some embodiments, device 500 has a touch-sensitive display 504, referred to below as touchscreen 504. As an alternative to or complement to touchscreen 504, device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, touchscreen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., touch). The one or more intensity sensors of touchscreen 504 (or touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of device 500 can respond to touches based on the intensity of the touch, meaning that touches of different intensities can invoke different user interface operations on device 500.
[0199] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is incorporated herein by reference in its entirety.
[0200] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) may be physical. Examples of physical input mechanisms include push-buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch straps, bangles, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow a user to wear device 500.
[0201] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include a reference retrieval system. Figure 1A , Figure 1B and Figure 3 Some or all of the components described herein. Device 500 has a bus 512 that operatively couples I / O portion 514 to one or more computer processors 516 and memory 518. I / O portion 514 may be connected to display 504, which may have touch-sensitive components 522 and optionally have an intensity sensor 524 (e.g., a contact intensity sensor). Furthermore, I / O portion 514 may be connected to communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 may optionally be a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 508 may optionally be a button.
[0202] In some examples, the input mechanism 508 is optionally a microphone. The personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.
[0203] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform, for example, the techniques described below, including processes 700, 900, 1100, 1300, and 1400. Figure 7A , Figure 7B , Figure 9 , Figure 11 , Figure 13 and Figure 14A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state storage such as flash memory, solid-state drives, etc. Personal electronic devices are not limited to... Figure 5B It can be the components and configurations, or it can include other components or additional components in a variety of configurations.
[0204] As used herein, the term "power indication" refers optionally to the power indication in devices 100, 300, and / or 500 ( Figure 1A , Figure 3 and Figures 5A to 5C A user-interactive graphical user interface object displayed on a screen. For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) optionally each constitute a functional representation.
[0205] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some specific implementations that include a cursor or other positional marker, the cursor acts as a "focus selector," such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the cursor is positioned on a touch-sensitive surface (e.g., a...). Figure 3 The touchpad 355 or Figure 4B When an input (e.g., a press input) is detected on the touch-sensitive surface 451 of the display, the specific user interface element is adjusted according to the detected input. This applies to touchscreen displays (e.g., those capable of enabling direct interaction with user interface elements on the touchscreen display) Figure 1A The touch-sensitive display system 112 or Figure 4AIn some embodiments of the touchscreen 112, a touch detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by touch) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, that particular user interface element is adjusted according to the detected input. In some embodiments, focus moves from one area of the user interface to another without corresponding movement of the cursor or movement of a touch on the touchscreen display (e.g., moving focus from one button to another using tab keys or arrow keys); in these embodiments, the focus selector moves according to the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user-controlled user interface element (or a touch on the touchscreen display) that delivers the user-expected interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of the focus selector (e.g., a cursor, touch, or selection box) above the corresponding button will indicate to the user that they expect to activate the corresponding button (rather than other user interface elements shown on the device's display).
[0206] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted away, before or after contact begins to move, before contact ends, before or after contact intensity is detected to increase and / or before or after contact intensity decreases). The characteristic intensity of the contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half maximum value of the contact intensity, the 90% maximum value of the contact intensity, etc. In some embodiments, the duration of the contact is used when determining the characteristic intensity (e.g., when the characteristic intensity is the average value of the contact intensity over time). In some implementations, the feature intensity is compared to a set of one or more intensity thresholds to determine whether the user has performed an action. For example, the set of one or more intensity thresholds may optionally include a first intensity threshold and a second intensity threshold. In this example, contact with a feature intensity not exceeding the first threshold results in a first action, contact with a feature intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a feature intensity exceeding the second threshold results in a third action. In some implementations, a comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more actions (e.g., whether to perform the corresponding action or abandon performing the corresponding action) rather than to determine whether to perform the first action or the second action.
[0207] In some implementations, a portion of the gesture is identified to determine the characteristic intensity. For example, a touch-sensitive surface optionally receives a series of swipe contacts that transition from a starting position to an ending position, where the contact intensity increases. In this example, the characteristic intensity of the contact at the ending position is optionally based only on a portion of the series of swipe contacts, rather than the entire swipe contact (e.g., only the portion of the swipe contact at the ending position). In some implementations, a smoothing algorithm is optionally applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of the following: unweighted moving average smoothing algorithm, triangular smoothing algorithm, median filter smoothing algorithm, and / or exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swipe contact to achieve the purpose of determining the characteristic intensity.
[0208] Optionally, the contact intensity on a touch-sensitive surface is characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device performs an operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device performs an operation different from the operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, when a contact with a characteristic intensity lower than the light press intensity threshold (e.g., and higher than the nominal contact detection intensity threshold, where contacts lower than the nominal contact detection intensity threshold are no longer detected) is detected, the device will move the focus selector based on the movement of the contact on the touch-sensitive surface without performing the operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise stated, these intensity thresholds are consistent across different groups of user interface figures.
[0209] An increase in contact intensity from below a light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact intensity from below a deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact intensity from below a contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact intensity from above a contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a contact being lifted off the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.
[0210] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or multiple contacts), wherein the corresponding press input is detected at least in part based on detecting that the intensity of the contact (or multiple contacts) increases to above a press input intensity threshold. In some embodiments, the corresponding operation is performed in response to detecting that the intensity of the corresponding contact increases to above a press input intensity threshold (e.g., a "downward stroke" of the corresponding press input). In some embodiments, the press input includes the intensity of the corresponding contact increasing to above a press input intensity threshold and the intensity of the contact subsequently decreasing to below the press input intensity threshold, and the corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases to below the press input threshold (e.g., an "upward stroke" of the corresponding press input).
[0211] In some implementations, the device employs intensity hysteresis to avoid unintended inputs sometimes referred to as "jitter," wherein the device defines or selects a hysteresis intensity threshold that has a predefined relationship with a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Therefore, in some implementations, a press input includes an increase in the intensity of the corresponding contact above the press input intensity threshold and a subsequent decrease in the intensity of that contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and a corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases below the hysteresis intensity threshold (e.g., an "upstroke" of the corresponding press input). Similarly, in some implementations, a press input is detected only when the device detects that the contact intensity increases from an intensity equal to or below the hysteresis intensity threshold to an intensity equal to or above the press input intensity threshold and optionally the contact intensity subsequently decreases to an intensity equal to or below the hysteresis intensity threshold, and a corresponding operation is performed in response to detecting a press input (e.g., an increase or decrease in contact intensity depending on the environment).
[0212] For ease of explanation, optionally, the description of an operation triggered in response to a press input associated with a press input strength threshold or in response to a gesture including a press input is provided in response to detecting any of the following conditions: the contact strength increases to above the press input strength threshold, the contact strength increases from below a hysteresis strength threshold to above the press input strength threshold, the contact strength decreases to below the press input strength threshold, and / or the contact strength decreases to below the hysteresis strength threshold corresponding to the press input strength threshold. Additionally, in the example where the operation is described as being performed in response to detecting a decrease in contact strength below the press input strength threshold, the operation is optionally performed in response to detecting a decrease in contact strength below a hysteresis strength threshold corresponding to and less than the press input strength threshold.
[0213] Now let’s turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as portable multifunction devices 100, 300 or 500).
[0214] Figure 5CAn exemplary diagram depicts a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500, and each device shares one or more data connections 510 (such as an internet connection, Wi-Fi connection, cellular connection, short-range communication connection, and / or any other such data connection or network) to facilitate real-time communication of audio and / or video data between the respective devices for a sustained period of time. In some embodiments, the exemplary communication session may include a shared data session, whereby data is transmitted from one or more electronic devices to other electronic devices to enable simultaneous output of corresponding content at the electronic devices. In some embodiments, the exemplary communication session may include a video conferencing session, whereby audio and / or video data are transmitted between devices 500A, 500B, and 500C, enabling users of the respective devices to communicate in real time using the electronic devices.
[0215] exist Figure 5C In this design, device 500A represents an electronic device associated with user A. Device 500A (via data connection 510) communicates with devices 500B and 500C, which are associated with users B and C, respectively. Device 500A includes a camera 501A for capturing video data of the communication session, and a display 504A (e.g., a touchscreen) for displaying content associated with the communication session. Device 500A also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0216] Device 500A displays a communication UI 520A via a display 504A. This communication UI is a user interface used to facilitate communication sessions (e.g., video conferencing sessions) between device 500B and device 500C. The communication UI 520A includes video feeds 525-1A and 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during a communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during a communication session.
[0217] The communication UI 520A includes a camera preview 550A, which is a representation of video data captured by camera 501A at device 500A. Camera preview 550A indicates to user A the expected video feed to be displayed at corresponding devices 500B and 500C.
[0218] The communication UI 520A includes one or more controls 555A for controlling one or more aspects of a communication session. For example, controls 555A may include controls for muting audio in the communication session, changing the camera view of the communication session (e.g., changing the camera used to capture video of the communication session, adjusting the zoom value), terminating the communication session, applying visual effects to the camera view of the communication session, and activating one or more modes associated with the communication session. In some embodiments, one or more controls 555A are optionally displayed in the communication UI 520A. In some embodiments, one or more controls 555A are displayed separately from the camera preview 550A. In some embodiments, one or more controls 555A are displayed to cover at least a portion of the camera preview 550A.
[0219] exist Figure 5C In this context, device 500B represents an electronic device associated with user B, who communicates with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B for capturing video data of the communication session, and a display 504B (e.g., a touchscreen) for displaying content associated with the communication session. Device 500B also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0220] Device 500B displays a communication UI 520B similar to the communication UI 520A of device 500A via touchscreen 504B. Communication UI 520B includes video feeds 525-1B and 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during a communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during a communication session. Communication UI 520B also includes: a camera preview 550B, which is a representation of video data captured at device 500B via camera 501B; and one or more controls 555B similar to controls 555A, which are used to control one or more aspects of the communication session. Camera preview 550B indicates to user B the expected video feed that user B will see displayed at the corresponding devices 500A and 500C.
[0221] exist Figure 5CIn this context, device 500C represents an electronic device associated with user C, who communicates with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C for capturing video data of the communication session, and a display 504C (e.g., a touchscreen) for displaying content associated with the communication session. Device 500C also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0222] Device 500C displays a communication UI 520C, similar to the communication UI 520A of device 500A and the communication UI 520B of device 500B, via a touchscreen 504C. The communication UI 520C includes video feeds 525-1C and 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during a communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during a communication session. The communication UI 520C also includes: a camera preview 550C, which is a representation of video data captured at device 500C via camera 501C; and one or more controls 555C similar to controls 555A and 555B, which are used to control one or more aspects of the communication session. The camera preview 550C indicates to user C the expected video feed to be displayed at the corresponding devices 500A and 500B.
[0223] Although Figure 5C The diagram depicts a communication session between three electronic devices, but such a session can be established between two or more electronic devices, and the number of devices participating in the session can change as electronic devices join or leave. For example, if one electronic device leaves the communication session, audio and video data from the device that has stopped participating are no longer represented on the participating devices. For instance, if device 500B stops participating in the communication session, there is no data connection 510 between devices 500A and 500C, and there is no data connection 510 between devices 500C and 500B. Furthermore, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and video and audio data are shared among all devices, enabling each device to output data transmitted from other devices.
[0224] Figure 5C The implementation scheme depicted in the diagram represents a communication session between multiple electronic devices, including Figures 6A to 6Q , Figures 8A to 8R , Figures 10A to 10J and Figures 12A to 12U The exemplary communication session is depicted in the image. In some implementations, Figures 6A to 6Q , Figures 8A to 8R , Figures 10A to 10J and Figures 12A to 12U The communication session depicted in the figure includes two or more electronic devices, even if other electronic devices involved in the communication session are not depicted in the figure.
[0225] Figures 6A to 6Q Exemplary user interfaces for managing real-time video communication sessions (e.g., video conferencing) according to some implementation schemes are shown. The user interfaces in these figures are used to illustrate including... Figure 7A and Figure 7B The process described below is the process in the middle.
[0226] Figures 6A to 6Q Device 600 is shown, which displays a user interface for managing real-time video communication sessions on display 601 (e.g., a display device or display generating component). Figures 6A to 6Q Various implementation schemes are described in which the device 600 automatically reconstructs the display portion of the camera's field of view based on conditions detected in the scene within the camera's field of view when the automatic viewfinder mode is enabled. The following section discusses... Figures 6A to 6Q One or more of the implementation schemes discussed may be related to... Figures 8A to 8R , Figures 10A to 10J and Figures 12A to 12U One or more combinations of implementation schemes are discussed in the implementation scheme.
[0227] Device 600 includes one or more cameras 602 (e.g., a front-facing camera) for capturing image data and optionally capturing depth data of the scene within the camera's field of view. In some embodiments, camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600 includes one or more features of device 100, 300, or 500.
[0228] exist Figure 6AIn the device 600, a video conferencing request interface 604-1 is displayed, depicting an incoming request from "John" to participate in a live video conference. The video conferencing request interface 604-1 includes a camera preview 606, an options menu 608, a viewfinder mode indicator 610, and a background blur indicator 611. The camera preview 606 is a real-time representation of the video feed from camera 602, which is enabled for output during the video conferencing session (if the incoming request is accepted). Figure 6A In the image, camera 602 is currently enabled, and camera preview 606 depicts a representation of "Jane," who is currently in front of device 600 and within the field of view of camera 602.
[0229] Options menu 608 includes various selectable options for controlling one or more aspects of the video conference. For example, mute option 608-1 can be selected to mute any audio transmission detected by device 600. Flip option 608-2 can be selected to switch the camera being used for the video conference between camera 602 and one or more different cameras (such as cameras on the opposite side of device 600 (e.g., rear cameras)). Accept option 608-3 can be selected to accept a request to participate in a live video conference. Deny option 608-4 can be selected to deny a request to participate in a live video conference.
[0230] Background blur capability indicates that 611 can be selected to enable or disable background blur mode for video conferencing sessions. Figure 6A In the described implementation, when device 600 receives or initiates a request to participate in a video conference, the background blur mode is disabled by default. Therefore, the background blur capability indicator 611 is depicted as having an unselected state, such as... Figure 6A The bolded indicator does not include the power indication. In some implementations, when device 600 receives or initiates a request to participate in a video conference, a background blur mode is enabled by default (e.g., the power indication is bolded). Regarding Figures 12A to 12U The background blur feature will be discussed in more detail.
[0231] The framing mode indicator 610 can be selected to enable or disable the automatic framing mode for video conferencing sessions. Figure 6A In the described implementation, when device 600 receives or initiates a request to participate in a video conference, the automatic framing mode is enabled by default. Therefore, the framing mode indication 610 is depicted as having a selected state, as indicated by… Figure 6A The bolded power indicator indicates this. In some implementations, when device 600 receives or initiates a request to participate in a video conference, the automatic framing mode is disabled by default (e.g., and there is no bolding of the power indicator).
[0232] When the automatic viewfinder mode is enabled, device 600 detects the conditions of the scene within the field of view of the enabled camera (e.g., camera 602) (e.g., the presence and / or position of one or more objects within the camera's field of view), and adjusts the field of view of the video output for the video conferencing session in real time (as shown in the camera preview) based on the conditions or changes in the scene detected within the camera's field of view (e.g., changes in the position and / or movement of objects during a video conferencing session) (e.g., without moving camera 602 or device 600). Various implementations of the automatic viewfinder mode have been discussed throughout this disclosure.
[0233] For example, in Figures 6A to 6Q In the implementation described herein, device 600 automatically adjusts (e.g., reconstructs) the display output video feed field of view to maintain the display of one or more objects (e.g., Jane) within the field of view of the camera (e.g., camera 602). Because device 600 automatically adjusts the display portion of the camera's field of view to include Jane's display, Jane can move around the scene while participating in a video conference without having to manually adjust the outgoing video feed angle to account for her movement or other changes in the scene. Therefore, participants in the video conference require less interaction with device 600 because device 600 automatically reconstructs the outgoing video feed, allowing remote participants receiving the video feed from device 600 to continuously view Jane as she moves around in her environment. Other benefits of the automatic framing mode will be described in the disclosure below. Figures 6A to 6Q , Figures 8A to 8R and Figures 10A to 10J The implementation scheme described herein discusses various features of the automatic viewfinder mode. One or more of these features can be combined with other features of the automatic viewfinder mode, as discussed herein.
[0234] Figure 6A The diagram depicts input 612 (e.g., tap input) on the viewfinder mode enable display 610 and input 614 on the accept option 608-3. In response to detecting input 612, device 600 disables the auto-view mode. In response to detecting input 614 after input 612, device 600 accepts a video conferencing call and joins the video conferencing session while the auto-view mode is disabled, as shown. Figure 6B As described in [the text]. If device 600 does not detect input 612, or if input 612 is an input that enables auto-framing mode (e.g., framing mode enable indicator 610 is not selected when input 612 is received), then device 600 accepts the video conferencing call in response to input 614, and joins the video conferencing session if auto-framing mode is enabled, as [the text continues]. Figure 6F As depicted in the text.
[0235] Figure 6BScene 615 is depicted, which is the physical environment within the field of view 620 of camera 602. Figure 6B In the embodiment depicted herein, Jane 622 is seated on sofa 621, with device 600 positioned in front of her (e.g., on a table), and door 618 in the background. Field of view 620 refers to the field of view of camera 602 (e.g., the maximum field of view of camera 602 or the wide-angle field of view of camera 602) that encompasses scene 615. Part 625 indicates a portion of field of view 620 currently being output (or selected for output) for a video conferencing session (e.g., part 625 represents a display portion of the camera's field of view). Therefore, part 625 indicates a portion of scene 615 currently represented in the display video feed depicted in camera preview 606. Field of view 620 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view, and part 625 is sometimes referred to herein as the video feed field of view.
[0236] When an incoming video conferencing request is accepted, the video conferencing request interface 604-1 redirects to the video conferencing interface 604, as shown below. Figure 6B As shown. The video conferencing interface 604 is similar to the video conferencing request interface 604-1, but is updated to depict, for example, an incoming video feed 623, which includes video data of a remote participant in the video conference received at device 600. Video feed 623 includes a representation 623-1 of John as a remote participant in a video conference with Jane 622. Compared to Figure 6A The camera preview 606 is now smaller in size and shifted toward the upper right corner of the display 601. The camera preview 606 includes Jane's representation 622-1 and a representation of the environment captured within portion 625 (e.g., video feed field of view). The options menu 608 has been updated to include a framing mode indication display 610. In some embodiments, for example, the framing mode indication display 610 is displayed in the camera preview 606, such as... Figure 6O As shown. Figure 6B As shown, the viewfinder mode indicator 610 is shown as unselected, thus indicating that the automatic viewfinder mode is disabled.
[0237] In some implementations, when the automatic viewfinder mode is disabled, device 600 outputs a predetermined portion of the available field of view of camera 602 as the video feed field of view. Examples of such implementations are depicted in... Figures 6B to 6D In this context, portion 625 represents a predetermined portion of the usable field of view of camera 602 located at the center of field of view 620. In some embodiments, when auto-view mode is disabled, device 600 outputs the entire field of view 620 as a video feed field of view. Examples of such embodiments are depicted in... Figure 10A middle.
[0238] exist Figure 6B In the video conference, Jane is using device 600 to participate with John. Similarly, John is using a device with one or more features of device 100, 300, 500, or 600 to participate with Jane. For example, John is using a tablet computer similar to device 600 (e.g., Figures 10H to 10J and Figures 12B to 12N John's tablet (600a). Therefore, John's device displays a video conferencing interface similar to video conferencing interface 604, except that John's device displays a camera preview showing the video feed captured from John's device (currently in...). Figure 6B (As depicted in video feed 623), and the incoming video feed on John's device shows the video feed output from Jane's device 600 (currently in...). Figure 6B (As depicted in the camera preview 606).
[0239] exist Figure 6C In this scene, device 600 moves relative to scene 615, and therefore field of view 620 and portion 625 pivot with device 600. Because automatic viewfinder mode is disabled, device 600 does not automatically adjust the video feed field of view to remain fixed in Jane's position within field of view 620. Instead, the video feed angle moves with device 600, and Jane 622 is no longer centered in the video feed field of view, as indicated by portion 625 and depicted as in camera preview 606, which shows the background of scene 615 and a portion of Jane's representation 622-1. Figure 6C In the embodiments depicted, the movement of device 600 is pivotal; however, the movement of field of view 620 and portion 625 may be caused by other movements, such as tilting, rotating and / or moving device 600 in a manner that prevents Jane 622 from remaining within portion 625 (e.g., forward, backward and / or left and right).
[0240] exist Figure 6D In this process, device 600 returns to its original position, and Jane 622 bends downward, moving out of portion 625. Similarly, because the automatic viewfinder mode is disabled, device 600 does not automatically adjust the video feed field of view to follow Jane's movement when Jane moves out of portion 625. While Jane 622 moves, the video feed field of view remains stationary, and Jane's representation 622-1 is mostly outside the viewfinder in camera preview 606.
[0241] exist Figure 6DIn this process, device 600 detects input 626 (e.g., a tap) on the viewfinder mode indication display 610. In response, device 600 bolds the viewfinder mode indication display 610 (to indicate its selected / enabled state) and enables automatic viewfinder mode, as shown below. Figure 6E As shown in the diagram. When automatic framing mode is enabled, device 600 automatically adjusts the display video feed field of view based on conditions detected within scene 615. Figure 6E In the implementation described herein, device 600 adjusts the display video feed field of view to center Jane's face. Accordingly, device 600 updates camera preview 606 to include a representation 622-1 of Jane centered in the viewfinder and a representation 621-1 of the sofa she sits on in the background. Field of view 620 remains fixed because the position of camera 602 remains unchanged. However, the position of Jane's face within field of view 620 does change. Therefore, device 600 adjusts (e.g., repositions) a portion of the display of field of view 620 so that Jane remains positioned within camera preview 606. This is in Figure 6E The repositioning of part 625 is indicated by centering it on Jane's face. Figure 6E In this context, portion 627 corresponds to the previous position of portion 625, and thus represents a portion of the field of view 620 previously displayed in camera preview 606 (before the adjustments made by enabling auto-view mode).
[0242] Figure 6F A scene 615 and device 600 are depicted when the auto-view mode is enabled in response to input 614, which accepts an incoming request to join a video conference, and the auto-view mode is enabled (or alternatively, in response to input 626, which enables the auto-view mode). Accordingly, device 600 displays Jane's representation 622-1 in the center of camera preview 606.
[0243] exist Figure 6G In the middle, equipment 600 is similar to the one mentioned above. Figure 6C The method of movement is under discussion. However, because... Figure 6G With automatic viewfinder mode enabled, device 600 automatically adjusts the video feed field of view (part 625) relative to field of view 620 to remain fixed on Jane's face. This has changed its position relative to device 600 and camera 602 in response to the pivoting of device 600. Therefore, camera preview 606 continues to display Jane's representation 622-1 at the center of the video feed field of view. The adjustment of the video feed field of view is represented by the change in the position of part 625 within field of view 620. For example, when compared to... Figure 6F At that time, the relative position of part 625 has shifted from the center position within the field of view 620 (in Figure 6G (represented by part 627) moved to Figure 6GThe displaced position depicted in the text. Figure 6G In the embodiments depicted, the movement of device 600 is pivotal; however, device 600 may automatically adjust the display video feed field of view in response to other movements, such as tilting, rotating, and / or moving device 600 in a manner that keeps Jane 622 within the field of view 620 (e.g., forward, backward, and / or left and right).
[0244] exist Figure 6H In this scene, device 600 returns to its original position, and Jane 622 has moved to a standing position near sofa 621. Device 600 detects the updated position of Jane 622 in scene 615 and updates the video feed field of view to maintain its position on Jane 622, as shown in camera preview 606. Therefore, portion 625 has moved from its previous position, represented by portion 627, to an updated position surrounding Jane's face, as shown in camera preview 606. Figure 6H As depicted in the text.
[0245] In some implementation schemes, from Figure 6G The field of view depicted in the camera preview 606 is... Figure 6H The field-of-view transition depicted in camera preview 606 is performed as a matching clip. For example, from Figure 6G Camera preview 606 in Figure 6H The transition in camera preview 606 is a match clip performed when Jane 622 has moved from her seated position on sofa 621 to a standing position closer to the sofa. The result of the match clip is that camera preview 606 appears to change from... Figure 6G The first camera view in the transition to Figure 6H Different camera views in (which may optionally have different camera views than) Figure 6G (The same zoom level as the camera view in the image). However, the actual field of view of camera 602 (e.g., field of view 620) has not changed. Instead, only a portion of the field of view (part 625) that is displayed changes position within field of view 620.
[0246] Figure 6I It is similar to Figure 6H The implementation scheme described in the document, but the camera preview 606 is compared to Figure 6H The camera preview shown has a larger, extended field of view. Specifically, Figure 6I The implementation scheme depicted in the diagram illustrates the process from... Figure 6G Camera preview in Figure 6I The camera preview jump cut transition in the video. The jump cut transition is achieved by switching from... Figure 6G Camera Preview 606 in the middle is converted to Figure 6I The camera preview 606 (which has a large (e.g., zoomed-out) field of view) is used to depict this. Therefore, Figure 6IThe video feed field of view (represented by portion 625) is the larger portion of the field of view 620. This is achieved through portion 625 (corresponding to...) Figure 6I (Camera preview in the image) and part 627 (corresponding to) Figure 6G The size difference between the cameras (in the camera preview) is shown.
[0247] Figure 6H and Figure 6I A specific implementation of the transition between different camera previews is shown. In some implementations, other transitions can be performed, such as by continuously moving (e.g., panning and / or zooming) the video feed field of view within the field of view 620 to follow Jane 622 as she moves around scene 615.
[0248] exist Figure 6J In the scene 615, Jane 622 has moved away from device 600 and is now behind sofa 621. In response to detecting a change in Jane 622's position, device 600 performs a transition (e.g., a jump cut transition), where camera preview 606 depicts a close-up view of Jane in scene 615. Figure 6J In the embodiments depicted, device 600 moves closer to Jane 622 (e.g., reduces the field of view) as Jane 622 moves away from the camera (e.g., moves a threshold distance). In some embodiments, device 600 pulls away from Jane 622 as she moves toward the camera (e.g., moves a threshold distance) (e.g., expands the display field of view). For example, if Jane 622 wants to move away from her... Figure 6J Move the position in the middle to her Figure 6I If the camera preview 606 is moved from its previous position, it will zoom out to... Figure 6I The camera preview depicted in the image.
[0249] exist Figure 6K In the middle, another object, Jack 628, enters scene 615. Device 600 continues to display with features such as Figure 6J The same camera preview video conferencing interface 604 depicted in the image. Figure 6K In the depicted implementation, device 600 detects Jack 628 within field of view 620, but maintains the same video feed field of view as Jack 628 moves around the scene or until Jack 628 moves to a specific location in the scene (e.g., closer to the center of the scene). In some implementations, when an additional object is detected within field of view 620, device 600 displays a prompt to adjust the camera preview. The following section discusses... Figure 6P and Figures 8A to 8J The implementation plan described in the document illustrates an example of this suggestion in more detail.
[0250] exist Figure 6LIn this embodiment, device 600 reconstructs camera preview 606 to include representation 628-1 of Jack 628, who is now standing next to Jane 622 in scene 615. In some implementations, device 600 automatically adjusts camera preview 606 in response to determining that Jack 628 has stopped moving around scene 615 and / or is exhibiting behavior indicating a desire to participate in a video conference. Examples of such behavior may include turning attention to device 600 and / or camera 602, focusing / looking at camera 602 or in the general direction of that camera, remaining still (e.g., for at least a certain amount of time), positioning oneself next to a participant in the video conference (e.g., Jane 622), facing device 600, speaking, etc. In some implementations, when automatic framing mode is enabled, device 600 automatically adjusts camera preview in response to detecting a change in the number of objects detected within field of view 620 (such as when Jack 628 enters scene 615). Figure 6L As depicted herein, part 625 represents the adjusted video feed field of view, and part 627 represents... Figure 6L The size of the video feed field of view was adjusted before. Compared to Figure 6K The video feed field of view in the middle, Figure 6L The adjusted field of view was widened and re-centered on Jane 622 and Jack 628, as depicted in camera preview 606.
[0251] exist Figure 6M In the above embodiments, Jane 622 begins to move away from Jack 628 and out of the video feed field of view represented by portion 625 and camera preview 606. When Jane's movement is detected, device 600 maintains (e.g., does not adjust) the field of view of camera preview 606. In some embodiments, as Jane moves away from Jack, device 600 resizes the video feed field of view so that both objects remain within the video feed field of view (camera preview). In some embodiments, after Jane has moved away from Jack, device 600 resizes the video feed field of view after Jane stops moving so that both objects remain within the camera preview.
[0252] exist Figure 6N In the middle, device 600 detected that Jane 622 was no longer present. Figure 6M Within part 625 (e.g., Figure 6N In response, the displayed video feed field of view is adjusted to zoom in on Jack 628. Thus, device 600 displays a video conferencing interface 604, where camera preview 606 has a zoomed-in view of Jack 628. Figure 6N A portion 625, which is smaller than portion 627, is depicted to indicate the change in the field of view of the displayed video feed.
[0253] In some implementations, when automatic framing mode is enabled, device 600 displays a prompt to adjust the video feed field of view to include one or more additional participants in response to detecting an additional object within the field of view 620. The following is about... Figures 60 to 6Q Examples of such implementation schemes are described.
[0254] Figure 6O Depicting something similar to Figure 6N The embodiment shown is an implementation where the framing mode indication 610 is displayed in the camera preview area instead of in the options menu 608. Device 600 displays a framing indicator 630 positioned around Jack's indication 628-1 to indicate that device 600 has detected the presence of a face (e.g., Jack's face) in the camera preview. In some embodiments, the framing indicator also indicates that an automatic framing mode is enabled and that the detected face is being tracked when it is detected within section 625, camera preview 606, and / or field of view 620. Jack 628 is currently the only participant in the video conference located in scene 615.
[0255] exist Figure 6P In the process, device 600 detects that Jane 622 has entered scene 615 within field of view 620. In response, device 600 updates the video conferencing interface 604 by displaying an added revelation 632 in the camera preview area. The added revelation 632 can be selected to adjust the displayed video feed field of view to include additional objects detected within field of view 620.
[0256] exist Figure 6P In the process, device 600 detects the input 634 added to the display representation 632, and in response, adjusts portion 625 so that camera preview 606 includes Jane's representation 622-1 and Jack's representation 628-1, as shown. Figure 6Q As depicted in [the document]. In some implementations, Jane 622 is added as a participant in the video conference. Device 600 also recognizes the presence of Jane's face and displays framing indicators 630 (in addition to those framing indicators around Jack's face) around a representation of Jane's face to indicate that framing mode is enabled and that Jane's face is being tracked when it is detected in section 625, camera preview 606, and / or field of view 620.
[0257] Figures 7A to 7BA flowchart illustrating a method for managing a real-time video communication session using an electronic device, according to some embodiments, is depicted. Method 700 is performed at a computer system (e.g., a smartphone, tablet computer) (e.g., 100, 300, 500, 600) that communicates with display generation components (e.g., 601) (e.g., a display controller, a touch-sensitive display system), one or more cameras (e.g., 602) (e.g., an infrared camera, a depth camera, a visible light camera), and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 700 are optionally combined, some operations are optionally changed in order, and some operations are optionally omitted.
[0258] As described below, method 700 provides an intuitive way to manage real-time video communication sessions. This method reduces the cognitive burden on users managing real-time video communication sessions, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to manage real-time video communication sessions faster and more efficiently, saving power and increasing the time interval between battery charges.
[0259] In method 700, the computer system (e.g., 600) displays (702) a communication request interface (e.g., via a display generation component (e.g., 601) via a display generation component (e.g., 702). Figure 6A (e.g., an interface for receiving or sending real-time video communication sessions (e.g., real-time video chat sessions, real-time video conferencing sessions)).
[0260] The computer system (e.g., 600) displays (704) a communication request interface (e.g., Figure 6A In embodiment 604-1), the communication request interface includes a first selectable graphical user interface object (e.g., 608-3) (e.g., an "Accept" enablement) associated with the process of joining a real-time video communication session. In some embodiments, the accept enablement can be selected to initiate a process for accepting an incoming request to join a real-time video communication session. In some embodiments, the first selectable graphical user interface object can be selected to initiate a "Cancel" enablement for canceling or terminating an outgoing request to join a real-time video communication session.
[0261] The computer system (e.g., 600) (e.g., simultaneously with 704) displays (706) a communication request interface (e.g., including a second selectable graphical user interface object (e.g., 610)). Figure 6AIn 604-1), the second selectable graphical user interface (e.g., "framing mode" indication, "background blur" indication, "dynamic video quality" indication) is associated with a process for selecting between using a first camera mode (e.g., auto-framing mode, background blur mode, dynamic video quality mode) for the one or more cameras and using a second camera mode (e.g., a mode different from the first camera mode (e.g., a mode in which auto-framing mode is disabled, a mode in which background blur mode is disabled, and / or a mode in which dynamic video quality mode is disabled)) for the one or more cameras during a real-time video communication session.
[0262] In some implementations, the framing mode indication can be selected to enable / disable a mode (e.g., an auto framing mode) for: 1) tracking the position and / or localization of one or more objects detected within the field of view of one or more cameras during a live video communication session, and 2) automatically adjusting the display view of the objects based on the tracking of the objects during a live video communication session.
[0263] In some implementations, the background blur capability can be selected to enable / disable a mode (e.g., background blur mode) in which visual effects (e.g., blurring, darkening, coloring, occlusion, desaturation, or other de-emphasis effects) are applied to the background portion of the camera field of view during a real-time video communication session (e.g., 606) (e.g., camera preview, output video feed of the camera field of view) (e.g., without applying visual effects to the portion of the camera field of view that includes the representation of objects (e.g., the foreground portion) (e.g., 622-1)).
[0264] In some embodiments, the dynamic video quality indicator can be selected to enable / disable a mode (e.g., dynamic video quality mode) for outputting (e.g., transmitting and optionally displaying) the camera field of view, where portions have different degrees of compression and / or video quality. For example, the portion of the camera field of view including detected faces (e.g., 622-1) is compressed less than the portion of the camera field of view excluding detected faces. In some embodiments, there is an inverse relationship between the degree of compression and video quality (e.g., greater compression produces lower video quality; less compression produces higher video quality). Thus, a video feed for a real-time video communication session can be transmitted (e.g., via a computer system (e.g., 600)) to a receiving device of a remote participant in the real-time video communication session, such that the portion including detected faces can be displayed at the receiving device at a higher video quality than the portion of the camera field of view excluding detected faces (due to reduced compression of the portion including detected faces and increased compression of the portion excluding detected faces). In some embodiments, the computer system varies the amount of compression as the video bandwidth changes (e.g., increases, decreases). For example, the compression of the portion of the camera's field of view that does not include the detected face (e.g., the camera feed) varies (e.g., increases or decreases with a corresponding change in bandwidth), while the compression of the portion of the camera's field of view that includes the detected face remains constant (or, in some embodiments, varies at a smaller rate or by a smaller amount than that of the portion of the camera's field of view that does not include the face).
[0265] In the interface displaying communication requests (e.g., Figure 6A When 604-1 is in the middle, the computer system (e.g., 600) receives (708) a set of one or more inputs (e.g., 612 and / or 614) via one or more input devices (e.g., 601) including a selection (e.g., 614) of a first selectable graphical user interface object (e.g., 608-3) (e.g., the set of one or more inputs includes a selection of an accept capability representation, and optionally includes a selection of a framing mode capability representation, a background blur capability representation, and / or a dynamic video quality capability representation).
[0266] In response to receiving one or more inputs (e.g., 612 and / or 614) including a selection (e.g., 608-3) of a first selectable graphical user interface object (e.g., 614), the computer system (e.g., 600) displays (710) a real-time video communication interface (e.g., 604) for a real-time video communication session via a display generation component (e.g., 601).
[0267] When a real-time video communication interface is displayed (e.g., 604), a computer system (e.g., 600) detects (712) changes (e.g., changes in the position of objects) in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602). In some embodiments, the scene includes a representation of objects and optionally includes one or more additional objects within the field of view of the one or more cameras.
[0268] In response to detecting a change (714) in the scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602), a computer system (e.g., 600) executes one or more of steps 716 and 718 in method 700.
[0269] Based on the determination that a first camera mode is selected (e.g., enabled) (e.g., if the first camera mode is disabled by default, the group of one or more inputs includes the selection of a second selectable graphical user interface object; if the first camera mode is enabled by default, the group of one or more inputs does not include the selection of a second selectable graphical user interface object) (e.g., the viewfinder mode enablement is selected when the enablement is selected), the computer system (e.g., 600) adjusts (716) (e.g., automatically, without user input) the representation of the field of view of the one or more cameras (e.g., 606) (e.g., the display field of view of the one or more cameras) (e.g., the representation of the field of view of the one or more cameras is automatically adjusted during the real-time video communication session) based on detected changes in the scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras (e.g., automatically adjusted during the real-time video communication session, based on the detected object position). When the first camera mode is selected, the representation of the field of view of one or more cameras is adjusted during real-time video communication based on detected changes in the scene within the field of view of one or more cameras. This enhances the video communication session experience by automatically adjusting the camera's field of view (e.g., to maintain the display of the object / user) without additional input from the user. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system). This, in turn, reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively.
[0270] In some implementations, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session (e.g., 606) includes: 1) determining that a first set of criteria is met, including that the scene is included at a first location (within the field of view 620 of the one or more cameras, in... Figure 6F The detected object (e.g., 622) (e.g., one or more users of a computer system) is displayed as a representation with a first field of view (e.g., Figure 6F 606) of the real-time video communication interface 604 (e.g., displaying the real-time video communication interface in a first display portion of the field of view of the first digital zoom level and the one or more cameras) (in some embodiments, the representation of the first field of view includes a representation of the object when the object is located at the first position); and 2) based on determining that a second set of criteria is met, including detecting the object at a second position different from the first position (e.g., Figure 6H (625 in the text), showing a representation having a second field of view that is different from the representation of the first field of view (e.g., Figure 6H The real-time video communication interface (606) of the camera (e.g., displaying the real-time video communication interface in a second display portion of the field of view of the second digital zoom level and / or the one or more cameras) (e.g., a representation of the field of view that is zoomed in, zoomed out, and / or panned in a direction relative to the representation of the first field of view) (in some embodiments, the representation of the second field of view includes a representation of the object when the object is located at a second position). In some embodiments, when the first camera mode is selected (e.g., enabled), the representation of the field of view changes automatically in response to a detected change in the object's position and / or in response to the detection of a second object entering or leaving the field of view of the one or more cameras (e.g., without changing the actual field of view of the one or more cameras). For example, the representation of the field of view is changed to track the position of the object and the display position and / or zoom level (e.g., digital zoom level) is adjusted to make the object appear more prominent (e.g., changing the digital zoom level to appear as if the object is being pushed closer when it moves away from the camera; changing the digital zoom level to appear as if the object is being pulled away when it moves toward the camera; changing the display portion of the field of view of the one or more cameras to appear as if the object is being translated in a particular direction when it moves in that direction).
[0271] Based on the determination to use a second camera mode (e.g., enabled) (e.g., when the receiving display is selected, the viewfinder mode display is in an unselected or deselected state), the computer system (e.g., 600) abandons (718) adjusting the representation of the field of view of the one or more cameras during the real-time video communication session (e.g., as...). Figure 6D(As depicted herein) (e.g., based on detected changes in the scene within the field of view of the one or more cameras) (e.g., when a first camera mode is disabled, the real-time video communication interface maintains the same (e.g., default) representation of the field of view, regardless of whether an object is located within the scene within the field of view of the one or more cameras, and regardless of where the object is located within the field of view of the one or more cameras). In some embodiments, abandoning the adjustment of the representation of the field of view of the one or more cameras during a real-time video communication session includes: 1) when (e.g., according to determination) an object has a first position within the scene within the field of view of the one or more cameras (e.g., as described herein). Figure 6B (as depicted in the text), displaying a real-time video communication interface with a representation having a first field of view (e.g., Figure 6B 606); and 2) when (e.g., according to determination) the object has a second position within the scene in the field of view of the one or more cameras (e.g., in Figure 6D (in the middle), displaying a real-time video communication interface with a representation having a first field of view (e.g., Figure 6D (606 in the original text). In some implementations, the representation of the first field of view is a standard or default representation of the field of view that does not change based on changes in the scene (e.g., changes in the position of an object relative to the one or more cameras or a second object entering or leaving the field of view of the one or more cameras).
[0272] In some embodiments, detected changes in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) include detected changes in a set of attention-based factors of one or more objects (e.g., 622, 628) within the scene (e.g., a first object directing its attention (e.g., focusing on, looking at) the one or more cameras (e.g., based on the first object's gaze position, head position, and / or body position)). In some embodiments, a computer system (e.g., 600) adjusts the representation of the field of view of the one or more cameras during a real-time video communication session based on detected changes in the scene within the field of view of the one or more cameras (e.g., 606), including: adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) detected changes in the set of attention-based factors of the one or more objects within the scene (e.g., such as...). Figure 6L(As depicted in the text). Adjusting the field of view representation of one or more cameras during real-time video communication based on detected changes in this set of attention-based factors for one or more objects in the scene enhances the video communication session experience by automatically adjusting the camera's field of view based on this set of attention-based factors for objects in the scene without additional input from the user. Performing operations when a set of conditions are met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system). This, in turn, reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively.
[0273] In some implementations, when automatic viewfinder mode is enabled, a computer system (e.g., 600) adjusts (e.g., reconstructs) the display portion of the field of view of one or more cameras (e.g., 602) based on one or more attention-based factors of objects (e.g., 622, 628) detected within the field of view (e.g., 620) of the one or more cameras (e.g., 606). For example, when a first object draws attention to the one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include a representation of the first object or to focus on the first object. Conversely, when the first object's attention shifts away from the one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms in) to exclude a representation of the first object (e.g., if other objects remain in the field of view of the one or more cameras) or to focus on another object.
[0274] In some implementations, the set of attention-based factors includes a first factor, which is based on the detected focal plane (e.g., as shown by the image) of a first object (e.g., 628) among the one or more objects in the scene. Figure 6L(As depicted in the text). Adjusting the representation of the field of view of the one or more cameras during real-time video communication based on the detected focal plane of a first object among the one or more objects in the scene enhances the video communication session experience by automatically adjusting the camera's field of view without additional user input when the focal plane of the first object meets a criterion. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently. In some implementations, the attention of the first object is determined based on the focal plane of the first object. For example, if the focal plane of the first object is aligned (e.g., coplanar) with the focal plane of the one or more cameras or the focal plane of another object participating in the real-time video communication session, the first object is considered to be paying attention to the one or more cameras. Therefore, the first object is considered an active participant in the real-time video communication session, and the computer system then adjusts the representation of the field of view of the one or more cameras to include the first object in the real-time video communication interface.
[0275] In some implementations, the set of attention-based factors includes a second factor based on whether a second object (e.g., 628) (e.g., an object other than the first object) among the one or more objects in the scene (e.g., 615) is determined to be looking at the one or more cameras (e.g., 602). The representation of adjusting the field of view of the one or more cameras during real-time video communication based on whether a second object in the scene is determined to be looking at the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera when the second object is looking at it, without additional input from the user. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some implementations, the attention of the second object is determined based on whether the second object is looking at the one or more cameras. If so, the second object is considered to be paying attention to the one or more cameras and is therefore considered an active participant in the real-time video communication session. Therefore, the computer system adjusts the representation of the field of view of the one or more cameras to include the second object in the real-time video communication interface.
[0276] In some embodiments, detected changes in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) include detected changes in the number (e.g., quantity, number) of objects (e.g., 622, 628) in the scene (e.g., quantity, number of objects) detected in the scene (e.g., detected changes in the number of objects in the scene that meet a first set of criteria (e.g., objects located within the field of view of the one or more cameras and optionally stationary) (e.g., one or more objects entering or leaving the scene within the field of view of the one or more cameras)). In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on detected changes in the scene within the field of view of the one or more cameras (e.g., 606) includes adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) detected changes in the number of objects detected in the scene (e.g., meeting a first set of criteria) (e.g., as...). Figure 6L and / or Figure 6N (As depicted in the text). Adjusting the representation of the field of view of one or more cameras during real-time video communication based on detected changes in the number of objects detected in the scene enhances the video communication session experience by automatically adjusting the camera's field of view as the number of objects in the scene changes without additional input from the user. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, when automatic framing mode is enabled, the computer system (e.g., 600) adjusts (e.g., reconstructs) the display portion of the field of view of the one or more cameras based on the number of objects detected within the field of view of the one or more cameras. For example, when the number of objects detected in the scene increases, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include additional objects (e.g., along with previously detected objects). Similarly, when the number of objects detected in a scene decreases, the representation of the field of view of one or more cameras changes (e.g., zooms in) to capture the objects remaining in the scene.
[0277] In some implementations, adjusting the representation of the field of view of one or more cameras (e.g., 606) during a real-time video communication session based on detected changes in the number of objects detected in the scene is based on the determination of whether objects in the field of view (e.g., 620) are stationary (e.g., relatively stationary; moving within the field of view of the one or more cameras by no more than a threshold amount of movement). Adjusting the representation of the field of view of one or more cameras during real-time video communication based on whether objects in the field of view are stationary enhances the video communication session experience by automatically adjusting the camera's field of view when objects in the scene are stationary without additional input from the user. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively.
[0278] In some implementations, when automatic framing mode is enabled, the computer system (e.g., 600) considers an object (e.g., 628) to be a participant in a real-time video communication session if the detected object's movement does not exceed a threshold amount of movement. When an object is considered an active participant, the computer system adjusts (e.g., reconstructs) the display portion of the field of view of one or more cameras (e.g., 606) to subsequently include a representation of the object (e.g., as shown in the image). Figure 6L (As depicted in the text). This prevents the computer system from automatically reconstructing the display portion of the field of view of one or more cameras based on irrelevant movement in the scene (e.g., movement caused by objects passing through the background or children jumping around in the field of view of one or more cameras), which would otherwise distract the participants / viewers of a real-time video communication session.
[0279] In some embodiments, before a computer system (e.g., 600) detects a change in the scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602), the representation of the field of view of the one or more cameras (e.g., 606) has a first representing field of view (e.g., the computer system is displaying a portion of the field of view of the one or more cameras before detecting the change in the scene). In some embodiments, the change in the scene within the field of view of the one or more cameras includes a third object (e.g., 622) from the field of view of the one or more cameras corresponding to (e.g., represented by; included in) the first representing field of view (e.g., ...). Figure 6G The first part of 606 (e.g., Figure 6G (625) to the field of view of the one or more cameras does not correspond to (e.g., not represented by it; not included in) the second part of the first represented field of view (e.g., Figure 6H The detected movement (e.g., a third object moving from a portion of the field of view of the one or more cameras that was displayed before the third object moved to a portion of the field of view of the one or more cameras that was not displayed before the third object moved) is considered. In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on (in some embodiments, in response to) the detected change in the scene in the field of view of the one or more cameras includes: adjusting the representation of the field of view from a first representation field of view to a second representation field of view (e.g., different from the first representation field of view) corresponding to (e.g., representing; displaying; including) a second portion of the field of view of the one or more cameras, based on determining that a fourth object is not detected in the scene in the first portion of the field of view of the one or more cameras (e.g., 628). Figure 6H The camera preview 606 depicted in the image. Based on the determination that a fourth object (e.g., 628) is detected in the scene within a first portion of the field of view of the one or more cameras (e.g., Jack 628 is located when Jane 622 leaves the frame). Figure 6M In part 625), the adjustment of the field of view representation from the first representation field of view to the second representation field of view is abandoned (e.g., the first representation field of view continues to be displayed) (e.g., in Figure 6M In this context, when Jane 622 leaves and Jack 628 remains, device 600 continues to display a camera preview 606 of the depicting portion 625. After another object (e.g., a third object) leaves the first portion of the field of view, the representation of the field of view of the one or more cameras is selectively adjusted from a first representation field of view to a second representation field of view during a real-time video communication session based on whether an object (e.g., a fourth object) is detected in the scene within the first portion of the field of view of the one or more cameras. This practice enhances the video communication session experience by automatically adjusting the camera's field of view based on whether an additional object remains in the first portion of the field of view when another object leaves, without requiring additional user input. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively. In some implementations, when automatic framing mode is enabled, the computer system does not track (e.g., follow; adjust the representation of the field of view in response to this) the movement of another object that leaves the display field of view while the object remains in the display field of view.
[0280] In some implementations, before detecting a change in the scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602), the representation of the field of view of those one or more cameras (e.g., 606) has a third representation of the field of view (e.g., in...). Figure 6F (For example, the computer system is displaying a portion of the field of view of the one or more cameras before detecting a change in the scene.) In some embodiments, a change in the scene within the field of view of the one or more cameras includes a fifth object (e.g., 622) corresponding to (e.g., represented by it; included in) a third portion of the field of view (e.g., ...) representing the third portion of the field of view. Figure 6F (625) to the field of view of the one or more cameras does not correspond to (e.g., not represented by it; not included in) the third representation of the fourth part of the field of view (e.g., Figure 6H The movement of 625 in the middle. In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on (in some embodiments, in response to) detected changes in the scene within the field of view of the one or more cameras includes: displaying a representation of the field of view of the one or more cameras in a real-time video communication interface (e.g., 604), the field of view having a fourth representation field of view (e.g., Figure 6H 606 (e.g., different from the third representation field of view; in some embodiments, including a subset of the third representation field of view), the field of view corresponds to a fourth portion of the field of view of the one or more cameras and includes a representation of a fifth object (e.g., 622-1) (e.g., replacing the display of the third representation field of view with a fourth representation field of view that includes a representation of the fifth object (and in some embodiments, including a subset of the third representation field of view)). Stopping the display of the third representation field of view and displaying a fourth representation field of view corresponding to the fourth portion of the field of view and including a representation of the fifth object enhances the video communication session experience by automatically adjusting the camera's field of view to maintain the display of the object as it moves to different positions in the scene without additional input from the user. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0281] In some implementations, adjusting the representation of the field of view of one or more cameras (e.g., 606) during a real-time video communication session based on detected changes in the scene (e.g., 615) within the field of view (e.g., 602) of one or more cameras (e.g., 606) also includes stopping the display of the representation of the field of view of the one or more cameras with a third representation field of view in the real-time video communication interface (e.g., ...). Figure 6F In some embodiments, when auto view mode is enabled and the computer system (e.g., 600) detects that an object (e.g., 622) has moved out of the viewfinder (e.g., 606) (e.g., out of the displayed representation field of view), the computer system switches to a different viewfinder that includes the object (e.g., a representation of a portion of the camera's field of view). In some embodiments, clipping (e.g., jump cut) includes a change in zoom level (e.g., zoom in or zoom out). For example, a change in the representation field of view is a jump cut, which includes a zoomed-out view including the user (e.g., when in single-person tracking mode). In some embodiments, clipping (e.g., match clipping) includes a change from a first area displaying the camera's field of view to a second area displaying the camera's field of view that includes the user but does not include the first area. In some embodiments, when a second object remains in the viewfinder after the first object has moved out of the viewfinder, the computer system displays a jump cut to a zoomed view of the second object (e.g., when auto view mode is enabled).
[0282] In some implementations, before detecting a change in the scene within the field of view of one or more cameras, the representation of the field of view of those cameras (e.g., 606) has a first zoom value (e.g., zoom setting (e.g., 1x, 0.5x, 0.7x)) (e.g., as shown in the image). Figure 6I (As depicted in the image). In some embodiments, changes in the scene (e.g., 615) within the field of view (e.g., 620) of the one or more cameras (e.g., 602) include a sixth object (e.g., 622) moving from a first position within the field of view of the one or more cameras (e.g., ...). Figure 6I (625) to a second position within the field of view of the one or more cameras (e.g., Figure 6JThe movement of 625 in the first position corresponds to the representation of the field of view (e.g., represented by it) and is a first distance from the one or more cameras, and the second position corresponds to the representation of the field of view (e.g., represented by it) and is a threshold distance from the one or more cameras (e.g., an object moves within the displayed viewfinder (e.g., toward the camera; away from the camera) to a predetermined distance from the one or more cameras). In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on detected changes in the scene in the field of view of the one or more cameras includes: displaying in the real-time video communication interface (e.g., 604) a representation of the field of view of the one or more cameras having a second zoom value different from the first zoom value (e.g., ...). Figure 6J (e.g., zooming in; zooming out; in some embodiments, including the entire portion of the field of view of the one or more cameras that was previously displayed in a representation of the field of view with a first zoom value but is instead displayed in a second zoom value) (e.g., jumping from the first zoom level to the second zoom level when an object moves within the viewfinder of the original display to a predetermined distance from the one or more cameras). Stopping the display of the representation of the field of view with the first zoom value and displaying a representation of the field of view with the second zoom value enhances the video communication session experience by automatically adjusting the zoom value of the camera's field of view representation to maintain the object's prominence without additional user input as the object moves through the scene at different distances from the camera. Performing operations when a set of conditions have been met without additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0283] In some implementations, adjusting the representation of the field of view of one or more cameras (e.g., 606) during a real-time video communication session based on detected changes in the scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602) includes stopping at the real-time video communication interface (e.g., ...). Figure 6IIn 606, a representation of the field of view of one or more cameras with a first zoom value is displayed. In some embodiments, as an object (e.g., 622) moves toward the camera at a first threshold distance from the camera, the representation of the field of view of the one or more cameras is switched to a second zoom value, which is a zoom-out view (e.g., a jump cut to a wide-angle view) of a previously displayed portion of the field of view of the one or more cameras. In some embodiments, as an object moves away from the camera at a second threshold distance from the camera, the representation of the field of view of the one or more cameras is switched to a second zoom value, which is a zoom-in view of a previously displayed portion of the field of view of the one or more cameras.
[0284] In some embodiments, the computer system (e.g., 600) simultaneously displays a second selectable graphical user interface object (e.g., 610) and a real-time video communication interface (e.g., 604), the real-time video communication interface including one or more other selectable controls for controlling the real-time video communication (e.g., 608) (e.g., an end call button for ending a real-time video communication session, a switch camera button for switching which camera is used in the real-time video communication session, a mute button for mute / unmute the audio of a user on a device in the real-time video communication session, an effects button for adding / removing visual effects to the real-time video communication session, an add user button for adding a user to the real-time video communication session, and / or a camera on / off button for turning a user's video on / off in the real-time video communication session). In some embodiments, the second selectable graphical user interface object (e.g., a "viewing mode" indicator) is continuously displayed during the real-time video communication session.
[0285] In some implementations, when a real-time video communication interface (e.g., 604) is displayed when a seventh object (e.g., 628) (e.g., a first participant in a real-time video communication session) is detected in a scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602), a computer system (e.g., 600) detects an eighth object (e.g., 622) (e.g., a second participant in a real-time video communication session) (e.g., an increase in the number of objects in the scene) within the field of view of the one or more cameras. In response to the detection of an eighth object in the scene within the field of view of the one or more cameras, the computer system displays a prompt (e.g., 632) via a display generation component (e.g., 601) (e.g., text indicating the addition of a second participant to the real-time video communication session, an indication for adding a second participant, an indication for the second participant (e.g., a framing indication in a potential preview (e.g., a blurred area) for indicating the identification of the detected additional object (e.g., 630)), a stacked camera preview window, or something similar to "About..." Figures 8A to 8ROther tips discussed) to adjust the representation of the field of view of the one or more cameras to include the representation of the eighth object (e.g., 622-1) in the real-time video communication interface (e.g., 604) (e.g., as...). Figure 6Q (As depicted in the text). In response to the detection of an eighth object in the scene, a prompt is displayed to adjust the representation of the field of view of one or more cameras to include the representation of the eighth object in the real-time video communication interface. This provides the user of the computer system with feedback that an additional object has been detected in the field of view of the one or more cameras, and reduces the amount of user input at the computer system by providing the option to automatically adjust the representation of the field of view to include the additional object without requiring the user to navigate to settings menus or other additional interfaces. Providing improved feedback and reducing the amount of input at the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0286] In some implementations, before detecting a change in the scene (e.g., 615) within the field of view (e.g., 620) of one or more cameras (e.g., 602), the representation of the field of view of those one or more cameras (e.g., 606) has a fifth representation of the field of view (e.g., in...). Figure 6K (e.g., in the middle) (e.g., a computer system (e.g., 600) is displaying a portion of the field of view of the one or more cameras before detecting a change in the scene). In some embodiments, changes in the scene within the field of view of the one or more cameras include movement of one or more objects (e.g., 628) detected in the scene. In some embodiments, adjusting the representation of the field of view of the one or more cameras during a real-time video communication session based on detected changes in the scene within the field of view of the one or more cameras includes: displaying a sixth representation field of view (e.g., a non-zero movement threshold) in the real-time video communication interface (e.g., 604) based on determining that the one or more objects have movement less than a threshold amount within at least a threshold time amount (e.g., a predetermined time amount (e.g., one second, two seconds, three seconds)). Figure 6LThe representation of the field of view of the one or more cameras (e.g., different from the fifth field of view) is adjusted after the one or more objects have moved less than a threshold amount within a predetermined time period. Based on the determination that the one or more objects have not moved less than a threshold amount for at least a threshold time period, the representation of the field of view of the one or more cameras with the fifth field of view is continued to be displayed in the real-time video communication interface (e.g., until the one or more objects have moved less than a threshold amount within at least a threshold time period) (e.g., maintaining the original viewfinder while one or more of these objects move). Based on whether one or more objects detected in the scene have moved less than a threshold amount for at least a threshold time period, the representation of the field of view of one or more cameras is selectively adjusted from the fifth to the sixth field of view during the real-time video communication session. This enhances the video communication session experience by automatically adjusting the camera's field of view when an attached object enters the scene with the intention of participating in the real-time video communication session, rather than adjusting the field of view when an object enters the scene with the intention of not participating. This also reduces the amount of computation performed by the computer system by eliminating irrelevant adjustments to the represented field of view whenever the number of participants in the scene changes. Performing operations when a set of conditions are met without requiring additional user input and reducing the amount of computation performed by the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system). This, in turn, reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently. In some implementations, when automatic framing mode is enabled, the computer system maintains the original display view until one or more of the objects remain stationary.
[0287] In some implementations, a computer system (e.g., 600, 600a) displays, via a display generation component (e.g., 601, 601a), a representation of a first portion of the field of view of one or more cameras (e.g., 606, 1006, 1056, 1208, 1218) of a corresponding device of a relevant participant in a real-time video communication session (e.g., a portion of the field of view of a remote participant's camera in the real-time video communication session, which includes the detected face of the remote participant (e.g., ...). Figure 12L Video feed 1220-1b; video feed 1210-1 including a portion of John's face; video feed 1023 including a portion of John's face; video feed 1053-2 including a portion of Jane's face) (e.g., the field of view of one or more cameras of a computer system including a portion of the face of a detected object (e.g., Figure 12LIn 1208-2; camera preview 1218 including a portion of John's face; camera preview 1006 including a portion of Jane's face; camera preview 1056 including a portion of John's face) and a representation of the second portion of the field of view of one or more cameras of the respective participants' respective devices (e.g., the field of view of a remote participant's camera excluding the detected portion of the remote participant's face (e.g., ... Figure 12L 1220-1a; video feed 1210-1 does not include a portion of John's face; video feed 1023 does not include a portion of John's face; video feed 1053-2 does not include a portion of Jane's face (e.g., the field of view of one or more cameras in a computer system does not include a portion of the face of a detected object (e.g., Figure 12L In 1208-1; camera preview 1218 excluding a portion of John's face; camera preview 1006 excluding a portion of Jane's face; camera preview 1056 excluding a portion of John's face), including determining that the first part of the field of view of the one or more cameras includes a corresponding type of detected feature (e.g., a face; multiple different faces) while the corresponding type of detected feature is not detected in the second part of the field of view of the one or more cameras of the corresponding participant's corresponding device (e.g., when a face (or multiple different faces) is detected in the first part of the field of view but not in the second part of the field of view, the second part is compressed to a greater extent than the first part (e.g., by the transmitting device (e.g., the device of the remote participant (e.g., 600a, 600); the computer system of the object (e.g., 600, 600a))), such that when a face is detected in the first part but not in the second part, the first part of the field of view can be displayed with a higher video quality than the second part of the field of view (e.g., at the receiving device (e.g., the computer system; the device of the remote participant)). Figure 12LIn the above, 1220-1b is displayed as having a higher video quality than 1220-1a; the portion of video feed 1210-1 including John's face is displayed as having a higher image quality than the portion of video feed 1210-1 excluding John's face; the portion of video feed 1053-2 including Jane's face is displayed as having a higher video quality than the portion of video feed 1053-2 excluding Jane's face; the portion of video feed 1023 including John's face is displayed as having a higher video quality than the portion of video feed 1023 excluding John's face), with a reduced compression (e.g., higher video quality) compared to the representation of the second portion of the field of view of one or more cameras of the corresponding participant's corresponding device, to display the representation of the first portion of the field of view of one or more cameras of the corresponding participant's corresponding device. Based on determining that a first portion of the field of view of one or more cameras includes detected features of the corresponding type, while a second portion of the field of view of the one or more cameras does not contain detected features of the corresponding type, a representation of the first portion of the field of view of one or more cameras of a corresponding participant in a real-time video communication session is displayed with a reduced compression level compared to the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device. This practice saves computational resources by conserving bandwidth and reducing the amount of image data processed for display and / or transmission in high image quality. The saved computational resources enhance the operability of the computer system and make the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently.
[0288] In some embodiments, the computer system (e.g., 600, 600a) enables a dynamic video quality mode for output (e.g., transmitted to a receiving device (e.g., 600a, 600), optionally simultaneously displayed at the sending device (e.g., 600, 600a)) of camera fields of view with different levels of video compression (e.g., 606, 1006, 1056, 1208, 1210-1, 1218, 1220-1, 1023, 1053-2). In some embodiments, the computer system compresses portions of the camera field of view that do not include one or more faces (e.g., Figure 12L In 1220-1a; video feed 1210-1 excluding a portion of John's face; video feed 1023 excluding a portion of John's face; video feed 1053-2 excluding a portion of Jane's face) the portion of the camera's field of view that includes one or more faces (e.g., Figure 12LVideo feeds 1220-1b, 1210-1 including a portion of John's face, 1023 including a portion of John's face, and 1053-2 including a portion of Jane's face are among the most common. In some embodiments, the computer system optionally displays compressed video feeds in a camera preview. In some embodiments, the computer system transmits video feeds with different levels of compression during a real-time video communication session, such that the receiving device (e.g., a remote participant) is able to display a video feed received from the transmitting device (e.g., the computer system) that has a higher video quality portion displayed simultaneously with a lower video quality portion, wherein the higher video quality portion of the video feed includes the face, and the lower video quality portion of the video feed does not include the face (e.g., 1220-1b is displayed as having a higher video quality portion than the lower video quality portion). Figure 12L In video feed 1220-1a, the portion including John's face is displayed with higher video quality than the portion excluding John's face in video feed 1210-1; in video feed 1053-2, the portion including Jane's face is displayed with higher video quality than the portion excluding Jane's face in video feed 1053-2; in video feed 1023, the portion including John's face is displayed with higher video quality than the portion excluding John's face in video feed 1023. Similarly, in some embodiments, the computer system receives compressed video data from a remote device (e.g., a device of a remote participant in a real-time video communication session) and displays video feeds from the remote device with different levels of compression, such that the video feeds from the remote device can be displayed together with a higher quality portion including the remote participant's face and a lower quality portion excluding the remote participant's face (displayed simultaneously with the higher quality portion) (e.g., 1220-1b is displayed with higher video quality than the portion including John's face in video feed 1220-1a). Figure 12L The video feed 1220-1a has higher video quality; the portion of video feed 1210-1 including John's face is displayed with higher video quality than the portion of video feed 1210-1 excluding John's face; the portion of video feed 1053-2 including Jane's face is displayed with higher video quality than the portion of video feed 1053-2 excluding Jane's face; the portion of video feed 1023 including John's face is displayed with higher video quality than the portion of video feed 1023 excluding John's face. In some embodiments, different degrees of compression can be applied to video feeds that detect multiple faces. For example, the video feed can have multiple higher quality (less compressed) portions, each corresponding to the location of one face among the detected faces.
[0289] In some implementations, the dynamic video quality mode is independent of the auto-framing mode and the background blur mode, allowing the dynamic video quality mode to be enabled and disabled separately from the auto-framing mode and the background blur mode. In some implementations, the dynamic video quality mode is implemented together with the auto-framing mode, such that the dynamic video quality mode is enabled when the auto-framing mode is enabled, and disabled when the auto-framing mode is disabled. In some implementations, the dynamic video quality mode is implemented together with the background blur mode, such that the dynamic video quality mode is enabled when the background blur mode is enabled, and disabled when the background blur mode is disabled.
[0290] In some implementations, after a feature of a corresponding type has moved from a first portion (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras on the corresponding participant's device to a second portion of the field of view of one or more cameras on the corresponding participant's device (e.g., detecting the movement of a feature of the corresponding type from the first portion of the field of view of one or more cameras to the second portion of the field of view of one or more cameras; and, in response to detecting the movement of a feature of the corresponding type from the first portion of the field of view of one or more cameras to the second portion of the field of view of one or more cameras), (Movement of a portion of the field of view of the one or more cameras), the computer system (e.g., 600, 600a) displays, via a display generation component (e.g., 601, 601a), a representation of the first portion of the field of view of the one or more cameras of the corresponding participant's device and a representation of the second portion of the field of view of the one or more cameras of the corresponding participant's device (e.g., a portion of the field of view including the detected face), including determining that the second portion of the field of view of the one or more cameras of the corresponding participant's device includes a corresponding type of detected feature while the corresponding type of detected feature is not detected in the representation of the first portion of the field of view of the one or more cameras of the corresponding participant's device, to display the representation of the first portion of the field of view of the one or more cameras of the corresponding participant's device with an increased compression level (e.g., lower video quality) compared to the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's device (e.g., as the face moves within the field of view of the one or more cameras, the compression level of the corresponding portion of the field of view of the one or more cameras changes, such that the face (e.g., a portion of the field of view including the face) is output with a lower compression level (e.g., preserved) than the portion of the field of view excluding the face. For example, transmission and optional display (e.g., as Jane's face moves, video feeds 1053-2 and / or 1220-1 are updated such that her face continues to be displayed at a higher video quality, and portions of her face not included in the video feed (even portions previously displayed at a higher quality) are displayed at a lower video quality; as John's face moves, video feeds 1023 and / or 1210-1 are updated such that his face continues to be displayed at a higher video quality, and portions of his face not included in the video feed (even portions previously displayed at a higher quality) are displayed at a lower video quality)).After a feature of a corresponding type has been moved from a first portion of the field of view of one or more cameras on a corresponding participant's device to a second portion of the field of view of one or more cameras on a corresponding participant's device, the representation of the first portion of the field of view of one or more cameras on a corresponding participant's device in a real-time video communication session is displayed with an increased degree of compression compared to the representation of the second portion of the field of view of one or more cameras on a corresponding participant's device, based on the determination that the second portion of the field of view of one or more cameras includes the detected feature of the corresponding type while the detected feature of the corresponding type was not detected in the first portion of the field of view of one or more cameras. This practice saves computational resources by saving bandwidth and reducing the amount of image data processed for displaying and / or transmitting in high image quality as faces move within a scene. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently.
[0291] In some implementations, the corresponding type of feature is a face (e.g., a face detected within the field of view of one or more cameras; the face of a remote participant (e.g., Jane's face in video feeds 1220-1 and / or 1053-2; John's face in video feeds 1023 and / or 1210-1); the face of an object (e.g., Jane's face in camera previews 606, 1006 and / or 1208; John's face in camera previews 1056 and / or 1218)). In some implementations, displaying a representation of a second portion of the field of view of one or more cameras on a participant's device (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) includes, based on determining that a first portion of the field of view of the one or more cameras includes a detected face, while the face is not detected in the second portion of the field of view of the one or more cameras on the participant's device, displaying a representation of the second portion of the field of view of the one or more cameras on the participant's device with a reduced video quality compared to the representation of the first portion of the field of view of the one or more cameras on the participant's device (e.g., 1220-1a is displayed with a lower quality than the representation of the first portion of the field of view of the one or more cameras on the participant's device). Figure 12LThe video quality of 1220-1b is lower than that of video feed 1210-1, where the portion excluding John's face is displayed is displayed with a lower video quality than that of video feed 1210-1, where John's face is included; the video quality of 1053-2, where Jane's face is not displayed is displayed with a lower video quality than that of video feed 1053-2, where Jane's face is included; the video quality of 1023, where John's face is not displayed is displayed with a lower video quality than that of video feed 1023, where John's face is included (e.g., due to compression reduction of the representation of the first portion of the field of view of the one or more cameras). Based on determining that the first portion of the field of view includes the detected face while the face is not detected in the second portion of the field of view of the one or more cameras of the corresponding participant's device, the representation of the second portion of the field of view of the one or more cameras of the corresponding participant's device is displayed with a reduced video quality compared to the representation of the first portion of the field of view. This practice saves computational resources by saving bandwidth and reducing the amount of image data processed for display and / or transmission in high image quality. Saving computing resources enhances the operability of computer systems and makes user-system interfaces more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with computer systems), which in turn reduces power consumption and extends the battery life of computer systems by enabling users to use the system more quickly and efficiently.
[0292] In some implementations, portions of the camera's field of view that do not include detected faces (e.g., portions of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that do not include John's face; and portions of 1006 and / or 1053-2 that do not include Jane's face) are output with lower image quality (e.g., transmitted and optionally displayed) than portions of the camera's field of view that include detected faces (e.g., portions of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 that include John's face; and portions of 1006 and / or 1053-2 that include Jane's face) (due to increased compression of portions that do not include detected faces). In some embodiments, when no face is detected in the field of view of one or more cameras, a computer system (e.g., 600, 600a) applies a uniform or substantially uniform degree of compression to a first and a second portion of the field of view of the one or more cameras, such that a video feed with uniform or substantially uniform video quality (e.g., both the first and second portions) can be output. In some embodiments, when multiple faces are detected in the camera's field of view (e.g., multiple participants in a real-time video communication session are detected), the computer system simultaneously applies reduced compression to the portion of the field of view corresponding to the detected faces, such that faces with higher image quality can be displayed simultaneously (e.g., at the receiving device). In some embodiments, the computer system applies increased compression to the representation of a second portion of the field of view of the one or more cameras, even if a face is detected in that second portion. For example, the computer system may determine that the face in the second portion is not a participant in the real-time video communication session (e.g., the person is a bystander in the background), and therefore does not reduce the compression of the second portion containing that face.
[0293] In some implementations, after a change in the bandwidth of the representation of the field of view of one or more cameras of a corresponding device (e.g., 600, 600a) for transmitting data of a corresponding participant occurs (e.g., is detected), when a feature of a corresponding type (e.g., a face) is detected in the first portion of the field of view of the corresponding device of the corresponding participant (e.g., 600, 600a) (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1), and in the second portion of the field of view of the corresponding device of the corresponding participant, a feature of a corresponding type (e.g., a face) is detected (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 including John's face; the portion of 1006 and / or 1053-2 including Jane's face), and in the third portion of the field of view of the corresponding device of the corresponding participant, a feature of a corresponding type (e.g., a face) is detected (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 including Jane's face), and in the third portion of the field of view of the corresponding device of the corresponding participant, a feature of a corresponding type (e.g., a face) is detected in the first portion of the field of view of the corresponding device of the corresponding participant (e.g., 606, 1006, 1023, 1053-2 including Jane's face), and in the second portion of the field of view of the corresponding device of the corresponding participant, a feature of a corresponding type (e.g., a face) is detected in the first When no feature of the corresponding type is detected in the two parts (e.g., the parts of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 excluding John's face; 1006 and / or 1053-2 excluding Jane's face), the compression degree (e.g., amount of compression) of the representation of the first part of the field of view of the corresponding device of the corresponding participant is changed by a smaller amount than the amount of change in the compression degree of the representation of the second part of the field of view of the corresponding device of the corresponding participant (e.g., when a face is detected in the first part of the field of view of the one or more cameras and no face is detected in the second part of the field of view of the one or more cameras, the rate of change of compression (in response to a change in bandwidth (e.g., a decrease in bandwidth)) is smaller for the first part of the field of view than for the second part of the field of view). When a feature of a corresponding type is detected in the first part but not in the second part, the compression degree of the representation of the first part of the field of view of one or more cameras of the corresponding participant is changed by a smaller amount than the change in the compression degree of the representation of the second part. This practice saves computational resources by conserving bandwidth for the representation of the first part of the field of view of one or more cameras that include the corresponding type of feature and by reducing the amount of image data processed for display and / or transmission in high image quality. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently.
[0294] In some implementations, after a change in the bandwidth of the representation of the field of view (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of one or more cameras of a corresponding device (e.g., 600, 600a) used to transmit data of a corresponding participant occurs (e.g., is detected), when no feature of the corresponding type is detected in the first part of the field of view of the one or more cameras of the corresponding participant's corresponding device (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 excluding John's face; the portion of 1006 and / or 1053-2 excluding Jane's face), and in the second part of the field of view of the one or more cameras of the corresponding participant's corresponding device, a feature of the corresponding type is detected. When a feature of the corresponding type is detected in the field of view (e.g., the portion of John's face in 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1; the portion of Jane's face in 1006 and / or 1053-2), the degree of compression (e.g., amount of compression) of the representation of the first portion of the field of view of the corresponding device of the corresponding participant is changed by a greater amount than the amount of change in the degree of compression of the representation of the second portion of the field of view of the corresponding device of the corresponding participant (e.g., when a face is detected in the second portion of the field of view of the one or more cameras, and no face is detected in the first portion of the field of view of the one or more cameras, the rate of compression change of the first portion of the field of view is greater than the rate of compression change of the second portion of the field of view in response to a change in bandwidth (e.g., a decrease in bandwidth)). When a feature of the corresponding type is not detected in the first part but is detected in the second part, the compression degree of the representation of the first part of the field of view of one or more cameras of the corresponding participant is changed by a greater amount than the change in the compression degree of the representation of the second part. This practice saves computational resources by conserving bandwidth for the second part of the representation of the field of view of one or more cameras that include the corresponding type of feature and by reducing the amount of image data processed for display and / or transmission in high image quality. Saving computational resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently.
[0295] In some implementations, in response to a change in the bandwidth of the representation of the field of view (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of one or more cameras of the corresponding device (e.g., 600, 600a) for transmitting the corresponding participant, (e.g., detected), the quality (e.g., video quality) of the representation of the second portion (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 excluding John's face; the portion of 1006 and / or 1053-2 excluding Jane's face) of the field of view of one or more cameras of the corresponding device of the corresponding participant is improved (e.g., due to a change in the amount of video compression). The change in quality of the representation of the second part of the field of view of one or more cameras of the corresponding participant's device is greater than the change in quality of the first part (e.g., the portion including John's face in 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1; the portion including Jane's face in 1006 and / or 1056-2). (In some embodiments, the representation of the first part is not changed in quality or has a nominal change in quality.) (For example, when a face is detected in the first part of the field of view of one or more cameras but not in the second part of the field of view of one or more cameras, the image quality of the second part changes more than the image quality of the first part in response to a change in bandwidth (e.g., a reduction in bandwidth). This practice of changing the quality of the representation of the second part of the field of view of one or more cameras of the corresponding participant's device by a greater amount than the change in the quality of the representation of the first part saves computational resources by conserving bandwidth for the representation of the first part of the field of view of one or more cameras and reducing the amount of image data processed for display and / or transmission in high image quality. Saving computing resources enhances the operability of computer systems and makes user-system interfaces more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with computer systems), which in turn reduces power consumption and extends the battery life of computer systems by enabling users to use the system more quickly and efficiently.
[0296] In some implementations, when a face is detected in a first portion of the field of view of one or more cameras (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 including John's face; 1006 and / or 1053-2 including Jane's face), and not detected in a second portion of the field of view (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 excluding John's face; 1006 and / or 1053-2 excluding Jane's face), a computer system (e.g., 600, 600a) detects a change in available bandwidth (e.g., an increase in bandwidth; a decrease in bandwidth) and, in response, adjusts (e.g., increases, decreases) the compression of the second portion of the representation of the field of view of the one or more cameras, without adjusting the compression of the first portion of the representation of the field of view of the one or more cameras. In some embodiments, when a bandwidth change is detected, the computer system adjusts the compression of the first portion at a lower rate than the adjustment for the second portion. In some embodiments, the method includes detecting (e.g., at the corresponding device of the corresponding participant) a change in the bandwidth used to transmit a representation of the field of view of the one or more cameras of the corresponding participant's corresponding device when a feature of a corresponding type (e.g., a face) is detected in the first portion of the field of view of the one or more cameras of the corresponding participant's corresponding device, and when no feature of the corresponding type is detected in the second portion of the field of view of the one or more cameras of the corresponding participant's corresponding device.
[0297] It should be noted that the above reference method 700 (for example, Figures 7A to 7B The details of the process also apply in a similar manner to the methods 900, 1100, 1300, and 1400 described below. For example, method 900, method 1100, method 1300, and / or method 1400 may optionally include one or more features of the various methods described above with reference to method 700. For the sake of brevity, these details will not be repeated below.
[0298] Figures 8A to 8R Exemplary user interfaces for managing real-time video communication sessions (e.g., video conferencing) according to some implementation schemes are shown. The user interfaces in these figures are used to illustrate the processes described herein, including... Figure 9 The process in.
[0299] Figures 8A to 8R A device 600 is shown that displays a user interface for managing real-time video communication sessions on a display 601, similar to the one described above. Figures 6A to 6Q The subject of discussion. Figures 8A to 8RVarious implementations are described in which the device 600, when in auto-view mode, prompts the user to adjust the displayed portion of the camera's field of view (e.g., camera preview) in response to the detection of another object in the scene. The following section discusses... Figures 8A to 8R One or more of the implementation schemes discussed may be related to... Figures 6A to 6Q , Figures 10A to 10J and Figures 12A to 12U One or more combinations of implementation schemes are discussed in the implementation scheme.
[0300] Figures 8A to 8J An exemplary embodiment is described in which device 600 displays a prompt to adjust the video feed field of view to include the additional participant in response to the detection of an additional object in scene 615. Figures 8A to 8D An implementation scheme is shown in which the prompt includes an option to display a stacked camera preview. Figures 8E to 8G An implementation scheme is shown in which prompts include displaying a camera preview with both shaded and unshaded areas. Figures 8H to 8J The illustration shows an implementation where prompts include options for switching between single-person and multi-person framing modes. Other implementations are also provided (such as those mentioned above). Figure 6P The described implementation includes a prompt that can be selected to adjust the display field of view to include a power indication of an additional object (e.g., adding power indication 632).
[0301] Figure 8A Describing something similar to the above about Figure 6O The implementation scheme discussed differs in that the video feed 623 now includes Pam's representation 623-2 instead of John's representation 623-1. Automatic framing mode is enabled, as indicated by the bold state of the framing mode indicator 610 and the display of the framing indicator 630.
[0302] exist Figure 8A In the video conference, Jack is using device 600 to participate with Pam. Similarly, Pam is using a device with one or more features of devices 100, 300, 500, or 600 to participate with Jack. For example, Pam is using a tablet similar to device 600. Therefore, Pam's device displays a video conference interface similar to video conference interface 604, except that the camera preview on Pam's device displays the video feed captured from Pam's device (currently in...). Figure 8A (As described in video feed 623), and the incoming video feed on Pam's device displays the video feed output from device 600 (currently in...). Figure 8A (As depicted in the camera preview 606).
[0303] exist Figure 8BIn this scenario, device 600 detects Jane 622 entering scene 615 within field of view 620. In response, device 600 updates the video conferencing interface 604 by displaying an auxiliary camera preview 806 located behind and offset from the camera preview 606. The stacked appearance of camera preview 606 and auxiliary camera preview 806 indicates that multiple video feed fields of view are available for the video conferencing, and that the user can change the displayed field of view. Device 600 indicates that camera preview 606 is the currently selected or enabled video feed field of view because it is located on top of auxiliary camera preview 806. Therefore, auxiliary camera preview 806 represents an option for adjusting the video feed field of view, which in this embodiment is an alternative field of view including both Jack 628 and Jane 622 (an additional object that has entered the scene). In some embodiments, different camera previews represent different zoom values, and therefore, different camera preview options can also be considered as different zoom controls / options.
[0304] Device 600 detects input 804 (e.g., a tap) on a stacked preview (e.g., on an auxiliary camera preview 606) and, in response, updates the video conferencing interface 604 by panning the position of the auxiliary camera preview 806 so that it is no longer behind the camera preview 806. Figure 8C As depicted in the text.
[0305] exist Figure 8C In the video conferencing interface 604, device 600 displays both camera preview 606 and auxiliary camera preview 806 separately (without stacking). Device 600 also displays a bold outline 807 to indicate the currently selected video feed field of view. Figure 8C The central image shows camera preview 606. Auxiliary camera preview 806 represents the available video feed field of view provided by device 600. Part 825 represents a portion of the field of view 620 displayed in auxiliary camera preview 806, while part 625 represents the portion of the field of view 620 currently displayed in camera preview 606. Although camera preview 806 shows a rendering of part 825, camera preview 8060 is not currently selected; therefore, device 600 is not currently outputting a view of part 825 of the video conference.
[0306] Camera preview 606 includes a portion of Jack's representation 628-1 and Jane's representation 622-1. Auxiliary camera preview 806 is a zoomed-out view (compared to the view in camera preview 606) including Jack's representation 628-1 and Jane's representation 622-1. As previously discussed, framing indicators 630 are depicted in both camera preview 606 and auxiliary camera preview 806 to indicate that the faces of Jane and Jack have been detected within the respective video feed fields of view.
[0307] When the camera preview option is Figure 8C When the non-stacked configuration is displayed as depicted, device 600 can maintain the availability of preview options or switch between available preview options in response to user input. For example, if device 600 detects input 811 on camera preview 606, the device continues to use (e.g., output) the video feed field of view represented by camera preview 606, and video conferencing interface 604 returns to... Figure 8B The view depicted in the image. If device 600 detects input 812 on auxiliary camera preview 806, device 600 switches (e.g., outputs) to the video feed field of view represented by auxiliary camera preview 806, and the camera preview returns to a stacked configuration where auxiliary camera preview 806 is on top and camera preview 606 is on the bottom, as shown in the image. Figure 8D As depicted in the diagram. In some embodiments, when device 600 switches from camera preview 606 to auxiliary camera preview 806, bold outline 807 moves from camera preview 606 to auxiliary camera preview 806 to indicate a switch from output camera preview 606 to output auxiliary camera preview 806.
[0308] exist Figure 8D In this context, portion 625 represents the portion of the field of view 620 currently being output to the video conference. Because auxiliary camera preview 806 was selected in response to input 812, portion 625 now corresponds to auxiliary camera preview 806. Portion 627 represents the previous video feed field of view, which now corresponds to camera preview 606.
[0309] exist Figure 8E In the scene 615, device 600 detects Jane 622 entering. Device 600 displays a camera preview 606 with an unblurred area 606-1 (represented by boundary 808 and without a shadow line) and a blurred area 606-2 (represented by boundary 809 and with a shadow line). The unblurred area 606-1 represents the current video feed field of view, while the blurred area 606-2 represents an additional video feed field of view that can be used for video conferencing but is not currently being output. Therefore, portion 625 corresponds to the field of view of the unblurred area, and portion 825 corresponds to the combined available field of view of the blurred and unblurred areas. The display of the blurred and unblurred areas in camera preview 606 indicates that the video feed field of view can be adjusted. The use of blur is described as a way to distinguish the current video feed field of view from the additional available video feed field of view. However, these areas can be distinguished by other visual cues and appearances, such as shadows, darkening, highlighting, or other visual masking to emphasize or de-emphasize the individual areas. Boundaries 808 and 809 are also used to visually distinguish these areas.
[0310] exist Figure 8EIn the diagram, the unblurred region 606-1 depicts Jack's unmasked representation 628-1a (specifically, Jack's face), which is being output for video conferencing. The blurred region 606-2 depicts the masked (e.g., blurred) representation of the available video feed field of view, included in portion 825 of the field of view 620 but not included in portion 625. For example, in... Figure 8E In the diagram, blurred region 606-2 depicts the occlusion representation 628-1b of Jack's body. In some embodiments, the blurred and / or unblurred regions include framing indicators when the device 600 detects a face in the corresponding area.
[0311] exist Figure 8F In the middle, Jane 622 has entered part 825 of the field of view 620, and device 600 displays Jane's occlusion representation 622-1b in the blurred area 606-2 of the camera preview. Device 600 detects input 813 on the blurred area 606-2 (or on the viewfinder indicator located around Jane's face in the blurred area 606-2), and in response, adjusts (e.g., expands) the video feed field of view to include the previously blurred area 606-2, as... Figure 8G As depicted in the diagram. In some implementations, the blurred / unblurred areas of the camera preview represent different zoom values in the video feed field of view. Therefore, the camera preview 606, which can be selected to switch to a field of view with different zoom values (e.g., by expanding the unblurred area), can also be considered a zoom control.
[0312] Figure 8G Part 625 in the text represents the extended portion of the field of view 620 that is now being output for video conferencing, and part 627 represents the portion corresponding to... Figure 8F The unblurred portion 606-1 in the field of view 620 is the previously displayed portion. The camera preview 606 now depicts the unmasked representation 628-1 of Jack and the unmasked representation 622-1 of Jane.
[0313] In some implementation schemes, when the conditions for triggering the adjustment are no longer met, the above-mentioned... Figures 8E to 8G The adjustment of the video feed field of view under discussion has been reversed. For example, if Jane leaves... Figure 8G If the viewfinder 625 is selected, then device 600 returns to [the previous state]. Figure 8E The state depicted in the image includes an unblurred area 606-1 and a blurred area 606-2.
[0314] Figures 8H to 8J Various interfaces of the implementation scheme in which the device 600 switches between single-person framing mode and multi-person framing mode settings are described. Figure 8HIn the context of automatic framing mode, device 600 detects Jack 628 in scene 615 and displays video conferencing interface 604, where framing mode option 830 is depicted in camera preview 606. Framing mode option 830 includes single-person option 830-1 and multi-person option 830-2. When single-person option 830-1 is selected, as... Figure 8H As depicted, the single-person framing mode setting is enabled and device 600 keeps the video feed field of view focused on a single user's face, even when another person is detected in the field of view 620. For example, in Figure 8I In this scene, although Jane 622 is now positioned next to Jack 628 in scene 615, device 600 maintains the video feed field of view characterized by Jack 628 (represented in camera preview 606 and section 625) instead of automatically adjusting the video feed field of view to include Jane.
[0315] exist Figure 8I In this context, device 600 detects input 832 on the multi-person option 830-2. In response, device 600 switches from single-person framing mode setting to multi-person framing mode setting. When multi-person framing mode setting is enabled, device 600 automatically adjusts the video feed field of view to include additional objects (or subsets thereof) detected in field of view 620. For example, when device 600 switches to multi-person framing mode setting, device 600 expands the video feed field of view to include representations 628-1 and 622-1 of both Jack and Jane, as shown below. Figure 8J As depicted in camera preview 606. Therefore, Figure 8J Part 625 represents the expanded video feed field of view resulting from enabling the multi-person framing mode setting, and part 627 represents the previous video feed field of view corresponding to the single-person framing mode setting. In some embodiments, framing mode option 830 corresponds to a video feed field of view with different zoom values, and therefore framing mode option 830 can also be regarded as a zoom control / option.
[0316] In some implementation schemes, Figures 8H to 8J The transformation described in the text can be related to, for example, the transformations described in the text. Figures 8E to 8G The discussed combination of camera previews with blurred and unblurred areas. For example, device 600 can display camera preview 606 with blurred and unblurred areas, similar to... Figure 8E The descriptions in the text, but also including similar ones Figure 8H The viewfinder mode option 830 is described herein. When the device 600 is in single-person viewfinder mode setting, the device 600 displays a camera preview 606 with both blurred and unblurred areas, regardless of whether anyone is detected in the blurred portion of the viewfinder (similar to...). Figure 8E and 8F(As depicted in the text). However, when the device 600 is in multi-person viewfinder mode setting, the device 600 can switch the camera preview 606 from a blurred and unblurred appearance to an unblurred appearance in response to detecting another person in the blurred area (similar to...). Figure 8F and Figure 8G (The transition in the middle). In a similar way, if a person is detected in the blurry area when the device 600 switches from single-person viewfinder mode to multi-person viewfinder mode (such as...). Figure 8F As shown in the diagram, device 600 adjusts the video feed field of view to include the previously blurred area, which includes people previously detected in the blurred area (similar to...). Figure 8F and Figure 8G The transformation described herein). In some implementations, device 600 can reverse the above transformation. For example, if device 600 is displaying a camera preview 606 in which both objects are in the field of view (similar to...). Figure 8J (As depicted in the text), and if device 600 detects the selection of single-person framing option 830-1, then device 600 can adjust the video feed field of view to return to a blurred / unblurred appearance, similar to... Figure 8F As depicted in the text.
[0317] Now for reference Figure 8K Device 600 displays a video conferencing interface 834, which depicts an incoming request to join a live video conference with John and two other remote participants. Video conferencing interface 834 is similar to video conferencing interface 604, except that multiple participants are active in the video conferencing session depicted in video conferencing interface 834. Therefore, the implementation described herein with respect to video conferencing interface 604 can be applied in a similar manner to video conferencing interface 834. Similarly, the implementation described herein with respect to video conferencing interface 834 can be applied in a similar manner to video conferencing interface 604, etc.
[0318] exist Figure 8K In the process, device 600 detects input 835 on accept option 608-3, and simultaneously enables automatic viewfinder mode (as indicated by the bold appearance of viewfinder mode indicator 610) and disables background blur mode (as indicated by the non-bold appearance of background blur indicator 611). In response, as... Figure 8L As described, device 600 accepts real-time video conferencing calls and joins video conferencing sessions with automatic framing mode enabled and background blur mode disabled.
[0319] Figure 8LA device 600 is depicted displaying a video conferencing interface 834 having a camera preview 606 (similar to camera preview 836), and incoming video feeds 840-1, 840-2, and 840-3 for each of the corresponding remote participants in a real-time video conferencing session. Camera preview 836 includes a representation 622-1 of Jane and a framing mode indication 610. In some embodiments, the framing mode indication 610 is selectable in camera preview 836 to enable or disable automatic framing mode. In some embodiments, the framing mode indication 610 is not selectable until camera preview 836 is displayed in a magnified state, such as... Figure 8M As depicted in [the document]. In some embodiments, device 600 displays a viewfinder mode power indicator 610 in camera preview 836 when auto viewfinder mode is enabled, and does not display the power indicator when auto viewfinder mode is disabled. In some embodiments, device 600 persistently displays the viewfinder mode power indicator 610 and indicates whether auto viewfinder mode is enabled by changing the appearance of the viewfinder mode power indicator (e.g., by bolding the power indicator when mode is enabled). In some embodiments, the viewfinder mode power indicator 610 is displayed in options menu 608.
[0320] exist Figure 8L In the scenario, Jane is using device 600 to participate in a video conference with Pam, John, and Jack. Similarly, Pam, John, and Jack each use a corresponding device, including one or more features of devices 100, 300, 500, or 600, to participate in a video conference with Jane and other corresponding participants. For example, John, Jack, and Pam each use a tablet similar to device 600. Therefore, the devices of the other participants (John, Jack, and Pam) each display a video conference interface similar to video conference interface 834, except that the camera preview on each corresponding device displays the video feed captured from that user's corresponding device (e.g., Pam's camera preview displays what is currently depicted in video feed 840-3, John's camera preview displays what is currently depicted in video feed 840-1, and Jack's camera preview displays what is currently depicted in video feed 840-2), and the incoming video feeds on the devices of the other participants (John, Jack, and Pam) include the video feed output from device 600 (currently depicted in...). Figure 8L (The content depicted in the camera preview 836) and video feeds output from other participants' devices.
[0321] exist Figure 8L In the process, device 600 detects input 837 on camera preview 836 and, in response, zooms in on camera preview 836, such as... Figure 8M As depicted in [the text]. Figure 8M In the video conferencing interface 834, device 600 displays the viewfinder mode indicator 610 in two locations. Viewfinder mode indicator 610-1 is displayed in the camera preview 836, and viewfinder mode indicator 610-2 is displayed in the options menu 608. In some embodiments, the viewfinder mode indicator 610 is displayed in only one location at any given time (e.g., in the options menu 608 or the camera preview 836).
[0322] In some implementations, when the automatic viewfinder mode is unavailable, the device 600 displays a viewfinder mode indication 610 with a changed appearance. For example, in Figure 8N In scene 615, the lighting conditions are poor, and in response to detecting poor lighting conditions, device 600 displays viewfinder mode indicator 610-1 and viewfinder mode indicator 610-2 with a grayed-out appearance to indicate that the automatic viewfinder mode is currently unavailable. When the lighting conditions improve, device 600 displays a viewfinder mode indicator 610-2 with a grayed-out appearance. Figure 8M The viewfinder modes shown are 610-1 and 610-2.
[0323] exist Figure 8N In the middle, device 600 detects input 839 (e.g., tap input or drag gesture) on option menu 608, and in response, displays Figure 8O The interface depicted in [the document]. In some implementations, device 600 responds to detection in [the context of...]. Figure 8N The input is displayed in other locations on the video conferencing interface 834 depicted therein (e.g., on the camera preview 836 or in other locations within the interface besides the options menu 608). Figure 8L The video conferencing interface depicted in [the document] is 834.
[0324] exist Figure 8O In the device 600, a video conferencing interface 834 is displayed, featuring an extended options menu 845, incoming video feeds 840-2 and 840-3, and a camera preview 836. The extended options menu 845 includes information and various options for video conferencing, including a framing mode option 845-1, which is similar to the framing mode display 610.
[0325] Now for reference Figure 8P The device 600 displays a video conferencing interface 834 with a magnified camera preview 836 and control options 850. In some implementations, in response to Figure 8O The input is displayed on the camera preview 836. Figure 8PThe interface is depicted in the diagram. Control options 850 include viewfinder mode option 850-1, 1X zoom option 850-2, and 0.5X zoom option 850-3. Viewfinder mode option 850-1 is similar to viewfinder mode indicator 610 and is displayed in bold to indicate that auto viewfinder mode is enabled. Because auto viewfinder mode is enabled, camera preview 836 includes viewfinder indicator 852 (similar to viewfinder indicator 630) positioned around Jane's face (indication 622-1). Zoom options 850-2 and 850-3 can be selected to manually change the digital zoom level of the video feed field of view. Because zoom options 850-2 and 850-3 manually adjust the digital zoom of camera preview 836, selecting a zoom option disables auto viewfinder mode, as discussed below.
[0326] exist Figure 8P In the process, device 600 detects input 853 on zoom option 850-2. In response, device 600 emphasizes (e.g., bolds and optionally magnifies) zoom option 850-2 and disables auto-view mode. Figure 8P In the implementation described, the 1X zoom level is the zoom setting prior to detecting input 853. Therefore, device 600 continues to display indication 622-1 at the 1X zoom level. Because the auto-viewing mode is disabled, the viewfinder mode option 850-1 is weakened (e.g., no longer bolded) and the viewfinder indicator 852 is no longer displayed. Figure 8Q middle.
[0327] exist Figure 8Q In the process, device 60 detects input 855 on zoom option 850-3 and, in response, adjusts the digital zoom level, such as... Figure 8R As indicated in the document. Therefore, device 600 emphasizes zoom option 850-3, downplays zoom option 850-2, and displays a camera preview 836 with a 0.5X digital zoom value (compared to...). Figure 8Q (The camera preview 836 in the video feed is zoomed out). Part 625 represents a portion of the field of view 620 that is displayed after the video feed field of view is zoomed out, and part 627 represents a portion of the field of view 620 that was previously displayed when zoom option 850-2 is selected.
[0328] Figure 9This is a flowchart illustrating a method for managing a real-time video communication session using an electronic device, according to some embodiments. Method 900 is performed at a computer system (e.g., a smartphone, tablet computer) (e.g., 100, 300, 500, 600) that communicates with display generation components (e.g., a display controller, a touch-sensitive display system), one or more cameras (e.g., 602) (e.g., a visible light camera, an infrared camera, a depth camera), and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 900 may be optionally combined, some operations may be optionally changed in order, and some operations may be optionally omitted.
[0329] As described below, method 900 provides an intuitive way to manage real-time video communication sessions. This method reduces the cognitive burden on users managing real-time video communication sessions, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to manage real-time video communication sessions faster and more efficiently, saving power and increasing the time interval between battery charges.
[0330] In method 900, the computer system (e.g., 600) displays (902) a real-time video communication interface (e.g., 604, 834) for a real-time video communication session (e.g., an interface for a real-time video communication session, such as a real-time video chat session, a real-time video conferencing session, etc.) via a display generation component (e.g., 601). In some embodiments, the real-time video communication interface includes a real-time preview of the computer system's user and a real-time representation of one or more participants (e.g., remote users) in the real-time video communication session.
[0331] A computer system (e.g., 600) displays a real-time video communication interface (e.g., 604, 834), which includes (904) representations (e.g., 623, 623-1, 840-1, 840-2, 840-3) of one or more participants (e.g., remote participants in the real-time video communication session) in addition to those visible via one or more cameras (e.g., 602) (e.g., 622, 628). In some embodiments, the participants visible via the one or more cameras are objects located within the field of view (e.g., 620) of the one or more cameras and represented (e.g., displayed) in the real-time video communication session (e.g., 606, 806) via a display generation component (e.g., 601).
[0332] A computer system (e.g., 600) displays a real-time video communication interface (e.g., 640, 834), which includes (904) (e.g., simultaneously with the representation of the one or more participants) a representation of the field of view of the one or more cameras (e.g., 606, 806, 836), which is visually associated (e.g., adjacent display; grouped display) with visual indications (e.g., 610, 610-1, 610-2, 630, 632) of options for changing (e.g., adjusting) the representation of the field of view of the one or more cameras during a real-time video communication session (e.g., changing the digital zoom level / value, expanding the display field of view, shrinking the display field of view). 806, 808, 809, 606-1, 606-2, 830, 830-1, 830-2, 845-1, 850, 850-1, 850-2, 850-3, 852) (e.g., prompts (e.g., text), selectable graphical user interface objects (e.g., zoom controls, framing mode indications, framing indicators, indications for selecting single-person framing modes, indications for selecting multi-person framing modes), camera preview representations (e.g., alternative camera previews), framing indicators (e.g., a representation of the field of view of the one or more cameras is displayed with a first digital zoom level and a first display portion of the field of view of the one or more cameras). Displaying a visual indication of the field of view of one or more cameras during a real-time video communication session, associated with visual instructions for changing the representation of the field of view of one or more cameras, provides feedback to the user of the computer system that alternative representations of the field of view of the one or more cameras are available. This reduces the amount of user input at the computer system by providing options for adjusting the representation of the field of view without requiring the user to navigate to settings menus or other additional interfaces. Providing improved feedback and reducing the amount of input at the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system). This, in turn, reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0333] In some embodiments, the representation of the field of view (e.g., 620) of the one or more cameras (e.g., 602) is output by a computer system (e.g., 600) or can be output to a preview (e.g., 606, 806, 836) of image data associated with one or more electronic devices (e.g., remote participants) of one or more participants in a real-time video communication session. In some embodiments, the representation of the field of view of the one or more cameras includes a representation (e.g., 622, 628) of objects participating in the real-time video communication session (e.g., participants; users of the computer system detected within the field of view (e.g., 620) of the one or more cameras during the real-time video communication conference) (e.g., 622-1, 628-1) (e.g., a camera preview of the user of the computer system for the real-time video communication session). In some embodiments, visual indications (e.g., 850-2, 850-3, 806, 606, 632) of options for changing the representation of the field of view of the one or more cameras can be selected to manually adjust the framing (e.g., digital zoom level) of the representation of the field of view of the one or more cameras during the real-time video communication session. In some embodiments, visual indications (e.g., 610, 610-1, 610-2, 845-1, 830-1, 830-2, 850-1) of options for changing the representation of the field of view of one or more cameras can be selected to enable or disable a mode for automatically adjusting the representation of the field of view of the one or more cameras during a real-time video communication session. In some embodiments, the representation of the field of view of the one or more cameras includes the representation of objects and optionally one or more additional objects.
[0334] When a real-time video communication interface (e.g., 604, 834) for a real-time video communication session is displayed, a computer system (e.g., 600) detects (908) a set of one or more inputs (e.g., 626, 634, 804, 811, 812, 813, 832, 850-2, 850-3) via one or more input devices (e.g., 601) that correspond to a request to initiate a process for adjusting (in some embodiments, manually; in some embodiments, automatically (e.g., without user input)) the representa...
Claims
1. A method for managing real-time video communication sessions, comprising: At the computer system that communicates with the display generation component, one or more cameras, and one or more input devices: The display generation component displays a real-time video communication interface for a real-time video communication session, the real-time video communication interface including: The representation of one or more remote participants in the real-time video communication session, in addition to those visible via the one or more cameras; and The representation of the field of view of the one or more cameras, the representation of the field of view being visually associated with a visual indication of an option to change the representation of the field of view of the one or more cameras during the real-time video communication session; When displaying the real-time video communication interface for the real-time video communication session, a set or more inputs are detected via the one or more input devices corresponding to a request to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and In response to the detection of the set of one or more inputs, a process is initiated for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.
2. The method according to claim 1, wherein: The visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session includes a set of one or more controls for adjusting the scaling level of the content displayed in the representation of the field of view of the one or more cameras.
3. The method according to claim 2, further comprising: A first input is detected via the one or more input devices for the set of one or more controls, the set of one or more controls being used to adjust the scaling level of the content displayed in the representation of the field of view of the one or more cameras; In response to detecting the first input to the set of one or more controls, the set of one or more controls is displayed such that the first control option and the second control option are displayed separately; When displaying the set of one or more controls, a second input corresponding to the selection of the first control option or the second control option is detected; as well as In response to the detection of the second input, the scaling level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session is adjusted based on the selection of the first control option or the second control option.
4. The method of claim 2, wherein, The set of one or more controls for adjusting the zoom level of the representation of the field of view of the one or more cameras includes a first zoom control having a first fixed position and a second zoom control having a second fixed position, the method further comprising: Detecting a third input corresponding to the selection of the first zoom control or the second zoom control via the one or more input devices; and While continuing to display the first zoom control with the first fixed position and the second zoom control with the second fixed position, and in response to detecting the third input, the zoom level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session is adjusted based on the selection of the first zoom control or the second zoom control.
5. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Detecting the number of objects within the field of view of the one or more cameras during the real-time video communication session; and The representation of the field of view of the one or more cameras is adjusted based on the number of objects detected in the field of view of the one or more cameras during the real-time video communication session.
6. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Based on one or more characteristics of the scene in the field of view of the one or more cameras, the scaling level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session is adjusted.
7. The method according to claim 1, wherein, The visual indication of the option to change the representation of the field of view of the one or more cameras during the real-time video communication session can be selected to enable a first camera mode or a second camera mode, the method further comprising: When displaying the real-time video communication interface for the real-time video communication session, changes in the scene in the field of view of the one or more cameras are detected, including changes in the number of objects in the scene; In response to detecting a change in the scene within the field of view of the one or more cameras: Based on determining that the first camera mode is enabled, the representation of the field of view of the one or more cameras is adjusted according to the change in the number of objects in the scene; and Based on determining that the second camera mode is enabled, the representation of the field of view of the one or more cameras is abandoned based on the change in the number of objects in the scene.
8. The method according to claim 1, wherein, The real-time video communication interface includes: A first representation of a first participant in the real-time video communication session, the first representation corresponding to a first portion of the field of view of the one or more cameras; and A second representation of a second participant in the real-time video communication session, different from the first participant, the second representation corresponding to a second portion of the field of view of the one or more cameras, wherein the second portion of the field of view of the one or more cameras is different from the first portion of the field of view of the one or more cameras, and wherein the second representation is separate from the first representation.
9. The method according to claim 8, wherein, In response to detecting input on the real-time video communication interface for the real-time video communication session, the first representation of the first participant and the second representation of the second participant are displayed.
10. The method according to claim 1, wherein, The representation of the field of view of the one or more cameras includes a graphical indication of whether the option to change the representation of the field of view of the one or more cameras is enabled during the real-time video communication session.
11. The method of claim 10, further comprising: When displaying a representation of the field of view of the one or more cameras, input for the representation of the field of view of the one or more cameras is detected; as well as In response to the input that detects a representation of the field of view for the one or more cameras, a selectable graphical user interface object is displayed via the display generation component. The selectable graphical user interface object can be selected to enable the option to change the representation of the field of view of the one or more cameras to different framing options for the field of view of the one or more cameras during the real-time video communication session.
12. The method according to claim 10, wherein: When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is available for use, the graphical indication of whether to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session has a first appearance, and When the option to change the representation of the field of view of the one or more cameras during the real-time video communication session is not available, the graphical indication of whether to enable the option to change the representation of the field of view of the one or more cameras during the real-time video communication session has a second appearance different from the first appearance.
13. The method according to claim 1, wherein, The representation of the field of view of the one or more cameras includes: A first display area representing the field of view of the one or more cameras, the first display area corresponding to a first portion of the field of view of the one or more cameras; and A second display area representing the field of view of the one or more cameras, the second display area corresponding to a second portion of the field of view of the one or more cameras that is different from the first portion of the field of view, wherein the second display area is visually distinct from the first display area.
14. The method of claim 13, further comprising: Based on determining that a first portion of the field of view of the one or more cameras is selected for the real-time video communication session, a first display area with a visually unobstructed appearance is displayed and a second display area representing the field of view with an obstructed appearance is displayed. as well as Based on determining that the second portion of the field of view of the one or more cameras is selected for the real-time video communication session, a second display area with a visually unobstructed appearance is displayed.
15. The method according to claim 13, wherein, The representation of the field of view of the one or more cameras includes graphic elements displayed to separate the first display area from the second display area.
16. The method of claim 13, further comprising: One or more indicators of one or more faces are displayed in the second portion of the field of view of the one or more cameras, wherein the one or more indicators can be selected to initiate a process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session.
17. The method of claim 16, wherein: When a face of a first object is detected in the field of view of the one or more cameras and a first portion of the field of view of the one or more cameras is selected for the real-time video communication session, the one or more indicators of the one or more faces in the second portion of the field of view of the one or more cameras are displayed, and In response to detecting one or more faces of an object other than the first object in the second portion of the field of view of the one or more cameras, display one or more indications of the one or more faces.
18. The method of claim 16, further comprising: When the first portion of the field of view of the one or more cameras is selected for the real-time video communication session: Detecting the selection of one or more indications of one or more faces in the second portion of the field of view of the one or more cameras; as well as In response to detecting a selection of one or more of the indicated means, a process is initiated to adjust the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session to include one or more representations of the one or more faces corresponding to the selected indication.
19. The method according to claim 1, wherein: The representation of the field of view of the one or more cameras includes a representation of a first object detected in the field of view of the one or more cameras, and The process of initiating an adjustment of the field of view of the content displayed in a representation of the field of view of the one or more cameras during the real-time video communication session includes: automatically adjusting the field of view of the content displayed in a representation of the field of view of the one or more cameras to include a representation of a second object detected in the field of view of the one or more cameras that satisfies a set of criteria, wherein the representation of the second object is displayed simultaneously with the representation of the first object.
20. The method according to claim 1, further comprising: When displaying a representation of the field of view of the one or more cameras having a first display state, and when detecting the face of an object in a first region of the field of view of the one or more cameras, detecting a second object in a second region of the field of view of the one or more cameras; In response to detecting the second object in the second region of the field of view of the one or more cameras, a second selectable graphical user interface object is displayed; Detect the selection of the second selectable graphical user interface object; as well as In response to detecting the selection of the second selectable graphical user interface object, the representation of the field of view of the one or more cameras is adjusted from having a first display state to having a second display state, the second display state being different from the first display state and including a representation of the second object.
21. The method according to claim 1, further comprising: When a mode for automatically adjusting the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session based on changes in the position of objects detected in the field of view of the one or more cameras is enabled: Detecting the selection of an option for changing the scaling level of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; as well as In response to detecting the selection of the option for changing the zoom level, the mode for automatically adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session is disabled, and the zoom level of the content displayed in the representation of the field of view of the one or more cameras is adjusted.
22. The method according to claim 1, wherein, The process for adjusting the field of view of the content displayed in the representation of the field of view of the one or more cameras during the real-time video communication session includes: Based on one or more conditions detected in the scene within the field of view of the one or more cameras, the content displayed in the representation of the field of view of the one or more cameras is panned during the real-time video communication session.
23. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions that, when executed by the one or more processors, cause the computer system to perform the method according to any one of claims 1 to 22.
24. A computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices, the computer system comprising: One or more processors; as well as A memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions that, when executed by the one or more processors, cause the computer system to perform the method according to any one of claims 1 to 22.
25. A computer system, comprising: Display generated components; One or more cameras; One or more input devices; and Apparatus for performing the method according to any one of claims 1 to 22.
26. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising instructions that, when executed by the one or more processors, cause the computer system to perform the method according to any one of claims 1 to 22.