User interface for wide-angle video conferencing
Through simplified user interface and camera mode selection, the problem of inefficiency of real-time video communication sessions in the prior art is solved, and device efficiency and user satisfaction are improved, especially energy saving in battery-driven devices.
Patent Information
- Application Number
- CN202410924550.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-24
- Filing Date
- 2022-01-28
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-01-28
AI Technical Summary
The user interface for the prior art for managing real-time video communication sessions is complex and inefficient, resulting in wasted user time and device energy, especially in battery-driven devices.
Provides a simplified user interface, which enables quick selection of camera mode and adjusts field of view representation by displaying communication between the generator components and the input device, reduces user operation steps, and improves interface efficiency and power utilization.
Improves the efficiency and user experience of real-time video communication sessions, saving user time and device energy, especially the battery life of battery-driven devices.
Smart Images

Figure CN118890430B_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 202280008837.3, filed on January 28, 2022, and entitled "User Interface for Wide-Angle Video Conferencing."
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Patent Application 63 / 143,881, filed on January 31, 2021, entitled “USER INTERFACES FOR WIDE ANGLE VIDEO CONFERENCE,” U.S. Provisional Patent Application 63 / 176,811, filed on April 19, 2021, entitled “USER INTERFACES FOR WIDE ANGLE VIDEO CONFERENCE,” U.S. Patent Application 17 / 482,977, filed on September 23, 2021, entitled “USER INTERFACES FOR WIDE ANGLE VIDEO CONFERENCE,” U.S. Patent Application 17 / 482,987, filed on September 23, 2021, entitled “USER INTERFACES FOR WIDE ANGLE VIDEO CONFERENCE,” and U.S. Provisional Patent Application 63 / 176,811, filed on April 19, 2021, entitled “USER INTERFACES FOR WIDE ANGLE VIDEO CONFERENCE.” CONFERENCE," the contents of each of which are hereby incorporated by reference herein in their entirety. Technical Field
[0004] The present disclosure relates generally to computer user interfaces and, more particularly, to techniques for managing real-time video communication sessions. Background Art
[0005] The computer system may include hardware and / or software for displaying an interface for a real-time video communication session. Summary of the Invention
[0006] However, some existing techniques for managing real-time video communication sessions using electronic devices are often cumbersome and inefficient. For example, some existing techniques use complex and time-consuming user interfaces that may include multiple button presses or keystrokes. These techniques require more time than necessary, resulting in wasted user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0007] Thus, the present technology provides electronic devices with faster, more efficient methods and interfaces for managing real-time video communication sessions. Such methods and interfaces optionally supplement or replace other methods for managing real-time video communication sessions. Such methods and interfaces reduce the cognitive burden on users and produce a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces conserve power and increase the time between battery charges.
[0008] This article describes an example method. An exemplary method includes, at a computer system in communication with a display generation component, one or more cameras, and one or more input devices: displaying, via the display generation component, a communication request interface, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for the one or more cameras during the real-time video communication session; while displaying the communication request interface, receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; while displaying the real-time video communication interface, detecting a change in a scene in a field of view of the one or more cameras; and in response to detecting the change in the field of view of the one or more cameras: in accordance with a determination that the first camera mode is selected for use, adjusting a representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras; and in accordance with a determination that the second camera mode is selected for use, forgoing adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0009] An example non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices is described herein. The one or more programs include instructions for: displaying, via a display generation component, a communication request interface, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for the one or more cameras during the real-time video communication session; while displaying the communication request interface, receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; while displaying the real-time video communication interface, detecting a change in a scene in a field of view of the one or more cameras; and in response to detecting the change in the field of view of the one or more cameras: in accordance with a determination that the first camera mode is selected for use, adjusting a representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras; and in accordance with a determination that the second camera mode is selected for use, forgoing adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0010] An example non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices is described herein. The one or more programs include instructions for: displaying, via a display generation component, a communication request interface, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for the one or more cameras during the real-time video communication session; while displaying the communication request interface, receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; while displaying the real-time video communication interface, detecting a change in a scene in a field of view of the one or more cameras; and in response to detecting the change in the field of view of the one or more cameras: in accordance with a determination that the first camera mode is selected for use, adjusting a representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras; and in accordance with a determination that the second camera mode is selected for use, forgoing adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0011] An exemplary computer system is described herein. An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying, via a display generation component, a communication request interface, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for the one or more cameras during the real-time video communication session; while displaying the communication request interface, receiving, via one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object; in response to receiving the set of one or more inputs including a selection of the first selectable graphical user interface object, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session; while displaying the real-time video communication interface, detecting a change in a scene in a field of view of the one or more cameras; and in response to detecting the change in the field of view of the one or more cameras: in accordance with a determination that the first camera mode is selected for use, adjusting a representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras; and in accordance with a determination that the second camera mode is selected for use, forgoing adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0012] An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; means for displaying, via the display generation component, a communication request interface, the communication request interface including: a first selectable graphical user interface object associated with a process for joining a real-time video communication session; and a second selectable graphical user interface object associated with a process for selecting between using a first camera mode for one or more cameras and using a second camera mode for one or more cameras during the real-time video communication session; means for receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object when displaying the communication request interface; and means for receiving, via the one or more input devices, a set of one or more inputs including a selection of the first selectable graphical user interface object. Apparatus for displaying a real-time video communication interface for a real-time video communication session via a display generating component in response to receiving the set of one or more inputs including a selection of a first selectable graphical user interface object; apparatus for detecting a change in a scene in the field of view of one or more cameras while displaying the real-time video communication interface; and apparatus for performing the following items in response to detecting a change in the field of view of one or more cameras: adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on a determination that a first camera mode is selected for use; and forgoing adjusting the representation of the field of view of one or more cameras during the real-time video communication session based on a determination that a second camera mode is selected for use.
[0013] An exemplary method includes: at a computer system in communication with a display generation component, one or more cameras, and one or more input devices: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session in addition to participants visible via the one or more cameras; and representations of the field of view of the one or more cameras, the representations being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; and while displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session; and in response to detecting the set of one or more inputs, initiating a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0014] An exemplary non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface comprising simultaneously displaying: representations of one or more participants in the real-time video communication session in addition to participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations visually associated with a visual indication of an option to change the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0015] An exemplary transient computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface comprising simultaneously displaying: representations of one or more participants in the real-time video communication session in addition to participants visible via the one or more cameras; and representations of the fields of view of the one or more cameras, the representations visually associated with a visual indication of an option to change the representations of the fields of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representations of the fields of view of the one or more cameras during the real-time video communication session.
[0016] An exemplary computer system includes: a display generation component; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including simultaneously displaying: representations of one or more participants in the real-time video communication session in addition to participants visible via the one or more cameras; and representations of the field of view of the one or more cameras, the representations visually associated with a visual indication of an option to change the representation of the field of view of the one or more cameras during the real-time video communication session; and, while displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session; and, in response to detecting the set of one or more inputs, initiating a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session.
[0017] An exemplary computer system comprises: a display generation component; one or more cameras; one or more input devices; a device for displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface comprising simultaneously displaying: representations of one or more participants in the real-time video communication session in addition to participants visible via the one or more cameras; and representations of the field of view of the one or more cameras, the representations being visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during the real-time video communication session; a device for detecting, via the one or more input devices, when displaying the real-time video communication interface for the real-time video communication session, a set of one or more inputs corresponding to a request to initiate a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session; and a device for initiating a process for adjusting the representation of the field of view of the one or more cameras during the real-time video communication session in response to detecting the set of one or more inputs.
[0018] An exemplary method includes: at a computer system in communication with a display generation component and one or more cameras: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including one or more representations of a field of view of the one or more cameras; capturing, via the one or more cameras, image data of the real-time video communication session while the real-time video communication session is active; and, based on determining that an amount of separation between a first participant and a second participant satisfies a separation criterion based on the image data of the real-time video communication session captured via the one or more cameras, simultaneously displaying, via the display generation component: a representation of a first portion of the field of view of the one or more cameras at a first area of the real-time video communication interface; and a representation of a second portion of the field of view of the one or more cameras at a second area of the real-time video communication interface that is different from the first area. a representation of a second portion of the field of view of one or more cameras at an area, wherein a representation of the first portion of the field of view of the one or more cameras and a representation of the second portion of the field of view of the one or more cameras are displayed, while a representation of a third portion of the field of view of the one or more cameras between the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras is not displayed; and based on determining that an amount of separation between the first participant and the second participant based on image data of the real-time video communication session captured via the one or more cameras does not meet the separation criteria, displaying a representation of a fourth portion of the field of view of the one or more cameras including the first participant and the second participant via the display generation component, while maintaining display of a portion of the field of view of the one or more cameras between the first participant and the second participant.
[0019] An exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the field of view of the one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and simultaneously displaying via the display generation component: a representation of a first portion of the field of view of the one or more cameras at a first area of the real-time video communication interface, and one or more representations of the field of view of the one or more cameras at a second area of the real-time video communication interface that is different from the first area, based on determining that an amount of separation between a first participant and a second participant satisfies a separation criterion based on the image data of the real-time video communication session captured via the one or more cameras. a representation of a second portion of the field of view of the one or more cameras, wherein a representation of the first portion of the field of view of the one or more cameras and a representation of the second portion of the field of view of the one or more cameras are displayed, without displaying a representation of a third portion of the field of view of the one or more cameras between the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras; and based on determining that an amount of separation between the first participant and the second participant based on image data of the real-time video communication session captured via the one or more cameras does not meet a separation criterion, displaying via the display generation component a representation of a fourth portion of the field of view of the one or more cameras including the first participant and the second participant, while maintaining display of a portion of the field of view of the one or more cameras between the first participant and the second participant.
[0020] An exemplary transient computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the field of view of the one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and, based on determining that an amount of separation between a first participant and a second participant satisfies a separation criterion based on the image data of the real-time video communication session captured via the one or more cameras, simultaneously displaying via the display generation component: a representation of a first portion of the field of view of the one or more cameras at a first area of the real-time video communication interface; and one or more representations of a second portion of the field of view of the one or more cameras at a second area of the real-time video communication interface that is different from the first area. a representation of a second portion of the field of view of the one or more cameras, wherein a representation of the first portion of the field of view of the one or more cameras and a representation of the second portion of the field of view of the one or more cameras are displayed, without displaying a representation of a third portion of the field of view of the one or more cameras between the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras; and based on determining that an amount of separation between the first participant and the second participant based on image data of the real-time video communication session captured via the one or more cameras does not meet a separation criterion, displaying via the display generation component a representation of a fourth portion of the field of view of the one or more cameras including the first participant and the second participant, while maintaining display of a portion of the field of view of the one or more cameras between the first participant and the second participant.
[0021] An exemplary computer system includes: a display generation component; one or more cameras; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for the following operations: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including one or more representations of the field of view of the one or more cameras; capturing image data of the real-time video communication session via the one or more cameras when the real-time video communication session is active; and based on determining that an amount of separation between a first participant and a second participant satisfies a separation criterion based on the image data of the real-time video communication session captured via the one or more cameras, simultaneously displaying via the display generation component: a representation of a first portion of the field of view of the one or more cameras at a first area of the real-time video communication interface; and one or more representations of a second portion of the field of view of the one or more cameras at a second area of the real-time video communication interface that is different from the first area. a representation of a second portion of the field of view of the one or more cameras, wherein a representation of the first portion of the field of view of the one or more cameras and a representation of the second portion of the field of view of the one or more cameras are displayed, without displaying a representation of a third portion of the field of view of the one or more cameras between the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras; and based on determining that an amount of separation between the first participant and the second participant based on image data of the real-time video communication session captured via the one or more cameras does not meet a separation criterion, displaying via the display generation component a representation of a fourth portion of the field of view of the one or more cameras including the first participant and the second participant, while maintaining display of a portion of the field of view of the one or more cameras between the first participant and the second participant.
[0022] An exemplary computer system comprising: a display generating component; one or more cameras; means for displaying, via the display generating component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface comprising one or more representations of a field of view of the one or more cameras; means for capturing, via the one or more cameras, image data of the real-time video communication session when the real-time video communication session is active; means for simultaneously displaying, via the display generating component, a representation of a first portion of the field of view of the one or more cameras at a first area of the real-time video communication interface, and a portion of the field of view of the real-time video communication interface that is different from the first area, based on determining that an amount of separation between a first participant and a second participant satisfies a separation criterion based on the image data of the real-time video communication session captured via the one or more cameras; and a representation of a second portion of the field of view of one or more cameras at a second area of a communication interface, wherein a representation of the first portion of the field of view of the one or more cameras and a representation of the second portion of the field of view of the one or more cameras are displayed, without displaying a representation of a third portion of the field of view of the one or more cameras between the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras; and a device for displaying a representation of a fourth portion of the field of view of the one or more cameras including the first participant and the second participant via a display generating component based on determining that an amount of separation between the first participant and the second participant based on image data of a real-time video communication session captured via the one or more cameras does not meet a separation criterion, while maintaining display of a portion of the field of view of the one or more cameras between the first participant and the second participant.
[0023] According to some embodiments, a method is described for execution at a computer system in communication with one or more output generating components and one or more input devices. The method includes: detecting, via the one or more input devices, a request to display a system interface; in response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of concurrently displayed controls for controlling different system functions of the computer system, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of concurrently displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for the predetermined amount of time, displaying the plurality of concurrently displayed controls without displaying the set of one or more media communication controls; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls; and when the corresponding media communication session has been active for the predetermined amount of time, in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.
[0024] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more output generating components and one or more input devices, the one or more programs including instructions for: detecting a request to display a system interface via the one or more input devices; in response to detecting the request to display the system interface, displaying a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system via the one or more output generating components, including: determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to the media communication session; settings that determine how media is handled by a computer system during a media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying multiple concurrently displayed controls without displaying the set of one or more media communication controls; while displaying a system interface having the set of one or more media communication controls, detecting, via one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.
[0025] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more output generating components and one or more input devices, the one or more programs including instructions for: detecting a request to display a system interface via the one or more input devices; in response to detecting the request to display the system interface, displaying a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system via the one or more output generating components, including: determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to the media communication session; settings that determine how media is handled by a computer system during a media communication session; and based on determining that the media communication session has not been active for a predetermined amount of time, displaying multiple concurrently displayed controls without displaying the set of one or more media communication controls; while displaying a system interface having the set of one or more media communication controls, detecting, via one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls; and when the corresponding media communication session has been active for a predetermined amount of time, in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session.
[0026] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: detecting a request to display a system interface via the one or more input devices; in response to detecting the request to display the system interface, displaying a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system via the one or more output generating components, including: determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls including a set of one or more media communication controls, wherein the media communication controls provide access to the system interface; Accessing media communication settings that determine how media is handled by a computer system during a media communication session; and displaying a plurality of simultaneously displayed controls without displaying the set of one or more media communication controls based on determining that the media communication session has not been active for a predetermined amount of time; detecting, via one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls while displaying a system interface having the set of one or more media communication controls; and adjusting the media communication settings for the corresponding media communication session in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls when the corresponding media communication session has been active for a predetermined amount of time.
[0027] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more input components; a device for detecting a request to display a system interface via the one or more input devices; a device for displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system in response to detecting the request to display the system interface, including: based on determining that a media communication session has been active for a predetermined amount of time, the plurality of simultaneously displayed controls include a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is processed by the computer system during the media communication session; and based on determining that the media communication session has not been active for the predetermined amount of time, displaying the plurality of simultaneously displayed controls without displaying the set of one or more media communication controls; a device for detecting, via the one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls when the system interface having the set of one or more media communication controls is displayed; and a device for adjusting the media communication settings for the corresponding media communication session in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls when the corresponding media communication session has been active for the predetermined amount of time.
[0028] According to some embodiments, a method is described for execution at a computer system in communication with one or more output generation components, one or more cameras, and one or more input devices. The method includes: displaying, via the one or more output generation components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of a field of view of the one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a procedure for changing the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and, while displaying the real-time video communication interface for the real-time video communication session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and, in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than the object displayed in the representation of the field of view of the one or more cameras.
[0029] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs for execution by one or more processors of a computer system in communication with one or more output generating components, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the one or more output generating components, a real-time video communications interface for a real-time video communications session, wherein displaying the real-time video communications interface includes simultaneously displaying: a representation of a field of view of the one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and a representation of one or more participants in the real-time video communications session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and while displaying the real-time video communications interface for the real-time video communications session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than the object displayed in the representation of the field of view of the one or more cameras.
[0030] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs for execution by one or more processors of a computer system in communication with one or more output generating components, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the one or more output generating components, a real-time video communications interface for a real-time video communications session, wherein displaying the real-time video communications interface includes simultaneously displaying: a representation of a field of view of the one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and a representation of one or more participants in the real-time video communications session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; and while displaying the real-time video communications interface for the real-time video communications session, detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than the object displayed in the representation of the field of view of the one or more cameras.
[0031] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more cameras; one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the one or more output generating components, a real-time video communication interface for a real-time video communication session, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of a field of view of one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing an option displayed in addition to the field of view of the one or more cameras during the real-time video communication session. the appearance of a portion of the representation of the field of view of the one or more cameras other than the objects displayed in the representation of the field of view of the one or more cameras; and a representation of one or more participants in a real-time video communications session that is different from the representation of the field of view of the one or more cameras of the computer system; and while displaying a real-time video communications interface for the real-time video communications session, detecting via one or more input devices a set of one or more inputs corresponding to a request to change the appearance of a portion of the representation of the field of view of the one or more cameras other than the objects displayed in the representation of the field of view of the one or more cameras; and in response to detecting the set of one or more inputs, changing the appearance of the portion of the representation of the field of view of the one or more cameras other than the objects displayed in the representation of the field of view of the one or more cameras.
[0032] According to some embodiments, a computer system is described. The computer system includes: one or more output generating components; one or more cameras; one or more input devices; and means for displaying a real-time video communication interface for a real-time video communication session via the one or more output generating components, wherein displaying the real-time video communication interface includes simultaneously displaying: a representation of a field of view of the one or more cameras of the computer system, wherein the representation of the field of view of the one or more cameras is visually associated with an indication of an option for initiating a process for changing the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras during the real-time video communication session; and a representation of one or more participants in the real-time video communication session, the representation being different from the representation of the field of view of the one or more cameras of the computer system; means for detecting, via the one or more input devices, a set of one or more inputs corresponding to a request to change the appearance of a portion of the representation of the field of view of the one or more cameras other than an object displayed in the representation of the field of view of the one or more cameras, while displaying the real-time video communication interface for the real-time video communication session; and means for changing the appearance of the portion of the representation of the field of view of the one or more cameras other than the object displayed in the representation of the field of view of the one or more cameras in response to detecting the set of one or more inputs.
[0033] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are optionally included in a transient computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0034] Thus, a faster, more efficient method and interface for managing real-time video communication sessions is provided for devices, thereby improving the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces may supplement or replace other methods for managing real-time video communication sessions. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.
[0036] Figure 1A is a block diagram illustrating a portable multifunction device with a touch-sensitive display according to some embodiments.
[0037] Figure 1B is a block diagram illustrating example components for event handling according to some embodiments.
[0038] Figure 2 A portable multifunction device with a touch screen according to some embodiments is shown.
[0039] Figure 3 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments.
[0040] Figure 4A An exemplary user interface for a menu of applications on a portable multifunction device is shown according to some embodiments.
[0041] Figure 4B An exemplary user interface is shown for a multifunction device having a touch-sensitive surface that is separate from the display according to some embodiments.
[0042] Figure 5A A personal electronic device according to some embodiments is shown.
[0043] Figure 5B is a block diagram illustrating a personal electronic device according to some embodiments.
[0044] Figure 5C An exemplary diagram of a communication session between electronic devices according to some embodiments is shown.
[0045] Figures 6A to 6Q An exemplary user interface for managing a real-time video communication session is shown according to some embodiments.
[0046] 7A to 7B Depicted is a flow diagram illustrating a method for managing a real-time video communication session according to some embodiments.
[0047] Figures 8A to 8R An exemplary user interface for managing a real-time video communication session is shown according to some embodiments.
[0048] Figure 9 is a flow chart illustrating a method for managing a real-time video communication session according to some embodiments.
[0049] 10A to 10J An exemplary user interface for managing a real-time video communication session is shown according to some embodiments.
[0050] Figure 11 is a flow chart illustrating a method for managing a real-time video communication session according to some embodiments.
[0051] Figures 12A to 12U An exemplary user interface for managing a real-time video communication session is shown according to some embodiments.
[0052] Figure 13is a flow chart illustrating a method for managing a real-time video communication session according to some embodiments.
[0053] Figure 14 is a flow chart illustrating a method for managing a real-time video communication session according to some embodiments. DETAILED DESCRIPTION
[0054] The following description sets forth exemplary methods, parameters, etc. However, it should be recognized that such description is not intended to limit the scope of the present disclosure, but is provided as a description of exemplary embodiments.
[0055] There is a need for electronic devices that provide efficient methods and interfaces for managing real-time video communication sessions. Such techniques can reduce the cognitive burden on users participating in video communication sessions, thereby increasing productivity. Furthermore, such techniques can reduce processor and battery power that would otherwise be wasted on redundant user input.
[0056] under, Figure 1A to Figure 1B 、 Figure 2 、 Figure 3 、 Figures 4A to 4B and Figures 5A to 5C A description of an exemplary device for performing techniques for managing a real-time video communication session is provided. Figures 6A to 6Q An exemplary user interface for managing a real-time video communication session is shown. 7A to 7B Depicted is a flow chart illustrating a method of managing a real-time video communication session according to some embodiments. Figures 6A to 6Q The user interface in the diagram is used to illustrate the process described below, including 7A to 7B in the process. Figures 8A to 8R An exemplary user interface for managing a real-time video communication session is shown. Figure 9 is a flow chart illustrating a method of managing a real-time video communication session according to some embodiments. Figures 8A to 8R The user interface in is used to illustrate the processes described below, which include Figure 9 in the process. 10A to 10J An exemplary user interface for managing a real-time video communication session is shown. Figure 11 is a flow chart illustrating a method of managing a real-time video communication session according to some embodiments. Figures 10A-10J The user interface is used to show the Figure 11 The process is described below. Figures 12A to 12U An exemplary user interface for managing a real-time video communication session is shown. Figure 13 and Figure 14 is a flow chart illustrating a method of managing a real-time video communication session according to some embodiments. Figures 12A to 12U The user interface in the Figure 13 and Figure 14 The process is described below.
[0057] In addition, in the method described herein where one or more steps depend on having met one or more conditions, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions of the steps in the method of determining the method have been met in different repetitions of the method. For example, if the method needs to perform the first step (if the condition is met), and perform the second step (if the condition is not met), then those of ordinary skill will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on having met one or more conditions can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require a system or computer-readable medium to declare that the system or computer-readable medium includes instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether a possible situation has been met without explicitly repeating the steps of the method until all conditions of the steps in the method of determining the method have been met. Those of ordinary skill in the art will also understand that, similar to the method with a contingent step, a system or computer-readable storage medium can repeat the steps of the method as needed multiple times to ensure that all contingent steps have been performed.
[0058] Although the following description uses the terms "first," "second," etc. to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch can be named a second touch and similarly, a second touch can be named a first touch without departing from the scope of the various described embodiments. In some embodiments, the first touch and the second touch are two separate references to the same touch. In some embodiments, both the first touch and the second touch are touches, but they are not the same touch.
[0059] The terms used in the description of the various embodiments described herein are only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms "a" and "the" are intended to also include plural forms unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" used herein refer to and encompass any and all possible combinations of one or more items in the associated listed items. It will also be understood that the terms "includes," "including," "comprises," and / or "comprising" when used in this specification specify the presence of stated features, integers, steps, operations, elements, and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their groupings.
[0060] The term "if" is optionally interpreted to mean "when," "upon," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined that," or "if [stated condition or event] is detected" are optionally interpreted to mean "upon determining," or "in response to determining," or "upon detecting [stated condition or event]," or "in response to detecting [stated condition or event]," depending on the context.
[0061] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as a PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, but are not limited to, the Apple Watch from Apple Inc. (Cupertino, California). Devices, iPod equipment, and Device. Optionally, other portable electronic devices are used, such as laptops or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or touchpads). In some embodiments, the electronic device is a computer system that communicates with the display generation component (e.g., via wireless communication, via wired communication). The display generation component is configured to provide visual output, such as display via a CRT display, display via an LED display, or display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separated from the computer system. As used herein, "display" content includes transmitting data (e.g., image data or video data) to an integrated or external display generation component via a wired or wireless connection to visually generate content to display content (e.g., video data rendered or decoded by display controller 156).
[0062] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick.
[0063] The device typically supports a variety of applications, such as one or more of the following: a drawing application, a rendering application, a word processing application, a website creation application, a disk editing application, a spreadsheet application, a gaming application, a telephony application, a video conferencing application, an email application, an instant messaging application, a fitness support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0064] Various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed on the device are optionally adjusted and / or varied for different applications and / or adjusted and / or varied within the respective applications. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally supports the various applications with a user interface that is intuitive and clear to the user.
[0065] Attention is now turned to embodiments of portable devices having touch-sensitive displays. Figure 1A1 is a block diagram illustrating a portable multifunction device 100 with a touch-sensitive display system 112 according to some embodiments. Touch-sensitive display 112 is sometimes referred to as a "touch screen" for convenience, and is sometimes referred to as or referred to as a "touch-sensitive display system." Device 100 includes memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and external ports 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact force sensors 165 for detecting the intensity of contacts on device 100 (e.g., a touch-sensitive surface, such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 for generating tactile output on device 100 (e.g., generating tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touch pad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0066] As used in this specification and claims, the term "intensity" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or to a surrogate (surrogate) for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some implementations, the force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size of the contact area detected on the touch-sensitive surface and / or its change, the capacitance of the touch-sensitive surface near the contact and / or its change, and / or the resistance of the touch-sensitive surface near the contact and / or its change are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the surrogate measurement of the contact force or pressure is used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measurement). In some embodiments, the surrogate measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of the contact as an attribute of the user input allows the user to access additional device functionality that would otherwise be inaccessible to the user on a smaller device with limited real estate, which is used to display an indication (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).
[0067] As used in this specification and claims, the term "tactile output" refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of a device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device that will be detected by a user using the user's sense of touch. For example, when a device or a component of the device is in contact with a surface that is touch-sensitive to a user (e.g., a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation that corresponds to a perceived change in a physical characteristic of the device or component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a "press click" or "release click" on a physical actuation button. In some cases, the user will feel a tactile sensation, such as a "press click" or "release click," even when the physical actuation button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when the smoothness of the touch-sensitive surface does not change, movement of the touch-sensitive surface may optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface. While such a user's interpretation of touch will be limited by the user's individualized sensory perceptions, many sensory perceptions of touch are common to most users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "press click," "release click," "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or a component thereof that would generate that sensory perception for a typical (or average) user.
[0068] It should be understood that device 100 is merely one example of a portable multifunction device and that device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. Figure 1A The various components shown in the EMBODIMENTS 100 are implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0069] Memory 102 optionally includes high-speed random access memory and optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.
[0070] The peripheral device interface 118 can be used to couple the input peripheral devices and output peripheral devices of the device to the CPU 120 and the memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in the memory 102 to perform various functions of the device 100 and process data. In some embodiments, the peripheral device interface 118, the CPU 120, and the memory controller 122 are optionally implemented on a single chip such as the chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0071] RF (radio frequency) circuitry 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuitry 108 converts electrical signals into / from electromagnetic signals and communicates with communication networks and other communication devices via the electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, memory, and the like. RF circuitry 108 optionally communicates with networks and other devices via wireless communications, such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). RF circuitry 108 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via a short-range communication radio. The wireless communication optionally uses any of a variety of communication standards, protocols and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), Evolution, Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11d), IEEE 802.11e, IEEE 802.11f, IEEE 802.11g, IEEE 802.11g), IEEE 802.11f, IEEE 802.11g ... 11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this document.
[0072] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts the audio data into electrical signals, and transmits the electrical signals to the speaker 111. The speaker 111 converts the electrical signals into sound waves audible to humans. The audio circuit 110 also receives electrical signals converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signals into audio data and transmits the audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., Figure 2 The headset jack provides an interface between the audio circuit 110 and a removable audio input / output peripheral device, such as an output-only headset or a headset with both output (e.g., a single or dual-ear headset) and input (e.g., a microphone).
[0073] The I / O subsystem 106 couples input / output peripherals on the device 100, such as the touch screen 112 and other input control devices 116, to a peripherals interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a tactile feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive / send electrical signals from / to other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some embodiments, the input controller 160 is optionally coupled to any (or none) of the following: a keyboard, an infrared port, a USB port, and a pointing device such as a mouse. One or more buttons (e.g., Figure 2 208) optionally includes an up / down button for volume control of the speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., Figure 2206 in). In some embodiments, the electronic device is a computer system that communicates (e.g., via wireless communication, via wired communication) with one or more input devices. In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body through the air, including movement of the user's body relative to an absolute reference (for example, the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (for example, movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (for example, a tap gesture involving moving the hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of a part of the user's body).
[0074] A quick press of the push button optionally releases the lock on the touch screen 112 or optionally initiates the process of unlocking the device using gestures on the touch screen, as described in U.S. patent application Ser. No. 11 / 322,549, filed Dec. 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. A long press of the push button (e.g., 206) optionally turns the device 100 on or off. The functions of one or more buttons are optionally user-customizable. The touch screen 112 is used to implement virtual or soft buttons and one or more soft keyboards.
[0075] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch screen 112 and / or sends electrical signals to the touch screen 112. The touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0076] The touch screen 112 has a touch-sensitive surface, sensor, or sensor group that accepts input from the user based on tactile and / or haptic contact. The touch screen 112 and display controller 156 (together with any associated modules and / or instruction sets in memory 102) detect contact on the touch screen 112 (and any movement or interruption of that contact) and convert the detected contact into interaction with a user interface object (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.
[0077] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies are used in other embodiments. The touch screen 112 and display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112 to detect contact and any movement or interruption thereof. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as in the Apple ® from Apple Inc. (Cupertino, California). and iPod The technology used in
[0078] The touch-sensitive display in some embodiments of the touch screen 112 is optionally similar to the multi-touch-sensitive touchpads described in the following U.S. Patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.), and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, the touch screen 112 displays visual output from the device 100, whereas a touch-sensitive touchpad does not provide visual output.
[0079] The touch-sensitive display in some embodiments of the touch screen 112 is described in the following applications: (1) U.S. patent application 11 / 381,313, filed May 2, 2006, “Multipoint Touch Surface Controller”; (2) U.S. patent application 10 / 840,862, filed May 6, 2004, “Multipoint Touchscreen”; (3) U.S. patent application 10 / 903,964, filed July 30, 2004, “Gestures For Touch Sensitive Input Devices”; (4) U.S. patent application 11 / 048,264, filed January 31, 2005, “Gestures For Touch Sensitive Input Devices”; (5) U.S. patent application 11 / 038,590, filed January 18, 2005, “Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices”; (6) U.S. patent application 11 / 228,758, filed September 16, 2005, “Virtual Input Device Placement On A TouchScreen User Interface”; (7) U.S. patent application 11 / 228,700, filed September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. patent application 11 / 228,737, filed September 16, 2005, “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”; and (9) U.S. patent application 11 / 367,749, filed March 3, 2006, “Multi-Functional Hand-Held Device”. All of these applications are incorporated herein by reference in their entirety.
[0080] The touch screen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user optionally uses any suitable object or appendage, such as a stylus, a finger, or the like, to contact the touch screen 112. In some embodiments, the user interface is designed to work primarily through finger-based contacts and gestures, which may not be as precise as stylus-based input due to the larger contact area of a finger on the touch screen. In some embodiments, the device converts rough finger-based input into precise pointer / cursor positions or commands for performing the user's desired actions.
[0081] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad is optionally a touch-sensitive surface that is separate from the touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.
[0082] Device 100 also includes a power system 162 for powering the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, power fault detection circuitry, a power converter or inverter, a power status indicator (e.g., a light emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0083] Device 100 optionally also includes one or more optical sensors 164 . Figure 1AAn optical sensor is shown coupled to the optical sensor controller 158 in the I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or complementary metal oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light from the environment projected through one or more lenses and converts the light into data representing an image. In conjunction with the imaging module 143 (also called a camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, facing away from the touch screen display 112 on the front of the device, enabling the touch screen display to be used as a viewfinder for still and / or video image acquisition. In some embodiments, the optical sensor is located on the front of the device, allowing the user to optionally capture an image of the user for video conferencing while viewing other video conference participants on the touch screen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor in the device housing), allowing a single optical sensor 164 to be used with the touch screen display for both video conferencing and still and / or video image acquisition.
[0084] Device 100 optionally also includes one or more depth camera sensors 175 . Figure 1A A depth camera sensor is shown coupled to a depth camera controller 169 in the I / O subsystem 106. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of objects (e.g., faces) within the scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with the imaging module 143 (also referred to as a camera module), the depth camera sensor 175 is optionally used to determine a depth map for different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located on the front of the device 100, enabling the user to optionally capture an image of the user with depth information for video conferencing while viewing other video conference participants on the touchscreen display, and to capture selfies with depth map data. In some embodiments, the depth camera sensor 175 is located on the rear of the device, or on both the rear and front of the device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor in the device housing), enabling the depth camera sensor 175 to be used in conjunction with the touchscreen display for both video conferencing and still and / or video image acquisition.
[0085] In some embodiments, a depth map (e.g., a depth map image) includes information (e.g., values) related to the distance of objects in a scene from a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines the position of its corresponding two-dimensional pixel in the Z axis of the viewpoint. In some embodiments, a depth map is composed of pixels, where each pixel is defined by a value (e.g., 0 to 255). For example, a value of "0" represents the pixel that is farthest from the viewpoint (e.g., a camera, an optical sensor, a depth camera sensor) in a "three-dimensional" scene, and a value of "255" represents the pixel that is closest to the viewpoint in a "three-dimensional" scene. In other embodiments, a depth map represents the distance between objects in a scene and the plane of the viewpoint. In some embodiments, a depth map includes information about the relative depths of various features of an object of interest in the field of view of a depth camera (e.g., the relative depths of the eyes, nose, mouth, ears of a user's face). In some embodiments, a depth map includes information that enables the device to determine the outline of an object of interest in the z direction.
[0086] Device 100 optionally also includes one or more contact intensity sensors 165 . Figure 1A A contact force sensor is shown coupled to force sensor controller 159 in I / O subsystem 106. Contact force sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other force sensors (e.g., sensors for measuring the force (or pressure) of a contact on a touch-sensitive surface). Contact force sensor 165 receives contact force information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact force sensor is juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact force sensor is located on the back of device 100, opposite touch screen display 112 located on the front of device 100.
[0087] Device 100 optionally also includes one or more proximity sensors 166 . Figure 1AA proximity sensor 166 is shown coupled to the peripherals interface 118. Alternatively, the proximity sensor 166 is optionally coupled to the input controller 160 in the I / O subsystem 106. The proximity sensor 166 is optionally implemented as described in the following U.S. patent applications: No. 11 / 241,839, entitled “Proximity Detector In Handheld Device”; No. 11 / 240,788, entitled “Proximity Detector In Handheld Device”; No. 11 / 620,702, entitled “Using Ambient Light Sensor To Augment Proximity Sensor Output”; No. 11 / 586,862, entitled “Automated Response To And Sensing Of User Activity In Portable Devices”; and No. 11 / 638,251, entitled “Methods And Systems For Automatic Configuration Of Peripherals,” which are hereby incorporated by reference in their entireties. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is on a phone call), the proximity sensor turns off and disables the touch screen 112.
[0088] Device 100 optionally also includes one or more tactile output generators 167 . Figure 1AA tactile output generator is shown coupled to a tactile feedback controller 161 in the I / O subsystem 106. The tactile output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components for converting electrical signals into tactile outputs on the device). The contact force sensor 165 receives tactile feedback generation instructions from the tactile feedback module 133 and generates tactile outputs on the device 100 that can be felt by the user of the device 100. In some embodiments, at least one tactile output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112) and optionally generates tactile outputs by moving the touch-sensitive surface vertically (e.g., inward / outward toward the surface of the device 100) or laterally (e.g., back and forth in the same plane as the surface of the device 100). In some embodiments, at least one tactile output generator sensor is located on the back of the device 100, opposite the touch screen display 112 located on the front of the device 100.
[0089] Device 100 optionally also includes one or more accelerometers 168 . Figure 1A An accelerometer 168 is shown coupled to the peripherals interface 118. Alternatively, the accelerometer 168 is optionally coupled to the input controller 160 in the I / O subsystem 106. The accelerometer 168 optionally implements as described in the following U.S. Patent Publication No. 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication No. 20060017692, entitled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on analysis of data received from one or more accelerometers. The device 100 optionally includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver in addition to the accelerometer 168 for obtaining information about the position and orientation (e.g., portrait or landscape) of the device 100.
[0090] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a global positioning system (GPS) module (or instruction set) 135, and an application (or instruction set) 136. In addition, in some embodiments, memory 102 ( Figure 1A ) or 370( Figure 3 ) storage device / global internal state 157, such as Figure 1A and Figure 3 . The device / global internal state 157 includes one or more of the following: active application state, which indicates which application (if any) is currently active; display state, which indicates what applications, views, or other information occupy various areas of the touch screen display 112; sensor state, which includes information obtained from the device's various sensors and input control devices 116; and position information relating to the device's position and / or posture.
[0091] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware components and software components.
[0092] The communication module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are suitable for coupling directly to other devices or indirectly through a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external ports are connected to (trademark of Apple Inc.) devices.
[0093] The contact / motion module 130 optionally detects contact with the touch screen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger down event), determining the strength of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there has been movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has ceased (e.g., detecting a finger up event or contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of a contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to a single point of contact (e.g., a single-finger contact) or multiple points of contact simultaneously (e.g., "multi-touch" / multiple-finger contact). In some embodiments, the contact / motion module 130 and display controller 156 detect contact on the touchpad.
[0094] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has "clicked" an icon). In some embodiments, at least a subset of the intensity thresholds are determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and can be adjusted without changing the physical hardware of the device 100). For example, a mouse "click" threshold for a touchpad or touchscreen can be set to any one of a large range of predefined thresholds without changing the touchpad or touchscreen display hardware. In addition, in some embodiments, a software setting is provided to the user of the device for adjusting one or more intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by utilizing a system-level click on an "intensity" parameter to adjust multiple intensity thresholds at once).
[0095] Contact / motion module 130 optionally detects gesture input by the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of the detected contacts). Thus, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and then detecting a finger lift (lift-off) event.
[0096] The graphics module 132 includes various known software components for rendering and displaying graphics on the touch screen 112 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual attributes) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0097] In some embodiments, the graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes specifying the graphics to be displayed from an application program or the like, along with coordinate data and other graphic attribute data, if necessary, and then generates screen image data for output to the display controller 156.
[0098] Haptic feedback module 133 includes various software components for generating instructions used by tactile output generator 167 to produce tactile output at one or more locations on device 100 in response to user interaction with device 100 .
[0099] Text input module 134, which is optionally a component of graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).
[0100] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for use in location-based dialing; to the camera 143 as picture / video metadata; and to applications that provide location-based services, such as a weather widget, a local yellow pages widget, and a map / navigation widget).
[0101] Application 136 optionally includes the following modules (or instruction sets), or a subset or superset thereof:
[0102] ● Contacts module 137 (sometimes called address book or contact list);
[0103] ● Telephone module 138;
[0104] ● Video conferencing module 139;
[0105] ● Email client module 140;
[0106] Instant messaging (IM) module 141;
[0107] ●Fitness support module 142;
[0108] A camera module 143 for still and / or video images;
[0109] Image management module 144;
[0110] ●Video player module;
[0111] Music player module;
[0112] ●Browser module 147;
[0113] ●Calendar module 148;
[0114] A widget module 149 , which optionally includes one or more of the following: a weather widget 149 - 1 , a stock market widget 149 - 2 , a calculator widget 149 - 3 , an alarm clock widget 149 - 4 , a dictionary widget 149 - 5 , and other widgets acquired by the user, and a user-created widget 149 - 6 ;
[0115] A widget creator module 150 for forming user-created widgets 149-6;
[0116] ●Search module 151;
[0117] ● Video and music player module 152, which merges the video player module and the music player module;
[0118] ●Note module 153;
[0119] ● Map module 154; and / or
[0120] ●Online video module 155.
[0121] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice replication.
[0122] In combination with the touch screen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the contacts module 137 is optionally used to manage an address book or contact list (e.g., stored in the application internal state 192 of the contacts module 137 in memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0123] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, phone module 138 is optionally used to enter a character sequence corresponding to a phone number, access one or more phone numbers in contacts module 137, modify an entered phone number, dial the corresponding phone number, conduct a conversation, and disconnect or hang up when the conversation is complete. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0124] In combination with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the optical sensor 164, the optical sensor controller 158, the contact / motion module 130, the graphics module 132, the text input module 134, the contact module 137 and the telephone module 138, the video conferencing module 139 includes executable instructions for initiating, conducting and terminating a video conference between a user and one or more other participants in accordance with user instructions.
[0125] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. In conjunction with image management module 144, email client module 140 makes it very easy to create and send emails with still images or video images captured by camera module 143.
[0126] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134, the instant messaging module 141 includes executable instructions for entering a character sequence corresponding to an instant message, modifying previously entered characters, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0127] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, the GPS module 135, the map module 154, and the music player module, the fitness support module 142 includes executable instructions for creating a workout (e.g., with time, distance, and / or calorie burn goals); communicating with fitness sensors (sports equipment); receiving fitness sensor data; calibrating sensors for monitoring fitness; selecting and playing music for a workout; and displaying, storing, and transmitting fitness data.
[0128] In combination with the touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132 and image management module 144, the camera module 143 includes executable instructions for the following operations: capturing still images or videos (including video streams) and storing them in the memory 102, modifying the characteristics of the still images or videos, or deleting the still images or videos from the memory 102.
[0129] In conjunction with touch screen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143, image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0130] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132 and the text input module 134, the browser module 147 includes executable instructions for browsing the Internet in accordance with user instructions, including searching for, linking to, receiving and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.
[0131] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132, the text input module 134, the email client module 140 and the browser module 147, the calendar module 148 includes executable instructions for creating, displaying, modifying and storing a calendar and data associated with the calendar (e.g., calendar entries, to-do items, etc.) in accordance with user instructions.
[0132] In conjunction with the RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, the widget module 149 is a mini-application that is optionally downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or created by a user (e.g., user-created widget 149-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., a Yahoo! widget).
[0133] In combination with the RF circuitry 108, touch screen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, the widget creator module 150 is optionally used by a user to create widgets (e.g., converting a user-specified portion of a web page into a widget).
[0134] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132 and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.
[0135] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch screen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (trademark of Apple Inc.).
[0136] In conjunction with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the note module 153 includes executable instructions for creating and managing notes, to-do lists, etc. according to user instructions.
[0137] In combination with the RF circuitry 108, touch screen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, the map module 154 is optionally used to receive, display, modify, and store maps and data associated with the maps (e.g., driving directions, data relating to stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.
[0138] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, email client module 140, and browser module 147, online video module 155 includes instructions for allowing a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on the touch screen or on an external display connected via external port 124), send an email with a link to a particular online video, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, instant messaging module 141 is used instead of email client module 140 to send a link to a particular online video. Additional descriptions of online video applications can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed on June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed on December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are hereby incorporated by reference in their entirety.
[0139] Each of the modules and applications described above corresponds to an executable instruction set for performing one or more of the functions described above and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) do not have to be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, a video player module is optionally combined with a music player module into a single module (e.g., Figure 1A In some embodiments, the memory 102 optionally stores a subset of the above modules and data structures. In addition, the memory 102 optionally stores additional modules and data structures not described above.
[0140] In some embodiments, device 100 is a device where operation of a predefined set of functions on the device is performed exclusively through a touch screen and / or a touchpad. By using a touch screen and / or a touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on device 100 is optionally reduced.
[0141] A predefined set of functions that are exclusively performed through the touch screen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by the user, navigates the device 100 from any user interface displayed on the device 100 to a main menu, home menu, or root menu. In such embodiments, the touchpad is used to implement a "menu button." In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touchpad.
[0142] Figure 1B is a block diagram illustrating exemplary components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370( Figure 3 ) includes an event classifier 170 (e.g., in the operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 137 to 151, 155, 380 to 390).
[0143] Event classifier 170 receives event information and determines the application 136-1 and the application view 191 of application 136-1 to which the event information is to be delivered. Event classifier 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, application 136-1 includes an application internal state 192 that indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0144] In some embodiments, the application internal state 192 includes additional information, such as one or more of the following: resumption information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by the application 136-1, a state queue for enabling the user to return to a previous state or view of the application 136-1, and a redo / undo queue of previous actions taken by the user.
[0145] Event monitor 171 receives event information from peripherals interface 118. The event information includes information about sub-events (e.g., a user touch on touch-sensitive display 112 as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via audio circuit 110). The information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display 112 or a touch-sensitive surface.
[0146] In some embodiments, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other embodiments, peripheral device interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or receiving an input for more than a predetermined duration).
[0147] In some embodiments, the event classifier 170 also includes a hit view determination module 172 and / or an active event identifier determination module 173.
[0148] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides software procedures for determining where within one or more views a sub-event has occurred. A view consists of controls and other elements that a user can see on the display.
[0149] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of the respective application) in which a touch is detected optionally correspond to programmatic levels within the application's programmatic or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as a hit view, and the set of events that are recognized as correct input is optionally determined based at least in part on the hit view of the initial touch that started the touch-based gesture.
[0150] Hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy where the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once a hit view is identified by hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.
[0151] Active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with one particular view, views higher in the hierarchy will still remain actively participating views.
[0152] Event dispatcher module 174 dispatches event information to event recognizers (e.g., event recognizer 180). In embodiments that include active event recognizer determination module 173, event dispatcher module 174 delivers the event information to the event recognizer determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores the event information in an event queue, which is retrieved by corresponding event receiver 182.
[0153] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or part of another module stored in memory 102, such as contact / motion module 130.
[0154] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events that occur within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes multiple event recognizers 180. In other embodiments, one or more of event recognizers 180 is part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handler 190 optionally utilizes or calls data updater 176, object updater 177, or GUI updater 178 to update the application's internal state 192. Alternatively, one or more of the application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0155] A corresponding event identifier 180 receives event information (e.g., event data 179) from event classifier 170 and identifies an event based on the event information. Event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event identifier 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0156] The event receiver 182 receives event information from the event classifier 170. The event information includes information about sub-events such as touches or touch movements. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device posture).
[0157] The event comparator 184 compares the event information with a predefined event or sub-event definition and determines the event or sub-event based on the comparison, or determines or updates the state of the event or sub-event. In some embodiments, the event comparator 184 includes an event definition 186. The event definition 186 includes the definition of an event (e.g., a predefined sequence of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, the sub-events in event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, the definition of event 1 (187-1) is a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off (touch end) of a predetermined duration. In another example, the definition of event 2 (187-2) is a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on a displayed object, movement of the touch on the touch-sensitive display 112, and lifting of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0158] In some embodiments, event definition 187 includes definitions of events for corresponding user interface objects. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects an event handler that is associated with the sub-event and the object that triggered the hit test.
[0159] In some embodiments, the definition of the corresponding event (187) also includes a delay action that delays the delivery of the event information until it has been determined that the sub-event sequence does or does not correspond to the event type of the event identifier.
[0160] When a corresponding event recognizer 180 determines that a sequence of sub-events does not match any event in event definitions 186, the corresponding event recognizer 180 enters the event impossible, event failed, or event ended state, after which subsequent sub-events of the touch-based gesture are ignored. In this case, other event recognizers (if any) that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.
[0161] In some embodiments, corresponding event recognizers 180 include metadata 183 with configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery for actively participating event recognizers. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact or can interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0162] In some embodiments, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferred sending) the sub-events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with the flag obtains the flag and performs a predefined process.
[0163] In some embodiments, event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and executes a predetermined process.
[0164] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video player module. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the location of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on a touch-sensitive display.
[0165] In some embodiments, event handler 190 includes or has access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0166] It should be understood that the above discussion of event handling for user touches on a touch-sensitive display also applies to other forms of user input utilizing input devices to operate the multifunction device 100, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses, optionally in conjunction with single or multiple keyboard presses or holddowns; contact movement on a touchpad, such as taps, drags, scrolls, etc.; stylus input; movement of the device; spoken commands; detected eye movement; biometric input; and / or any combination thereof, are optionally used as input corresponding to sub-events defining the event to be recognized.
[0167] Figure 2A portable multifunction device 100 with a touch screen 112 is shown according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment and other embodiments described below, a user can select one or more of the graphics by, for example, making gestures on the graphics using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of the one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or rolling of the finger that has made contact with the device 100 (from right to left, from left to right, up and / or down). In some embodiments or in some cases, inadvertent contact with a graphic does not select the graphic. For example, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.
[0168] The device 100 optionally also includes one or more physical buttons, such as a "home" or menu button 204. As previously described, the menu button 204 is optionally used to navigate to any application 136 in a set of applications that are optionally executed on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen 112.
[0169] In some embodiments, the device 100 includes a touch screen 112, a menu button 204, a push button 206 for turning the device on / off and for locking the device, one or more volume adjustment buttons 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and holding it in the depressed state for a predefined time interval; to lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In an alternative embodiment, the device 100 also accepts voice input for activating or deactivating certain functions via the microphone 113. The device 100 also optionally includes one or more contact force sensors 165 for detecting the intensity of contact on the touch screen 112, and / or one or more tactile output generators 167 for generating tactile output for the user of the device 100.
[0170] Figure 33 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments. The device 300 does not have to be portable. In some embodiments, the device 300 is a laptop, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a gaming system, or a control device (e.g., a home controller or an industrial controller). The device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, a memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuits (sometimes referred to as a chipset) that interconnect system components and control communications between system components. The device 300 includes an input / output (I / O) interface 330 having a display 340, which is typically a touch screen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 for generating tactile output on the device 300 (e.g., similar to the above referenced device). Figure 1A The tactile output generator 167 described above), sensor 359 (e.g., optical sensor, acceleration sensor, proximity sensor, touch sensor and / or contact intensity sensor (similar to the above reference Figure 1A The memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 370 optionally includes one or more storage devices located remotely from the CPU 310. In some embodiments, the memory 370 stores data related to the portable multifunction device 100 ( Figure 1A ) or a subset thereof. In addition, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while portable multifunction device 100( Figure 1A )'s memory 102 optionally does not store these modules.
[0171] Figure 3Each element in the above-mentioned elements in is optionally stored in one or more memory devices of the memory device mentioned previously.Each module in the above-mentioned modules corresponds to the instruction set for performing the above-mentioned functions.Above-mentioned modules or computer programs (for example, instruction sets or including instructions) need not be realized with independent software programs (such as computer programs (for example, including instructions)), processes or modules, and therefore the various subsets of these modules are optionally combined or otherwise rearranged in various embodiments.In some embodiments, memory 370 optionally stores the subset of above-mentioned modules and data structures.In addition, memory 370 optionally stores additional modules and data structures not described above.
[0172] Attention is now turned to an embodiment of a user interface, optionally implemented on, for example, portable multifunction device 100 .
[0173] Figure 4A An exemplary user interface for an applications menu on portable multifunction device 100 is shown according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof:
[0174] Signal strength indicators 402 for wireless communications such as cellular and Wi-Fi signals;
[0175] ●Time 404;
[0176] Bluetooth indicator 405;
[0177] Battery status indicator 406;
[0178] A tray 408 with icons for commonly used applications, such as:
[0179] o An icon 416 labeled “Phone” for the phone module 138 , which optionally includes an indicator 414 of the number of missed calls or voicemails;
[0180] o An icon 418 of the email client module 140 labeled “Mail,” which optionally includes an indicator 410 of the number of unread emails;
[0181] o An icon 420 labeled "Browser" of the browser module 147; and
[0182] o An icon 422 labeled “iPod” for the video and music player module 152 (also referred to as the iPod (trademark of Apple Inc.) module 152); and
[0183] Icons for other applications, such as:
[0184] o Icon 424 labeled “Messages” of IM module 141;
[0185] o Icon 426 labeled “Calendar” of calendar module 148;
[0186] o Icon 428 labeled “Photos” of the image management module 144;
[0187] o An icon 430 labeled “Camera” of the camera module 143;
[0188] o Icon 432 labeled “Online Video” of the online video module 155;
[0189] o Icon 434 labeled “Stock Market” of the stock market widget 149 - 2 ;
[0190] o Icon 436 labeled “Map” of the map module 154;
[0191] o Icon 438 labeled “Weather” of weather widget 149 - 1 ;
[0192] o Icon 440 labeled “Clock” of the alarm clock widget 149 - 4 ;
[0193] o An icon 442 labeled “Fitness Support” of the fitness support module 142 ;
[0194] o An icon 444 labeled "Notes" of the notes module 153; and
[0195] o An icon 446 of a settings application or module labeled “Settings” that provides access to settings for the device 100 and its various applications 136 .
[0196] It should be pointed out that Figure 4A The icon labels shown are exemplary only. For example, icon 422 for video and music player module 152 is labeled "Music" or "Music Player." Other labels are optionally used for various application icons. In some embodiments, the label for a respective application icon includes the name of the application corresponding to the respective application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to the particular application icon.
[0197] Figure 4B A touch-sensitive surface 451 (eg, touch screen display 112) is shown having a touch-sensitive surface 451 (eg, touch screen display 112) that is separate from a display 450 (eg, touch screen display 112). Figure 3 tablet or touchpad 355) of the device (e.g., Figure 3Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting intensity of contacts on touch-sensitive surface 451 and / or one or more tactile output generators 357 for generating tactile output for a user of device 300.
[0198] Although some of the examples below will be given with reference to input on a touch screen display 112 (where the touch-sensitive surface and display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, such as Figure 4B In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a main axis (e.g., Figure 4B 453) corresponding to the main axis (for example, Figure 4B According to these embodiments, the device detects a position corresponding to a corresponding position on the display (e.g., Figure 4B , 460 corresponds to 468 and 462 corresponds to 470 ) at contact with touch-sensitive surface 451 (e.g., Figure 4B 460 and 462 in FIG. 4. Thus, when the touch-sensitive surface (e.g., Figure 4B 451) and a display of a multi-function device (e.g., Figure 4B When the user interface 450 in FIG. 1 is separated, the user input detected by the device on the touch-sensitive surface (e.g., contacts 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0199] In addition, although the following examples are primarily given with reference to finger inputs (e.g., finger contacts, single-finger tap gestures, finger swipe gestures), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). For another example, a tap gesture is optionally replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., instead of detecting the contact, followed by ceasing to detect the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or mice and finger contacts are optionally used simultaneously.
[0200] Figure 5AAn exemplary personal electronic device 500 is shown. The device 500 includes a body 502. In some embodiments, the device 500 may include a body 502 relative to the devices 100 and 300 (e.g., Figures 1A to 4B ) some or all of the features described in . In some embodiments, device 500 has a touch-sensitive display screen 504, referred to hereinafter as touch screen 504. As an alternative to or in addition to touch screen 504, device 500 has a display and a touch-sensitive surface. As with devices 100 and 300, in some embodiments, touch screen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of applied contact (e.g., touch). The one or more intensity sensors of touch screen 504 (or touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of device 500 can respond to touches based on the intensity of the touch, which means that touches of different intensities can invoke different user interface operations on device 500.
[0201] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed on May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed on November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.
[0202] In some embodiments, the device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) can be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, the device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) can allow the device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watchbands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear the device 500.
[0203] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, the device 500 may include a reference Figure 1A 、 Figure 1B and Figure 3 Some or all of the components described. Device 500 has a bus 512 that operatively couples an I / O portion 514 to one or more computer processors 516 and a memory 518. The I / O portion 514 can be connected to a display 504, which can have a touch-sensitive component 522 and optionally a strength sensor 524 (e.g., a contact strength sensor). In addition, the I / O portion 514 can be connected to a communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, the input mechanism 506 is optionally a rotatable input device or a depressible input device and a rotatable input device. In some examples, the input mechanism 508 is optionally a button.
[0204] In some examples, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to I / O portion 514.
[0205] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform the techniques described below, including processes 700, 900, 1100, 1300, and 1400 ( Figure 7A 、 Figure 7B 、 Figure 9 、 Figure 11 、 Figure 13 and Figure 14). Computer-readable storage media can be any medium that can tangibly contain or store computer-executable instructions for use by or in conjunction with instruction execution systems, devices, and apparatuses. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. Personal electronic device 500 is not limited to Figure 5B components and configurations, but may include other components or additional components in a variety of configurations.
[0206] As used herein, the term "indicator" refers to an indication that is optionally provided on device 100, 300, and / or 500 ( Figure 1A 、 Figure 3 and Figures 5A to 5C ) is a user-interactive graphical user interface object displayed on a display screen of a computer. For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) optionally each constitute an affordance.
[0207] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 3 Touchpad 355 or Figure 4B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 451 in FIG, the particular user interface element is adjusted according to the detected input. In the case of a touch screen display (e.g., a touch screen display) that enables direct interaction with user interface elements on the touch screen display Figure 1A touch-sensitive display system 112 or Figure 4AIn some implementations of the touch screen 112 in FIG, 2 , a contact detected on the touch screen acts as a “focus selector” such that when input (e.g., a press input by the contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted according to the detected input. In some implementations, the focus moves from one area of the user interface to another area of the user interface without corresponding movement of a cursor or movement of a contact on the touch screen display (e.g., by using a tab key or arrow keys to move the focus from one button to another); in these implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on the touch screen display) that is controlled by the user to deliver the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touch screen), the position of a focus selector (e.g., a cursor, contact, or selection box) over a corresponding button will indicate that the user intends to activate the corresponding button (rather than other user interface elements shown on the device display).
[0208] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected during a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted off, before or after contact begins to move, before contact ends, before or after contact is detected to increase in intensity, and / or before or after contact is detected to decrease in intensity). The characteristic intensity of a contact is optionally based on one or more of the following: the maximum value of the intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the value at the top 10% of the intensity of the contact, the half-maximum value of the intensity of the contact, the 90% maximum value of the intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is the average value of the intensity of the contact over time). In some embodiments, the feature strength is compared to a set of one or more strength thresholds to determine whether the user has performed an operation. For example, the set of one or more strength thresholds optionally includes a first strength threshold and a second strength threshold. In this example, a contact whose feature strength does not exceed the first threshold results in a first operation, a contact whose feature strength exceeds the first strength threshold but does not exceed the second strength threshold results in a second operation, and a contact whose feature strength exceeds the second threshold results in a third operation. In some embodiments, a comparison between the feature strength and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to abandon the corresponding operation) rather than to determine whether to perform the first operation or the second operation.
[0209] In some embodiments, a portion of a gesture is identified for determining the characteristic strength. For example, the touch-sensitive surface optionally receives a continuous swipe contact that transitions from a starting position and reaches an end position where the contact strength increases. In this example, the characteristic strength of the contact at the end position is optionally based only on a portion of the continuous swipe contact, rather than the entire swipe contact (e.g., only the portion of the swipe contact at the end position). In some embodiments, a smoothing algorithm is optionally applied to the intensity of the swipe contact before determining the characteristic strength of the contact. For example, the smoothing algorithm optionally includes one or more of the following: an unweighted sliding average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow peaks or dips in the intensity of the swipe contact to achieve the purpose of determining the characteristic strength.
[0210] The intensity of the contact on the touch-sensitive surface is optionally characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device will perform an operation typically associated with clicking a button of a physical mouse or trackpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device will perform an operation different from the operation typically associated with clicking a button of a physical mouse or trackpad. In some embodiments, when a contact is detected with a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold, contacts below the nominal contact detection intensity threshold are no longer detected), the device will move the focus selector in accordance with the movement of the contact on the touch-sensitive surface, without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally speaking, unless otherwise stated, these intensity thresholds are consistent between different groups of user interface illustrations.
[0211] An increase in contact feature intensity from an intensity below a light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact feature intensity from an intensity below a deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact feature intensity from an intensity below a contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact feature intensity from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a contact lifted from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.
[0212] In some embodiments described herein, one or more operations are performed in response to detecting a gesture that includes a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or multiple contacts), wherein the corresponding press input is detected at least in part based on detecting an increase in the intensity of the contact (or multiple contacts) to above a press input intensity threshold. In some embodiments, the corresponding operation is performed in response to detecting an increase in the intensity of the corresponding contact to above the press input intensity threshold (e.g., a "down stroke" of the corresponding press input). In some embodiments, the press input includes an increase in the intensity of the corresponding contact to above the press input intensity threshold and a subsequent decrease in the intensity of the contact to below the press input intensity threshold, and the corresponding operation is performed in response to detecting a subsequent decrease in the intensity of the corresponding contact to below the press input threshold (e.g., an "up stroke" of the corresponding press input).
[0213] In some embodiments, the device employs intensity hysteresis to avoid unexpected inputs, sometimes referred to as "jitter," where the device defines or selects a hysteresis intensity threshold that has a predefined relationship to a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable proportion of the press input intensity threshold). Thus, in some embodiments, a press input includes an increase in the intensity of the corresponding contact to above the press input intensity threshold and a subsequent decrease in the intensity of the contact to below a hysteresis intensity threshold corresponding to the press input intensity threshold, and a corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases to below the hysteresis intensity threshold (e.g., an "upstroke" of the corresponding press input). Similarly, in some embodiments, a press input is detected only when the device detects that the contact intensity increases from an intensity equal to or below the hysteresis intensity threshold to an intensity equal to or above the press input intensity threshold and, optionally, that the contact intensity subsequently decreases to an intensity equal to or below the hysteresis intensity, and a corresponding operation is performed in response to detecting the press input (e.g., an increase in contact intensity or a decrease in contact intensity, depending on the circumstances).
[0214] For ease of explanation, a description of an operation performed in response to a press input associated with a press input intensity threshold or in response to a gesture including a press input is optionally triggered in response to detecting any of the following: contact intensity increasing above the press input intensity threshold, contact intensity increasing from an intensity below a hysteresis intensity threshold to an intensity above the press input intensity threshold, contact intensity decreasing below the press input intensity threshold, and / or contact intensity decreasing below a hysteresis intensity threshold corresponding to the press input intensity threshold. Additionally, in examples where an operation is described as being performed in response to detecting that the intensity of the contact decreases below the press input intensity threshold, the operation is optionally performed in response to detecting that the intensity of the contact decreases below a hysteresis intensity threshold that corresponds to and is less than the press input intensity threshold.
[0215] Attention is now turned to an embodiment of a user interface ("UI") and associated processes implemented on an electronic device, such as portable multifunction device 100, device 300, or device 500.
[0216] Figure 5CAn exemplary diagram of a communication session between electronic devices 500A, 500B, and 500C is depicted. Devices 500A, 500B, and 500C are similar to electronic device 500, and each device shares one or more data connections 510 (such as an internet connection, a Wi-Fi connection, a cellular connection, a short-range communication connection, and / or any other such data connection or network) with each other to facilitate real-time communication of audio data and / or video data between the respective devices for a period of time. In some embodiments, the exemplary communication session may include a shared data session, whereby data is transmitted from one or more electronic devices in the electronic device to other electronic devices so that corresponding content can be output simultaneously at the electronic device. In some embodiments, the exemplary communication session may include a video conferencing session, whereby audio data and / or video data is transmitted between devices 500A, 500B, and 500C so that users of the respective devices can communicate in real time using the electronic device.
[0217] exist Figure 5C , device 500A represents an electronic device associated with user A. Device 500A communicates (via data connection 510) with devices 500B and 500C, which are associated with user B and user C, respectively. Device 500A includes a camera 501A for capturing video data of the communication session, and a display 504A (e.g., a touch screen) for displaying content associated with the communication session. Device 500A also includes other components, such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0218] Device 500A displays a communication UI 520A via display 504A, which is a user interface for facilitating a communication session (e.g., a video conferencing session) between device 500B and device 500C. Communication UI 520A includes video feed 525-1A and video feed 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during the communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during the communication session.
[0219] Communication UI 520A includes camera preview 550A, which is a representation of video data captured at device 500A via camera 501 A. Camera preview 550A represents to user A the intended video feed of user A displayed at respective devices 500B and 500C.
[0220] The communication UI 520A includes one or more controls 555A for controlling one or more aspects of the communication session. For example, the controls 555A may include controls for muting the audio of the communication session, changing the camera view of the communication session (e.g., changing the camera used to capture video of the communication session, adjusting the zoom value), terminating the communication session, applying visual effects to the camera view of the communication session, or activating one or more modes associated with the communication session. In some embodiments, the one or more controls 555A are optionally displayed in the communication UI 520A. In some embodiments, the one or more controls 555A are displayed separately from the camera preview 550A. In some embodiments, the one or more controls 555A are displayed so as to overlay at least a portion of the camera preview 550A.
[0221] exist Figure 5C , device 500B represents an electronic device associated with user B, who communicates with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B for capturing video data of the communication session, and a display 504B (e.g., a touch screen) for displaying content associated with the communication session. Device 500B also includes other components, such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0222] Device 500B displays a communication UI 520B similar to the communication UI 520A of device 500A via touch screen 504B. Communication UI 520B includes video feed 525-1B and video feed 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during the communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during the communication session. Communication UI 520B also includes a camera preview 550B, which is a representation of the video data captured at device 500B via camera 501B; and one or more controls 555B, similar to controls 555A, for controlling one or more aspects of the communication session. Camera preview 550B represents to user B the intended video feed of user B displayed at respective devices 500A and 500C.
[0223] exist Figure 5C, device 500C represents an electronic device associated with user C, who communicates with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C for capturing video data of the communication session, and a display 504C (e.g., a touch screen) for displaying content associated with the communication session. Device 500C also includes other components, such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0224] Device 500C displays a communication UI 520C similar to communication UI 520A of device 500A and communication UI 520B of device 500B via touch screen 504C. Communication UI 520C includes video feed 525-1C and video feed 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during the communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during the communication session. Communication UI 520C also includes a camera preview 550C, which is a representation of video data captured at device 500C via camera 501C; and one or more controls 555C, similar to controls 555A and 555B, for controlling one or more aspects of the communication session. Camera preview 550C shows user C the expected video feed of user C displayed at respective devices 500A and 500B.
[0225] Although Figure 5C The diagram depicted in represents a communication session between three electronic devices, but the communication session can be established between two or more electronic devices, and the number of devices participating in the communication session can change as electronic devices join or leave the communication session. For example, if one of the electronic devices leaves the communication session, the audio data and video data from the device that ceased participating in the communication session are no longer displayed on the participating devices. For example, if device 500B ceases participating in the communication session, there is no data connection 510 between devices 500A and 500C, and there is no data connection 510 between devices 500C and 500B. In addition, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and video data and audio data are shared among all devices, allowing each device to output data transmitted from other devices.
[0226] Figure 5C The embodiment depicted in FIG represents a diagram of a communication session between multiple electronic devices, including Figures 6A to 6Q 、 Figures 8A to 8R 、 10A to 10J and Figures 12A to 12U In some embodiments, Figures 6A to 6Q 、 Figures 8A to 8R 、 10A to 10J and Figures 12A to 12U A communication session depicted in FIG. 1 includes two or more electronic devices even if other electronic devices participating in the communication session are not depicted in the figure.
[0227] Figures 6A to 6Q An exemplary user interface for managing a real-time video communication session (e.g., a video conference) according to some embodiments is shown. The user interface in these figures is used to illustrate a user interface including Figure 7A and Figure 7B The process is described below.
[0228] Figures 6A to 6Q A device 600 is shown displaying a user interface for managing a real-time video communication session on a display 601 (eg, a display device or display generating component). Figures 6A to 6Q Various embodiments are depicted in which the device 600 automatically reconstructs a displayed portion of a camera's field of view based on conditions detected in a scene within the camera's field of view when an auto-framing mode is enabled. Figures 6A to 6Q One or more of the embodiments discussed may be combined with Figures 8A to 8R 、 10A to 10J and Figures 12A to 12U Combinations of one or more of the discussed embodiments.
[0229] Device 600 includes one or more cameras 602 (e.g., a front-facing camera) for capturing image data and, optionally, depth data of a scene within the camera's field of view. In some embodiments, camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600 includes one or more features of device 100, 300, or 500.
[0230] exist Figure 6A, device 600 displays video conference request interface 604-1, which depicts an incoming request from "John" to participate in a real-time video conference. Video conference request interface 604-1 includes camera preview 606, options menu 608, framing mode enable indication 610, and background blur enable indication 611. Camera preview 606 is a real-time representation of the video feed from camera 602, which is enabled for output of the video conference session (if the incoming request is accepted). Figure 6A , camera 602 is currently enabled and camera preview 606 depicts a representation of “Jane” who is currently located in front of device 600 and within the field of view of camera 602 .
[0231] The options menu 608 includes various selectable options for controlling one or more aspects of the video conference. For example, a mute option 608-1 can be selected to mute any audio transmission detected by the device 600. A flip option 608-2 can be selected to switch the camera being used for the video conference between the camera 602 and one or more different cameras (such as a camera on the opposite side of the device 600 from the camera 602 (e.g., a rear camera)). An accept option 608-3 can be selected to accept the request to participate in the real-time video conference. A reject option 608-4 can be selected to reject the request to participate in the real-time video conference.
[0232] The background blur enable indication 611 can be selected to enable or disable the background blur mode for the video conferencing session. Figure 6A In the depicted embodiment, when device 600 receives or initiates a request to participate in a video conference, background blur mode is disabled by default. Therefore, background blur affordance 611 is depicted as having an unselected state, as shown in FIG. Figure 6A In some embodiments, when device 600 receives or initiates a request to participate in a video conference, background blur mode is enabled by default (e.g., and the affordance is bolded). Figures 12A to 12U The background blur feature is discussed in more detail.
[0233] The framing mode enable indicator 610 can be selected to enable or disable the automatic framing mode for the video conferencing session. Figure 6A In the depicted embodiment, when device 600 receives or initiates a request to participate in a video conference, the automatic framing mode is enabled by default. Therefore, the framing mode enable indication 610 is depicted as having a selected state, as indicated by Figure 6A In some embodiments, when device 600 receives or initiates a request to participate in a video conference, automatic framing mode is disabled by default (e.g., and there is no bolding of the affordance).
[0234] When auto-framing mode is enabled, device 600 detects conditions of a scene within the field of view of an enabled camera (e.g., camera 602) (e.g., the presence and / or position of one or more objects within the camera's field of view) and adjusts the field of view of the video output for the video conferencing session (as represented in the camera preview) in real time (e.g., without moving camera 602 or device 600) based on the conditions or changes in the scene detected in the scene within the camera's field of view (e.g., changes in the position and / or movement of objects during the video conferencing session). Various embodiments of auto-framing mode are discussed throughout this disclosure.
[0235] For example, in Figures 6A to 6Q In the embodiment depicted in , device 600 automatically adjusts (e.g., reframes) the display output video feed field of view to maintain the display of one or more objects (e.g., Jane) within the field of view of a camera (e.g., camera 602). Because device 600 automatically adjusts the displayed portion of the camera's field of view to include Jane's display, Jane is able to move around the scene while participating in the video conference without having to manually adjust the viewing angle of the outgoing video feed to account for her movement or other changes in the scene. Thus, participants in the video conference need to interact less with device 600 because device 600 automatically reframes the outgoing video feed so that remote participants receiving the video feed from device 600 can continuously view Jane as she moves around her environment. Other benefits of the automatic framing mode will be described in the disclosure below. About Figures 6A to 6Q 、 Figures 8A to 8R and 10A to 10J Various features of the auto-framing mode are discussed in connection with the embodiments depicted in
[0014] . One or more of these features may be combined with other features of the auto-framing mode, as discussed herein.
[0236] Figure 6A Depicted are input 612 (e.g., a tap input) on framing mode enable representation 610 and input 614 on accept option 608-3. In response to detecting input 612, device 600 disables auto-framing mode. In response to detecting input 614 after input 612, device 600 accepts the video conference call and joins the video conference session with auto-framing mode disabled, as shown. Figure 6B If device 600 does not detect input 612, or if input 612 is an input to enable auto-framing mode (e.g., framing mode enable indication 610 is not selected when input 612 is received), device 600 accepts the video conference call in response to input 614 and joins the video conference session with auto-framing mode enabled, as shown. Figure 6F As described in .
[0237] Figure 6BDepicted is a scene 615, which is the physical environment within the field of view 620 of the camera 602. Figure 6B , Jane 622 is seated on sofa 621, device 600 is positioned in front of her (e.g., on a table), and door 618 is in the background. Field of view 620 represents the field of view of camera 602 (e.g., the maximum field of view of camera 602 or the wide-angle field of view of camera 602), which encompasses scene 615. Portion 625 indicates the portion of field of view 620 currently being output (or selected for output) for the video conferencing session (e.g., portion 625 represents the displayed portion of the camera's field of view). Thus, portion 625 indicates the portion of scene 615 currently being displayed in the displayed video feed depicted in camera preview 606. Field of view 620 is sometimes referred to herein as the available field of view, the entire field of view, or the camera's field of view, and portion 625 is sometimes referred to herein as the video feed field of view.
[0238] When the incoming video conference request is accepted, the video conference request interface 604-1 switches to the video conference interface 604, as shown in FIG. Figure 6B Video conference interface 604 is similar to video conference request interface 604-1, but is updated to depict, for example, incoming video feed 623 including video data of remote participants of the video conference received at device 600. Video feed 623 includes representation 623-1 of John as a remote participant in the video conference with Jane 622. Figure 6A , the camera preview 606 is now reduced in size and shifted toward the upper right corner of the display 601. The camera preview 606 includes a representation of Jane 622-1 and a representation of the environment captured within portion 625 (e.g., the video feed field of view). The options menu 608 is updated to include a view mode enable indication 610. In some embodiments, for example, the view mode enable indication 610 is displayed in the camera preview 606, as shown in FIG. Figure 6O As shown in Figure 6B As shown, the framing mode enable indication 610 is shown in an unselected state, thereby indicating that the automatic framing mode is disabled.
[0239] In some embodiments, when the auto-view mode is disabled, the device 600 outputs a predetermined portion of the available field of view of the camera 602 as the video feed field of view. Examples of such embodiments are depicted in Figures 6B to 6D , where portion 625 represents a predetermined portion of the available field of view of camera 602 that is located in the center of field of view 620. In some embodiments, when the auto-view mode is disabled, device 600 outputs the entire field of view 620 as the video feed field of view. Examples of such embodiments are depicted in Figure 10A middle.
[0240] exist Figure 6B In FIG, Jane is using device 600 to participate in a video conference with John. Similarly, John is using a device that includes one or more features of device 100, 300, 500, or 600 to participate in a video conference with Jane. For example, John is using a tablet computer similar to device 600 (e.g., Figures 10H to 10J and Figures 12B to 12N Thus, John's device displays a video conferencing interface similar to video conferencing interface 604, except that the camera preview on John's device displays the video feed captured from John's device (currently in Figure 6B ), and the incoming video feed on John's device displays the video feed output from Jane's device 600 (currently in Figure 6B (as depicted in the camera preview 606 in FIG. 1 ).
[0241] exist Figure 6C , device 600 moves relative to scene 615, and therefore field of view 620 and portion 625 pivot with device 600. Because auto-framing mode is disabled, device 600 does not automatically adjust the video feed field of view to remain fixed on Jane's position within field of view 620. Instead, the perspective of the video feed moves with device 600, and Jane 622 is no longer in the center of the video feed field of view, as indicated by portion 625 and as depicted by camera preview 606, which shows the background of scene 615 and a portion of Jane's representation 622-1. Figure 6C In the embodiment depicted in , the movement of device 600 is pivoting, however, movement of field of view 620 and portion 625 may be caused by other movements, such as tilting, rotating, and / or moving device 600 (e.g., forward, backward, and / or left to right) in a manner such that Jane 622 does not remain within portion 625.
[0242] exist Figure 6D , device 600 returns to its original position, and Jane 622 bends downward, moving out of portion 625. Again, because auto-framing mode is disabled, device 600 does not automatically adjust the video feed field of view to follow Jane's movement as she moves out of portion 625. As Jane 622 moves, the video feed field of view remains stationary, and Jane's representation 622-1 is mostly outside the framing frame in camera preview 606.
[0243] exist Figure 6D, device 600 detects input 626 (e.g., a tap input) on framing mode enable indication 610. In response, device 600 bolds framing mode enable indication 610 (to indicate its selected / enabled state) and enables auto-framing mode, as shown in FIG. Figure 6E When the auto-viewing mode is enabled, the device 600 automatically adjusts the displayed video feed field of view based on conditions detected within the scene 615. Figure 6E , the device 600 adjusts the displayed video feed field of view to be centered on Jane's face. Accordingly, the device 600 updates the camera preview 606 to include a representation 622-1 of Jane centered in the frame and a representation 621-1 of the sofa on which she is sitting in the background. The field of view 620 remains fixed because the position of the camera 602 remains unchanged. However, the position of Jane's face within the field of view 620 does change. Therefore, the device 600 adjusts (e.g., repositions) the displayed portion of the field of view 620 so that Jane remains positioned within the camera preview 606. This is shown in FIG. Figure 6E is represented by the repositioning of portion 625 so that it is centered on Jane's face. Figure 6E , portion 627 corresponds to the previous position of portion 625 and, therefore, represents the portion of field of view 620 that was previously displayed in camera preview 606 (before adjustments resulting from enabling auto-framing mode).
[0244] Figure 6F Depicted is a scene 615 and device 600 when auto-framing mode is enabled in response to input 614 accepting an incoming request to join a video conference while auto-framing mode is enabled (or alternatively in response to input 626 enabling auto-framing mode). Accordingly, device 600 displays a representation 622-1 of Jane in the center of camera preview 606.
[0245] exist Figure 6G In the embodiment, the device 600 is similar to the above description of Figure 6C However, because Figure 6G 620 to remain fixed on Jane's face, which has changed position relative to the device 600 and camera 602 in response to the pivoting of the device 600. Thus, the camera preview 606 continues to display Jane's representation 622-1 in the center of the video feed field of view. The adjustment of the video feed field of view is represented by the change in position of portion 625 within the field of view 620. For example, when compared to Figure 6F , the relative position of portion 625 has changed from the central position within the field of view 620 (at Figure 6G ) is moved to Figure 6GThe displaced position depicted in . Figure 6G In the embodiment depicted in FIG, the movement of device 600 is pivoting, however, device 600 may automatically adjust the displayed video feed field of view in response to other movements, such as tilting, rotating, and / or moving device 600 (e.g., forward, backward, and / or left to right) in a manner that maintains Jane 622 within field of view 620.
[0246] exist Figure 6H , device 600 has returned to its original position and Jane 622 has moved to a standing position near sofa 621. Device 600 detects the updated position of Jane 622 in scene 615 and updates the video feed field of view to maintain its position on Jane 622, as shown in camera preview 606. Thus, portion 625 has moved from its previous position represented by portion 627 to an updated position around Jane's face, as shown in FIG. Figure 6H Described in.
[0247] In some embodiments, from Figure 6G The field of view depicted in the camera preview 606 is Figure 6H A transformation of the field of view depicted in the camera preview 606 in is performed to match the clip. For example, from Figure 6G Camera preview in 606 to Figure 6H The transition of the camera preview 606 in FIG is a match cut performed when Jane 622 has moved from her position sitting on the sofa 621 to her position standing near the sofa. The result of the match cut is that the camera preview 606 looks from Figure 6G The first camera view in the transform is Figure 6H Different camera views in (the different camera views optionally have Figure 6G 6). However, the actual field of view of camera 602 (e.g., field of view 620) has not changed. Instead, only the portion of the field of view that is displayed (portion 625) has changed position within field of view 620.
[0248] Figure 6I is similar to Figure 6H An embodiment of the embodiment depicted in FIG, but wherein the camera preview 606 is compared to Figure 6H The camera preview shown in has a larger, more distant field of view. Specifically, Figure 6I The embodiment depicted in FIG shows the Figure 6G The camera preview in Figure 6I Jump cut transitions from the camera preview in Figure 6G The camera preview in 606 is converted to Figure 6I , which has a larger (e.g., more distant) field of view. Figure 6IThe video feed field of view in FIG (represented by portion 625) is a larger portion of the field of view 620. This is shown by portion 625 (corresponding to Figure 6I Camera preview in ) and section 627 (corresponding to Figure 6G to illustrate the size difference between the camera preview in .
[0249] Figure 6H and Figure 6I A specific embodiment of transitions between different camera previews is shown. In some embodiments, other transitions can be performed, such as by continuously moving (e.g., panning and / or zooming) the video feed field of view within the field of view 620 to follow Jane 622 as she moves around the scene 615.
[0250] exist Figure 6J , Jane 622 has left device 600 and is behind sofa 621. In response to detecting the change in position of Jane 622, device 600 performs a transition (e.g., a jump cut transition) in which camera preview 606 depicts a zoomed-in view of Jane in scene 615. Figure 6J In the embodiment depicted in FIG, device 600 zooms in on Jane 622 (e.g., reduces the field of view) as she moves away from the camera (e.g., moves a threshold distance). In some embodiments, device 600 zooms out from Jane 622 (e.g., expands the display field of view) as she moves toward the camera (e.g., moves a threshold distance). For example, if Jane 622 were to move from her Figure 6J Move to her position in Figure 6I , the camera preview 606 will zoom out to the previous position in Figure 6I The camera preview depicted in .
[0251] exist Figure 6K In the example, another object Jack 628 walks into the scene 615. The device 600 continues to display the image with Figure 6J The same camera preview as depicted in the video conferencing interface 604. Figure 6K In the depicted embodiment, the device 600 detects Jack 628 within the field of view 620, but maintains the same video feed field of view as Jack 628 moves around the scene or until Jack 628 moves to a specific location in the scene (e.g., closer to the center of the scene). In some embodiments, when additional objects are detected within the field of view 620, the device 600 displays a prompt to adjust the camera preview. Figure 6P and Figures 8A to 8J The embodiments depicted in discuss examples of this prompt in more detail.
[0252] exist Figure 6L, the device 600 reconstructs the camera preview 606 to include a representation 628-1 of Jack 628, who is now standing next to Jane 622 in the scene 615. In some embodiments, the device 600 automatically adjusts the camera preview 606 in response to determining that Jack 628 stops moving around the scene 615 and / or that he is exhibiting behavior that indicates a desire to participate in the video conference. Examples of such behavior may include turning attention to the device 600 and / or camera 602, focusing / looking at the camera 602 or in the general direction of the camera, remaining still (e.g., for at least a particular amount of time), positioning next to a participant in the video conference (e.g., Jane 622), facing the device 600, speaking, etc. In some embodiments, when the auto-framing mode is enabled, the device 600 automatically adjusts the camera preview in response to detecting a change in the number of objects detected within the field of view 620, such as when Jack 628 enters the scene 615. Figure 6L As depicted in FIG. 6 , portion 625 represents the adjusted video feed field of view, and portion 627 represents the adjusted video feed field of view. Figure 6L Adjust the size of the video feed field of view before. Figure 6K The video feed field of view in Figure 6L The adjusted field of view is zoomed out and re-centered on Jane 622 and Jack 628 , as depicted in camera preview 606 .
[0253] exist Figure 6M , Jane 622 begins to move away from Jack 628 and leaves the video feed field of view represented by portion 625 and camera preview 606. When Jane's movement is detected, device 600 maintains (e.g., does not adjust) the field of view of camera preview 606. In some embodiments, as Jane moves away from Jack, device 600 resizes the video feed field of view so that both objects remain within the video feed field of view (camera preview). In some embodiments, after Jane has moved away from Jack, device 600 resizes the video feed field of view after Jane stops moving so that both objects are within the camera preview.
[0254] exist Figure 6N In the example, device 600 detects that Jane 622 is no longer in Figure 6M within portion 625 (e.g., Figure 6N 627 in the video feed), and in response, adjusts the displayed video feed field of view to zoom in on Jack 628. Thus, device 600 displays video conferencing interface 604 with camera preview 606 having a zoomed-in view of Jack 628. Figure 6N Portion 625 is depicted as having a smaller size than portion 627 to indicate the change in the field of view of the displayed video feed.
[0255] In some embodiments, when the auto-framing mode is enabled, the device 600 displays one or more prompts to adjust the video feed field of view to include the additional participant in response to detecting the additional object within the field of view 620. Figures 6O to 6Q Examples of such implementations are described.
[0256] Figure 6O Depicts something like Figure 6N , but in which the framing mode enabling indication 610 is displayed in the camera preview area rather than in the options menu 608. The device 600 displays a framing indicator 630 positioned around Jack's representation 628-1 to indicate that the device 600 has detected the presence of a face (e.g., Jack's face) in the camera preview. In some embodiments, the framing indicator also indicates that auto-framing mode is enabled and that a face is being tracked when it is detected within portion 625, camera preview 606, and / or field of view 620. Jack 628 is currently the only participant in the video conference located in scene 615.
[0257] exist Figure 6P , device 600 detects that Jane 622 has entered scene 615 within field of view 620. In response, device 600 updates video conferencing interface 604 by displaying an add affordance 632 in the camera preview area. Add affordance 632 can be selected to adjust the displayed video feed field of view to include the additional object detected within field of view 620.
[0258] exist Figure 6P , device 600 detects input 634 on adding affordance 632 and, in response, adjusts portion 625 so that camera preview 606 includes Jane's representation 622-1 and Jack's representation 628-1, as shown. Figure 6Q In some embodiments, Jane 622 is added as a participant in the video conference. Device 600 also recognizes the presence of Jane's face and displays framing indicators 630 around the representation of Jane's face (in addition to those around the representation of Jack's face) to indicate that framing mode is enabled and is tracking Jane's face when Jane's face is detected within portion 625, camera preview 606, and / or field of view 620.
[0259] 7A to 7BA flowchart illustrating a method for managing a real-time video communication session using an electronic device according to some embodiments is depicted. Method 700 is performed at a computer system (e.g., a smartphone, a tablet) (e.g., 100, 300, 500, 600) that communicates with a display generation component (e.g., 601) (e.g., a display controller, a touch-sensitive display system, one or more cameras (e.g., 602) (e.g., an infrared camera, a depth camera, a visible light camera), and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0260] As described below, method 700 provides an intuitive way to manage real-time video communication sessions. The method reduces the cognitive burden on users managing real-time video communication sessions, thereby creating a more efficient human-computer interface. For battery-powered computing devices, this enables users to manage real-time video communication sessions faster and more efficiently, conserving power and increasing the time between battery charges.
[0261] In method 700, a computer system (eg, 600) displays (702) a communication request interface (eg, Figure 6A 604-1 in) (e.g., an interface for incoming or outgoing real-time video communication sessions (e.g., real-time video chat sessions, real-time video conferencing sessions)).
[0262] The computer system (eg, 600) displays (704) a communication request interface (eg, Figure 6A ), the communication request interface includes a first selectable graphical user interface object (e.g., 608-3) associated with a process for joining the real-time video communication session (e.g., an "Accept" affordance). In some embodiments, the Accept affordance can be selected to initiate a process for accepting an incoming request to join the real-time video communication session. In some embodiments, the first selectable graphical user interface object can be selected to initiate a "Cancel" affordance for canceling or terminating an outgoing request to join the real-time video communication session.
[0263] The computer system (eg, 600) displays (eg, simultaneously with 704) 706 a communication request interface (eg, 610) including a second selectable graphical user interface object (eg, 610). Figure 6A604-1 in), the second selectable graphical user interface (e.g., a “framing mode” enable indication, a “background blur” enable indication, a “dynamic video quality” enable indication) is associated with a process for selecting between using a first camera mode (e.g., an automatic framing mode, a background blur mode, a dynamic video quality mode) for the one or more cameras and using a second camera mode (e.g., a mode different from the first camera mode (e.g., a mode in which the automatic framing mode is disabled, a mode in which the background blur mode is disabled, and / or a mode in which the dynamic video quality mode is disabled)) for the one or more cameras during a real-time video communication session.
[0264] In some embodiments, the framing mode indicator can be selected to enable / disable a mode (e.g., an auto-framing mode) for: 1) tracking the position and / or location of one or more objects detected within the field of view of one or more cameras during a real-time video communication session, and 2) automatically adjusting the displayed view of the object based on the tracking of the object during the real-time video communication session.
[0265] In some embodiments, the background blur enable representation can be selected to enable / disable a mode (e.g., background blur mode) in which a visual effect (e.g., a blur, darken, colorize, mask, desaturate, or otherwise de-emphasize effect) is applied to a background portion (e.g., 606) of a camera's field of view (e.g., a camera preview, an output video feed of the camera's field of view) during a real-time video communication session (e.g., without applying the visual effect to a portion of the camera's field of view that includes a representation of an object (e.g., a foreground portion) (e.g., 622-1)).
[0266] In some embodiments, a dynamic video quality indicator can be selected to enable / disable a mode (e.g., a dynamic video quality mode) for outputting (e.g., transmitting and optionally displaying) a camera field of view in which portions have different degrees of compression and / or video quality. For example, a portion of the camera field of view that includes a detected face (e.g., 622-1) is compressed less than a portion of the camera field of view that does not include a detected face. In some embodiments, there is an inverse relationship between the degree of compression and video quality (e.g., greater compression produces lower video quality; less compression produces higher video quality). Thus, a video feed for a real-time video communication session can be transmitted (e.g., by a computer system (e.g., 600)) to a receiving device of a remote participant in the real-time video communication session such that the portion that includes the detected face can be displayed at the receiving device with a higher video quality than the portion of the camera field of view that does not include the detected face (due to reduced compression of the portion of the camera field of view that includes the detected face and increased compression of the portion of the camera field of view that does not include the detected face). In some embodiments, the computer system changes the amount of compression as the video bandwidth changes (e.g., increases, decreases). For example, the degree of compression of a portion of the camera's field of view that does not include a detected face (e.g., the camera feed) varies (e.g., increases or decreases with a corresponding change in bandwidth), while the degree of compression of a portion of the camera's field of view that includes a detected face remains constant (or, in some embodiments, changes at a smaller rate or by a smaller amount than the portion of the camera's field of view that does not include a face).
[0267] When the communication request interface is displayed (for example, Figure 6A 1 ), the computer system (e.g., 600) receives (708) via one or more input devices (e.g., 601) a set of one or more inputs (e.g., 612 and / or 614) including a selection (e.g., 614) of a first selectable graphical user interface object (e.g., 608-3) (e.g., the set of one or more inputs including a selection of an accept enable indication and, optionally, a selection of a framing mode enable indication, a background blur enable indication, and / or a motion video quality enable indication).
[0268] In response to receiving the set of one or more inputs (e.g., 612 and / or 614) including a selection (e.g., 614) of a first selectable graphical user interface object (e.g., 608-3), the computer system (e.g., 600) displays (710) a real-time video communication interface (e.g., 604) for a real-time video communication session via a display generation component (e.g., 601).
[0269] While displaying the real-time video communication interface (e.g., 604), the computer system (e.g., 600) detects (712) a change (e.g., a change in position of an object) in a scene (e.g., 615) in a field of view (e.g., 620) of the one or more cameras (e.g., 602). In some embodiments, the scene includes a representation of the object and, optionally, one or more additional objects in the field of view of the one or more cameras.
[0270] In response to detecting a change (714) in the scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras (e.g., 602), the computer system (e.g., 600) performs one or more of steps 716 and 718 of method 700.
[0271] Based on determining that the first camera mode is selected for use (e.g., enabled) (e.g., if the first camera mode is disabled by default, the set of one or more inputs includes selection of a second selectable graphical user interface object; if the first camera mode is enabled by default, the set of one or more inputs does not include selection of the second selectable graphical user interface object) (e.g., the viewfinder mode enable indication is in a selected state when the accept enable indication is selected), the computer system (e.g., 600) adjusts (716) (e.g., automatically, without user input) a representation of the field of view (e.g., 606) (e.g., the displayed field of view of the one or more cameras) during the real-time video communication session based on a detected change in the field of view (e.g., 620) of the one or more cameras (e.g., 602) (e.g., automatically adjusting the representation of the field of view of the one or more cameras (e.g., based on the detected object position) during the real-time video communication session). When the first camera mode is selected for use, the representation of the field of view of the one or more cameras is adjusted during the real-time video communication based on detected changes in the scene in the field of view of the one or more cameras, which enhances the video communication session experience by automatically adjusting the field of view of the cameras (e.g., to maintain display of an object / user) without requiring additional input from the user. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0272] In some embodiments, adjusting the representation of the field of view of the one or more cameras during the real-time video communication session (e.g., 606) includes: 1) adjusting the representation of the field of view of the one or more cameras based on determining that a first set of criteria is satisfied, including that the scene is included in the first position (within the field of view of the one or more cameras 620, at Figure 6F ), displaying a representation having a first field of view (e.g., 622 ) of the detected object (e.g., one or more users of the computer system). Figure 6F 606) of the real-time video communication interface 604 (e.g., displaying the real-time video communication interface at a first digital zoom level and a first display portion of the field of view of the one or more cameras) (in some embodiments, the representation of the first field of view includes a representation of the object when the object is at a first location); and 2) based on determining that a second set of criteria is met, including detecting the object at a second location different from the first location (e.g., Figure 6H 625 ), displaying a representation having a second field of view different from the representation of the first field of view (e.g., Figure 6H In some embodiments, the representation of the field of view automatically changes in response to a detected change in the position of the object and / or in response to detecting a second object entering or leaving the field of view of the one or more cameras (e.g., without changing the actual field of view of the one or more cameras). For example, the representation of the field of view changes to track the position of the object and adjusts the display position and / or zoom level (e.g., digital zoom level) to more prominently display the object (e.g., changing the digital zoom level to appear to zoom in on the object as the object moves away from the camera; changing the digital zoom level to appear to zoom out from the object as the object moves toward the camera; changing the displayed portion of the field of view of the one or more cameras to appear to pan in a particular direction as the object moves in that direction).
[0273] Based on determining that the second camera mode is selected for use (e.g., enabled) (e.g., when the accept affordance is selected, the view mode affordance is in an unselected or deselected state), the computer system (e.g., 600) forgoes (718) adjusting the representation of the field of view of the one or more cameras during the real-time video communication session (e.g., as Figure 6D) (e.g., based on a detected change in the scene in the field of view of the one or more cameras) (e.g., when the first camera mode is disabled, the real-time video communication interface maintains the same (e.g., default) representation of the field of view regardless of whether the object is within the scene in the field of view of the one or more cameras and regardless of where the object is within the field of view of the one or more cameras). In some embodiments, forgoing adjusting the representation of the field of view of the one or more cameras during the real-time video communication session includes: 1) when (e.g., based on a determination) that the object has a first position within the scene in the field of view of the one or more cameras (e.g., as Figure 6B ), displaying a real-time video communication interface having a representation of the first field of view (e.g., Figure 6B and 2) when (e.g., based on determining) the object has a second position within the scene in the field of view of the one or more cameras (e.g., in Figure 6D ), displaying a real-time video communication interface having a representation of the first field of view (e.g., Figure 6D In some embodiments, the representation of the first field of view is a standard or default representation of the field of view that does not change based on changes in the scene (e.g., a change in the position of an object relative to the one or more cameras or a second object entering or leaving the field of view of the one or more cameras).
[0274] In some embodiments, the detected change in the scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras (e.g., 602) includes a detected change in a set of attention-based factors of one or more objects (e.g., 622, 628) in the scene (e.g., a first object directing their attention (e.g., focusing on, looking toward) the one or more cameras (e.g., based on the first object's gaze position, head position, and / or body position). In some embodiments, the computer system (e.g., 600) adjusts the representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras (e.g., 606), including: adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) the detected change in the set of attention-based factors of the one or more objects in the scene (e.g., Figure 6L). Adjusting a representation of the field of view of the one or more cameras during a real-time video communication based on detected changes in the set of attention-based factors for one or more objects in the scene enhances the video communication session experience by automatically adjusting the camera's field of view based on the set of attention-based factors for the objects in the scene without requiring additional input from the user. Performing an operation when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0275] In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) adjusts (e.g., reframes) a displayed portion of the field of view of one or more cameras (e.g., 602) based on one or more attention-based factors of objects (e.g., 622, 628) detected within the field of view (e.g., 620) of the one or more cameras (e.g., 606). For example, when a first subject turns their attention toward one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include a representation of the first subject or focuses on the first subject. Conversely, when the first subject's attention shifts away from the one or more cameras, the representation of the field of view of the one or more cameras changes (e.g., zooms in) to not include a representation of the first subject (e.g., if other objects remain in the field of view of the one or more cameras) or focuses on another object.
[0276] In some embodiments, the set of attention-based factors includes a first factor that is based on a detected focal plane (e.g., as Figure 6L). Adjusting the representation of the field of view of the one or more cameras during the real-time video communication based on the detected focal plane of a first object among the one or more objects in the scene enhances the video communication session experience by automatically adjusting the camera's field of view when the first object's focal plane meets criteria without requiring additional input from the user. Performing an operation when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, the attention of the first object is determined based on the focal plane of the first object. For example, if the focal plane of the first object is aligned (e.g., coplanar) with the focal plane of the one or more cameras or the focal plane of another object participating in the real-time video communication session, the first object is considered to be paying attention to the one or more cameras. Therefore, the first object is considered to be an active participant in the real-time video communication session, and the computer system then adjusts the representation of the field of view of the one or more cameras to include the first object in the real-time video communication interface.
[0277] In some embodiments, the set of attention-based factors includes a second factor based on whether a second object (e.g., 628) (e.g., an object other than the first object) among the one or more objects in the scene (e.g., 615) is determined to be looking at the one or more cameras (e.g., 602). Adjusting the representation of the field of view of the one or more cameras during the real-time video communication based on whether the second object in the scene is determined to be looking at the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera when the second object is looking at the camera without requiring additional input from the user. Performing operations when a set of conditions are met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, the attention of the second object is determined based on whether the second object is looking at the one or more cameras. If so, the second object is considered to be paying attention to the one or more cameras and, therefore, is considered an active participant in the real-time video communication session. Thus, the computer system adjusts the representation of the field of view of the one or more cameras to include the second object in the real-time video communication interface.
[0278] In some embodiments, the detected change in the scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras (e.g., 602) includes a detected change in the number (e.g., amount, number) of objects (e.g., 622, 628) in the scene (e.g., a detected change in the number of objects detected in the scene that meet a first set of criteria (e.g., objects are located in the field of view of the one or more cameras and, optionally, are stationary) (e.g., one or more objects enter or leave the scene in the field of view of the one or more cameras). In some embodiments, adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on the detected change in the scene in the field of view of the one or more cameras (e.g., 606) includes adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on (in some embodiments, in response to) the detected change in the number of objects detected in the scene (e.g., meeting the first set of criteria) (e.g., Figure 6L and / or Figure 6N ). Adjusting the representation of the field of view of one or more cameras during real-time video communication based on a detected change in the number of objects detected in the scene enhances the video communication session experience by automatically adjusting the camera's field of view as the number of objects in the scene changes without requiring additional input from the user. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) adjusts (e.g., reframes) the displayed portion of the field of view of the one or more cameras based on the number of objects detected within the field of view of the one or more cameras. For example, as the number of objects detected in the scene increases, the representation of the field of view of the one or more cameras changes (e.g., zooms out) to include the additional objects (e.g., along with previously detected objects). Similarly, as the number of objects detected in the scene decreases, the representation of the field of view of the one or more cameras changes (eg, zooms in) to capture the objects that remain in the scene.
[0279] In some embodiments, adjusting the representation of the field of view of the one or more cameras (e.g., 606) during a real-time video communication session based on a detected change in the number of objects detected in the scene is based on a determination of whether objects in the field of view (e.g., 620) are stationary (e.g., relatively stationary; not moving by more than a threshold amount of movement within the field of view of the one or more cameras). Adjusting the representation of the field of view of the one or more cameras during a real-time video communication based on whether objects in the field of view are stationary enhances the video communication session experience by automatically adjusting the camera's field of view when objects in the scene are stationary without requiring additional input from the user. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0280] In some embodiments, when the auto-framing mode is enabled, the computer system (e.g., 600) deems a detected object (e.g., 628) to be a participant in the real-time video communication session when the object does not move more than a threshold amount of movement. When the object is deemed to be an active participant, the computer system adjusts (e.g., reframes) the displayed portion of the field of view of one or more cameras (e.g., 606) to subsequently include a representation of the object (e.g., as Figure 6L This prevents the computer system from automatically reconstructing the displayed portion of the field of view of the camera or cameras based on extraneous motion in the scene (e.g., such as motion caused by objects passing in the background or children jumping around in the field of view of the camera or cameras), which would otherwise be distracting to participants / viewers of the real-time video communication session.
[0281] In some embodiments, before a computer system (e.g., 600) detects a change in a scene (e.g., 615) in a field of view (e.g., 620) of one or more cameras (e.g., 602), a representation (e.g., 606) of the field of view of the one or more cameras has a first representation field of view (e.g., the computer system is displaying a portion of the field of view of the one or more cameras before detecting the change in the scene). In some embodiments, the change in the field of view of the one or more cameras includes a third object (e.g., 622) being removed from a portion of the field of view of the one or more cameras corresponding to (e.g., represented by; included in) the first representation field of view (e.g., Figure 6G 606) (e.g., Figure 6G 625) to a second portion of the field of view of the one or more cameras that does not correspond to (e.g., is not represented by; is not included in) the first representation field of view (e.g., Figure 6H In some embodiments, adjusting the representation of the field of view of one or more cameras during a real-time video communication session based on (in some embodiments, in response to) a detected change in the scene in the field of view of the one or more cameras includes: based on determining that the fourth object is not detected in the scene in the first portion of the field of view of the one or more cameras (e.g., 628), adjusting the representation of the field of view from a first representation field of view to a second representation field of view (e.g., different from the first representation field of view) corresponding to (e.g., representing; displaying; including) a second portion of the field of view of the one or more cameras (e.g., displaying Figure 6H 6). Based on determining that a fourth object (e.g., 628) is detected in the scene in the first portion of the field of view of the one or more cameras (e.g., when Jane 622 leaves the frame, Jack 628 is located Figure 6M 625 in the section), abandoning the adjustment of the representation of the field of view from the first representation field of view to the second representation field of view (e.g., continuing to display the first representation field of view) (e.g., in Figure 6M , when Jane 622 leaves and Jack 628 remains, device 600 continues to display camera preview 606 depicting portion 625. After another object (e.g., a third object) leaves the first portion of the field of view, selectively adjusting the representation of the field of view of the one or more cameras from a first representation field of view to a second representation field of view during the real-time video communication session based on whether an object (e.g., a fourth object) is detected in the scene in the first portion of the field of view of the one or more cameras enhances the video communication session experience by automatically adjusting the field of view of the camera based on whether additional objects remain in the first portion of the field of view when the other object leaves, without requiring further input from the user. Performing an operation when a set of conditions have been met without requiring further user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, when auto-framing mode is enabled, the computer system does not track (e.g., follow; adjust the representation of the field of view in response to) the movement of another object that leaves the display field of view while the object remains in the display field of view.
[0282] In some embodiments, prior to detecting a change in a scene (e.g., 615) in the field of view (e.g., 620) of one or more cameras (e.g., 602), the representation (e.g., 606) of the field of view of the one or more cameras has a third representation field of view (e.g., Figure 6F In some embodiments, the change in the scene in the field of view of the one or more cameras includes the fifth object (e.g., 622) being displayed from (e.g., represented by; included in) a third portion of the field of view of the one or more cameras corresponding to (e.g., represented by; included in) a third portion of the field of view of the one or more cameras (e.g., Figure 6F 625) to a fourth portion of the field of view of the one or more cameras that does not correspond to (e.g., is not represented by; is not included in) the third representation field of view (e.g., Figure 6H In some embodiments, adjusting the representation of the field of view of the one or more cameras during the real-time video communications session based on (in some embodiments, in response to) a detected change in the scene in the field of view of the one or more cameras includes: displaying a representation of the field of view of the one or more cameras in the real-time video communications interface (e.g., 604) having a fourth representation field of view (e.g., Figure 6H 606) (e.g., different from the third representation field of view; in some embodiments, including a subset of the third representation field of view), the field of view corresponding to a fourth portion of the field of view of the one or more cameras and including a representation of a fifth object (e.g., 622-1) (e.g., replacing display of the third representation field of view with a fourth representation field of view, the fourth representation field of view including a representation of the fifth object (and in some embodiments, including a subset of the third representation field of view)). Ceasing to display the third representation field of view and displaying a fourth representation field of view corresponding to the fourth portion of the field of view and including a representation of the fifth object enhances the video communication session experience by automatically adjusting the field of view of the camera to maintain display of the object as the object moves to different locations in the scene without requiring additional input from the user. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0283] In some embodiments, adjusting a representation of a field of view of one or more cameras (e.g., 606) during a real-time video communications session based on a detected change in a scene (e.g., 615) in the field of view (e.g., 602) of the one or more cameras (e.g., 606) further includes ceasing to display, in the real-time video communications interface, a representation of the field of view of the one or more cameras having a third representation of the field of view (e.g., 615). Figure 6F 606 in). In some embodiments, when auto-framing mode is enabled and the computer system (e.g., 600) detects that an object (e.g., 622) moves out of the framing frame (e.g., 606) (e.g., out of the displayed representation field of view), the computer system switches to a different framing frame that includes the object (e.g., displaying a representation of a portion of the camera field of view). In some embodiments, a clip (e.g., a jump cut) includes a change in zoom level (e.g., zooming in or out). For example, a change in the representation field of view is a jump cut that includes a zoomed-out view that includes the user (e.g., when in single-person tracking mode). In some embodiments, a clip (e.g., a match clip) includes a change from displaying a first area of the camera field of view to displaying a second area of the camera field of view that includes the user but does not include the first area. In some embodiments, when a second object remains in the framing frame after the first object moves out of the framing frame, the computer system displays a jump cut to a zoomed view of the second object (e.g., when auto-framing mode is enabled).
[0284] In some embodiments, prior to detecting a change in the scene in the field of view of the one or more cameras, the representation of the field of view of the one or more cameras (e.g., 606) has a first zoom value (e.g., a zoom setting (e.g., 1x, 0.5x, 0.7x)) (e.g., as Figure 6I In some embodiments, the change in the scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras (e.g., 602) includes a change in the sixth object (e.g., 622) from a first position (e.g., 620) within the field of view of the one or more cameras. Figure 6I 625) to a second position within the field of view of the one or more cameras (e.g., Figure 6JThe first position corresponds to (e.g., represented by) a representation of the field of view and is a first distance from the one or more cameras, and the second position corresponds to (e.g., represented by) a representation of the field of view and is a threshold distance from the one or more cameras (e.g., an object moves within the displayed viewfinder (e.g., toward a camera; away from a camera) to a predetermined distance from the one or more cameras). In some embodiments, adjusting the representation of the field of view of the one or more cameras during the real-time video communication session based on a detected change in the scene in the field of view of the one or more cameras includes: displaying in the real-time video communication interface (e.g., 604) a representation of the field of view of the one or more cameras having a second zoom value different from the first zoom value (e.g., Figure 6J 606 of the video communication session) (e.g., zooming in on the view; zooming out on the view; in some embodiments, including the entire portion of the field of view of the one or more cameras that was previously displayed in the representation of the field of view with the first zoom value but is instead displayed at the second zoom value) (e.g., jumping from the first zoom level to the second zoom level when the object moves to a predetermined distance from the one or more cameras within the originally displayed viewfinder). Ceasing to display the representation of the field of view with the first zoom value and displaying the representation of the field of view with the second zoom value enhances the video communication session experience by automatically adjusting the zoom value of the representation of the field of view of the camera to maintain prominent display of the object as the object moves in the scene to different distances from the camera without requiring additional input from the user. Performing operations when a set of conditions have been met without requiring additional user input enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0285] In some embodiments, adjusting a representation (e.g., 606) of a field of view of one or more cameras (e.g., 602) based on a detected change in a scene (e.g., 615) in the field of view (e.g., 620) of the one or more cameras during a real-time video communication session includes stopping the real-time video communication interface (e.g., Figure 6IIn some embodiments, when an object (e.g., 622) moves toward a camera to a first threshold distance from the camera, the representation of the field of view of the one or more cameras transitions to a second zoom value that is a zoomed-out view of the previously displayed portion of the field of view of the one or more cameras (e.g., a jump cut to a wide angle view). In some embodiments, when an object (e.g., 622) moves away from the camera to a second threshold distance from the camera, the representation of the field of view of the one or more cameras transitions to a second zoom value that is a zoomed-in view of the previously displayed portion of the field of view of the one or more cameras.
[0286] In some embodiments, the computer system (e.g., 600) simultaneously displays a second selectable graphical user interface object (e.g., 610) and a real-time video communication interface (e.g., 604) that includes one or more other selectable controls (e.g., 608) for controlling the real-time video communication (e.g., an end call button for ending the real-time video communication session, a switch camera button for switching which camera is used for the real-time video communication session, a mute button for muting / unmuting the audio of users of devices in the real-time video communication session, an effects button for adding / removing visual effects to / from the real-time video communication session, an add user button for adding a user to the real-time video communication session, and / or a camera on / off button for turning on / off a user's video in the real-time video communication session). In some embodiments, the second selectable graphical user interface object (e.g., a "view mode" affordance) is continuously displayed during the real-time video communication session.
[0287] In some embodiments, when a real-time video communication interface (e.g., 604) is displayed when a seventh object (e.g., 628) (e.g., a first participant in a real-time video communication session) is detected in a scene (e.g., 615) in a field of view (e.g., 620) of one or more cameras (e.g., 602), the computer system (e.g., 600) detects an eighth object (e.g., 622) (e.g., a second participant in the real-time video communication session) in the scene in the field of view of the one or more cameras (e.g., detecting an increase in the number of objects in the scene). In response to detecting the eighth object in the scene in the field of view of the one or more cameras, the computer system displays a prompt (e.g., 632) (e.g., text to add a second participant to the real-time video communication session, an affordance for adding a second participant, an indication of the second participant (e.g., a framing indication (e.g., 630) in a potential preview (e.g., a blurred area) for showing recognition of the detected additional object), a stacked camera preview window, or as described with respect to Figures 8A to 8ROther prompts discussed herein) to adjust the representation of the field of view of the one or more cameras to include the representation of the eighth object (e.g., 622-1) in the real-time video communication interface (e.g., 604) (e.g., as Figure 6Q In response to detecting an eighth object in the scene, displaying a prompt to adjust the representation of the field of view of the one or more cameras to include the representation of the eighth object in the real-time video communication interface provides feedback to a user of the computer system that an additional object has been detected in the field of view of the one or more cameras, and reduces the number of user inputs at the computer system by providing an option for automatically adjusting the representation of the field of view to include the additional object without requiring the user to navigate a settings menu or other additional interface to adjust the representation of the field of view. Providing improved feedback and reducing the number of inputs at the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0288] In some embodiments, prior to detecting a change in a scene (e.g., 615) in the field of view (e.g., 620) of one or more cameras (e.g., 602), the representation (e.g., 606) of the field of view of the one or more cameras has a fifth representation field of view (e.g., Figure 6K In some embodiments, the change in the scene in the field of view of the one or more cameras comprises movement of one or more objects (e.g., 628) detected in the scene. In some embodiments, adjusting the representation of the field of view of the one or more cameras during the real-time video communications session based on the detected change in the field of view of the one or more cameras comprises: displaying in the real-time video communications interface (e.g., 604) a sixth representation of the field of view (e.g., 628) based on a determination that the one or more objects have less than a threshold amount of movement (e.g., a non-zero movement threshold) for at least a threshold amount of time (e.g., a predetermined amount of time (e.g., one second, two seconds, three seconds)). Figure 6L606) (e.g., different from the fifth representation field of view) (e.g., adjusting the representation of the field of view of the one or more cameras after the one or more objects have less than a threshold amount of movement for a predetermined amount of time). Based on determining that the one or more objects have not had less than a threshold amount of movement for at least a threshold amount of time, continuing to display the representation of the field of view of the one or more cameras with the fifth representation field of view in the real-time video communication interface (e.g., until the one or more objects have less than a threshold amount of movement for at least a threshold amount of time) (e.g., maintaining the originally displayed framing frame while one or more of the objects move). Selectively adjusting the representation of the field of view of the one or more cameras from the fifth representation field of view to the sixth representation field of view during the real-time video communication session based on whether one or more objects detected in the scene have less than a threshold amount of movement for at least a threshold amount of time enhances the video communication session experience by automatically adjusting the camera's field of view when additional objects enter the scene with the intent to participate in the real-time video communication session, while not adjusting the field of view when objects enter the scene with the intent not to participate. This also reduces the amount of computation performed by the computer system by eliminating extraneous adjustments to the representation field of view each time the number of participants in the scene changes. Performing an operation when a set of conditions have been met without requiring additional user input and reducing the amount of computation performed by the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, when the auto-framing mode is enabled, the computer system maintains the original display view until one or more of the objects are stationary.
[0289] In some embodiments, a computer system (e.g., 600, 600a) displays, via a display generation component (e.g., 601, 601a), a representation of a first portion of a field of view of one or more cameras (e.g., 606, 1006, 1056, 1208, 1218) of respective devices of respective participants in the real-time video communications session (e.g., a portion of the field of view of a camera of a remote participant in the real-time video communications session that includes a detected face of the remote participant (e.g., Figure 12L 1220-1b in; a portion of video feed 1210-1 including John's face; a portion of video feed 1023 including John's face; a portion of video feed 1053-2 including Jane's face)) (e.g., a portion of the field of view of one or more cameras of the computer system including the face of the detected object (e.g., Figure 12L1208-2 in the camera preview 1218; a portion of the camera preview 1006 that includes John's face; a portion of the camera preview 1056 that includes John's face) and a representation of a second portion of the field of view of one or more cameras of the respective devices of the respective participants (e.g., a portion of the field of view of the remote participant's camera that does not include the detected face of the remote participant (e.g., Figure 12L 1220-1a in; a portion of video feed 1210-1 that does not include John's face; a portion of video feed 1023 that does not include John's face; a portion of video feed 1053-2 that does not include Jane's face) (e.g., a portion of the field of view of one or more cameras of the computer system that does not include the face of the detected object (e.g., Figure 12L 1208-1 in; the portion of camera preview 1218 that does not include John's face; the portion of camera preview 1006 that does not include Jane's face; the portion of camera preview 1056 that does not include John's face)), including based on determining that a first portion of the field of view of the one or more cameras includes a corresponding type of detected feature (e.g., a face; a plurality of different faces) while a corresponding type of detected feature is not detected in a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., when a face (or a plurality of different faces) is detected in the first portion of the field of view but not in the second portion of the field of view, the second portion is compressed to a greater extent than the first portion (e.g., by the sending device (e.g., the remote participant's device (e.g., 600a, 600); the subject's computer system (e.g., 600, 600a))) so that when a face is detected in the first portion but not in the second portion, the first portion of the field of view can be displayed (e.g., at the receiving device (e.g., the computer system; the remote participant's device)) with a higher video quality than the second portion of the field of view (e.g., at the receiving device (e.g., the computer system; the remote participant's device)) Figure 12L1220-1b is displayed as having a higher video quality than 1220-1a; the portion of video feed 1210-1 that includes John's face is displayed as having a higher image quality than the portion of video feed 1210-1 that does not include John's face; the portion of video feed 1053-2 that includes Jane's face is displayed as having a higher video quality than the portion of video feed 1053-2 that does not include Jane's face; the portion of video feed 1023 that includes John's face is displayed as having a higher video quality than the portion of video feed 1023 that does not include John's face)) displays a representation of a first portion of the field of view of one or more cameras of the respective devices of the respective participants with a reduced degree of compression (e.g., higher video quality) as compared to a representation of a second portion of the field of view of the one or more cameras of the respective devices of the respective participants. Based on determining that a first portion of the field of view of the one or more cameras includes a detected feature of the corresponding type while a second portion of the field of view of the one or more cameras is not detected, a representation of a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant in the real-time video communication session is displayed at a reduced level of compression compared to a representation of a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant, which conserves computing resources by conserving bandwidth and reducing the amount of image data processed for display and / or transmission at high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0290] In some embodiments, a computer system (e.g., 600, 600a) enables a dynamic video quality mode for outputting (e.g., transmitting to a recipient device (e.g., 600a, 600), optionally simultaneously displaying at a sending device (e.g., 600, 600a)) portions of the camera field of view (e.g., 606, 1006, 1056, 1208, 1210-1, 1218, 1220-1, 1023, 1053-2) having different degrees of video compression. In some embodiments, the computer system compresses portions of the camera field of view that do not include one or more faces (e.g., Figure 12L 1220-1a in; the portion of video feed 1210-1 that does not include John's face; the portion of video feed 1023 that does not include John's face; the portion of video feed 1053-2 that does not include Jane's face) is larger than the portion of the camera's field of view that includes one or more faces (e.g., Figure 12LIn some embodiments, the computer system optionally displays the compressed video feeds in the camera preview. In some embodiments, the computer system transmits video feeds with different degrees of compression during a real-time video communication session, enabling a receiving device (e.g., a remote participant) to display a video feed received from a sending device (e.g., a computer system) having a higher video quality portion displayed simultaneously with a lower video quality portion, wherein the higher video quality portion of the video feed includes a face and the lower video quality portion of the video feed does not include a face (e.g., 1220-1b is displayed with a higher video quality portion than the video feed). Figure 12L 1 is displayed with a higher video quality than the portion of video feed 1220-1a in the real-time video communication session; the portion of video feed 1210-1 that includes John's face is displayed with a higher video quality than the portion of video feed 1210-1 that does not include John's face; the portion of video feed 1053-2 that includes Jane's face is displayed with a higher video quality than the portion of video feed 1053-2 that does not include Jane's face; the portion of video feed 1023 that includes John's face is displayed with a higher video quality than the portion of video feed 1023 that does not include John's face). Similarly, in some embodiments, a computer system receives compressed video data from a remote device (e.g., a device of a remote participant in a real-time video communication session) and displays video feeds from the remote device with different degrees of compression such that the video feed of the remote device can be displayed with a higher video quality portion that includes the remote participant's face and a lower video quality portion that does not include the remote participant's face (displayed simultaneously with the higher quality portion) (e.g., 1220-1b is displayed with a higher video quality than the portion of video feed 1220-1b that does not include the remote participant's face). Figure 12L (e.g., the portion of video feed 1053-2 that includes Jane's face is displayed with a higher video quality than the portion of video feed 1053-2 that does not include Jane's face; the portion of video feed 1023 that includes John's face is displayed with a higher video quality than the portion of video feed 1023 that does not include John's face.) In some embodiments, different degrees of compression can be applied to a video feed in which multiple faces are detected. For example, a video feed can have multiple higher quality (less compressed) portions, each corresponding to the location of one of the detected faces.
[0291] In some embodiments, the dynamic video quality mode is independent of the auto-framing mode and the background blur mode, such that the dynamic video quality mode can be enabled and disabled separately from the auto-framing mode and the background blur mode. In some embodiments, the dynamic video quality mode is implemented with the auto-framing mode, such that the dynamic video quality mode is enabled when the auto-framing mode is enabled and the dynamic video quality mode is disabled when the auto-framing mode is disabled. In some embodiments, the dynamic video quality mode is implemented with the background blur mode, such that the dynamic video quality mode is enabled when the background blur mode is enabled and the dynamic video quality mode is disabled when the background blur mode is disabled.
[0292] In some embodiments, after a feature of the corresponding type has moved from a first portion (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of the corresponding device of the corresponding participant to a second portion (e.g., detecting movement of the feature of the corresponding type from the first portion (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of the corresponding device of the corresponding participant to a second portion (e.g., detecting movement of the feature of the corresponding type from the first portion (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of the ... the computer system (e.g., 600, 600a), via display generation components (e.g., 601, 601a), displaying a representation of a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant and a representation of a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., a portion of the field of view that includes the detected face), including displaying the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant at an increased degree of compression (e.g., lower video quality) compared to the representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant based on a determination that the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant includes the corresponding type of detected feature while the corresponding type of detected feature is not detected in the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., as the face moves within the field of view of the one or more cameras, the degree of compression of the corresponding portion of the field of view of the one or more cameras changes such that the face (e.g., the portion of the field of view that includes the face) is output (e.g., retained) at a lower degree of compression than the portion of the field of view that does not include the face) ( For example, transmitted and optionally displayed) (e.g., as Jane's face moves, video feed 1053-2 and / or 1220-1 is updated so that her face continues to be displayed at a higher video quality, and portions of the video feed that do not include her face (even portions previously displayed at a higher quality) are displayed at a lower video quality; as John's face moves, video feed 1023 and / or 1210-1 is updated so that his face continues to be displayed at a higher video quality, and portions of the video feed that do not include his face (even portions previously displayed at a higher quality) are displayed at a lower video quality)).After a feature of the corresponding type has moved from a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant to a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant, based on a determination that the second portion of the field of view of the one or more cameras includes a detected feature of the corresponding type while no detected feature of the corresponding type is detected in the first portion of the field of view of the one or more cameras, a representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant in the real-time video communication session is displayed at an increased degree of compression compared to a representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant, which conserves computing resources by conserving bandwidth and reducing the amount of image data processed for display and / or transmission at high image quality as a face moves within the scene. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0293] In some embodiments, the corresponding type of feature is a face (e.g., a face detected within the field of view of one or more cameras; a face of a remote participant (e.g., Jane's face in video feed 1220-1 and / or 1053-2; John's face in video feed 1023 and / or 1210-1); a face of an object (e.g., Jane's face in camera previews 606, 1006, and / or 1208; John's face in camera previews 1056 and / or 1218)). In some embodiments, displaying a representation of a second portion of the field of view of one or more cameras of the corresponding device of the corresponding participant (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) includes displaying a representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant at a reduced video quality compared to the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., 1220-1a is displayed as having a reduced video quality compared to the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant based on determining that the first portion of the field of view of the one or more cameras includes a detected face while the face is not detected in the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant. Figure 12LThe portion of video feed 1053-2 that does not include John's face is displayed with a lower video quality than the portion of video feed 1053-1b that includes John's face; the portion of video feed 1053-2 that does not include Jane's face is displayed with a lower video quality than the portion of video feed 1053-2 that includes Jane's face; and the portion of video feed 1023 that does not include John's face is displayed with a lower video quality than the portion of video feed 1023 that includes John's face (e.g., due to reduced compression of the representation of the first portion of the field of view of the one or more cameras). Displaying a representation of the second portion of the field of view of the one or more cameras of the respective device of the respective participant at a reduced video quality compared to the representation of the first portion of the field of view based on a determination that the first portion of the field of view includes a detected face while the face is not detected in the second portion of the field of view of the one or more cameras of the respective device of the respective participant conserves computing resources by conserving bandwidth and reducing the amount of image data processed for display and / or transmission at high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0294] In some embodiments, portions of the camera's field of view that do not include detected faces (e.g., portions of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that do not include John's face; portions of 1006 and / or 1053-2 that do not include Jane's face) are output (e.g., transmitted and optionally displayed) at a lower image quality than portions of the camera's field of view that include detected faces (e.g., portions of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 that include John's face; portions of 1006 and / or 1053-2 that include Jane's face) (due to increased compression of the portions that do not include detected faces). In some embodiments, when no faces are detected in the field of view of one or more cameras, the computer system (e.g., 600, 600a) applies a uniform or substantially uniform degree of compression to a first portion and a second portion of the field of view of the one or more cameras, so that a video feed (e.g., both the first portion and the second portion) can be output with uniform or substantially uniform video quality. In some embodiments, when multiple faces are detected in the camera field of view (e.g., multiple participants in a real-time video communication session are detected), the computer system simultaneously applies reduced compression to the portions of the field of view corresponding to the detected faces, so that the faces can be displayed simultaneously (e.g., at a recipient device) with higher image quality. In some embodiments, the computer system applies increased compression to the representation of the second portion of the field of view of the one or more cameras, even if a face is detected in the second portion. For example, the computer system may determine that the face in the second portion is not a participant in the real-time video communication session (e.g., the person is a bystander in the background), and therefore does not reduce the degree of compression of the second portion containing the face.
[0295] In some embodiments, after a change in the bandwidth used to transmit a representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of the corresponding device of the corresponding participant (e.g., 600, 600a) occurs (e.g., is detected), when a feature (e.g., a face) of a corresponding type is detected in a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1053-2 that includes Jane's face), and in a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant. When a corresponding type of feature is not detected in the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face), the degree of compression (e.g., the amount of compression) of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed by a lesser amount than the amount of change in the degree of compression of the representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., when a face is detected in the first portion of the field of view of the one or more cameras and no face is detected in the second portion of the field of view of the one or more cameras, the rate of change of compression is smaller for the first portion of the field of view than for the second portion of the field of view (in response to a change in bandwidth (e.g., a reduction in bandwidth)). When a feature of the corresponding type is detected in the first portion and a feature of the corresponding type is not detected in the second portion, changing the degree of compression of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant by a smaller amount than the amount of change in the degree of compression of the representation of the second portion saves computing resources by conserving bandwidth for the first portion of the representation of the field of view of the one or more cameras that includes the corresponding type of feature and reducing the amount of image data processed for display and / or transmission with high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0296] In some embodiments, after a change in the bandwidth used to transmit a representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of the corresponding device of the corresponding participant (e.g., 600, 600a) occurs (e.g., is detected), when no feature of the corresponding type is detected in a first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., the portion of 621-1; 1220-1a; 1208-1; 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face), and in a second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant. When a feature of the corresponding type is detected in the portion (e.g., the portion including John's face of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1; the portion including Jane's face of 1006 and / or 1053-2), the degree of compression (e.g., the amount of compression) of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed by an amount that is greater than the amount of change in the degree of compression of the representation of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant (e.g., when a face is detected in the second portion of the field of view of the one or more cameras and no face is detected in the first portion of the field of view of the one or more cameras, the rate of change of compression of the first portion of the field of view is greater than the rate of change of compression of the second portion of the field of view (in response to a change in bandwidth (e.g., a reduction in bandwidth)). Changing the degree of compression of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant by an amount greater than the amount of the change in the degree of compression of the representation of the second portion when a feature of the corresponding type is not detected in the first portion and a feature of the corresponding type is detected in the second portion conserves computing resources by conserving bandwidth for the second portion of the representation of the field of view of the one or more cameras that includes the feature of the corresponding type and reducing the amount of image data processed for display and / or transmission at high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0297] In some embodiments, in response to a change in the bandwidth used to transmit the representation (e.g., 606, 1006, 1023, 1053-2, 1056, 1208, 1210-1, 1218, 1220-1) of the field of view of one or more cameras of the respective devices of the respective participants (e.g., 600, 600a) occurring (e.g., being detected), the quality (e.g., video quality) of the representation of the second portion (e.g., the portion of the field of view of the one or more cameras of the respective devices of the respective participants that does not include John's face) of the one or more cameras of the respective devices of the respective participants (e.g., the portion of 1006 and / or 1053-2 that does not include Jane's face) is reduced (e.g., due to the change in the amount of video compression). In some embodiments, the quality of the representation of the first portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant is changed by an amount that is greater than the amount of change in the quality of the representation of the first portion (e.g., the portion of 622-1; 1220-1b; 1208-2; 1218 and / or 1210-1 including John's face; the portion of 1006 and / or 1056-2 including Jane's face) (in some embodiments, the representation of the first portion does not change in quality or has a nominal amount of change in quality) (e.g., when a face is detected in the first portion of the field of view of the one or more cameras and no face is detected in the second portion of the field of view of the one or more cameras, the image quality of the second portion changes more than the image quality of the first portion in response to a change in bandwidth (e.g., a reduction in bandwidth). Changing the quality of the second portion of the field of view of the one or more cameras of the corresponding device of the corresponding participant by an amount that is greater than the amount of change in the quality of the representation of the first portion saves computing resources by conserving bandwidth for the first portion of the representation of the field of view of the one or more cameras and reducing the amount of image data processed for display and / or transmission at high image quality. Saving computing resources enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0298] In some embodiments, when a face is detected in a first portion of the field of view of one or more cameras (e.g., 622-1; 1220-1b; 1208-2; the portion of 1218 and / or 1210-1 that includes John's face; the portion of 1006 and / or 1053-2 that includes Jane's face) and not detected in a second portion of the field of view (e.g., 621-1; 1220-1a; 1208-1; the portion of 1218 and / or 1210-1 that does not include John's face; the portion of 1006 and / or 1053-2 that does not include Jane's face), the computer system (e.g., 600, 600a) detects a change in available bandwidth (e.g., an increase in bandwidth; a decrease in bandwidth) and, in response, adjusts (e.g., increases, decreases) the compression of the second portion of the representation of the field of view of the one or more cameras without adjusting the compression of the first portion of the representation of the field of view of the one or more cameras. In some embodiments, when a bandwidth change is detected, the computer system adjusts the compression of the first portion at a slower rate than the adjustment of the second portion. In some embodiments, the method includes detecting (e.g., at the respective device of the respective participant) a change in the bandwidth used to transmit the representation of the field of view of the one or more cameras of the respective device of the respective participant when a feature of the respective type (e.g., a face) is detected in a first portion of the field of view of the one or more cameras of the respective device of the respective participant, and when the feature of the respective type is not detected in a second portion of the field of view of the one or more cameras of the respective device of the respective participant.
[0299] Note that the above reference method 700 (e.g., 7A to 7B ) also apply in a similar manner to methods 900, 1100, 1300, and 1400 described below. For example, method 900, method 1100, method 1300, and / or method 1400 optionally include one or more features of the various methods described above with reference to method 700. For the sake of brevity, these details are not repeated below.
[0300] Figures 8A to 8R An exemplary user interface for managing a real-time video communication session (e.g., a video conference) according to some embodiments is shown. The user interfaces in these figures are used to illustrate the processes described herein, including Figure 9 in the process.
[0301] Figures 8A to 8R A device 600 is shown displaying a user interface for managing a real-time video communication session on a display 601, similar to that described above with respect to FIG. Figures 6A to 6Q discussed. Figures 8A to 8RVarious embodiments are depicted in which the device 600 prompts a user to adjust a displayed portion of the camera's field of view (e.g., the camera preview) in response to detecting another object in the scene when the auto-framing mode is enabled. Figures 8A to 8R One or more of the embodiments discussed may be combined with Figures 6A to 6Q 、 10A to 10J and Figures 12A to 12U Combinations of one or more of the discussed embodiments.
[0302] Figures 8A to 8J An example embodiment is depicted in which device 600 displays a prompt to adjust the field of view of the video feed to include the additional participant in response to detecting the additional object in scene 615 . Figures 8A to 8D An embodiment is shown in which the prompt includes a display stacked camera preview option. Figures 8E to 8G An embodiment is shown in which the prompting includes displaying a camera preview with obscured areas and unobscured areas. Figures 8H to 8J An embodiment is shown in which the prompt includes displaying an option to switch between a single-person view mode and a multi-person view mode. Other embodiments are also provided (such as those described above with respect to Figure 6P ) in which the device 600 display includes a prompt that can be selected to adjust the display field of view to include an enable representation of an additional object (e.g., add enable representation 632).
[0303] Figure 8A Describes a similar Figure 6O The embodiment discussed above is different in that video feed 623 now includes Pam's representation 623-2 rather than John's representation 623-1. Auto-framing mode is enabled, as indicated by the bolding of framing mode enable representation 610 and the display of framing indicator 630.
[0304] exist Figure 8A , Jack is participating in a video conference with Pam using device 600. Similarly, Pam is participating in a video conference with Jack using a device that includes one or more features of device 100, 300, 500, or 600. For example, Pam is using a tablet computer similar to device 600. Thus, Pam's device displays a video conference interface similar to video conference interface 604, except that the camera preview on Pam's device displays the video feed captured from Pam's device (currently in Figure 8A ), and the incoming video feed on Pam's device displays the video feed output from device 600 (currently in Figure 8A (as depicted in the camera preview 606 in FIG. 1 ).
[0305] exist Figure 8B, device 600 detects Jane 622 entering scene 615 within field of view 620. In response, device 600 updates video conferencing interface 604 by displaying secondary camera preview 806 positioned behind and offset from camera preview 606. The stacked appearance of camera preview 606 and secondary camera preview 806 indicates that multiple video feed fields of view are available for video conferencing and that the user can change the displayed field of view. Device 600 indicates that camera preview 606 is the currently selected or enabled video feed field of view because it is positioned on top of secondary camera preview 806. Thus, secondary camera preview 806 represents an option for adjusting the video feed field of view, which in this embodiment is an alternative field of view that includes both Jack 628 and Jane 622 (the additional object that has entered the scene). In some embodiments, different camera previews represent different zoom values, and therefore, different camera preview options can also be considered different zoom controls / options.
[0306] Device 600 detects input 804 (e.g., a tap input) on the stacked preview (e.g., on secondary camera preview 606) and, in response, updates video conferencing interface 604 by translating the position of secondary camera preview 806 so that it is no longer behind camera preview 806, as shown. Figure 8C As described in .
[0307] exist Figure 8C , the device 600 displays both the camera preview 606 and the secondary camera preview 806 separately (not stacked) in the video conferencing interface 604. The device 600 also displays a bold outline 807 to indicate the currently selected video feed field of view, which is Figure 8C 6. Secondary camera preview 806 represents the available video feed field of view provided by device 600. Portion 825 represents the portion of field of view 620 displayed in secondary camera preview 806, while portion 625 represents the portion of field of view 620 currently displayed in camera preview 606. Although camera preview 806 shows a rendering of portion 825, camera preview 806 is not currently selected, and therefore, device 600 is not currently outputting a view of portion 825 of the video conference.
[0308] The camera preview 606 includes a portion of Jack's representation 628-1 and Jane's representation 622-1. The secondary camera preview 806 is a zoomed-out view (compared to the view in the camera preview 606) that includes Jack's representation 628-1 and Jane's representation 622-1. As previously discussed, framing indicators 630 are depicted in the camera preview 606 and the secondary camera preview 806 to indicate that Jane's and Jack's faces are detected within the respective video feed fields of view.
[0309] When the camera preview option is Figure 8C , device 600 can maintain or switch between available preview options in response to user input. For example, if device 600 detects input 811 on camera preview 606, the device continues to use (e.g., outputs) the video feed field of view represented by camera preview 606, and video conferencing interface 604 returns to Figure 8B If device 600 detects input 812 on secondary camera preview 806, device 600 switches to (e.g., outputs) the video feed field of view represented by secondary camera preview 806, and the camera previews return to a stacked configuration with secondary camera preview 806 at the top and camera preview 606 at the bottom, as shown. Figure 8D In some embodiments, when the device 600 switches from the camera preview 606 to the auxiliary camera preview 806, the bold outline 807 moves from the camera preview 606 to the auxiliary camera preview 806 to indicate the switch from outputting the camera preview 606 to outputting the auxiliary camera preview 806.
[0310] exist Figure 8D , portion 625 represents the portion of field of view 620 currently being output for the video conference. Because secondary camera preview 806 was selected in response to input 812, portion 625 now corresponds to secondary camera preview 806. Portion 627 represents the previous video feed field of view, which now corresponds to camera preview 606.
[0311] exist Figure 8E , device 600 detects Jane 622 entering scene 615. Device 600 displays camera preview 606 with an unblurred area 606-1 (indicated by border 808 and without hatching) and a blurred area 606-2 (indicated by border 809 and with hatching). Unblurred area 606-1 represents the current video feed field of view, while blurred area 606-2 represents an additional video feed field of view that is available for video conferencing but is not currently being output. Thus, portion 625 corresponds to the field of view of the unblurred area, and portion 825 corresponds to the available field of view of the combined blurred and unblurred areas. The display of the blurred and unblurred areas in camera preview 606 indicates that the video feed field of view can be adjusted. The use of blur is described as one way to distinguish the current video feed field of view from the additional available video feed field of view. However, these areas can be distinguished by other visual indications and appearances, such as shading, darkening, highlighting, or other visual masking to emphasize or deemphasize individual areas. Boundaries 808 and 809 are also used to visually distinguish these areas.
[0312] exist Figure 8E, unblurred region 606-1 depicts an unobstructed representation 628-1a of Jack (specifically, Jack's face) that is being output for the video conference. Blurred region 606-2 depicts an obstructed (e.g., blurred) representation of the available video feed field of view that is included in portion 825 of field of view 620 but not included in portion 625. For example, in Figure 8E , the blurred area 606-2 depicts an obscured representation 628-1b of Jack's body. In some embodiments, the blurred area and / or the unblurred area includes a framing indicator when the device 600 detects a face in the corresponding area.
[0313] exist Figure 8F , Jane 622 has entered portion 825 of field of view 620, and device 600 displays an obscured representation 622-1b of Jane in blurred region 606-2 of the camera preview. Device 600 detects input 813 on blurred region 606-2 (or on a framing indicator around Jane's face in blurred region 606-2), and in response, adjusts (e.g., expands) the video feed field of view to include the previously blurred region 606-2, as shown. Figure 8G In some embodiments, the blurred / unblurred areas of the camera preview represent different zoom values of the video feed field of view. Thus, the camera preview 606, which can be selected to switch to a field of view with a different zoom value (e.g., by extending the unblurred area), can also be considered a zoom control.
[0314] Figure 8G The portion 625 in FIG. 1 represents the expanded portion of the field of view 620 that is now being output for video conferencing, and the portion 627 represents the portion corresponding to FIG. Figure 8F The camera preview 606 now depicts an unobstructed representation 628-1 of Jack and an unobstructed representation 622-1 of Jane.
[0315] In some embodiments, when the conditions triggering the adjustment are no longer met, Figures 8E to 8G The adjustment of the field of view of the video feed in question is reversed. For example, if Jane leaves Figure 8G 625 in the viewfinder, the device 600 returns to Figure 8E , where the camera preview 606 includes an unblurred area 606 - 1 and a blurred area 606 - 2 .
[0316] Figures 8H to 8J Depicts various interfaces for an embodiment in which the device 600 switches the automatic framing mode between a single-person framing mode setting and a multi-person framing mode setting. Figure 8HIn FIG. 6 , when the automatic framing mode is enabled, the device 600 detects Jack 628 in the scene 615 and displays the video conferencing interface 604, wherein the framing mode options 830 are depicted in the camera preview 606. The framing mode options 830 include a single-person option 830-1 and a multi-person option 830-2. When the single-person option 830-1 is in the selected state, as shown in FIG. Figure 8H As depicted in , the single person view mode setting is enabled and the device 600 keeps the video feed field of view focused on the face of a single user, even when another person is detected in the field of view 620. For example, in Figure 8I , although Jane 622 is now positioned next to Jack 628 in scene 615, device 600 maintains the video feed field of view (represented in camera preview 606 and portion 625) featuring Jack 628 rather than automatically adjusting the video feed field of view to include Jane.
[0317] exist Figure 8I , device 600 detects input 832 on multi-person option 830-2. In response, device 600 switches from a single-person view mode setting to a multi-person view mode setting. When the multi-person view mode setting is enabled, device 600 automatically adjusts the video feed field of view to include the additional objects (or a subset thereof) detected in field of view 620. For example, when device 600 switches to the multi-person view mode setting, device 600 expands the video feed field of view to include representations 628-1 and 622-1 of both Jack and Jane, as shown in FIG. Figure 8J As depicted in the camera preview 606. Thus, Figure 8J Portion 625 in the diagram represents the expanded video feed field of view resulting from the multi-person framing mode setting being enabled, and portion 627 represents the previous video feed field of view corresponding to the single-person framing mode setting. In some embodiments, framing mode options 830 correspond to video feed fields of view with different zoom values, and thus framing mode options 830 can also be considered zoom controls / options.
[0318] In some embodiments, Figures 8H to 8J The transformations described in can be compared to Figures 8E to 8G The discussed camera preview combination with blurred areas and unblurred areas. For example, the device 600 may display a camera preview 606 with blurred areas and unblurred areas, similar to Figure 8E as described in, but also includes Figure 8H When the device 600 is in the single person framing mode setting, the device 600 displays the camera preview 606 with both blurred and unblurred areas, regardless of whether any person is detected in the blurred portion of the framing frame (similar to Figure 8E and 8FHowever, when the device 600 is in a multi-person viewfinder mode setting, the device 600 may transition the camera preview 606 from a blurred and unblurred appearance to a unblurred appearance (similar to Figure 8F and Figure 8G In a similar manner, if a person is detected in the blurred area when the device 600 switches from the single-person view mode to the multi-person view mode (e.g. Figure 8F ), the device 600 adjusts the video feed field of view to include the previously blurred area, which includes the person previously detected in the blurred area (similar to Figure 8F and Figure 8G In some embodiments, the device 600 can reverse the above transformation. For example, if the device 600 is displaying a camera preview 606 with two objects in the field of view (similar to Figure 8J ), and the device 600 detects selection of the single-person view option 830-1, the device 600 may adjust the video feed field of view to return to a blurred / unblurred appearance, similar to Figure 8F As described in .
[0319] Now refer to Figure 8K , device 600 displays video conferencing interface 834, which depicts an incoming request to join a real-time video conference with John and two other remote participants. Video conferencing interface 834 is similar to video conferencing interface 604, except that multiple participants are active in the video conferencing session depicted in video conferencing interface 834. Therefore, the embodiments described herein with respect to video conferencing interface 604 can be applied in a similar manner to video conferencing interface 834. Similarly, the embodiments described herein with respect to video conferencing interface 834 can be applied in a similar manner to video conferencing interface 604, etc.
[0320] exist Figure 8K , device 600 detects input 835 on accepting option 608-3, while enabling auto-framing mode (as indicated by the bolded appearance of framing mode affordance 610) and disabling background blur mode (as indicated by the unbolded appearance of background blur affordance 611). Figure 8L As depicted in FIG, device 600 accepts a real-time video conference call and joins the video conference session with auto-framing mode enabled and background blur mode disabled.
[0321] Figure 8L6, and an image of the camera preview 836. The device 600 is depicted displaying a video conferencing interface 834 having a camera preview 606 (similar to the camera preview 836), and incoming video feeds 840-1, 840-2, and 840-3 for each of the respective remote participants of the real-time video conferencing session. The camera preview 836 includes a representation 622-1 of Jane and a framing mode enable indication 610. In some embodiments, the framing mode enable indication 610 is selectable in the camera preview 836 to enable or disable the automatic framing mode. In some embodiments, the framing mode enable indication 610 is not selectable until the camera preview 836 is displayed in a zoomed-in state, such as Figure 8M . In some embodiments, the device 600 displays the framing mode enable indication 610 in the camera preview 836 when the automatic framing mode is enabled, and does not display the enable indication when the automatic framing mode is disabled. In some embodiments, the device 600 persistently displays the framing mode enable indication 610 and indicates whether the automatic framing mode is enabled by changing the appearance of the framing mode enable indication (e.g., bolding the enable indication when the mode is enabled). In some embodiments, the framing mode enable indication 610 is displayed in the options menu 608.
[0322] exist Figure 8L , Jane is participating in a video conference with Pam, John, and Jack using device 600. Similarly, Pam, John, and Jack each use a respective device that includes one or more features of device 100, 300, 500, or 600 to participate in a video conference with Jane and the other respective participants. For example, John, Jack, and Pam each use a tablet computer similar to device 600. Thus, the devices of the other participants (John, Jack, and Pam) each display a video conference interface similar to video conference interface 834, except that the camera preview on each respective device displays the video feed captured from that user's respective device (e.g., Pam's camera preview displays what is currently depicted in video feed 840-3, John's camera preview displays what is currently depicted in video feed 840-1, and Jack's camera preview displays what is currently depicted in video feed 840-2), and the incoming video feeds on the devices of the other participants (John, Jack, and Pam) include the video feed output from device 600 (currently in Figure 8L ) and video feeds output from other participants' devices.
[0323] exist Figure 8L , device 600 detects input 837 on camera preview 836 and, in response, zooms in on camera preview 836, as shown in FIG. Figure 8M As depicted in Figure 8M , device 600 displays framing mode enable indications 610 in two locations in video conferencing interface 834. Framing mode enable indication 610-1 is displayed in camera preview 836, and framing mode enable indication 610-2 is displayed in options menu 608. In some embodiments, framing mode enable indication 610 is displayed in only one location (e.g., in options menu 608 or camera preview 836) at any given time.
[0324] In some embodiments, when the automatic framing mode is not available, device 600 displays a framing mode enabled indication 610 having a changed appearance. Figure 8N In the embodiment of the present invention, the lighting conditions in scene 615 are poor, and in response to detecting the poor lighting conditions, device 600 displays framing mode enable indication 610-1 and framing mode enable indication 610-2 having a grayed-out appearance to indicate that the automatic framing mode is currently unavailable. When the lighting conditions improve, device 600 displays the framing mode enable indication 610-1 and the framing mode enable indication 610-2 having a grayed-out appearance to indicate that the automatic framing mode is currently unavailable. Figure 8M The framing mode affordances of the appearance shown in FIG. 610 are representations 610 - 1 and 610 - 2 .
[0325] exist Figure 8N , the device 600 detects input 839 (e.g., a tap input or a drag gesture) on the options menu 608 and, in response, displays Figure 8O In some embodiments, the device 600 responds to detecting Figure 8N 608) in the video conferencing interface 834 (e.g., on the camera preview 836 or in a location other than the options menu 608 in the interface). Figure 8L Video conferencing interface 834 depicted in .
[0326] exist Figure 8O , device 600 displays video conferencing interface 834 with expanded options menu 845, incoming video feeds 840-2 and 840-3, and camera preview 836. Expanded options menu 845 includes information and various options for the video conference, including a view mode option 845-1, which is similar to view mode affordance 610.
[0327] Now refer to Figure 8P , the device 600 displays a video conferencing interface 834 with an enlarged camera preview 836 and control options 850. In some embodiments, in response to Figure 8O The camera preview 836 shows the input Figure 8P. The control options 850 include a framing mode option 850-1, a 1X zoom option 850-2, and a 0.5X zoom option 850-3. The framing mode option 850-1 is similar to the framing mode enable representation 610 and is displayed in a bold state to indicate that the automatic framing mode is enabled. Because the automatic framing mode is enabled, the camera preview 836 includes a framing indicator 852 (similar to the framing indicator 630) positioned around the face of Jane's representation 622-1. Zoom options 850-2 and 850-3 can be selected to manually change the digital zoom level of the video feed field of view. Because the zoom options 850-2 and 850-3 manually adjust the digital zoom of the camera preview 836, selecting a zoom option disables the automatic framing mode, as discussed below.
[0328] exist Figure 8P , device 600 detects input 853 on zoom option 850-2. In response, device 600 emphasizes (e.g., bolds and optionally magnifies) zoom option 850-2 and disables auto-framing mode. Figure 8P In the embodiment depicted in FIG, the 1X zoom level is the zoom setting before input 853 is detected. Therefore, device 600 continues to display representation 622-1 at the 1X zoom level. Because the automatic framing mode is disabled, the framing mode option 850-1 is de-emphasized (e.g., no longer bolded) and the framing indicator 852 is no longer displayed on the display. Figure 8Q middle.
[0329] exist Figure 8Q , device 60 detects input 855 on zoom option 850-3 and, in response, adjusts the digital zoom level, as shown in FIG. Figure 8R Thus, the device 600 emphasizes the zoom option 850-3, deemphasizes the zoom option 850-2, and displays the camera preview 836 with a 0.5X digital zoom value (compared to Figure 8Q 850 - 2 is selected). Portion 625 represents the portion of field of view 620 that is displayed after the video feed field of view is zoomed out, and portion 627 represents the portion of field of view 620 that was previously displayed when zoom option 850 - 2 was selected.
[0330] Figure 99 is a flowchart illustrating a method for managing a real-time video communication session using an electronic device according to some embodiments. Method 900 is performed at a computer system (e.g., a smartphone, a tablet) (e.g., 100, 300, 500, 600) that communicates with a display generation component (e.g., a display controller, a touch-sensitive display system), one or more cameras (e.g., 602) (e.g., a visible light camera, an infrared camera, a depth camera), and one or more input devices (e.g., a touch-sensitive surface). Some operations in method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0331] As described below, method 900 provides an intuitive way to manage real-time video communication sessions. The method reduces the cognitive burden on users managing real-time video communication sessions, thereby creating a more efficient human-computer interface. For battery-powered computing devices, this enables users to manage real-time video communication sessions faster and more efficiently, conserving power and increasing the time between battery charges.
[0332] In method 900, a computer system (e.g., 600) displays (902) a real-time video communication interface (e.g., 604, 834) for a real-time video communication session (e.g., an interface for a real-time video communication session (e.g., a real-time video chat session, a real-time video conferencing session, etc.) via a display generation component (e.g., 601). In some embodiments, the real-time video communication interface includes a real-time preview of a user of the computer system and real-time representations of one or more participants (e.g., remote users) of the real-time video communication session.
[0333] A computer system (e.g., 600) displays a real-time video communication interface (e.g., 604, 834) that includes (904) representations (e.g., 623, 623-1, 840-1, 840-2, 840-3) of one or more participants (e.g., remote participants of the real-time video communication session) in addition to participants (e.g., 622, 628) visible via one or more cameras (e.g., 602). In some embodiments, the participants visible via the one or more cameras are objects that are within the field of view of the one or more cameras (e.g., 620) and are represented (e.g., displayed) in the real-time video communication session (e.g., 606, 806) via a display generation component (e.g., 601).
[0334] A computer system (e.g., 600) displays a real-time video communication interface (e.g., 640, 834) that includes (904) (e.g., simultaneously with the representation of the one or more participants) a representation of the field of view of the one or more cameras (e.g., 606, 806, 836) that is visually associated with (e.g., displayed adjacent to; displayed in a group with) a visual indication (e.g., 610, 610-1, 610-2, 630, 632) of an option to change (e.g., adjust) the representation of the field of view of the one or more cameras (e.g., change digital zoom level / value, expand displayed field of view, shrink displayed field of view) during the real-time video communication session. , 806, 808, 809, 606-1, 606-2, 830, 830-1, 830-2, 845-1, 850, 850-1, 850-2, 850-3, 852) (e.g., prompts (e.g., text), selectable graphical user interface objects (e.g., zoom controls, framing mode enable representations, framing indications, enable representations for selecting a single-person framing mode, enable representations for selecting a multi-person framing mode), representations of camera previews (e.g., alternative camera previews), framing indications) (e.g., displaying a representation of the field of view of the one or more cameras at a first digital zoom level and a first displayed portion of the field of view of the one or more cameras). Displaying a representation of the field of view of one or more cameras visually associated with a visual indication of an option for changing the representation of the field of view of the one or more cameras during a real-time video communication session provides feedback to a user of the computer system that alternative representations of the field of view of the one or more cameras are available for selection, and reduces the number of user inputs at the computer system by providing the option to adjust the representation of the field of view without requiring the user to navigate a settings menu or other additional interface to adjust the representation of the field of view. Providing improved feedback and reducing the number of inputs at the computer system enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently.
[0335] In some embodiments, the representation of the field of view (e.g., 620) of the one or more cameras (e.g., 602) is a preview (e.g., 606, 806, 836) of image data output by the computer system (e.g., 600) or capable of being output to one or more electronic devices associated with one or more participants (e.g., remote participants) of the real-time video communication session. In some embodiments, the representation of the field of view of the one or more cameras includes representations (e.g., 622-1, 628-1) of objects (e.g., 622, 628) participating in the real-time video communication session (e.g., participants; users of the computer system detected within the field of view of the one or more cameras (e.g., 620) during the real-time video communication conference) (e.g., camera previews for users of the computer system for the real-time video communication session). In some embodiments, a visual indication (e.g., 850-2, 850-3, 806, 606, 632) of an option for changing the representation of the field of view of the one or more cameras can be selected to manually adjust the framing (e.g., digital zoom level) of the representation of the field of view of the one or more cameras during the real-time video communication session. In some embodiments, a visual indication of an option for changing the representation of the field of view of one or more cameras (e.g., 610, 610-1, 610-2, 845-1, 830-1, 830-2, 850-1) is selectable to enable or disable a mode for automatically adjusting the representation of the field of view of the one or more cameras during a real-time video communication session, where in some embodiments, the representation of the field of view of the one or more cameras includes a representation of an object and, optionally, one or more additional objects.
[0336] While displaying a real-time video communications interface (e.g., 604, 834) for a real-time video communications session, a computer system (e.g., 600) detects (908) via one or more input devices (e.g., 601) a set of one or more inputs (e.g., 626, 634, 804, 811, 812, 813, 832, 850-2, 850-3) corresponding to a request to initiate a process for adjusting (in some embodiments, manually; in some embodiments, automatically (e.g., without user input)) a representation of the field of view of the one or more cameras (e.g., 606, 806, 836) during the real-time video communications session.
[0337] In response to detecting the set of one or more inputs (e.g., 626, 634, 804, 811, 812, 813, 832, 850-2, 850-3), the computer system (e.g., 600) initiates (910) a process for adjusting the representa...
Claims
1. A method for managing a real-time video communication session, comprising: At a computer system in communication with one or more output generating components and one or more input devices: detecting, via the one or more input devices, a request to display a system interface; In response to detecting the request to display the system interface, displaying, via the one or more output generating components, a system interface including a plurality of simultaneously displayed controls for controlling different system functions of the computer system, comprising: Based on determining that a media communication session has been active for a predetermined amount of time, the plurality of concurrently displayed controls includes a set of one or more media communication controls, wherein the media communication controls provide access to media communication settings that determine how media is handled by the computer system during the media communication session; and displaying the plurality of concurrently displayed controls without displaying the set of one or more media communication controls based on determining that a media communication session has not been active within the predetermined amount of time; while displaying the system interface having the set of one or more media communication controls, detecting, via the one or more input devices, a set of one or more inputs including inputs directed to the set of one or more media communication controls; and When the corresponding media communication session has been active for the predetermined amount of time, in response to detecting the set of one or more inputs including inputs directed to the set of one or more media communication controls, adjusting the media communication settings for the corresponding media communication session. 2 . The method of claim 1 , wherein displaying the set of one or more media communication controls comprises displaying a control selected from the group consisting of a camera control and a microphone control.
3. The method according to claim 1, wherein: The set of one or more media communication controls includes a camera control that changes a communication setting of a camera when selected, and Displaying the set of one or more media communication controls includes: Based on determining that the camera is in an enabled state, displaying the camera control having a first visual appearance indicative of the enabled state of the camera; and Based on determining that the camera is in a disabled state, the camera control is displayed having a second visual appearance different from the first visual appearance that indicates the disabled state of the camera.
4. The method according to claim 1, wherein: The set of one or more media communication controls includes a microphone control that, when selected, changes a communication setting of a microphone, and Displaying the set of one or more media communication controls includes: Based on determining that the microphone is in an enabled state, displaying the microphone control having a third visual appearance indicating the enabled state of the microphone; and Based on determining that the microphone is in a disabled state, displaying the microphone control having a fourth visual appearance different from the third visual appearance that indicates the disabled state of the microphone.
5. The method of claim 1 , wherein displaying the system interface having the plurality of simultaneously displayed controls including the set of one or more media communication controls comprises: based on determining that one or more media communication settings are enabled, displaying the set of one or more media communication controls having a visual appearance indicative of an enabled state of the set of one or more media communication settings; as well as Based on determining that one or more media communication settings are disabled, displaying the set of one or more media communication controls having a visual appearance indicative of a disabled state of the one or more media communication settings.
6. The method of claim 1 , wherein displaying the system interface having the plurality of simultaneously displayed controls including the set of one or more media communication controls comprises: based on determining that a first media communication setting associated with the first media communication control is enabled, displaying the first media communication control having a first visual appearance indicative of an enabled state of the first media communication setting; as well as Based on determining that a second media communication setting associated with the first media communication control, different from the first media communication setting, is enabled, displaying the first media communication control having a second visual appearance different from the first visual appearance indicating an enabled state of the second media communication setting.
7. The method of claim 1 , wherein the set of one or more media communication controls includes a first camera control option that is selectable to initiate a process for changing the appearance of a representation of a background portion of a camera's field of view.
8. A method according to claim 1, wherein the set of one or more media communication controls includes a second camera control option that can be selected to enable a mode for automatically adjusting the representation of the field of view of the one or more cameras based on a change in the position of an object detected in the field of view of the one or more cameras.
9. The method of claim 8, wherein the second camera control option is disabled when the one or more cameras are incompatible with the mode for automatically adjusting the representation of the field of view of the one or more cameras.
10. The method of claim 1, wherein the set of one or more media communication controls comprises controls selected from the group consisting of: a first microphone control option selectable to enable a voice isolation mode of the microphone; a second microphone control option selectable to enable a music emphasis mode of the microphone; as well as A third microphone control option that can be selected to enable a mode that uses the microphone to filter out background noise.
11. The method according to claim 1 , further comprising: receiving input directed to the set of one or more media communication controls; as well as In response to receiving the input directed to the set of one or more media communication controls: changing media communication settings in a first manner for a first application running at the computer system; as well as Media communication settings are changed in the first manner for a second application running at the computer system that is different from the first application.
12. The method of claim 1 , wherein displaying the system interface having the plurality of simultaneously displayed controls including the set of one or more media communication controls comprises: displaying the set of one or more media communication controls concurrently with the camera controls based on determining that the camera has been active for the predetermined amount of time; as well as Based on determining that the camera has not been active for the predetermined amount of time, the set of one or more media communication controls is displayed without displaying the camera control.
13. The method of claim 1 , wherein displaying the system interface having the plurality of simultaneously displayed controls including the set of one or more media communication controls comprises: displaying the set of one or more media communication controls concurrently with the microphone audio control based on determining that the microphone has been active for the predetermined amount of time; as well as Based on determining that the microphone has not been active for the predetermined amount of time, the set of one or more media communication controls is displayed without displaying the microphone audio control.
14. The method of claim 1 , wherein displaying the system interface having the plurality of simultaneously displayed controls including the set of one or more media communication controls comprises displaying representations of one or more applications associated with the media communication session that has been active for the predetermined amount of time.
15. A method according to claim 1, wherein displaying the system interface having the multiple simultaneously displayed controls including the set of one or more media communication controls includes displaying a graphical user interface object, which is selectable to display a user interface including the media communication settings, which determine how media is handled by the computer system during a media communication session.
16. The method of claim 1 , wherein displaying the system interface having the plurality of concurrently displayed controls including the set of one or more media communication controls comprises: displaying a first set of media communication controls for a first media communication device, wherein the first set of media communication controls provides access to media communication settings that determine how media is handled using the first media communication device during a media communication session provided via a first application; as well as A second set of media communication controls for the first media communication device is displayed, wherein the second set of media communication controls provides access to media communication settings that determine how media is handled using the first media communication device during a media communication session provided via a second application that is different from the first application.
17. The method according to claim 1, further comprising: When the media communication session has not been active for said predetermined amount of time: receiving, via the one or more input devices, input corresponding to a request to display a settings user interface; as well as In response to the input corresponding to the request to display the settings user interface, a settings user interface is displayed via the one or more output generating components, the settings user interface including the media communication settings that determine how media is handled by the computer system during a media communication session.
18. The method of claim 17, wherein the settings user interface includes a selectable option for enabling a setting selected from the group consisting of: a first media communication setting for changing the appearance of a representation of a background portion of a camera's field of view, wherein the first media communication setting corresponds to an application running at the computer system that uses the camera; a second media communication setting for changing an appearance of a representation of a background portion of a camera's field of view, wherein the second media communication setting corresponds to a first category of applications running at the computer system that use the camera and does not correspond to a second category of applications running at the computer system that use the camera; as well as A third media communication setting is provided for changing an audio setting of a microphone, wherein the third media communication setting corresponds to an application running at the computer system that uses the microphone.
19. The method of claim 17, wherein: the settings user interface including a plurality of controls corresponding to a mode for automatically adjusting a representation of a field of view of the one or more cameras based on a change in position of an object detected in the field of view of the one or more cameras; The plurality of controls includes a first control selectable to enable a mode of a first application running at the computer system that uses the camera; and The plurality of controls includes a second control selectable to enable a mode of a second application running at the computer system that uses the camera, different from the first application.
20. The method of claim 1, wherein displaying the system interface having the plurality of concurrently displayed controls including the set of one or more media communication controls comprises: displaying the set of one or more media communication controls in a first area of the system interface and displaying a set of one or more system controls in a second area of the system interface based on determining that the media communication session has been active for the predetermined amount of time; as well as Based on determining that a media communication session has not been active for the predetermined amount of time, the set of one or more system controls is displayed at least in part in the first area of the system interface. The method of claim 20 , wherein the first area of the system interface is non-user configurable.
22. The method of claim 1, wherein the media communication session is a communication session between the computer system and one or more electronic devices.
23. A computer system comprising: one or more output generating components; one or more input devices; one or more processors; and A memory storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 22.
24. A non-transitory computer-readable storage medium storing one or more programs that, when executed by one or more processors of a computer system in communication with one or more output generating components and one or more input devices, cause the one or more processors to perform the method according to any one of claims 1-22.
25. A computer program product storing one or more programs which, when executed by one or more processors of a computer system in communication with one or more output generating components and one or more input devices, cause the one or more processors to perform the method according to any one of claims 1 to 22.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Acceleration-based theft detection system for portable electronic devices
US20050190059A1
Methods and apparatuses for operating a portable device based on an accelerometer
US20060017692A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1