WIDE-ANGLE VIDEO CONFERENCE

DE602022024528T2Active Publication Date: 2025-11-05APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602022024528
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-22
Filing Date
2022-09-23
Publication Date
2025-11-05
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

Existing techniques for managing live video communication sessions are cumbersome and inefficient, often requiring multiple key presses or keystrokes, wasting user time and device energy, particularly in battery-operated devices.

Method used

A method and system for managing live video communication sessions using a computer system with display generation components, cameras, and input devices, which detects user inputs and modifies the displayed image based on the position of the input relative to the cameras, allowing for faster and more efficient interaction.

Benefits of technology

The solution reduces cognitive burden on users, conserves power in battery-operated devices, and increases the time between battery charges by providing a more efficient human-machine interface.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

FIELD

[0001] The present disclosure relates generally to computer user interfaces, and more specifically to techniques for managing a live video communication session and / or managing digital content.BACKGROUND

[0002] Computer systems can include hardware and / or software for displaying an interface for a live video communication session. Japanese Patent Application H09-233 84A relates to video conferencing and describes an image input device connected with a communication control part, wherein an image dividing part divides the data of the image captured by a camera for every previously fixed area, an image specifying part recognises specified images such as the face and the whole of a person, etc., of an object from the divided images and an image magnifying part automatically magnifies the specified images and outputs the images.BRIEF SUMMARY

[0003] Some techniques for managing a live video communication session using electronic devices, however, are generally cumbersome and inefficient. For example, some existing techniques use a complex and time-consuming user interface, which may include multiple key presses or keystrokes. Existing techniques require more time than necessary, wasting user time and device energy. This latter consideration is particularly important in battery-operated devices.

[0004] Accordingly, the present technique provides electronic devices with faster, more efficient methods and interfaces for managing a live video communication session and / or managing digital content. Such methods and interfaces optionally complement or replace other methods for managing a live video communication session and / or managing digital content. Such methods and interfaces reduce the cognitive burden on a user and produce a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power and increase the time between battery charges.

[0005] The invention is set out in appended claims 1-12. In accordance with some embodiments, a method performed at a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices is described. The method comprises: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field-of-view of the one or more cameras; while displaying the live video communication interface, detecting, via the one or more input devices, one or more user inputs including a user input directed to a surface in a scene that is in the field-of-view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras that is modified based on a position of the surface relative to the one or more cameras.

[0006] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field-of-view of the one or more cameras; while displaying the live video communication interface, detecting, via the one or more input devices, one or more user inputs including a user input directed to a surface in a scene that is in the field-of-view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras that is modified based on a position of the surface relative to the one or more cameras.

[0007] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field-of-view of the one or more cameras; while displaying the live video communication interface, detecting, via the one or more input devices, one or more user inputs including a user input directed to a surface in a scene that is in the field-of-view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras that is modified based on a position of the surface relative to the one or more cameras.

[0008] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field-of-view of the one or more cameras; while displaying the live video communication interface, detecting, via the one or more input devices, one or more user inputs including a user input directed to a surface in a scene that is in the field-of-view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras that is modified based on a position of the surface relative to the one or more cameras.

[0009] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system comprises: means for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene that is in a field-of-view captured by the one or more cameras; and means, while displaying the live video communication interface, for obtaining, via the one or more cameras, image data for the field-of-view of the one or more cameras, the image data including a first gesture; and means, responsive to obtaining the image data for the field-of-view of the one or more cameras, for: in accordance with a determination that the first gesture satisfies a first set of criteria, displaying, via the display generation component, a representation of a second portion of the scene that is in the field-of-view of the one or more cameras, the representation of the second portion of the scene including different visual content from the representation of the first portion of the scene; and in accordance with a determination that the first gesture satisfies a second set of criteria different from the first set of criteria, continuing to display, via the display generation component, the representation of the first portion of the scene.

[0010] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene that is in a field-of-view captured by the one or more cameras; and while displaying the live video communication interface, obtaining, via the one or more cameras, image data for the field-of-view of the one or more cameras, the image data including a first gesture; and in response to obtaining the image data for the field-of-view of the one or more cameras: in accordance with a determination that the first gesture satisfies a first set of criteria, displaying, via the display generation component, a representation of a second portion of the scene that is in the field-of-view of the one or more cameras, the representation of the second portion of the scene including different visual content from the representation of the first portion of the scene; and in accordance with a determination that the first gesture satisfies a second set of criteria different from the first set of criteria, continuing to display, via the display generation component, the representation of the first portion of the scene.

[0011] In accordance with some embodiments, a method performed at a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices is described. The method comprises: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0012] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0013] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0014] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0015] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system comprises: means for detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; means, responsive to detecting the set of one or more user inputs, for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0016] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0017] In accordance with some embodiments, a method performed at a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices is described. The method comprises: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0018] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0019] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0020] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0021] In accordance with some embodiments, a computer system that is configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system comprises: means for detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; means, responsive to detecting the set of one or more user inputs, for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0022] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface of a live video communication session that includes a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including: a first representation of a field-of-view of the one or more first cameras of the first computer system; a second representation of the field-of-view of the one or more first cameras of the first computer system, the second representation of the field-of-view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is in the field-of-view of the one or more first cameras of the first computer system; a first representation of a field-of-view of one or more second cameras of a second computer system; and a second representation of the field-of-view of the one or more second cameras of the second computer system, the second representation of the field-of-view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is in the field-of-view of the one or more second cameras of the second computer system.

[0023] In accordance with some embodiments, a method is described. The method comprises: at a first computer system that is in communication with a first display generation component and one or more sensors: while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0024] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a first display generation component and one or more sensors, the one or more programs including instructions for: while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0025] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a first display generation component and one or more sensors, the one or more programs including instructions for: while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0026] In accordance with some embodiments, a computer system configured to communicate with a first display generation component and one or more sensors is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0027] In accordance with some embodiments, a computer system configured to communicate with a first display generation component and one or more sensors is described. The computer system comprises: means for, while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0028] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a first display generation component and one or more sensors, the one or more programs including instructions for: while the first computer system is in a live video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment that is in a field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in a position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system that is different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.

[0029] In accordance with some embodiments, a method is described. The method comprises: at a computer system that is in communication with a display generation component: displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0030] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0031] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0032] In accordance with some embodiments, a computer system configured to communicate with a display generation component is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0033] In accordance with some embodiments, a computer system configured to communicate with a display generation component is described. The computer system comprises: means for displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; means for, while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and means for, in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0034] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical mark and a physical background, and displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is in the field of view of the one or more cameras; while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras, obtaining data that includes a new physical mark in the physical environment; and in response to obtaining data representing the new physical mark in the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background that is in the field of view of the one or more cameras.

[0035] In accordance with some embodiments, a method is described. The method comprises: at a computer system that is in communication with a display generation component and one or more cameras: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0036] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0037] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0038] In accordance with some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0039] In accordance with some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system comprises: means for displaying, via the display generation component, an electronic document; means for detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and means for, in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0040] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting that includes physical marks on a physical surface that is in a field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting that includes physical marks on the physical surface that is in the field of view of the one or more cameras and is separate from the computer system, displaying, in the electronic document, digital text corresponding to the handwriting that is in the field of view of the one or more cameras.

[0041] In accordance with some embodiments, a method performed at a first computer system that is in communication with a display generation component, one or more cameras, and one or more input devices is described. The method comprises: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0042] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system that is in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0043] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0044] In accordance with some embodiments, a first computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0045] In accordance with some embodiments, a first computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system comprises: means for detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and means, responsive to detecting the one or more first user inputs, for: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0046] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a first computer system that is that is in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface that is in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: in accordance with a determination that a first set of one or more criteria is met, concurrently displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented as a view of the surface by a second computer system.

[0047] In accordance with some embodiments, a method is described. The method comprises: at a computer system that is in communication with a display generation component and one or more input devices: detecting, via the one or more input devices, a request to use a feature on the computer system; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0048] In accordance with some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to use a feature on the computer system; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0049] In accordance with some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to use a feature on the computer system; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0050] In accordance with some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system comprises: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, via the one or more input devices, a request to use a feature on the computer system; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0051] In accordance with some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system comprises: means for detecting, via the one or more input devices, a request to use a feature on the computer system; and means for, in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: means for, in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and means for, in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0052] In accordance with some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to use a feature on the computer system; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature that includes a virtual demonstration of the feature, including: in accordance with a determination that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance; and in accordance with a determination that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance that is different from the first appearance.

[0053] Executable instructions for performing these functions are, optionally, included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are, optionally, included in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0054] Thus, devices are provided with faster, more efficient methods and interfaces for managing a live video communication session, thereby increasing the effectiveness, efficiency, and user satisfaction with such devices. Such methods and interfaces may complement or replace other methods for managing a live video communication session.DESCRIPTION OF THE FIGURES

[0055] For a better understanding of the various described embodiments, reference should be made to the Description of Embodiments below, in conjunction with the following drawings in which like reference numerals refer to corresponding parts throughout the figures. FIG. 1A is a block diagram illustrating a portable multifunction device with a touch-sensitive display in accordance with some embodiments. FIG. 1B is a block diagram illustrating exemplary components for event handling in accordance with some embodiments. FIG. 2 illustrates a portable multifunction device having a touch screen in accordance with some embodiments. FIG. 3 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface in accordance with some embodiments. FIG. 4A illustrates an exemplary user interface for a menu of applications on a portable multifunction device in accordance with some embodiments. FIG. 4B illustrates an exemplary user interface for a multifunction device with a touch-sensitive surface that is separate from the display in accordance with some embodiments. FIG. 5A illustrates a personal electronic device in accordance with some embodiments. FIG. 5B is a block diagram illustrating a personal electronic device in accordance with some embodiments. FIG. 5C illustrates an exemplary diagram of a communication session between electronic devices, in accordance with some embodiments. FIGS. 6A-6AY illustrate exemplary user interfaces for managing a live video communication session, in accordance with some embodiments. FIG. 7 depicts a flow diagram illustrating a method for managing a live video communication session. FIG. 8 depicts a flow diagram illustrating a method for managing a live video communication session. FIGS. 9A-9T illustrate exemplary user interfaces for managing a live video communication session. FIG. 10 depicts a flow diagram illustrating a method for managing a live video communication session. FIGS. 11A-11P illustrate exemplary user interfaces for managing digital content. FIG. 12 is a flow diagram illustrating a method of managing digital content. FIGS. 13A-13K illustrate exemplary user interfaces for managing digital content. FIG. 14 is a flow diagram illustrating a method of managing digital content. FIG. 15 depicts a flow diagram illustrating a method for managing a live video communication session. FIGS. 16A-16Q illustrate exemplary user interfaces for managing a live video communication session, in accordance with some embodiments. FIG. 17 is a flow diagram illustrating a method for managing a live video communication session. FIGS. 18A-18N illustrate exemplary user interfaces for displaying a tutorial for a feature on a computer system. FIG. 19 is a flow diagram illustrating a method for displaying a tutorial for a feature on a computer system. DESCRIPTION OF EMBODIMENTS

[0056] The following description sets forth exemplary methods, parameters, and the like. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure but is instead provided as a description of exemplary embodiments.

[0057] There is a need for electronic devices that provide efficient methods and interfaces for managing a live video communication session and / or managing digital content. For example, there is a need for electronic devices to improve the sharing of content. Such techniques can reduce the cognitive burden on a user who shares content during live video communication session and / or manages digital content in an electronic document, thereby enhancing productivity. Further, such techniques can reduce processor and battery power otherwise wasted on redundant user inputs.

[0058] Below, FIGS. 1A-1B, 2, 3, 4A-4B, and 5A-5C provide a description of exemplary devices for performing the techniques for managing a live video communication session and / or managing digital content. FIGS. 6A-6AY illustrate exemplary user interfaces for managing a live video communication session. FIGS. 7-8, and 15 are flow diagrams illustrating methods of managing a live video communication session. The user interfaces in FIGS. 6A-6AY are used to illustrate the processes described below, including the processes in FIGS. 7-8, and 15. FIGS. 9A-9T illustrate exemplary user interfaces for managing a live video communication. FIG. 10 is a flow diagram illustrating methods of managing a live video communication. The user interfaces in FIGS. 9A-9T are used to illustrate the processes described below, including the process in FIG. 10. FIGS. 11A-11P illustrate exemplary user interfaces for managing digital content. FIG. 12 is a flow diagram illustrating methods of managing digital content. The user interfaces in FIGS. 11A-11P are used to illustrate the processes described below, including the process in FIG. 12. FIGS. 13A-13K illustrate exemplary user interfaces for managing digital content. FIG. 14 is a flow diagram illustrating methods of managing digital content. The user interfaces in FIGS. 13A-13K are used to illustrate the processes described below, including the process in FIG. 14. FIGS. 16A-16O illustrate exemplary user interfaces for managing a live video communication session. FIG. 17 is a flow diagram illustrating methods for managing a live video communication session. The user interfaces in FIGS. 16A-16Q are used to illustrate the processes described below, including the process in FIG. 17. FIGS. 18A-18N illustrate exemplary user interfaces for displaying a tutorial for a feature on a computer system. FIG. 19 is a flow diagram illustrating methods for displaying a tutorial for a feature on a computer system. The user interfaces in FIGS. 18A-18N are used to illustrate the processes described below, including the process in FIG. 19.

[0059] The processes described below enhance the operability of the devices and make the user-device interfaces more efficient (e.g., by helping the user to provide proper inputs and reducing user mistakes when operating / interacting with the device) through various techniques, including by providing improved visual feedback to the user, reducing the number of inputs needed to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, performing an operation when a set of conditions has been met without requiring further user input, improving efficiency in managing digital content, improving collaboration between users in a live communication session, improving the live communication session experience, and / or additional techniques. These techniques also reduce power usage and improve battery life of the device by enabling the user to use the device more quickly and efficiently.

[0060] In addition, in methods described herein where one or more steps are contingent upon one or more conditions having been met, it should be understood that the described method can be repeated in multiple repetitions so that over the course of the repetitions all of the conditions upon which steps in the method are contingent have been met in different repetitions of the method. For example, if a method requires performing a first step if a condition is satisfied, and a second step if the condition is not satisfied, then a person of ordinary skill would appreciate that the claimed steps are repeated until the condition has been both satisfied and not satisfied, in no particular order. Thus, a method described with one or more steps that are contingent upon one or more conditions having been met could be rewritten as a method that is repeated until each of the conditions described in the method has been met. This, however, is not required of system or computer readable medium claims where the system or computer readable medium contains instructions for performing the contingent operations based on the satisfaction of the corresponding one or more conditions and thus is capable of determining whether the contingency has or has not been satisfied without explicitly repeating steps of a method until all of the conditions upon which steps in the method are contingent have been met. A person having ordinary skill in the art would also understand that, similar to a method with contingent steps, a system or computer readable storage medium can repeat the steps of a method as many times as are needed to ensure that all of the contingent steps have been performed.

[0061] Although the following description uses terms "first," "second," etc. to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch could be termed a second touch, and, similarly, a second touch could be termed a first touch, without departing from the scope of the various described embodiments. In some embodiments, the first touch and the second touch are two separate references to the same touch. In some embodiments, the first touch and the second touch are both touches, but they are not the same touch.

[0062] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0063] The term "if" is, optionally, construed to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" is, optionally, construed to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]," depending on the context.

[0064] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communications device, such as a mobile telephone, that also contains other functions, such as PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, without limitation, the iPhone ®< , iPod Touch ®< , and iPad ®< devices from Apple Inc. of Cupertino, California. Other portable electronic devices, such as laptops or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touchpads), are, optionally, used. It should also be understood that, in some embodiments, the device is not a portable communications device, but is a desktop computer with a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). In some embodiments, the electronic device is a computer system that is in communication (e.g., via wireless communication, via wired communication) with a display generation component. The display generation component is configured to provide visual output, such as display via a CRT display, display via an LED display, or display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes causing to display the content (e.g., video data rendered or decoded by display controller 156) by transmitting, via a wired or wireless connection, data (e.g., image data or video data) to an integrated or external display generation component to visually produce the content.

[0065] In the discussion that follows, an electronic device that includes a display and a touch-sensitive surface is described. It should be understood, however, that the electronic device optionally includes one or more other physical user-interface devices, such as a physical keyboard, a mouse, and / or a joystick.

[0066] The device typically supports a variety of applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a gaming application, a telephone application, a video conferencing application, an e-mail application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0067] The various applications that are executed on the device optionally use at least one common physical user-interface device, such as the touch-sensitive surface. One or more functions of the touch-sensitive surface as well as corresponding information displayed on the device are, optionally, adjusted and / or varied from one application to the next and / or within a respective application. In this way, a common physical architecture (such as the touch-sensitive surface) of the device optionally supports the variety of applications with user interfaces that are intuitive and transparent to the user.

[0068] Attention is now directed toward embodiments of portable devices with touch-sensitive displays. FIG. 1A is a block diagram illustrating portable multifunction device 100 with touch-sensitive display system 112 in accordance with some embodiments. Touch-sensitive display 112 is sometimes called a "touch screen" for convenience and is sometimes known as or called a "touch-sensitive display system." Device 100 includes memory 102 (which optionally includes one or more computer-readable storage mediums), memory controller 122, one or more processing units (CPUs) 120, peripherals interface 118, RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, input / output (I / O) subsystem 106, other input control devices 116, and external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 for detecting intensity of contacts on device 100 (e.g., a touch-sensitive surface such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 for generating tactile outputs on device 100 (e.g., generating tactile outputs on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate over one or more communication buses or signal lines 103.

[0069] As used in the specification and claims, the term "intensity" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or to a substitute (proxy) for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four distinct values and more typically includes hundreds of distinct values (e.g., at least 256). Intensity of a contact is, optionally, determined (or measured) using various approaches and various sensors or combinations of sensors. For example, one or more force sensors underneath or adjacent to the touch-sensitive surface are, optionally, used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., a weighted average) to determine an estimated force of a contact. Similarly, a pressure-sensitive tip of a stylus is, optionally, used to determine a pressure of the stylus on the touch-sensitive surface. Alternatively, the size of the contact area detected on the touch-sensitive surface and / or changes thereto, the capacitance of the touch-sensitive surface proximate to the contact and / or changes thereto, and / or the resistance of the touch-sensitive surface proximate to the contact and / or changes thereto are, optionally, used as a substitute for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the substitute measurements for contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the substitute measurements). In some implementations, the substitute measurements for contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of a user input allows for user access to additional device functionality that may otherwise not be accessible by the user on a reduced-size device with limited real estate for displaying affordances (e.g., on a touch-sensitive display) and / or receiving user input (e.g., via a touch-sensitive display, a touch-sensitive surface, or a physical / mechanical control such as a knob or a button).

[0070] As used in the specification and claims, the term "tactile output" refers to physical displacement of a device relative to a previous position of the device, physical displacement of a component (e.g., a touch-sensitive surface) of a device relative to another component (e.g., housing) of the device, or displacement of the component relative to a center of mass of the device that will be detected by a user with the user's sense of touch. For example, in situations where the device or the component of the device is in contact with a surface of a user that is sensitive to touch (e.g., a finger, palm, or other part of a user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in physical characteristics of the device or the component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is, optionally, interpreted by the user as a "down click" or "up click" of a physical actuator button. In some cases, a user will feel a tactile sensation such as an "down click" or "up click" even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movements. As another example, movement of the touch-sensitive surface is, optionally, interpreted or sensed by the user as "roughness" of the touch-sensitive surface, even when there is no change in smoothness of the touch-sensitive surface. While such interpretations of touch by a user will be subject to the individualized sensory perceptions of the user, there are many sensory perceptions of touch that are common to a large majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., an "up click," a "down click," "roughness"), unless otherwise stated, the generated tactile output corresponds to physical displacement of the device or a component thereof that will generate the described sensory perception for a typical (or average) user.

[0071] It should be appreciated that device 100 is only one example of a portable multifunction device, and that device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. The various components shown in FIG. 1A are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0072] Memory 102 optionally includes high-speed random access memory and optionally also includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0073] Peripherals interface 118 can be used to couple input and output peripherals of the device to CPU 120 and memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or sets of instructions stored in memory 102 to perform various functions for device 100 and to process data. In some embodiments, peripherals interface 118, CPU 120, and memory controller 122 are, optionally, implemented on a single chip, such as chip 104. In some other embodiments, they are, optionally, implemented on separate chips.

[0074] RF (radio frequency) circuitry 108 receives and sends RF signals, also called electromagnetic signals. RF circuitry 108 converts electrical signals to / from electromagnetic signals and communicates with communications networks and other communications devices via the electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, and so forth. RF circuitry 108 optionally communicates with networks, such as the Internet, also referred to as the World Wide Web (WWW), an intranet and / or a wireless network, such as a cellular telephone network, a wireless local area network (LAN) and / or a metropolitan area network (MAN), and other devices by wireless communication. The RF circuitry 108 optionally includes well-known circuitry for detecting near field communication (NFC) fields, such as by a short-range communication radio. The wireless communication optionally uses any of a plurality of communications standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), long term evolution (LTE), near field communication (NFC), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, a protocol for e-mail (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this document.

[0075] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between a user and device 100. Audio circuitry 110 receives audio data from peripherals interface 118, converts the audio data to an electrical signal, and transmits the electrical signal to speaker 111. Speaker 111 converts the electrical signal to human-audible sound waves. Audio circuitry 110 also receives electrical signals converted by microphone 113 from sound waves. Audio circuitry 110 converts the electrical signal to audio data and transmits the audio data to peripherals interface 118 for processing. Audio data is, optionally, retrieved from and / or transmitted to memory 102 and / or RF circuitry 108 by peripherals interface 118. In some embodiments, audio circuitry 110 also includes a headset jack (e.g., 212, FIG. 2). The headset jack provides an interface between audio circuitry 110 and removable audio input / output peripherals, such as output-only headphones or a headset with both output (e.g., a headphone for one or both ears) and input (e.g., a microphone).

[0076] I / O subsystem 106 couples input / output peripherals on device 100, such as touch screen 112 and other input control devices 116, to peripherals interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, depth camera controller 169, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive / send electrical signals from / to other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, and so forth. In some embodiments, input controller(s) 160 are, optionally, coupled to any (or none) of the following: a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 208, FIG. 2) optionally include an up / down button for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., 206, FIG. 2). In some embodiments, the electronic device is a computer system that is in communication (e.g., via wireless communication, via wired communication) with one or more input devices. In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a trackpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independently of an input element that is a part of the device) and is based on detected motion of a portion of the user's body through the air including motion of the user's body relative to an absolute reference (e.g., an angle of the user's arm relative to the ground or a distance of the user's hand relative to the ground), relative to another portion of the user's body (e.g., movement of a hand of the user relative to a shoulder of the user, movement of one hand of the user relative to another hand of the user, and / or movement of a finger of the user relative to another finger or portion of a hand of the user), and / or absolute motion of a portion of the user's body (e.g., a tap gesture that includes movement of a hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes a predetermined speed or amount of rotation of a portion of the user's body).

[0077] A quick press of the push button optionally disengages a lock of touch screen 112 or optionally begins a process that uses gestures on the touch screen to unlock the device, as described in U.S. Patent Application 11 / 322,549, "Unlocking a Device by Performing Gestures on an Unlock Image," filed December 23, 2005, U.S. Pat. No. 7,657,849. A longer press of the push button (e.g., 206) optionally turns power to device 100 on or off. The functionality of one or more of the buttons are, optionally, user-customizable. Touch screen 112 is used to implement virtual or soft buttons and one or more soft keyboards.

[0078] Touch-sensitive display 112 provides an input interface and an output interface between the device and a user. Display controller 156 receives and / or sends electrical signals from / to touch screen 112. Touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively termed "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user-interface objects.

[0079] Touch screen 112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from the user based on haptic and / or tactile contact. Touch screen 112 and display controller 156 (along with any associated modules and / or sets of instructions in memory 102) detect contact (and any movement or breaking of the contact) on touch screen 112 and convert the detected contact into interaction with user-interface objects (e.g., one or more soft keys, icons, web pages, or images) that are displayed on touch screen 112. In an exemplary embodiment, a point of contact between touch screen 112 and the user corresponds to a finger of the user.

[0080] Touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies are used in other embodiments. Touch screen 112 and display controller 156 optionally detect contact and any movement or breaking thereof using any of a plurality of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with touch screen 112. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone ®< and iPod Touch ®< from Apple Inc. of Cupertino, California.

[0081] A touch-sensitive display in some embodiments of touch screen 112 is, optionally, analogous to the multi-touch sensitive touchpads described in the following U.S. Patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman), and / or U.S. Patent Publication 2002 / 0015024A1. However, touch screen 112 displays visual output from device 100, whereas touch-sensitive touchpads do not provide visual output.

[0082] A touch-sensitive display in some embodiments of touch screen 112 is described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, "Multipoint Touch Surface Controller," filed May 2, 2006; (2) U.S. Patent Application No. 10 / 840,862, "Multipoint Touchscreen," filed May 6, 2004; (3) U.S. Patent Application No. 10 / 903,964, "Gestures For Touch Sensitive Input Devices," filed July 30, 2004; (4) U.S. Patent Application No. 11 / 048,264, "Gestures For Touch Sensitive Input Devices," filed January 31, 2005; (5) U.S. Patent Application No. 11 / 038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices," filed January 18, 2005; (6) U.S. Patent Application No. 11 / 228,758, "Virtual Input Device Placement On A Touch Screen User Interface," filed September 16, 2005; (7) U.S. Patent Application No. 11 / 228,700, "Operation Of A Computer With A Touch Screen Interface," filed September 16, 2005; (8) U.S. Patent Application No. 11 / 228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," filed September 16, 2005; and (9) U.S. Patent Application No. 11 / 367,749, "Multi-Functional Hand-Held Device," filed March 3, 2006.

[0083] Touch screen 112 optionally has a video resolution in excess of 100 dpi. In some embodiments, the touch screen has a video resolution of approximately 160 dpi. The user optionally makes contact with touch screen 112 using any suitable object or appendage, such as a stylus, a finger, and so forth. In some embodiments, the user interface is designed to work primarily with finger-based contacts and gestures, which can be less precise than stylus-based input due to the larger area of contact of a finger on the touch screen. In some embodiments, the device translates the rough finger-based input into a precise pointer / cursor position or command for performing the actions desired by the user.

[0084] In some embodiments, in addition to the touch screen, device 100 optionally includes a touchpad for activating or deactivating particular functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad is, optionally, a touch-sensitive surface that is separate from touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.

[0085] Device 100 also includes power system 162 for powering the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)) and any other components associated with the generation, management and distribution of power in portable devices.

[0086] Device 100 optionally also includes one or more optical sensors 164. FIG. 1A shows an optical sensor coupled to optical sensor controller 158 in I / O subsystem 106. Optical sensor 164 optionally includes charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) phototransistors. Optical sensor 164 receives light from the environment, projected through one or more lenses, and converts the light to data representing an image. In conjunction with imaging module 143 (also called a camera module), optical sensor 164 optionally captures still images or video. In some embodiments, an optical sensor is located on the back of device 100, opposite touch screen display 112 on the front of the device so that the touch screen display is enabled for use as a viewfinder for still and / or video image acquisition. In some embodiments, an optical sensor is located on the front of the device so that the user's image is, optionally, obtained for video conferencing while the user views the other video conference participants on the touch screen display. In some embodiments, the position of optical sensor 164 can be changed by the user (e.g., by rotating the lens and the sensor in the device housing) so that a single optical sensor 164 is used along with the touch screen display for both video conferencing and still and / or video image acquisition.

[0087] Device 100 optionally also includes one or more depth camera sensors 175. FIG. 1A shows a depth camera sensor coupled to depth camera controller 169 in I / O subsystem 106. Depth camera sensor 175 receives data from the environment to create a three dimensional model of an object (e.g., a face) within a scene from a viewpoint (e.g., a depth camera sensor). In some embodiments, in conjunction with imaging module 143 (also called a camera module), depth camera sensor 175 is optionally used to determine a depth map of different portions of an image captured by the imaging module 143. In some embodiments, a depth camera sensor is located on the front of device 100 so that the user's image with depth information is, optionally, obtained for video conferencing while the user views the other video conference participants on the touch screen display and to capture selfies with depth map data. In some embodiments, the depth camera sensor 175 is located on the back of device, or on the back and the front of the device 100. In some embodiments, the position of depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and the sensor in the device housing) so that a depth camera sensor 175 is used along with the touch screen display for both video conferencing and still and / or video image acquisition.

[0088] In some embodiments, a depth map (e.g., depth map image) contains information (e.g., values) that relates to the distance of objects in a scene from a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines the position in the viewpoint's Z-axis where its corresponding two-dimensional pixel is located. In some embodiments, a depth map is composed of pixels wherein each pixel is defined by a value (e.g., 0 - 255). For example, the "0" value represents pixels that are located at the most distant place in a "three dimensional" scene and the "255" value represents pixels that are located closest to a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor) in the "three dimensional" scene. In other embodiments, a depth map represents the distance between an object in a scene and the plane of the viewpoint. In some embodiments, the depth map includes information about the relative depth of various features of an object of interest in view of the depth camera (e.g., the relative depth of eyes, nose, mouth, ears of a user's face). In some embodiments, the depth map includes information that enables the device to determine contours of the object of interest in a z direction.

[0089] Device 100 optionally also includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled to intensity sensor controller 159 in I / O subsystem 106. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is collocated with, or proximate to, a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the back of device 100, opposite touch screen display 112, which is located on the front of device 100.

[0090] Device 100 optionally also includes one or more proximity sensors 166. FIG. 1A shows proximity sensor 166 coupled to peripherals interface 118. Alternately, proximity sensor 166 is, optionally, coupled to input controller 160 in I / O subsystem 106. Proximity sensor 166 optionally performs as described in U.S. Patent Application Nos. 11 / 241,839, "Proximity Detector In Handheld Device"; 11 / 240,788, "Proximity Detector In Handheld Device"; 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals". In some embodiments, the proximity sensor turns off and disables touch screen 112 when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).

[0091] Device 100 optionally also includes one or more tactile output generators 167. FIG. 1A shows a tactile output generator coupled to haptic feedback controller 161 in I / O subsystem 106. Tactile output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components and / or electromechanical devices that convert energy into linear motion such as a motor, solenoid, electroactive polymer, piezoelectric actuator, electrostatic actuator, or other tactile output generating component (e.g., a component that converts electrical signals into tactile outputs on the device). Contact intensity sensor 165 receives tactile feedback generation instructions from haptic feedback module 133 and generates tactile outputs on device 100 that are capable of being sensed by a user of device 100. In some embodiments, at least one tactile output generator is collocated with, or proximate to, a touch-sensitive surface (e.g., touch-sensitive display system 112) and, optionally, generates a tactile output by moving the touch-sensitive surface vertically (e.g., in / out of a surface of device 100) or laterally (e.g., back and forth in the same plane as a surface of device 100). In some embodiments, at least one tactile output generator sensor is located on the back of device 100, opposite touch screen display 112, which is located on the front of device 100.

[0092] Device 100 optionally also includes one or more accelerometers 168. FIG. 1A shows accelerometer 168 coupled to peripherals interface 118. Alternately, accelerometer 168 is, optionally, coupled to an input controller 160 in I / O subsystem 106. Accelerometer 168 optionally performs as described in U.S. Patent Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and U.S. Patent Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer". In some embodiments, information is displayed on the touch screen display in a portrait view or a landscape view based on an analysis of data received from the one or more accelerometers. Device 100 optionally includes, in addition to accelerometer(s) 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for obtaining information concerning the location and orientation (e.g., portrait or landscape) of device 100.

[0093] In some embodiments, the software components stored in memory 102 include operating system 126, communication module (or set of instructions) 128, contact / motion module (or set of instructions) 130, graphics module (or set of instructions) 132, text input module (or set of instructions) 134, Global Positioning System (GPS) module (or set of instructions) 135, and applications (or sets of instructions) 136. Furthermore, in some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) stores device / global internal state 157, as shown in FIGS. 1A and 3. Device / global internal state 157 includes one or more of: active application state, indicating which applications, if any, are currently active; display state, indicating what applications, views or other information occupy various regions of touch screen display 112; sensor state, including information obtained from the device's various sensors and input control devices 116; and location information concerning the device's location and / or attitude.

[0094] Operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.

[0095] Communication module 128 facilitates communication with other devices over one or more external ports 124 and also includes various software components for handling data received by RF circuitry 108 and / or external port 124. External port 124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted for coupling directly to other devices or indirectly over a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, or similar to and / or compatible with, the 30-pin connector used on iPod ®< (trademark of Apple Inc.) devices.

[0096] Contact / motion module 130 optionally detects contact with touch screen 112 (in conjunction with display controller 156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). Contact / motion module 130 includes various software components for performing various operations related to detection of contact, such as determining if contact has occurred (e.g., detecting a finger-down event), determining an intensity of the contact (e.g., the force or pressure of the contact or a substitute for the force or pressure of the contact), determining if there is movement of the contact and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger-dragging events), and determining if the contact has ceased (e.g., detecting a finger-up event or a break in contact). Contact / motion module 130 receives contact data from the touch-sensitive surface. Determining movement of the point of contact, which is represented by a series of contact data, optionally includes determining speed (magnitude), velocity (magnitude and direction), and / or an acceleration (a change in magnitude and / or direction) of the point of contact. These operations are, optionally, applied to single contacts (e.g., one finger contacts) or to multiple simultaneous contacts (e.g., "multitouch" / multiple finger contacts). In some embodiments, contact / motion module 130 and display controller 156 detect contact on a touchpad.

[0097] In some embodiments, contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by a user (e.g., to determine whether a user has "clicked" on an icon). In some embodiments, at least a subset of the intensity thresholds are determined in accordance with software parameters (e.g., the intensity thresholds are not determined by the activation thresholds of particular physical actuators and can be adjusted without changing the physical hardware of device 100). For example, a mouse "click" threshold of a trackpad or touch screen display can be set to any of a large range of predefined threshold values without changing the trackpad or touch screen display hardware. Additionally, in some implementations, a user of the device is provided with software settings for adjusting one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting a plurality of intensity thresholds at once with a system-level click "intensity" parameter).

[0098] Contact / motion module 130 optionally detects a gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of detected contacts). Thus, a gesture is, optionally, detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger-down event followed by detecting a finger-up (liftoff) event at the same position (or substantially the same position) as the finger-down event (e.g., at the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger-down event followed by detecting one or more finger-dragging events, and subsequently followed by detecting a finger-up (liftoff) event.

[0099] Graphics module 132 includes various known software components for rendering and displaying graphics on touch screen 112 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual property) of graphics that are displayed. As used herein, the term "graphics" includes any object that can be displayed to a user, including, without limitation, text, web pages, icons (such as user-interface objects including soft keys), digital images, videos, animations, and the like.

[0100] In some embodiments, graphics module 132 stores data representing graphics to be used. Each graphic is, optionally, assigned a corresponding code. Graphics module 132 receives, from applications etc., one or more codes specifying graphics to be displayed along with, if necessary, coordinate data and other graphic property data, and then generates screen image data to output to display controller 156.

[0101] Haptic feedback module 133 includes various software components for generating instructions used by tactile output generator(s) 167 to produce tactile outputs at one or more locations on device 100 in response to user interactions with device 100.

[0102] Text input module 134, which is, optionally, a component of graphics module 132, provides soft keyboards for entering text in various applications (e.g., contacts 137, e-mail 140, IM 141, browser 147, and any other application that needs text input).

[0103] GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to telephone 138 for use in location-based dialing; to camera 143 as picture / video metadata; and to applications that provide location-based services such as weather widgets, local yellow page widgets, and map / navigation widgets).

[0104] Applications 136 optionally include the following modules (or sets of instructions), or a subset or superset thereof: Contacts module 137 (sometimes called an address book or contact list); Telephone module 138; Video conference module 139; E-mail client module 140; Instant messaging (IM) module 141; Workout support module 142; Camera module 143 for still and / or video images; Image management module 144; Video player module; Music player module; Browser module 147; Calendar module 148; Widget modules 149, which optionally include one or more of: weather widget 149-1, stocks widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6; Widget creator module 150 for making user-created widgets 149-6; Search module 151; Video and music player module 152, which merges video player module and music player module; Notes module 153; Map module 154; and / or Online video module 155.

[0105] Examples of other applications 136 that are, optionally, stored in memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice replication.

[0106] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, contacts module 137 are, optionally, used to manage an address book or contact list (e.g., stored in application internal state 192 of contacts module 137 in memory 102 or memory 370), including: adding name(s) to the address book; deleting name(s) from the address book; associating telephone number(s), e-mail address(es), physical address(es) or other information with a name; associating an image with a name; categorizing and sorting names; providing telephone numbers or e-mail addresses to initiate and / or facilitate communications by telephone 138, video conference module 139, e-mail 140, or IM 141; and so forth.

[0107] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, telephone module 138 are optionally, used to enter a sequence of characters corresponding to a telephone number, access one or more telephone numbers in contacts module 137, modify a telephone number that has been entered, dial a respective telephone number, conduct a conversation, and disconnect or hang up when the conversation is completed. As noted above, the wireless communication optionally uses any of a plurality of communications standards, protocols, and technologies.

[0108] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contacts module 137, and telephone module 138, video conference module 139 includes executable instructions to initiate, conduct, and terminate a video conference between a user and one or more other participants in accordance with user instructions.

[0109] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, e-mail client module 140 includes executable instructions to create, send, receive, and manage e-mail in response to user instructions. In conjunction with image management module 144, e-mail client module 140 makes it very easy to create and send e-mails with still or video images taken with camera module 143.

[0110] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the instant messaging module 141 includes executable instructions to enter a sequence of characters corresponding to an instant message, to modify previously entered characters, to transmit a respective instant message (for example, using a Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for telephony-based instant messages or using XMPP, SIMPLE, or IMPS for Internet-based instant messages), to receive instant messages, and to view received instant messages. In some embodiments, transmitted and / or received instant messages optionally include graphics, photos, audio files, video files and / or other attachments as are supported in an MMS and / or an Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephony-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0111] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, workout support module 142 includes executable instructions to create workouts (e.g., with time, distance, and / or calorie burning goals); communicate with workout sensors (sports devices); receive workout sensor data; calibrate sensors used to monitor a workout; select and play music for a workout; and display, store, and transmit workout data.

[0112] In conjunction with touch screen 112, display controller 156, optical sensor(s) 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, camera module 143 includes executable instructions to capture still images or video (including a video stream) and store them into memory 102, modify characteristics of a still image or video, or delete a still image or video from memory 102.

[0113] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and camera module 143, image management module 144 includes executable instructions to arrange, modify (e.g., edit), or otherwise manipulate, label, delete, present (e.g., in a digital slide show or album), and store still and / or video images.

[0114] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions to browse the Internet in accordance with user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.

[0115] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, e-mail client module 140, and browser module 147, calendar module 148 includes executable instructions to create, display, modify, and store calendars and data associated with calendars (e.g., calendar entries, to-do lists, etc.) in accordance with user instructions.

[0116] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget modules 149 are mini-applications that are, optionally, downloaded and used by a user (e.g., weather widget 149-1, stocks widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or created by the user (e.g., user-created widget 149-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! Widgets).

[0117] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, the widget creator module 150 are, optionally, used by a user to create widgets (e.g., turning a user-specified portion of a web page into a widget).

[0118] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, search module 151 includes executable instructions to search for text, music, sound, image, video, and / or other files in memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.

[0119] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that allow the user to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, and executable instructions to display, present, or otherwise play back videos (e.g., on touch screen 112 or on an external, connected display via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (trademark of Apple Inc.).

[0120] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, notes module 153 includes executable instructions to create and manage notes, to-do lists, and the like in accordance with user instructions.

[0121] In conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 are, optionally, used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data on stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.

[0122] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, e-mail client module 140, and browser module 147, online video module 155 includes instructions that allow the user to access, browse, receive (e.g., by streaming and / or download), play back (e.g., on the touch screen or on an external, connected display via external port 124), send an e-mail with a link to a particular online video, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 141, rather than e-mail client module 140, is used to send a link to a particular online video. Additional description of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," filed June 20, 2007, and U.S. Patent Application No. 11 / 968,067, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," filed December 31, 2007.

[0123] Each of the above-identified modules and applications corresponds to a set of executable instructions for performing one or more functions described above and the methods described in this application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), procedures, or modules, and thus various subsets of these modules are, optionally, combined or otherwise rearranged in various embodiments. For example, video player module is, optionally, combined with music player module into a single module (e.g., video and music player module 152, FIG. 1A). In some embodiments, memory 102 optionally stores a subset of the modules and data structures identified above. Furthermore, memory 102 optionally stores additional modules and data structures not described above.

[0124] In some embodiments, device 100 is a device where operation of a predefined set of functions on the device is performed exclusively through a touch screen and / or a touchpad. By using a touch screen and / or a touchpad as the primary input control device for operation of device 100, the number of physical input control devices (such as push buttons, dials, and the like) on device 100 is, optionally, reduced.

[0125] The predefined set of functions that are performed exclusively through a touch screen and / or a touchpad optionally include navigation between user interfaces. In some embodiments, the touchpad, when touched by the user, navigates device 100 to a main, home, or root menu from any user interface that is displayed on device 100. In such embodiments, a "menu button" is implemented using a touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device instead of a touchpad.

[0126] FIG. 1B is a block diagram illustrating exemplary components for event handling in accordance with some embodiments. In some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) includes event sorter 170 (e.g., in operating system 126) and a respective application 136-1 (e.g., any of the aforementioned applications 137-151, 155, 380-390).

[0127] Event sorter 170 receives event information and determines the application 136-1 and application view 191 of application 136-1 to which to deliver the event information. Event sorter 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates the current application view(s) displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event sorter 170 to determine which application(s) is (are) currently active, and application internal state 192 is used by event sorter 170 to determine application views 191 to which to deliver event information.

[0128] In some embodiments, application internal state 192 includes additional information, such as one or more of: resume information to be used when application 136-1 resumes execution, user interface state information that indicates information being displayed or that is ready for display by application 136-1, a state queue for enabling the user to go back to a prior state or view of application 136-1, and a redo / undo queue of previous actions taken by the user.

[0129] Event monitor 171 receives event information from peripherals interface 118. Event information includes information about a sub-event (e.g., a user touch on touch-sensitive display 112, as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or a sensor, such as proximity sensor 166, accelerometer(s) 168, and / or microphone 113 (through audio circuitry 110). Information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display 112 or a touch-sensitive surface.

[0130] In some embodiments, event monitor 171 sends requests to the peripherals interface 118 at predetermined intervals. In response, peripherals interface 118 transmits event information. In other embodiments, peripherals interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or for more than a predetermined duration).

[0131] In some embodiments, event sorter 170 also includes a hit view determination module 172 and / or an active event recognizer determination module 173.

[0132] Hit view determination module 172 provides software procedures for determining where a sub-event has taken place within one or more views when touch-sensitive display 112 displays more than one view. Views are made up of controls and other elements that a user can see on the display.

[0133] Another aspect of the user interface associated with an application is a set of views, sometimes herein called application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of a respective application) in which a touch is detected optionally correspond to programmatic levels within a programmatic or view hierarchy of the application. For example, the lowest level view in which a touch is detected is, optionally, called the hit view, and the set of events that are recognized as proper inputs are, optionally, determined based, at least in part, on the hit view of the initial touch that begins a touch-based gesture.

[0134] Hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies a hit view as the lowest view in the hierarchy which should handle the sub-event. In most circumstances, the hit view is the lowest level view in which an initiating sub-event occurs (e.g., the first sub-event in the sequence of sub-events that form an event or potential event). Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.

[0135] Active event recognizer determination module 173 determines which view or views within a view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that include the physical location of a sub-event are actively involved views, and therefore determines that all actively involved views should receive a particular sequence of sub-events. In other embodiments, even if touch sub-events were entirely confined to the area associated with one particular view, views higher in the hierarchy would still remain as actively involved views.

[0136] Event dispatcher module 174 dispatches the event information to an event recognizer (e.g., event recognizer 180). In embodiments including active event recognizer determination module 173, event dispatcher module 174 delivers the event information to an event recognizer determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores in an event queue the event information, which is retrieved by a respective event receiver 182.

[0137] In some embodiments, operating system 126 includes event sorter 170. Alternatively, application 136-1 includes event sorter 170. In yet other embodiments, event sorter 170 is a stand-alone module, or a part of another module stored in memory 102, such as contact / motion module 130.

[0138] In some embodiments, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events that occur within a respective view of the application's user interface. Each application view 191 of the application 136-1 includes one or more event recognizers 180. Typically, a respective application view 191 includes a plurality of event recognizers 180. In other embodiments, one or more of event recognizers 180 are part of a separate module, such as a user interface kit or a higher level object from which application 136-1 inherits methods and other properties. In some embodiments, a respective event handler 190 includes one or more of: data updater 176, object updater 177, GUI updater 178, and / or event data 179 received from event sorter 170. Event handler 190 optionally utilizes or calls data updater 176, object updater 177, or GUI updater 178 to update the application internal state 192. Alternatively, one or more of the application views 191 include one or more respective event handlers 190. Also, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in a respective application view 191.

[0139] A respective event recognizer 180 receives event information (e.g., event data 179) from event sorter 170 and identifies an event from the event information. Event recognizer 180 includes event receiver 182 and event comparator 184. In some embodiments, event recognizer 180 also includes at least a subset of: metadata 183, and event delivery instructions 188 (which optionally include sub-event delivery instructions).

[0140] Event receiver 182 receives event information from event sorter 170. The event information includes information about a sub-event, for example, a touch or a touch movement. Depending on the sub-event, the event information also includes additional information, such as location of the sub-event. When the sub-event concerns motion of a touch, the event information optionally also includes speed and direction of the sub-event. In some embodiments, events include rotation of the device from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation (also called device attitude) of the device.

[0141] Event comparator 184 compares the event information to predefined event or sub-event definitions and, based on the comparison, determines an event or sub-event, or determines or updates the state of an event or sub-event. In some embodiments, event comparator 184 includes event definitions 186. Event definitions 186 contain definitions of events (e.g., predefined sequences of sub-events), for example, event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in an event (187) include, for example, touch begin, touch end, touch movement, touch cancellation, and multiple touching. In one example, the definition for event 1 (187-1) is a double tap on a displayed object. The double tap, for example, comprises a first touch (touch begin) on the displayed object for a predetermined phase, a first liftoff (touch end) for a predetermined phase, a second touch (touch begin) on the displayed object for a predetermined phase, and a second liftoff (touch end) for a predetermined phase. In another example, the definition for event 2 (187-2) is a dragging on a displayed object. The dragging, for example, comprises a touch (or contact) on the displayed object for a predetermined phase, a movement of the touch across touch-sensitive display 112, and liftoff of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.

[0142] In some embodiments, event definition 187 includes a definition of an event for a respective user-interface object. In some embodiments, event comparator 184 performs a hit test to determine which user-interface object is associated with a sub-event. For example, in an application view in which three user-interface objects are displayed on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user-interface objects is associated with the touch (sub-event). If each displayed object is associated with a respective event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects an event handler associated with the sub-event and the object triggering the hit test.

[0143] In some embodiments, the definition for a respective event (187) also includes delayed actions that delay delivery of the event information until after it has been determined whether the sequence of sub-events does or does not correspond to the event recognizer's event type.

[0144] When a respective event recognizer 180 determines that the series of sub-events do not match any of the events in event definitions 186, the respective event recognizer 180 enters an event impossible, event failed, or event ended state, after which it disregards subsequent sub-events of the touch-based gesture. In this situation, other event recognizers, if any, that remain active for the hit view continue to track and process sub-events of an ongoing touch-based gesture.

[0145] In some embodiments, a respective event recognizer 180 includes metadata 183 with configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery to actively involved event recognizers. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact, or are enabled to interact, with one another. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to varying levels in the view or programmatic hierarchy.

[0146] In some embodiments, a respective event recognizer 180 activates event handler 190 associated with an event when one or more particular sub-events of an event are recognized. In some embodiments, a respective event recognizer 180 delivers event information associated with the event to event handler 190. Activating an event handler 190 is distinct from sending (and deferred sending) sub-events to a respective hit view. In some embodiments, event recognizer 180 throws a flag associated with the recognized event, and event handler 190 associated with the flag catches the flag and performs a predefined process.

[0147] In some embodiments, event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver event information to event handlers associated with the series of sub-events or to actively involved views. Event handlers associated with the series of sub-events or with actively involved views receive the event information and perform a predetermined process.

[0148] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates the telephone number used in contacts module 137, or stores a video file used in video player module. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates a new user-interface object or updates the position of a user-interface object. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends it to graphics module 132 for display on a touch-sensitive display.

[0149] In some embodiments, event handler(s) 190 includes or has access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of a respective application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.

[0150] It shall be understood that the foregoing discussion regarding event handling of user touches on touch-sensitive displays also applies to other forms of user inputs to operate multifunction devices 100 with input devices, not all of which are initiated on touch screens. For example, mouse movement and mouse button presses, optionally coordinated with single or multiple keyboard presses or holds; contact movements such as taps, drags, scrolls, etc. on touchpads; pen stylus inputs; movement of the device; oral instructions; detected eye movements; biometric inputs; and / or any combination thereof are optionally utilized as inputs corresponding to sub-events which define an event to be recognized.

[0151] FIG. 2 illustrates a portable multifunction device 100 having a touch screen 112 in accordance with some embodiments. The touch screen optionally displays one or more graphics within user interface (UI) 200. In this embodiment, as well as others described below, a user is enabled to select one or more of the graphics by making a gesture on the graphics, for example, with one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, right to left, upward and / or downward), and / or a rolling of a finger (from right to left, left to right, upward and / or downward) that has made contact with device 100. In some implementations or circumstances, inadvertent contact with a graphic does not select the graphic. For example, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.

[0152] Device 100 optionally also include one or more physical buttons, such as "home" or menu button 204. As described previously, menu button 204 is, optionally, used to navigate to any application 136 in a set of applications that are, optionally, executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touch screen 112.

[0153] In some embodiments, device 100 includes touch screen 112, menu button 204, push button 206 for powering the device on / off and locking the device, volume adjustment button(s) 208, subscriber identity module (SIM) card slot 210, headset jack 212, and docking / charging external port 124. Push button 206 is, optionally, used to turn the power on / off on the device by depressing the button and holding the button in the depressed state for a predefined time interval; to lock the device by depressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also accepts verbal input for activation or deactivation of some functions through microphone 113. Device 100 also, optionally, includes one or more contact intensity sensors 165 for detecting intensity of contacts on touch screen 112 and / or one or more tactile output generators 167 for generating tactile outputs for a user of device 100.

[0154] FIG. 3 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface in accordance with some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communications interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. Communication buses 320 optionally include circuitry (sometimes called a chipset) that interconnects and controls communications between system components. Device 300 includes input / output (I / O) interface 330 comprising display 340, which is typically a touch screen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and touchpad 355, tactile output generator 357 for generating tactile outputs on device 300 (e.g., similar to tactile output generator(s) 167 described above with reference to FIG. 1A), sensors 359 (e.g., optical, acceleration, proximity, touch-sensitive, and / or contact intensity sensors similar to contact intensity sensor(s) 165 described above with reference to FIG. 1A). Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices remotely located from CPU(s) 310. In some embodiments, memory 370 stores programs, modules, and data structures analogous to the programs, modules, and data structures stored in memory 102 of portable multifunction device 100 (FIG. 1A), or a subset thereof. Furthermore, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores drawing module 380, presentation module 382, word processing module 384, website creation module 386, disk authoring module 388, and / or spreadsheet module 390, while memory 102 of portable multifunction device 100 (FIG. 1A) optionally does not store these modules.

[0155] Each of the above-identified elements in FIG. 3 is, optionally, stored in one or more of the previously mentioned memory devices. Each of the above-identified modules corresponds to a set of instructions for performing a function described above. The above-identified modules or computer programs (e.g., sets of instructions or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), procedures, or modules, and thus various subsets of these modules are, optionally, combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subset of the modules and data structures identified above. Furthermore, memory 370 optionally stores additional modules and data structures not described above.

[0156] Attention is now directed towards embodiments of user interfaces that are, optionally, implemented on, for example, portable multifunction device 100.

[0157] FIG. 4A illustrates an exemplary user interface for a menu of applications on portable multifunction device 100 in accordance with some embodiments. Similar user interfaces are, optionally, implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof: Signal strength indicator(s) 402 for wireless communication(s), such as cellular and Wi-Fi signals; Time 404; Bluetooth indicator 405; Battery status indicator 406; Tray 408 with icons for frequently used applications, such as: ∘ Icon 416 for telephone module 138, labeled "Phone," which optionally includes an indicator 414 of the number of missed calls or voicemail messages; ∘ Icon 418 for e-mail client module 140, labeled "Mail," which optionally includes an indicator 410 of the number of unread e-mails; ∘ Icon 420 for browser module 147, labeled "Browser;" and ∘ Icon 422 for video and music player module 152, also referred to as iPod (trademark of Apple Inc.) module 152, labeled "iPod;" and Icons for other applications, such as: ∘ Icon 424 for IM module 141, labeled "Messages;" ∘ Icon 426 for calendar module 148, labeled "Calendar;" ∘ Icon 428 for image management module 144, labeled "Photos;" ∘ Icon 430 for camera module 143, labeled "Camera;" ∘ Icon 432 for online video module 155, labeled "Online Video;" ∘ Icon 434 for stocks widget 149-2, labeled "Stocks;" ∘ Icon 436 for map module 154, labeled "Maps;" ∘ Icon 438 for weather widget 149-1, labeled "Weather;" ∘ Icon 440 for alarm clock widget 149-4, labeled "Clock;" ∘ Icon 442 for workout support module 142, labeled "Workout Support;" ∘ Icon 444 for notes module 153, labeled "Notes;" and ∘ Icon 446 for a settings application or module, labeled "Settings," which provides access to settings for device 100 and its various applications 136.

[0158] It should be noted that the icon labels illustrated in FIG. 4A are merely exemplary. For example, icon 422 for video and music player module 152 is labeled "Music" or "Music Player." Other labels are, optionally, used for various application icons. In some embodiments, a label for a respective application icon includes a name of an application corresponding to the respective application icon. In some embodiments, a label for a particular application icon is distinct from a name of an application corresponding to the particular application icon.

[0159] FIG. 4B illustrates an exemplary user interface on a device (e.g., device 300, FIG. 3) with a touch-sensitive surface 451 (e.g., a tablet or touchpad 355, FIG. 3) that is separate from the display 450 (e.g., touch screen display 112). Device 300 also, optionally, includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting intensity of contacts on touch-sensitive surface 451 and / or one or more tactile output generators 357 for generating tactile outputs for a user of device 300.

[0160] Although some of the examples that follow will be given with reference to inputs on touch screen display 112 (where the touch-sensitive surface and the display are combined), in some embodiments, the device detects inputs on a touch-sensitive surface that is separate from the display, as shown in FIG. 4B. In some embodiments, the touch-sensitive surface (e.g., 451 in FIG. 4B) has a primary axis (e.g., 452 in FIG. 4B) that corresponds to a primary axis (e.g., 453 in FIG. 4B) on the display (e.g., 450). In accordance with these embodiments, the device detects contacts (e.g., 460 and 462 in FIG. 4B) with the touch-sensitive surface 451 at locations that correspond to respective locations on the display (e.g., in FIG. 4B, 460 corresponds to 468 and 462 corresponds to 470). In this way, user inputs (e.g., contacts 460 and 462, and movements thereof) detected by the device on the touch-sensitive surface (e.g., 451 in FIG. 4B) are used by the device to manipulate the user interface on the display (e.g., 450 in FIG. 4B) of the multifunction device when the touch-sensitive surface is separate from the display. It should be understood that similar methods are, optionally, used for other user interfaces described herein.

[0161] Additionally, while the following examples are given primarily with reference to finger inputs (e.g., finger contacts, finger tap gestures, finger swipe gestures), it should be understood that, in some embodiments, one or more of the finger inputs are replaced with input from another input device (e.g., a mouse-based input or stylus input). For example, a swipe gesture is, optionally, replaced with a mouse click (e.g., instead of a contact) followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is, optionally, replaced with a mouse click while the cursor is located over the location of the tap gesture (e.g., instead of detection of the contact followed by ceasing to detect the contact). Similarly, when multiple user inputs are simultaneously detected, it should be understood that multiple computer mice are, optionally, used simultaneously, or a mouse and finger contacts are, optionally, used simultaneously.

[0162] FIG. 5A illustrates exemplary personal electronic device 500. Device 500 includes body 502. In some embodiments, device 500 can include some or all of the features described with respect to devices 100 and 300 (e.g., FIGS. 1A-4B). In some embodiments, device 500 has touch-sensitive display screen 504, hereafter touch screen 504. Alternatively, or in addition to touch screen 504, device 500 has a display and a touch-sensitive surface. As with devices 100 and 300, in some embodiments, touch screen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors for detecting intensity of contacts (e.g., touches) being applied. The one or more intensity sensors of touch screen 504 (or the touch-sensitive surface) can provide output data that represents the intensity of touches. The user interface of device 500 can respond to touches based on their intensity, meaning that touches of different intensities can invoke different user interface operations on device 500.

[0163] Exemplary techniques for detecting and processing touch intensity are found, for example, in related applications: International Patent Application Serial No. PCT / US2013 / 040061, titled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," filed May 8, 2013, published as WIPO Publication No. WO / 2013 / 169849, and International Patent Application Serial No. PCT / US2013 / 069483, titled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," filed November 11, 2013, published as WIPO Publication No. WO / 2014 / 105276.

[0164] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508, if included, can be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms, if included, can permit attachment of device 500 with, for example, hats, eyewear, earrings, necklaces, shirts, jackets, bracelets, watch straps, chains, trousers, belts, shoes, purses, backpacks, and so forth. These attachment mechanisms permit device 500 to be worn by a user.

[0165] FIG. 5B depicts exemplary personal electronic device 500. In some embodiments, device 500 can include some or all of the components described with respect to FIGS. 1A, 1B, and 3. Device 500 has bus 512 that operatively couples I / O section 514 with one or more computer processors 516 and memory 518. I / O section 514 can be connected to display 504, which can have touch-sensitive component 522 and, optionally, intensity sensor 524 (e.g., contact intensity sensor). In addition, I / O section 514 can be connected with communication unit 530 for receiving application and operating system data, using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication techniques. Device 500 can include input mechanisms 506 and / or 508. Input mechanism 506 is, optionally, a rotatable input device or a depressible and rotatable input device, for example. Input mechanism 508 is, optionally, a button, in some examples.

[0166] Input mechanism 508 is, optionally, a microphone, in some examples. Personal electronic device 500 optionally includes various sensors, such as GPS sensor 532, accelerometer 534, directional sensor 540 (e.g., compass), gyroscope 536, motion sensor 538, and / or a combination thereof, all of which can be operatively connected to I / O section 514.

[0167] Memory 518 of personal electronic device 500 can include one or more non-transitory computer-readable storage mediums, for storing computer-executable instructions, which, when executed by one or more computer processors 516, for example, can cause the computer processors to perform the techniques described below, including processes 700, 800, 1000, 1200, 1400, 1500, 1700, and 1900 (FIGS. 7-8, 10, 12, 14, 15, 17, and 19). A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with the instruction execution system, apparatus, or device. In some examples, the storage medium is a transitory computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can include, but is not limited to, magnetic, optical, and / or semiconductor storages. Examples of such storage include magnetic disks, optical discs based on CD, DVD, or Blu-ray technologies, as well as persistent solid-state memory such as flash, solid-state drives, and the like. Personal electronic device 500 is not limited to the components and configuration of FIG. 5B, but can include other or additional components in multiple configurations.

[0168] As used here, the term "affordance" refers to a user-interactive graphical user interface object that is, optionally, displayed on the display screen of devices 100, 300, and / or 500 (FIGS. 1A, 3, and 5A-5C). For example, an image (e.g., icon), a button, and text (e.g., hyperlink) each optionally constitute an affordance.

[0169] As used herein, the term "focus selector" refers to an input element that indicates a current part of a user interface with which a user is interacting. In some implementations that include a cursor or other location marker, the cursor acts as a "focus selector" so that when an input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 in FIG. 3 or touch-sensitive surface 451 in FIG. 4B) while the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted in accordance with the detected input. In some implementations that include a touch screen display (e.g., touch-sensitive display system 112 in FIG. 1A or touch screen 112 in FIG. 4A) that enables direct interaction with user interface elements on the touch screen display, a detected contact on the touch screen acts as a "focus selector" so that when an input (e.g., a press input by the contact) is detected on the touch screen display at a location of a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted in accordance with the detected input. In some implementations, focus is moved from one region of a user interface to another region of the user interface without corresponding movement of a cursor or movement of a contact on a touch screen display (e.g., by using a tab key or arrow keys to move focus from one button to another button); in these implementations, the focus selector moves in accordance with movement of focus between different regions of the user interface. Without regard to the specific form taken by the focus selector, the focus selector is generally the user interface element (or contact on a touch screen display) that is controlled by the user so as to communicate the user's intended interaction with the user interface (e.g., by indicating, to the device, the element of the user interface with which the user is intending to interact). For example, the location of a focus selector (e.g., a cursor, a contact, or a selection box) over a respective button while a press input is detected on the touch-sensitive surface (e.g., a touchpad or touch screen) will indicate that the user is intending to activate the respective button (as opposed to other user interface elements shown on a display of the device).

[0170] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is, optionally, based on a predefined number of intensity samples, or a set of intensity samples collected during a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) relative to a predefined event (e.g., after detecting the contact, prior to detecting liftoff of the contact, before or after detecting a start of movement of the contact, prior to detecting an end of the contact, before or after detecting an increase in intensity of the contact, and / or before or after detecting a decrease in intensity of the contact). A characteristic intensity of a contact is, optionally, based on one or more of: a maximum value of the intensities of the contact, a mean value of the intensities of the contact, an average value of the intensities of the contact, a top 10 percentile value of the intensities of the contact, a value at the half maximum of the intensities of the contact, a value at the 90 percent maximum of the intensities of the contact, or the like. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an operation has been performed by a user. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact with a characteristic intensity that does not exceed the first threshold results in a first operation, a contact with a characteristic intensity that exceeds the first intensity threshold and does not exceed the second intensity threshold results in a second operation, and a contact with a characteristic intensity that exceeds the second threshold results in a third operation. In some embodiments, a comparison between the characteristic intensity and one or more thresholds is used to determine whether or not to perform one or more operations (e.g., whether to perform a respective operation or forgo performing the respective operation), rather than being used to determine whether to perform a first operation or a second operation.

[0171] FIG. 5C depicts an exemplary diagram of a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500, and each share with each other one or more data connections 510 such as an Internet connection, Wi-Fi connection, cellular connection, short-range communication connection, and / or any other such data connection or network so as to facilitate real time communication of audio and / or video data between the respective devices for a duration of time. In some embodiments, an exemplary communication session can include a shared-data session whereby data is communicated from one or more of the electronic devices to the other electronic devices to enable concurrent output of respective content at the electronic devices. In some embodiments, an exemplary communication session can include a video conference session whereby audio and / or video data is communicated between devices 500A, 500B, and 500C such that users of the respective devices can engage in real time communication using the electronic devices.

[0172] In FIG. 5C, device 500A represents an electronic device associated with User A. Device 500A is in communication (via data connections 510) with devices 500B and 500C, which are associated with User B and User C, respectively. Device 500A includes camera 501A, which is used to capture video data for the communication session, and display 504A (e.g., a touchscreen), which is used to display content associated with the communication session. Device 500A also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.

[0173] Device 500A displays, via display 504A, communication UI 520A, which is a user interface for facilitating a communication session (e.g., a video conference session) between device 500B and device 500C. Communication UI 520A includes video feed 525-1A and video feed 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and communicated from device 500B to devices 500A and 500C during the communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and communicated from device 500C to devices 500A and 500B during the communication session.

[0174] Communication UI 520A includes camera preview 550A, which is a representation of video data captured at device 500A via camera 501A. Camera preview 550A represents to User A the prospective video feed of User A that is displayed at respective devices 500B and 500C.

[0175] Communication UI 520A includes one or more controls 555A for controlling one or more aspects of the communication session. For example, controls 555A can include controls for muting audio for the communication session, changing a camera view for the communication session (e.g., changing which camera is used for capturing video for the communication session, adjusting a zoom value), terminating the communication session, applying visual effects to the camera view for the communication session, activating one or more modes associated with the communication session. In some embodiments, one or more controls 555A are optionally displayed in communication UI 520A. In some embodiments, one or more controls 555A are displayed separate from camera preview 550A. In some embodiments, one or more controls 555A are displayed overlaying at least a portion of camera preview 550A.

[0176] In FIG. 5C, device 500B represents an electronic device associated with User B, which is in communication (via data connections 510) with devices 500A and 500C. Device 500B includes camera 501B, which is used to capture video data for the communication session, and display 504B (e.g., a touchscreen), which is used to display content associated with the communication session. Device 500B also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.

[0177] Device 500B displays, via touchscreen 504B, communication UI 520B, which is similar to communication UI 520A of device 500A. Communication UI 520B includes video feed 525-1B and video feed 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and communicated from device 500A to devices 500B and 500C during the communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and communicated from device 500C to devices 500A and 500B during the communication session. Communication UI 520B also includes camera preview 550B, which is a representation of video data captured at device 500B via camera 501B, and one or more controls 555B for controlling one or more aspects of the communication session, similar to controls 555A. Camera preview 550B represents to User B the prospective video feed of User B that is displayed at respective devices 500A and 500C.

[0178] In FIG. 5C, device 500C represents an electronic device associated with User C, which is in communication (via data connections 510) with devices 500A and 500B. Device 500C includes camera 501C, which is used to capture video data for the communication session, and display 504C (e.g., a touchscreen), which is used to display content associated with the communication session. Device 500C also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.

[0179] Device 500C displays, via touchscreen 504C, communication UI 520C, which is similar to communication UI 520A of device 500A and communication UI 520B of device 500B. Communication UI 520C includes video feed 525-1C and video feed 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and communicated from device 500B to devices 500A and 500C during the communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and communicated from device 500A to devices 500B and 500C during the communication session. Communication UI 520C also includes camera preview 550C, which is a representation of video data captured at device 500C via camera 501C, and one or more controls 555C for controlling one or more aspects of the communication session, similar to controls 555A and 555B. Camera preview 550C represents to User C the prospective video feed of User C that is displayed at respective devices 500A and 500B.

[0180] While the diagram depicted in FIG. 5C represents a communication session between three electronic devices, the communication session can be established between two or more electronic devices, and the number of devices participating in the communication session can change as electronic devices join or leave the communication session. For example, if one of the electronic devices leaves the communication session, audio and video data from the device that stopped participating in the communication session is no longer represented on the participating devices. For example, if device 500B stops participating in the communication session, there is no data connection 510 between devices 500A and 500C, and no data connection 510 between devices 500C and 500B. Additionally, device 500A does not include video feed 525-1A and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and the video and audio data is shared among all devices such that each device is capable of outputting data communicated from the other devices.

[0181] The embodiment depicted in FIG. 5C represents a diagram of a communication session between multiple electronic devices, including the example communication sessions depicted in FIGS. 6A-6AY, 9A-9T, 11A-11P, 13A-13K, and 16A-16Q. In some embodiments, the communication session depicted in FIGS. 6A-6AY, 9A-9T, 13A-13K, and 16A-16Q includes two or more electronic devices, even if the other electronic devices participating in the communication session are not depicted in the figures.

[0182] Attention is now directed towards embodiments of user interfaces ("UI") and associated processes that are implemented on an electronic device, such as portable multifunction device 100, device 300, or device 500.

[0183] FIGS. 6A-6AY illustrate exemplary user interfaces for managing a live video communication session, in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes in FIGS. 7-8 and FIG. 15.

[0184] FIGS. 6A-6AY illustrate exemplary user interfaces for managing a live video communication session from the perspective of different users (e.g., users participating in the live video communication session from different devices, different types of devices, devices having different applications installed, and / or devices having different operating system software).

[0185] With reference to FIG. 6A, device 600-1 corresponds to user 622 (e.g., "John"), who is a participant of the live video communication session in some embodiments. Device 600-1 includes a display (e.g., touch-sensitive display) 601 and a camera 602 (e.g., front-facing camera) having a field of view 620. In some embodiments, camera 602 is configured to capture image data and / or depth data of a physical environment within field-of-view 620. Field-of-view 620 is sometimes referred to herein as the available field-of-view, entire field-of-view, or the camera field-of-view. In some embodiments, camera 602 is a wide angle camera (e.g., a camera that includes a wide angle lens or a lens that has a relatively short focal length and wide field-of-view). In some embodiments, device 600-1 includes multiple cameras. Accordingly, while description is made herein to device 600-1 using camera 602 to capture image data during a live video communication session, it will be appreciated that device 600-1 can use multiple cameras to capture image data.

[0186] With reference to FIG. 6A, device 600-2 corresponds to user 623 (e.g., "Jane"), who is a participant of the live video communication session in some embodiments. Device 600-2 includes a display (e.g., touch-sensitive display) 683 and a camera 682 (e.g., front-facing camera) having a field-of-view 688. In some embodiments, camera 682 is configured to capture image data and / or depth data of a physical environment within field-of-view 688. Field-of-view 688 is sometimes referred to herein as the available field-of-view, entire field-of-view, or the camera field-of-view. In some embodiments, camera 682 is a wide angle camera (e.g., a camera that includes a wide angle lens or a lens that has a relatively short focal length and wide field-of-view). In some embodiments, device 600-2 includes multiple cameras. Accordingly, while description is made herein to device 600-2 using camera 682 to capture image data during a live video communication session, it will be appreciated that device 600-2 can use multiple cameras to capture image data.

[0187] As shown, user 622 ("John") is positioned (e.g., seated) in front of desk 621 (and device 600-1) in environment 615. In some examples, user 622 is positioned in front of desk 621 such that user 622 is captured within field-of-view 620 of camera 602. In some embodiments, one or more objects proximate user 622 are positioned such that the objects are captured within field-of-view 620 of camera 602. In some embodiments, both user 622 and objects proximate user 622 are captured within field-of-view 620 simultaneously. For example, as shown, drawing 618 is positioned in front of user 622 (relative to camera 602) on surface 619 such that both user 622 and drawing 618 are captured in field-of-view 620 of camera 602 and displayed in representation 622-1 (displayed by device 600-1) and representation 622-2 (displayed by device 600-2).

[0188] Similarly, user 623 ("Jane") is positioned (e.g., seated) in front of desk 686 (and device 600-2) in environment 685. In some examples, user 623 is positioned in front of desk 686 such that user 623 is captured within field-of-view 688 of camera 682. As shown, user 623 is displayed in representation 623-1 (displayed by device 600-1) and representation 623-2 (displayed by device 600-2).

[0189] Generally, during operation, devices 600-1, 600-2 capture image data, which is in turn exchanged between devices 600-1, 600-2 and used by devices 600-1, 600-2 to display various representations of content during the live video communication session. While each of devices 600-1, 600-2 are illustrated, described examples are largely directed to the user interfaces displayed on and / or user inputs detected by device 600-1. It should be understood that, in some examples, electronic device 600-2 operates in an analogous manner as electronic device 600-1 during the live video communication session. In some examples devices 600-1, 600-2 display similar user interfaces and / or cause similar operations to be performed as those described below.

[0190] As will be described in further detail below, in some examples such representations include images that have been modified during the live video communication session to provide improved perspective of surfaces and / or objects within a field-of-view (also referred to herein as "field of view") of cameras of devices 600-1, 600-2. Images may be modified using any known image processing technique including but not limited to image rotation and / or distortion correction (e.g., image skew). Accordingly, although image data may be captured from a camera having a particular location relative to a user, representations may provide a perspective showing a user (and / or surfaces or objects in an environment of the user) from a perspective different than that of the camera capturing the image data. The embodiments of FIGS. 6A-6AY disclose displaying elements and detecting inputs (including hand gestures) at device 600-1 to control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2). In some embodiments, device 600-2 displays similar elements and detects similar inputs (including hand gestures) at device 600-2 to control how image data captured by camera 602 is displayed (at either device 600-1 and / or device 600-2).

[0191] With reference to FIG. 6A, device 600-1 displays, on display 601, video conference interface 604-1. Video conference interface 604-1 includes representation 622-1 which in turn includes an image (e.g., frame of a video stream) of a physical environment (e.g., a scene) within the field-of-view 620 of camera 602. In some examples, the image of representation 622-1 includes the entire field-of-view 620. In other examples, the image of representation 622-1 includes a portion (e.g., a cropped portion or subset) of the entire field-of-view 620. As shown, in some examples, the image of representation 622-1 includes user 622 and / or a surface 619 proximate user 622 on which drawing 618 is located.

[0192] Video conference interface 604-1 further includes representation 623-1 which in turn includes an image of a physical environment within the field-of-view 688 of camera 682. In some examples, the image of representation 623-1 includes the entire field-of-view 688. In other examples, the image of representation 623-1 includes a portion (e.g., a cropped portion or subset) of the entire field-of-view 688. As shown, in some examples, the image of representation 623-1 includes user 623. As shown, representation 623-1 is displayed at a larger magnitude than representation 622-1. In this manner, user 622 may better observe and / or interact with user 623 during the live communication session.

[0193] Device 600-2 displays, on display 683, video conference interface 604-2. Video conference interface 604-2 includes representation 622-2 which in turn includes an image of the physical environment within the field-of-view 620 of camera 602. Video conference interface 604-2 further includes representation 623-2 which in turn includes an image of a physical environment within the field-of-view 688 of camera 682. As shown, representation 622-2 is displayed at a larger magnitude than representation 623-2. In this manner, user 623 may better observe and / or interact with user 622 during the live communication session.

[0194] At FIG. 6A, device 600-1 displays interface 604-1. While displaying interface 604-1, device 600-1 detects input 612a (e.g., swipe input) corresponding to a request to display a settings interface. In response to detecting input 612a, device 600-1 displays settings interface 606, as depicted in FIG. 6B. As shown, settings interface 606 is overlaid on interface 604-1 in some embodiments.

[0195] In some embodiments, settings interface 606 includes one or more affordances for controlling settings of device 600-1 (e.g., volume, brightness of display, and / or Wi-Fi settings). For example, settings interface 606 includes a view affordance 607-1, which when selected causes device 600-1 to display a view menu, as shown in FIG. 6B.

[0196] As shown in FIG. 6B, while displaying settings interface 606, device 600-1 detects input 612b. Input 612b is a tap gesture on view affordance 607-1 in some embodiments. In response to detecting input 612b, device 600-1 displays view menu 616-1, as shown in FIG. 6C.

[0197] Generally, view menu 616-1 includes one or more affordances which may be used to manage (e.g., control) the manner in which representations are displayed during a live video communication session. By way of example, selection of a particular affordance may cause device 600-1 to display, or cease displaying, representations in an interface (e.g., interface 604-1 or interface 604-2).

[0198] View menu 616-1, for instance, includes a surface view affordance 610, which when selected, causes device 600-1 to display a representation including a modified image of a surface. In some embodiments, when surface view affordance 610 is selected, the user interfaces transition directly to the user interfaces of FIG. 6M. Additionally or alternatively, FIGS. 6D-6L (described below) illustrate other user interfaces that can be displayed prior to the user interfaces in FIG. 6M and other inputs to initiate the process of displaying the user interfaces as shown in FIG. 6M. For example, while displaying view menu 616-1, device 600-1 detects input 612c corresponding to a selection of surface view affordance 610. In some examples, input 612c is a touch input. In response to detecting input 612c, device 600-1 displays representation 624-1, as shown in FIG. 6M. Further in response to detecting input 612c, device 600-2 displays representation 624-2. As described, in some embodiments, an image is modified during the live video communication session to provide an image having a particular perspective. Accordingly, in some examples, representation 624-1 is provided by generating an image from image data captured by camera 602, modifying the image (or a portion of the image), and displaying representation 624-1 with the modified image. In some embodiments, the image is modified using any known image processing techniques, including but not limited to image rotation and / or distortion correction (e.g., image skewing). The image of representation 624-2 is also provided in this manner in some embodiments.

[0199] In some embodiments, the image of representation 624-1 is modified to provide a desired perspective (e.g., a surface view). In some embodiments, the image of representation 624-1 is modified based on a position of surface 619 relative to camera 602. By way of example, device 600-1 can rotate the image of representation 624-1 a predetermined amount (e.g., 45 degrees, 90 degrees, or 180 degrees) such that surface 619 can be more intuitively viewed in representation 624-1. As shown in FIG. 6M, for example, in which camera 602 captures surface 619 from a perspective facing the user 622, the image of representation 624-1 is rotated 180 degrees to provide a perspective of the image from that of user 622. Accordingly, during the live video communication session, devices 600-1, 600-2 display surface 619 (and by extension drawing 618) from the perspective of user 622 during the live communication session. The image of representation 624-2 is also provided in this manner in some examples.

[0200] In some embodiments, to ensure that user 623 maintains a view of user 622 while representation 624-2 includes a modified image of surface 619, device 600-2 maintains display of representation 622-2. As shown in FIG. 6M, maintaining display of representation 622-2 in this manner can include adjusting a size and / or position of representation 622-2 in interface 604-2. Optionally, in some embodiments, device 600-2 ceases display of representation 622-2 to provide a larger size of representation 624-2. Optionally, in some embodiments, device 600-1 ceases display of representation 622-1 to provide a larger size of representation 624-1.

[0201] Representations 624-1, 624-2 include an image of drawing 618 that is modified with respect to the position (e.g., location and / or orientation) of drawing 618 relative to camera 602. For example, as depicted in FIG. 6A, prior to modification, the image is shown as having a particular orientation (e.g., upside down) in representations 622-1, 622-1. As a result of modifying the image, the image of drawing 618 is rotated and / or skewed such that the perspective of representations 624-1, 624-2 appears to be from the perspective of user 622. In this manner, the modified image of drawing 618 provides a perspective that is different from the perspective of representations 624-1, 624-2, so as to give user 623 (and / or user 622) a more natural and direct view of drawing 618. Accordingly, drawing 618 may be more readily and intuitively viewed by user 623 during the live video communication session.

[0202] As described, a representation including a modified image of a surface is provided in response to selection of a surface image affordance (e.g., surface view affordance 610). In some examples, a representation including a modified view of a surface is provided in response to detecting other types of inputs.

[0203] With reference to FIG. 6D, in some examples, a representation including a modified image of a surface is provided in response to one or more gestures. As an example, device 600-1 can detect a gesture using camera 602, and in response to detecting the gesture, determine whether the gesture satisfies a set of criteria (e.g., a set of gesture criteria). In some embodiments, the criteria include a requirement that the gesture is a pointing gesture, and optionally, a requirement that the pointing gesture has a particular orientation and / or is directed at a surface and / or object. For example with reference to FIG. 6D, device 600-1 detects gesture 612d and determines that the gesture 612d is a pointing gesture directed at drawing 618. In response, device 600-1 displays a representation including a modified image of surface 619, as described with reference to FIG. 6M.

[0204] In some embodiments, the set of criteria includes a requirement that a gesture be performed for at least a threshold amount of time. For example, with reference to FIG. 6E, in response to detecting a gesture, device 600-1 overlays graphical object 626 on representation 622-1 indicating that device 600-1 has detected that the user is currently performing a gesture, such as 612d. As shown, in some embodiments, device 600-1 enlarges representation 622-1 to assist user 622 in better viewing the detected gesture and / or graphical object 626.

[0205] In some embodiments, graphical object 626 includes timer 628 indicating an amount of time gesture 612d has been detected (e.g., a numeric timer, a ring that is filled over time, and / or a bar that is filled over time). In some embodiments, timer 628 also (or alternatively) indicates a threshold amount of time gesture 612d is to continue to be provided to satisfy the set of criteria. In response to gesture 612 satisfying the threshold amount of time (e.g., 0.5 second, 2 seconds, and / or 5 seconds), device 600-1 displays representation 624-1 including a modified image of a surface (FIG. 6M), as described.

[0206] In some examples, graphical object 626 indicates the type of gesture currently detected by device 600-1. In some examples, graphical object 626 is an outline of a hand performing the detected type of gesture and / or an image of the detected type of gesture. Graphical object 626 can, for instance, include a hand performing a pointing gesture in response to device 600-1 detecting that user 622 is performing a pointing gesture. Additionally or alternatively, the graphical object 626 can, optionally, indicate a zoom level (e.g., zoom level at which the representation of the second portion of the scene is or will be displayed).

[0207] In some examples, a representation having an image that is modified is provided in response to one or more speech inputs. For example, during the live communication session, device 600-1 receives a speech input, such as speech input 614 ("Look at my drawing.") in FIG. 6D. In response, device 600-1 displays representation 624-1 including a modified image of a surface (FIG. 6M), as described.

[0208] In some examples, speech inputs received by device 600-1 can include references to any surface and / or object recognizable by device 600-1, and in response, device 600-1 provides a representation including a modified image of the referenced object or surface. For example, device 600-1 can receive a speech input that references a wall (e.g., a wall behind user 622). In response, device 600-1 provides a representation including a modified image of the wall.

[0209] In some embodiments, speech inputs can be used in combination with other types of inputs, such as gestures (e.g., gesture 612d). Accordingly, in some embodiments, device 600-1 displays a modified image of a surface (or object) in response to detecting both a gesture and a speech input corresponding to a request to provide a modified image of the surface.

[0210] In some embodiments, a surface view affordance is provided in other manners. With reference to FIG. 6F, for instance, video conference interface 604-1 includes options menu 608. Options menu 608 includes a set of affordances that can be used to control device 600-1 during a live video communication session, including view affordance 607-2.

[0211] While displaying options menu 608, device 600-1 detects an input 612f corresponding to a selection of view affordance 607-2. In response to detecting input 612f, device 600-1 displays view menu 616-2, as shown in FIG. 6G. View menu 616-2 can be used to control the manner in which representations are displayed during a live video communication session, as described with respect to FIG. 6C.

[0212] While options menu 608 is illustrated as being persistently displayed in video conference interface 604-1 throughout the figures, options menu 608 can be hidden and / or redisplayed at any point during the live video communications session by device 600-1. For example, options menu 608 can be displayed and / or removed from display in response to detecting one or more inputs and / or a period of inactivity by a user.

[0213] While detecting an input directed to a surface has been described as causing device 600-1 to display a representation including a modified image of a surface (for example, in response to detecting input 612c of FIG. 6C, device 600-1 displays representation 624-1, as shown in FIG. 6M), in some embodiments, detecting an input directed to a surface can cause device 600-1 to enter a preview mode (e.g., FIGS. 6H-6J), for instance, prior to displaying representation 624-1.

[0214] FIG. 6H illustrates an example in which device 600-1 is operating in a preview mode. Generally, the preview mode can be used to selectively provide portions, or regions, of an image of a representation to one or more other users during a live video communications session.

[0215] In some embodiments, prior to operating in the preview mode, device 600-1 detects an input (e.g., input 612c) directed to a surface view affordance 610. In response, device 600-1 initiates a preview mode. While operating in a preview mode, device 600-1 displays a preview interface 674-1. Preview interface 647-1 includes a left scroll affordance 634-2, a right scroll affordance 634-1, and preview 636.

[0216] In some embodiments, selection of the left scroll affordance causes device 600-1 to change (e.g., replace) preview 636. For example, selection of the left scroll affordance 634-2 or the right scroll affordance 634-1 causes device 600-1 to cycle through various images (image of a user, unmodified image of a surface, and / or modified image of surface 619) such that a user can select a particular perspective to be shared upon exiting the preview mode, for instance, in response to detecting an input directed to preview 636. Additionally or alternatively, these techniques can be used to cycle through and / or select a particular surface (e.g., vertical and / or horizontal surface) and / or particular portion (e.g., cropped portion or subset) in the field-of-view.

[0217] As shown, in some embodiments, preview user interface 674-1 is displayed at device 600-1 and is not displayed at device 600-2. For example, device 600-2 displays video conference interface 604-2 (including representation 622-2) while device 600-1 displays preview interface 674-1. As such, preview user interface 674-1 allows user 622 to select a view prior to sharing the view with user 623.

[0218] FIG. 6I illustrates an example in which device 600-1 is operating in a preview mode. As depicted, while the device 600-1 is operating in the preview mode, device 600-1 displays preview interface 674-2. In some embodiments, preview interface 674-2 includes representation 676 having regions 636-1, 636-2. In some embodiments, representation 676 includes an image that is the same or substantially similar to an image included in representation 622-1. Optionally, as shown, the size of representation 676 is larger than representation 622-1 of FIG. 6A. The position of representation 676 is different than the position of representation 622-1. Adjusting the size and / or position of a representation in preview interface 674-2 as compared the size and / or position of a representation including a similar or same image in video conference interface 604-1 allows user 622 to better view an image prior sharing that image with user 623.

[0219] In some embodiments, region 636-1 and region 636-2 correspond to respective portions of representation 676. For example, as shown, region 636-1 corresponds to an upper portion of representation 676 (e.g., a portion including an upper body of user 622), and region 636-2 corresponds to a lower portion of representation 676 (e.g., a portion including a lower body of user 622 and / or drawing 618).

[0220] In some embodiments, region 636-1 and region 636-2 are displayed as distinct regions (e.g., non-overlapping regions). In some embodiments, region 636-1 and region 636-2 overlap. Additionally or alternatively, one or more graphical objects 638-1 (e.g., lines, boxes, and / or dashes) can distinguish (e.g., visually distinguish) region 636-1 from region 636-2.

[0221] In some embodiments, preview interface 674-2 includes one or more graphical objects to indicate whether a region is active or inactive. In the example of FIG. 6I, preview interface 674-2 includes graphical objects 641a, 641b. The appearance (e.g., shape, size, and / or color) of graphical objects 641a, 641b indicates whether a respective region is active and / or inactive in some embodiments.

[0222] When active, a region is shared with one or more other users of a live video communication session. For example, with reference to FIG. 6I, graphical user interface object 641 indicates that region 636-1 is active. As a result, image data corresponding to region 636-1 is displayed by device 600-2 in representation 622-2. In some examples, device 600-1 shares only image data for active regions. In some embodiments, device 600-1 shares all image data, and instructs device 600-2 to display an image based on only the portion of image data corresponding to the active region 636-1.

[0223] While displaying interface 674-2, device 600-1 detects an input 612i at a location corresponding to region 636-2. Input 612i is a touch input in some embodiments. In response to detecting input 612i, device 600-1 activates region 636-2. As a result, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, region 636-1 remains active in response to input 612i (e.g., user 623 can see user 622, for example, in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates region 636-1 in response to input 612i (e.g., user 623 can no longer see user 622, for example, in representation 622-2).

[0224] While the example of FIG. 6I is described with respect to a preview mode having a representation including two regions 636-1, 636-2, in some embodiments other numbers of regions can be used. For example, with reference to FIG. 6J, device 600-1 is operating in a preview mode in which preview interface 674-3 includes a representation 676 that includes regions 636a-636i.

[0225] In some embodiments, a plurality of regions are active (and / or can be activated). For example, as shown, device 600-1 displays regions 636a-636i, of which regions 636a-f are active. As a result, device 600-2 displays representation 622-2.

[0226] In some embodiments, device 600-1 modifies an image of a surface having any type of orientation, including any angle (e.g., between zero to ninety degrees) with respect to gravity. For example, device 600-1 can modify an image of a surface when the surface is a horizontal surface (e.g., a surface that is in a plane that is within the range of 70 to 110 degrees of the direction of gravity). As another example, device 600-1 can modify an image of a surface when the surface is a vertical surface (e.g., a surface that is in a plane that up to 30 degrees of the direction of gravity).

[0227] While displaying interface 674-3, device 600-1 detects input 612j at a location corresponding to region 636h. In response to detecting input 612j, device 600-1 activates region 636-2. As a result, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, regions 636a-f remain active in response to input 612j (e.g., user 623 can see user 622, for example, in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates regions 636a-f in response to input 612j (e.g., user 623 can no longer see user 622, for example, in representation 622-2).

[0228] FIGS. 6K-6L illustrate example animations that can be displayed by device 600-1 and / or device 600-2. As discussed in FIGS. 6A-6I, device 600-1 can display representations including modified images. In some embodiments, device 600-1 and / or device 600-2 displays an animation to transition between views and / or show modifications to images over time. The animation can include, for instance, panning, rotating, and / or otherwise modifying an image to provide the modified image. Additionally or alternatively, the animation occurs in response to detecting an input directed at a surface (e.g., a selection of surface view affordance 610, a gesture, and / or a speech input).

[0229] FIG. 6K illustrates an example animation in which device 600-2 pans and rotates an image of representation 642a. During the animation, the image of representation 642a is panned down to view surface 619 at a more "overhead" perspective. The animation also includes rotating the image of representation 642a such that surface 619 is viewed from the perspective of user 622. While four frames of the animation are shown, the animation can include any number of frames. Optionally, in some embodiments, device 600-1 pans and rotates an image of a representation (e.g., representation 622-1).

[0230] FIG. 6L illustrates an example in which device 600-2 magnifies and rotates an image of representation 642a. During the animation, representation 642a is magnified until a desired zoom level is attained. The animation also includes rotating the representation 642a until an image of drawing 618 is oriented to a perspective of user 622, as described. While four frames of the animation are shown, the animation can include any number of frames. Optionally, in some embodiments, device 600-1 magnifies and rotates an image of a representation (e.g., representation 622-1).

[0231] FIGS. 6N-6R illustrate examples in which a modified image of a surface is further modified during a live communication session.

[0232] FIG. 6N illustrates an example of a live communication session in which a user provides various inputs. For example, while displaying interface 678, device 600-1 detects an input 677 corresponding to a rotation of device 600-1. As depicted in FIG. 6O, in response to detecting input 677, device 600-1 modifies interface 678 to compensate for the rotation (e.g., of camera 602). As shown in FIG. 6O, device 600-1 arranges representations 623-1 and 624-1 of interface 678 in a vertical configuration. Additionally, representation 624-1 is rotated according to the rotation of device 600-1 such that the perspective of representation 624-1 is maintained in the same orientation relative to the user 622. Additionally, the perspective of representation 624-2 is maintained in the same orientation relative to the user 623.

[0233] With further reference to FIG. 6N, in some examples, device 600-1 displays control affordances 648-1, 648-2 to modify the image of representation 624-1. Control affordances 648-1, 648-2 can be displayed in response to one or more inputs, for instance, corresponding to a selection of an affordance of options menu 608 (e.g., FIG. 6B).

[0234] As shown, in some embodiments, device 600-1 displays representation 624-1 including a modified image of a surface. Rotation affordance 648-1, when selected, causes device 600-1 to rotate the image of representation 624-1. For example, while displaying interface 678, device 600-1 detects input 650a corresponding to a selection of rotation affordance 648-1. In response to input 650a, device 600-1 modifies the orientation of the image of representation 624-1 from a first orientation (shown in FIG. 6N) to a second orientation (shown in FIG. 6O). In some embodiments, the image of representation 624-1 is rotated by a predetermined amount (e.g., 90 degrees).

[0235] Zoom affordance 648-2, when selected, modifies the zoom level of the image of representation 624-1. For example, as depicted in FIG. 6N, the image of representation 624-1 is displayed at a first zoom level (e.g., "1X"). While displaying zoom affordance 648-2, device 600-1 detects input 650b corresponding to a selection of zoom affordance 648-2. In response to input 650b, device 600-1 modifies a zoom level of the image of representation 624-1 from the first zoom level (e.g., "1X") to a second zoom level (e.g., "2X"), as shown in FIG. 6Q.

[0236] Additionally or alternatively, in some embodiments, video conference interface 604-1 includes an option to display a magnified view of at least a portion of the image of representation 624-1, as shown in FIG. 6R. For instance, while displaying representation 624-1, device 600-1 can detect an input 654 (e.g., a gesture directed to a surface and / or object) corresponding to a request to display a magnified view of a portion of the image of representation 624-1. In response to detecting input 654, device 600-1 displays magnified portion 652-1 at a greater zoom level than second portion 652-2 of representation 624-1. In some embodiments, the portion of the image of representation 624-1 that is magnified is determined based on a location of input 654. In some embodiments, in response to detecting input 650c (FIGS. 6R and 6Q), device 600-1 ceases to display control affordances 648-1, 648-2.

[0237] FIGS. 6S-6AC illustrate examples in which a device modifies an image of a representation in response to user input. As described in more detail below, device 600-1 can modify images of representations (e.g., representation 622-1) in video conference interface 604-1 in response to non-touch user input, including gestures and / or audio input, thereby improving the manner in which a user interacts with a device to manage and / or modify representations during a live video communication session.

[0238] FIGS. 6S-6T illustrate an example in which a device obscures at least a portion of an image of a representation in response to a gesture. As illustrated in FIG. 6S, device 600-1 detects gesture 656a corresponding to a request to modify at least a portion of an image of representation 622-1. In some examples, gesture 656a is a gesture in which user 622 points in an upward direction near the mouth of user 622 (e.g., a "shh" gesture). As shown in FIG. 6T, in response, device 600-1 replaces representation 622-1 with representation 622-1' that includes a modified image including a modified portion 658-1 (e.g., background of physical environment of user 622). In some examples, modifying portion 658-1 in this manner includes blurring, greying, or otherwise obscuring portion 658-1. In some examples, device 600-1 does not modify portion 658-2 in response to gesture 656a.

[0239] FIGS. 6U-6V illustrate an example in which a device magnifies a portion of the image of a representation in response to detecting a gesture. As shown in FIG. 6U, in some embodiments, device 600-1 detects pointing gesture 656b corresponding to a request to magnify at least a portion of representation 622-1. As shown, pointing gesture 656b is directed at object 660.

[0240] As depicted in FIG. 6V, in response to pointing gesture 656b, device 600-1 replaces representation 622-1 with representation 622-1' that includes a modified image by magnifying a portion of the image of representation 622-1 including object 660. In some embodiments, the magnification is based on the location of object 660 (e.g., relative to camera 602) and / or size of object 660.

[0241] FIGS. 6W-6X illustrate an example in which a device magnifies a portion of a view of a representation in response to detecting a gesture. As shown in FIG. 6W, in some embodiments, device 600-1 detects framing gesture 656c corresponding to a request to magnify at least a portion of representation 622-1. As shown, framing gesture 656c is directed at object 660 due to framing gesture 656c at least partially framing, surrounding, and / or outlining object 660.

[0242] As depicted in FIG. 6X, in response to framing gesture 656c, device 600-1 modifies the image of representation 622-1 by magnifying a portion of the image of representation 622-1 including object 660. In some embodiments, the magnification is based on the location of object 660 (e.g., relative to camera 602) and / or size of object 660. Additionally or alternatively, after magnifying a portion of the image of representation 622-1, device 600-1 can track a movement of framing gesture 656c. In response, device 600-1 can pan to a different portion of the image.

[0243] FIGS. 6Y-6Z illustrate an example in which a device pans an image of a representation in response to detecting a gesture. As shown in FIG. 6Y, device 600-1 detects pointing gesture 656d corresponding to a request to pan (e.g., horizontally pan) a view of the image of representation 622-1 in a particular direction. As shown, pointing gesture 656d is directed to the left of user 622.

[0244] As shown in FIG. 6Z, in response to pointing gesture 656d, device 600-1 replaces representation 622-1 with representation 622-1' that includes a modified image that is based on panning the image of representation 622-1 in a direction of pointing gesture 656d (e.g., to the left of user 622).

[0245] While in some embodiments, as shown in FIG. 6Z, a portion of user 622 (e.g., the right shoulder of user 622) can be excluded from the image of representation 622-1' due to a panning operation, in some embodiments, device 600-1 can adjust a zoom level of the image of representation 622-1' when panning so as to ensure user 622 remains fully in the image.

[0246] FIGS. 6AA-6AB illustrate an example in which a device modifies a zoom level of a representation in response to detecting a pinch and / or spread gesture. As shown in FIG. 6AA, in some embodiments, device 600-1 detects spread gesture 656e in which user 622 increases the distance between the thumb and index finger of the right hand of user 622.

[0247] As depicted in FIG. 6AB, in response to spread gesture 656e, device 600-1 replaces representations 622-1 with 622-1' by magnifying a portion of the image of representation 622-1. In some embodiments, the magnification is based on a location of spread gesture 656e (e.g., relative to camera 602) and / or a magnitude of spread gesture 656e. In some embodiments, the portion of the image is magnified according to a predetermined zoom level.

[0248] With reference to FIG. 6AA, in some embodiments, in response to detecting spread gesture 656e, device 600-1 displays zoom indicator 662 indicating a zoom level of the image of representation 622-1'. Once user 622 has completed the spread gesture 656e and device 600-1 has magnified the portion of representation 622-1', device 600-1 updates display of zoom indicator 662 to indicate the current zoom level of the image of representation 622-1'. In some embodiments, zoom indicator 662 is updated dynamically as user 622 performs gesture 656e.

[0249] While description is made herein with respect to increasing a zoom level of an image in response to a spread gesture 656e, in some examples, a zoom level of an image is decreased in response to a gesture (e.g., another type of gesture, such as a pinch gesture).

[0250] FIG. 6AC illustrates various gestures that can be used to modify an image of a representation. In some embodiments, for instance, a user can use gestures to indicate a zoom level. By way of example, gesture 664 can be used to indicate that a zoom level of an image of a representation should be at "1X", and in response to detecting gesture 664, device 600-1 can modify an image of a representation to have a "1X" zoom level. Similarly, gesture 666 can be used to indicate that a zoom level of an image of a representation should be at "2X" and in response to detecting gesture 666, device 600-1 can modify an image of a representation to have a "2X" zoom level. While two zoom levels (e.g., a "1X" and a "2X" zoom level) are described for FIG. 6AC, in some embodiments, device 600-1 can modify an image of a representation to other zoom levels (e.g., 0.5X, 3X, 5X, or 10X) using the same gesture or a different gesture. In some embodiments, device 600-1 can modify an image of a representation to three or more different zoom levels. In some embodiments, the zoom levels are discrete or continuous.

[0251] As another example, a gesture in which user 622 curls their fingers can be used to adjust a zoom level. For instance, gesture 668 (e.g., a gesture in which fingers of a user's hand are curled in a direction 668b away from a camera, for example, when the back of the hand 668a is oriented toward the camera) can be used to indicate that a zoom level of an image should be increased (e.g., zoomed in). Gesture 670 (e.g., a gesture in which fingers of a user's hand are curled in a direction 670b toward a camera, for example, when the palm of the hand 668a is oriented toward the camera) can be used to indicate that a zoom level of an image should be decreased (e.g., zoomed out).

[0252] FIGS. 6AD-6AE illustrate examples in which a user participates in a live video communication session using two devices.

[0253] As an example, as shown in FIG. 6AD, user 623 is using an additional device 600-3 during the live video communication session. In some embodiments, devices 600-2, 600-3 concurrently display representations including images that have different views. For example, while device 600-3 displays representation 622-2, device 600-2 displays representation 624-2.

[0254] In some embodiments, device 600-2 is positioned in front of user 623 on desk 686 in a manner that corresponds to the position of surface 619 relative to user 622. Accordingly, user 623 can view representation 624-2 (including an image of surface 619) in a manner analogous to that of user 622 viewing surface 619 in the physical environment.

[0255] As shown in FIG. 6AE, during the live communication session, user 623 can modify the image displayed in representation 624-2 by moving device 600-2. In response to user 623 changing an orientation of device 600-2, device 600-2 modifies an image of representation 624-2, for instance, in a manner corresponding to the change in orientation of device 600-2. For example, in response to user 623 tilting device 600-2, device 600-2 pans upward to display other portions of surface 619. In this manner, user 623 can change an orientation of device 600-2 (in any direction) to view various portions of surface 619 that are not otherwise displayed when device 600-2 is in a different orientation.

[0256] FIGS. 6AF-6AL illustrate embodiments for accessing the various user interfaces illustrated and described with reference to FIGS. 6A-6AE. In the embodiments depicted in FIGS. 6AF-6AL, the interfaces are illustrated using a laptop (e.g., John's device 6100-1 and / or Jane's device 6100-2). It should be appreciated that the embodiments illustrated in FIGS. 6AF-6AL can be implemented using a different device, such as a tablet (e.g., John's tablet 600-1 and / or Jane's device 600-2). Similarly, the embodiments illustrated in FIGS. 6A-6AE can be implemented using a different device such as John's device 6100-1 and / or Jane's device 6100-2. Therefore, various operations or features described above with respect to FIGS. 6A-6AE are not repeated below for the sake of brevity. For example, the applications, interfaces (e.g., 604-1 and / or 604-2), and displayed elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, and / or 624-2) discussed with respect to FIGS. 6A-6AE are similar to the applications, interfaces (e.g., 6121 and / or 6131), and displayed elements (e.g., 6124, 6132, 6122, 6134, 6116, 6140, and / or 6142) discussed with respect to FIGS. 6AF-6AL. Accordingly, details of these applications, interfaces, and displayed elements may not be repeated below for the sake of brevity.

[0257] FIG. 6AF depicts John's device 6100-1, which includes display 6101, one or more cameras 6102, and keyboard 6103 (which, in some embodiments, includes a trackpad). John's device 6100-1 displays, via display 6101, a home screen that includes camera application icon 6108 and video conferencing application icon 6110. Camera application icon 6108 corresponds to a camera application operating on John's device 6100-1 that can be used to access camera 6102. Video conferencing application icon 6110 corresponds to a video conferencing application operating on John's device 6100-1 that can be used to initiate and / or participate in a live video communication session (e.g., a video call and / or a video chat) similar to that discussed above with reference to FIGS. 6A-6AE. John's device 6100-1 also displays dock 6104, which includes various application icons, including a subset of icons that are displayed in dynamic region 6106. The icons displayed in dynamic region 6106 represent applications that are active (e.g., launched, open, and / or in use) on John's device 6100-1. In FIG. 6AF, neither the camera application nor the video conferencing application are currently active. Therefore, icons representing the camera application or video conferencing application are not displayed in dynamic region 6106, and John's device 6100-1 is not participating in a live video communication session.

[0258] In FIG. 6AF, John's device 6100-1 detects input as indicated by cursor 6112 (e.g., a cursor input caused by clicking a mouse, tapping on a trackpad, and / or other such input) selecting camera application icon 6108. In response, John's device 6100-1 launches the camera application and displays camera application window 6114, as shown in FIG. 6AG. In the embodiment depicted in FIG. 6AG, the camera application is being used to access camera 6102 to generate surface view 6116, which is similar to representation 624-1 depicted in FIG. 6M, for example, and described above. In some embodiments, the camera application can have different modes (e.g., user selectable modes) such as, for example, an expanded field-of-view mode (which provides an expanded field-of-view of camera 6102) and the surface view mode (which provides the surface view illustrated in FIG. 6AG). Accordingly, surface view 6116 represents a view of image data obtained using camera 6102 and modified (e.g., magnified, rotated, cropped, and / or skewed) by the camera application to produce surface view 6116 shown in FIG. 6AG. Additionally, because John's laptop launched the camera application, camera application icon 6108-1 is displayed in dynamic region 6106 of dock 6104, indicating that the camera application is active. In some embodiments, application icons (e.g., 6108-1) are displayed having an animated effect (e.g., bouncing) when they are added to the dynamic region of the dock.

[0259] In FIG. 6AG, John's device 6100-1 detects input 6118 selecting video conferencing application icon 6110. In response, John's device 6100-1 launches the video conferencing application, displays video conferencing application icon 6110-1 in dynamic region 6106, and displays video conferencing application window 6120, as shown in FIG. 6AH. Video conferencing application window 6120 includes video conferencing interface 6121, which is similar to interface 604-1, and includes video feed 6122 of Jane (similar to representation 623-1) and video feed 6124 of John (similar to representation 622-1). In some embodiments, John's device 6100-1 displays video conferencing application window 6120 with video conferencing interface 6121 after detecting one or more additional inputs after input 6118. For example, such inputs can be inputs to initiate a video call with Jane's laptop or to accept a request to participate in a video call with Jane's laptop.

[0260] In FIG. 6AH, John's device 6100-1 displays video conferencing application window 6120 partially overlaid on camera application window 6114. In some embodiments, John's device 6100-1 can bring camera application window 6114 to the front or foreground (e.g., partially overlaid on video conferencing application window 6120) in response to detecting a selection of camera application icon 6108, a selection of icon 6108-1, and / or an input on camera application window 6114. Similarly, video conferencing application window 6120 can be brought to the front or foreground (e.g., partially overlaying camera application window 6114) in response to detecting a selection of video conferencing application icon 6110, a selection of icon 6110-1, and / or an input on video conferencing application window 6120.

[0261] In FIG. 6AH, John's device 6100-1 is shown participating in a live video communication session with Jane's device 6100-2. Accordingly, Jane's device 6100-2 is depicted displaying video conferencing application window 6130, which is similar to video conferencing application window 6120 on John's device 6100-1. Video conferencing application window 6130 includes video conferencing interface 6131, which is similar to interface 604-2, and includes video feed 6132 of John (similar to representation 622-2) and video feed 6134 of Jane (similar to representation 623-2).

[0262] In the embodiment depicted in FIG. 6AH, the video conferencing application is being used to access camera 6102 to generate video feed 6124 and video feed 6132. Accordingly, video feeds 6124 and 6132 represent a view of image data obtained using camera 6102 and modified (e.g., magnified and / or cropped) by the video conferencing application to produce the image (e.g., video) shown in video feed 6124 and video feed 6132. In some embodiments, the camera application and the video conferencing application can use different cameras to provide respective video feeds.

[0263] Video conferencing application window 6120 includes menu option 6126, which can be selected to display different options for sharing content in the live video communication session. In FIG. 6AH, John's device 6100-1 detects input 6128 selecting menu option 6126 and, in response, displays share menu 6136, as shown in FIG. 6AI. Share menu 6136 includes share options 6136-1, 6136-2, and 6136-3. Share option 6136-1 is an option that can be selected to share content from the camera application. Share option 6136-2 is an option that can be selected to share content from the desktop of John's device 6100-1. Share option 6136-3 is an option that can be selected to share content from a presentation application. In response to detecting input 6138 on share option 6136-1, John's device 6100-1 begins sharing content from the camera application, as shown in FIG. 6AJ and FIG. 6AK.

[0264] In FIG. 6AJ, John's device 6100-1 updates video conferencing interface 6121 to include surface view 6140, which is shared with Jane's device 6100-2 in the live video communication session. In the embodiment depicted in FIG. 6AJ, John's device 6100-1 shares the video feed generated using the camera application (shown as surface view 6116 in camera application window 6114), and displays the representation of the video feed as surface view 6140 in the video conferencing application window 6120. Additionally, John's laptop emphasizes the display of surface view 6140 in video conferencing interface 6121 (e.g., by displaying the surface view with a larger size than other video feeds) and reduces the displayed size of Jane's video feed 6122. In FIG. 6AJ, John's device 6100-1 displays surface view 6140 concurrently with John's video feed 6124 and Jane's video feed 6122 in video conferencing application window 6120. In some embodiments, the display of John's video feed 6124 and / or Jane's video feed 6122 in video conferencing application window 6120 is optional.

[0265] Jane's device 6100-2 updates video conferencing interface 6131 to show surface video feed 6142, which is the surface view (from the camera application) being shared by John's device 6100-1. As shown in FIG. 6AJ, Jane's device 6100-2 adds surface video feed 6142 to video conferencing interface 6131 to show the surface video feed concurrently with Jane's video feed 6134 and John's video feed 6132, which has optionally been resized to accommodate the addition of surface video feed 6142. In some embodiments, Jane's device 6100-2 replaces John's video feed 6132 and / or Jane's video feed 6134 with surface video feed 6142.

[0266] FIG. 6AK illustrates an alternate embodiment depicting the sharing of content from the camera application in response to detecting input 6138 on share option 6136-1. In FIG. 6AK, John's laptop displays camera application window 6114 with surface view 6116 (optionally minimizing or hiding video conferencing application window 6120). John's device 6100-1 also displays John's video feed 6115 (similar to John's video feed 6124) and Jane's video feed 6117 (similar to Jane's video feed 6122), indicating that John's laptop is sharing surface view 6116 with Jane's device 6100-2 in a live video communication session (e.g., the video chat provided by the video conferencing application). In some embodiments, the display of John's video feed 6115 and / or Jane's video feed 6117 is optional. Similar to the embodiment shown in FIG. 6AJ, Jane's device 6100-2 shows surface video feed 6142, which is the surface view (from the camera application) being shared by John's device 6100-1.

[0267] FIG. 6AL illustrates a schematic view representing the field-of-view of camera 6102, and the portions of the field-of-view that are being used for the video conferencing application and camera application, for the embodiments depicted in FIGS. 6AF-6AK. For example, in FIG. 6AL, a profile view of John's laptop 6100 is shown in John's physical environment. Dashed line 6145-1 and dotted line 6147-2 represent the outer dimensions of the field-of-view of camera 6102, which in some embodiments is a wide angle camera. The collective field-of-view of camera 6102 is indicated by shaded regions 6144, 6146, and 6148. The portion of the camera field-of-view that is being used for the camera application (e.g., for surface view 6116) is indicated by dotted lines 6147-1 and 6147-2 and shaded regions 6146 and 6148. In other words, surface view 6116 (and surface view 6140) is generated by the camera application using the portion of the camera's field-of-view represented by shaded regions 6146 and 6148 that are between dotted lines 6147-1 and 6147-2. The portion of the camera field-of-view that is being used for the video conferencing application (e.g., for John's video feed 6124) is indicated by dashed lines 6145-1 and 6145-2 and shaded regions 6144 and 6146. In other words, John's video feed 6124 is generated by the video conferencing application using the portion of the camera's field-of-view represented by shaded regions 6144 and 6146 that are between dashed lines 6145-1 and 6145-2. Shaded region 6146 represents an overlap of the portion of the camera field-of-view that is being used to generate the video feeds for the respective camera and video conferencing applications.

[0268] FIGS. 6AM-6AY illustrate embodiments for controlling and / or interacting with the various user interfaces and views illustrated and described with reference to FIGS. 6A-6AL. In the embodiments depicted in FIGS. 6AM-6AY, the interfaces are illustrated using a tablet (e.g., John's tablet 600-1 and / or Jane's device 600-2) and computer (e.g., Jane's computer 600-4). The embodiments illustrated in FIGS. 6AM-6AY are optionally implemented using a different device, such as a laptop (e.g., John's device 6100-1 and / or Jane's device 6100-2). Similarly, the embodiments illustrated in FIGS. 6A-6AL are optionally implemented using a different device, such as Jane's computer 6100-2. Therefore, various operations or features described above with respect to FIGS. 6A-6AL are not repeated below for the sake of brevity.

[0269] Additionally, the applications, interfaces (e.g., 604-1, 604-2, 6121, and / or 6131) and field-of-views (e.g., 620, 688, 6145-1, and 6147-2) provided by one or more cameras (e.g., 602, 682, and / or 6102) discussed with respect to FIGS. 6A-6AL are similar to the applications, interfaces (e.g., 604-4) and field-of-views (e.g., 620) provided by camera (e.g., 602) discussed with respect to FIGS. 6AM-6AY. Accordingly, details of these applications, interfaces, and field-of-views may not be repeated below for the sake of brevity. Additionally, the options and requests (e.g., inputs and / or hand gestures) detected by device 600-1 to control the views associated with displayed elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, 624-2, 6121, and / or 6131) discussed with respect to FIGS. 6A-6AL are optionally detected by device 600-2 and / or device 600-4 to control the views associated with displayed elements (e.g., 622-1, 622-4, 623-1, 623-4, 6214, and / or 6216) discussed with respect to FIGS. 6AM-6AY (e.g., user 623 optionally provides the input to cause device 600-1 and / or device 600-2 to provide representation 624-1 including a modified image of a surface). Additionally, devices 600-1 and 600-2 in FIGS. 6AM-6AY are described and depicted as being in a landscape orientation. In some embodiments, device 600-1 and / or device 600-2 are in a portrait orientation, similar to device 600-1 in FIG. 6O. Accordingly, details of these the options and requests detected by device 600-2 may not be repeated below for the sake of brevity.

[0270] FIGS. 6AM-6AJ illustrate and describe exemplary user interfaces for controlling a view of a physical environment. The user interfaces in these figures are used to illustrate the processes described below, including the processes in FIG. 15. At FIG. 6AM, device 600-1 and device 600-4 display interfaces 604-1 and 604-4, respectively. Interface 604-1 includes representation 622-1 and interface 604-4 includes representation 622-4. Representations 622-1 and 622-4 include images of image data from a portion of field-of-view 620, specifically shaded region 6206. As illustrated, representations 622-1 and 622-4 include an image of a head and upper torso of user 622 and do not include an image of drawing 618 on desk 621. Interfaces 604-1 and 604-4 include representations 623-1 and 623-4, respectively, that include an image of user 223 that is in the field-of-view 6204 of camera 6202. Interfaces 604-1 and 604-4 further include options menu 609 (similar to options menu 608 discussed with respect to FIGS. 6A-6AE to control image data captured by 602 and / or captured by camera 6202, including FIGS. 6F-6G) allowing devices 600-1 and 600-4 to manage how image data is displayed.

[0271] At FIG. 6AN, user 623 brings device 600-2 near device 600-4 during a live video communication session. As depicted, in response to detecting device 600-2 (e.g., via wireless communication), device 600-4 displays add notification 6210a. Similarly, in response to detecting device 600-4, device 600-2, via display 683 (e.g., a touch-sensitive display), displays add notification 6210b. In some embodiments, devices 600-2 and 600-4 use specific device criteria to trigger the display of add notifications 6210a and 6210b. In some embodiments, the specific device criteria includes a criterion for a specific position (e.g., location, orientation, and / or angle) of device 600-2 that, when satisfied, triggers the display of add notifications 6210a and / or 6210b. In such embodiments, the specific position (e.g., location, orientation, and / or angle) of device 600-2 includes a criterion that device 600-2 has a specific angle or is within a range of angles (e.g., an angle or range of angles that indicate that the device is horizontal and / or lying flat on desk 686) and / or display 683 facing up (e.g., as opposed to facing down toward desk 686). In some embodiments, the specific device criteria include a criterion that device 600-2 is near device 600-4 (e.g., is within a threshold distance of device 600-4). In some embodiments, device 600-2 is in wireless communication with device 600-4 to communicate a location and / or proximity of device 600-2 (e.g., using location data and / or short-range wireless communications, such as Bluetooth and / or NFC). In some embodiments, the specific device criteria includes a criterion that device 600-2 and device 600-4 are associated with (e.g., are being used by and / or are logged into by) the same user. In some embodiments, the specific device criteria includes a criterion that device 600-2 has a particular state (e.g., unlocked and / or the display is powered on, as opposed to locked and / or the display is powered off).

[0272] At FIG. 6AN, connect notifications 6210a-6210b includes an indication of including device 600-4 in the live video communication session. For instance, add notifications 6210a-6210b includes an indication of adding a representation, for display on device 600-2, that includes an image of field-of-view 620 captured by camera 602. In some embodiments, the add notifications 6210a-6210b includes an indication of adding a representation, for display on device 600-1, that includes an image that is of field-of-view 6204 captured by camera 6202.

[0273] At FIG. 6AN, add notifications 6210a and 6210b include accept affordances 6212a and 6212b that, when selected, add (e.g., connect) device 600-2 to the live video communication session. Notifications 6210a and 6210b include decline affordances 6213a and 6213b that, when selected, dismiss notifications 6210a and 6210b, respectively, without adding device 600-2 to the live video communication session. While displaying accept affordance 6212b, device 600-2 detects input 6250an (e.g., tap, mouse click, or other selection input) directed at accept affordance 6212b. In response to detecting input 6250an, device 600-2 displays interface 604-2, as depicted in FIG. 6AO.

[0274] At FIG. 6AO, interface 604-2 is similar to interface 604-2 described herein (e.g., in reference to FIGS. 6A-6AE) and video conferencing interface 6131 as described herein (e.g., in reference to FIGS. 6AH-6AK) but has a different state. For example, interface 604-2 of FIG. 6AO does not include representations 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and options menu 609. In some embodiments, interface 604-2 of FIG. 6AO includes one or more of representations 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and / or options menu 609.

[0275] At FIG. 6AO, interface 604-2 includes adjustable view 6214 of a video feed captured by camera 602 (similar to John's video feed 6132 and representation 622-2, but including a different portion of the field of view 620). Adjustable view 6214 is associated with a portion of the field-of-view 620 corresponding to shaded region 6217. In some embodiments, interface 604-2 of FIG. 6AO includes representations 622-4 and 623-4 and / or option menu 609. In some embodiments, representations 622-4 and 623-4 and / or option menu 609 are moved from interface 604-4 to interface 604-2 in response to input detected at device 600-2 and / or device 600-4 so as to be concurrently displayed with adjustable view 6214. In such embodiments, display 6201 acts as a secondary display (e.g., extended display) of display 604-1 and / or vice versa.

[0276] At FIG. 6AO, in response to detection of input 6250an at FIG. 6AN, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as depicted in FIG. 6AO. Interface 604-1 of FIG. 6AO is similar to interface 604-1 of FIG. 6AN but has a different state (e.g., representations 623-1 and 622-1 are smaller in size and in different positions). Interface 604-1 includes adjustable view 6216, which is similar to adjustable view 6214 displayed at device 600-2 (e.g., adjustable view 6216 is associated with a portion of the field-of-view 620 corresponding to shaded region 6217). Adjustable view 6216 is updated to include similar images as adjustable view 6214 when inputs (e.g., movements of device 600-2) described herein are detected by device 600-2. Displaying adjustable view 6216 allows user 622 to see what portion of field-of-view 620 user 624 is currently viewing since, as described in greater detail below, user 623 optionally controls what view within field-of-view 620 is displayed.

[0277] At FIG. 6AO, while displaying interface 604-2, device 600-2 detects movement 6218ao of device 600-2. In response to detecting movement 6218ao, device 600-2 displays interface 602-4 of FIG. 6AP. Additionally, in response to detecting movement 6218ao, device 600-2 causes device 600-1 to display interface 604-1 of FIG. 6AP.

[0278] At FIG. 6AP, interface 602-4 includes an updated adjustable view 6214. Adjustable view 6214 of FIG. 6AP is a different view within field-of-view 620 as compared to adjustable view 6214 of FIG. 6AO. For example, shaded region 6217 of FIG. 6AP has moved with respect to shaded region 6217 of FIG. 6AO. Notably, camera 602 has not moved. In some embodiments, movement 6218ao of device 600-2 corresponds to (e.g., is proportional to) the amount of change in adjustable view 6214. For example, in some embodiments, the magnitude of the angle in which device 600-2 rotates (e.g., with respect to gravity) corresponds to the amount of change in adjustable view 6214 (e.g., the amount the image data is panned to include a new angle of view). In some embodiments, the direction of a movement (e.g., movement 6218ao) of device 600-2 (e.g., tilting down and / or rotating down) corresponds to the direction of change in adjustable view 6214 (e.g., pans down). In some embodiments, the acceleration and / or speed of a movement (e.g., movement 6218) corresponds to the speed in which adjustable view 6214 changes. In some embodiments, device 600-2 (and / or device 600-1) displays a gradual transition (e.g., a series views) from adjustable view 6214 in FIG. 6AO to adjustable view 6214 in FIG. 6AP. Additionally or alternatively, as depicted in FIG. 6AP, device 600-2 is lying flat on desk 686. In some embodiments, in response to detecting a specific position or a position within a predefined range of positions (e.g., horizontal and / or display up), device 600-2 displays the adjustable view 6214 of FIG. 6AP. As depicted, movement 6218ao in FIG. 6AO does not cause device 600-2 to update representations 622-4 and 623-4 (and / or representations 623-1 and 622-1 on device 600-1) in FIG. 6AP.

[0279] At FIG. 6AP, image of drawing 618 in adjustable view 6214 is at a different perspective than the perspective of the image of drawing 618 in adjustable view 6214 of FIG. 6AO. For example, adjustable view 6214 of FIG. 6AP includes a top-view perspective whereas adjustable view 6214 of FIG. 6AO includes a perspective that includes a combination of a side view and a top view. In some embodiments, the image of the drawing included in adjustable view 6214 of FIG. 6AP is based on image data that has been modified (e.g., skewed and / or magnified) using similar techniques described in reference to FIGS. 6A-6AL. In some embodiments, the image of drawing included in adjustable view 6214 of FIG. 6AO is based on image data that has not been modified (e.g., skewed and / or magnified) and / or has been modified in a different manner (e.g., at a lesser degree) than image of drawing 618 in adjustable view 6214 of FIG. 6AP (e.g., less skewed and / or less magnified as compared to the amount of skew and / or amount of magnification applied in FIG. 6AP). Providing a top-view perspective provides greater ease in collaborating and sharing content as it gives user 623 a view of the drawing that would be similar to the view user 623 would have if user 623 was sitting across from user 622 looking down at surface 619 of desk 621.

[0280] At FIG. 6AP, adjustable view 6216 of interface 604-1 has also been updated in a similar manner. In some embodiments, the images of adjustable view 6216 and / or adjustable view 6214 are modified based on a position of surface 619 relative to camera 602, as described in reference to FIGS. 6A-6AL. In such embodiments, device 600-1 and / or device 600-2 rotate the image of adjustable view 6214 by an amount (e.g., 45 degrees, 90 degrees, or 180 degrees) such that the image of drawing 618 can be more intuitively viewed in adjustable view 6216 and / or adjustable view 6214 (e.g., the image of drawing 618 is displayed such that the house is right-side up as opposed to upside down).

[0281] At FIG. 6AP, user 623 applies digit marks to adjustable view 6214 using stylist 6220. For example, while displaying adjustable view 6214 of FIG. 6AP, device 600-2 detects an input corresponding to a request to add digital marks to adjustable view 6214 (e.g., using stylist 6220). In response to detecting the input corresponding to the request to add digital marks to adjustable view 6214, device 600-2 displays interface 604-2, as depicted in FIG. 6AO. Additionally or alternatively, in response to detecting the input corresponding to the request to add digital marks to adjustable view 6214, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as depicted in FIG. 6AQ.

[0282] At FIG. 6AQ, interface 602-4 includes digital sun 6222 in adjustable view 6214 and interface 602-1 includes digital sun 6223 in adjustable view 6214. Displaying a digital sun at both devices allow users 623 and 622 to collaborate over the video communication session. Additionally, as depicted, digital sun 6222 has a position with respect to image of drawing 618. As described in greater detail below, digital sun 6222 maintains its position with respect to image of drawing 618 even if device 600-1 detects further movement and / or if drawing 618 moves on surface 619. In some embodiments, device 600-2 stores data corresponding to the relationship between digital marks (e.g., digital sun 6223) and objects (e.g., the house) detected in image data so as to determine where (and / or if) digital sun 6222 should be displayed. In some embodiments, device 600-2 stores data corresponding to the relationship between digital marks (e.g., digital sun 6223) and the position of device 600-2 so as to determine where (and / or if) digital sun 6222 should be displayed. In some embodiments, device 600-2 detects digital marks applied to other views in field-of-view 620. For example, digital marks can be applied in an image of a head of a user, such as the image of the head of user 622 in adjustable view 6214 of FIG. 6AR.

[0283] At FIG. 6AQ, interface 604-2 includes control affordance 648-1 (similar to control affordance 648-1 in FIG. 6N) to modify the image in adjustable view 6214. Rotation affordance 648-1, when selected, causes device 600-1 (and / or device 600-2) to rotate the image of adjustable view 6214, similar to how control affordance 648-1 modifies the image of representation 624-1 in FIG. 6N.

[0284] At FIG. 6AQ, in some embodiments, interface 604-2 includes a zoom affordance similar to zoom affordance 648-2 in FIG. 6N. In such embodiments, the zoom affordance modifies the image in adjustable view 6214, similar to how zoom affordance 648-2 modifies the image of representation 624-1 in FIG. 6N. Control affordances 648-1, 648-2 can be displayed in response to one or more inputs, for instance, corresponding to a selection of an affordance of options menu 609 (e.g., FIG. 6AM).

[0285] At FIG. 6AQ, in some embodiments, digital sun 6222 is projected onto a physical surface of drawing 618, similar to how markup 956 is projected onto surface 908b that is described in FIGS. 9K-9N. In such embodiments, an electronic device (e.g., a projector and / or a light emitting projector) is used to project an image and / or rendering of digital sun 6222 within physical environment 915. For example, an electronic device can cause a projection of a digital sun to be displayed next to drawing 618 based on the relative location of digital sun 6222 with respect to drawing 618 using the techniques described with respect to FIGS. 9K-9N.

[0286] At FIG. 6AQ, while displaying digital sun 6222 in adjustable view 6214, device 600-2 detects movement 6218aq (e.g., rotation and / or lifting). In response to detecting movement 6218aq, device 600-2 displays interface 604-2, as depicted in FIG. 6AR. In response to detecting movement 6218aq, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as depicted in FIG. 6AR.

[0287] At FIG. 6AR, interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). Adjustable view 6214 of FIG. 6AR is a different view within field-of-view 620 as compared to adjustable view 6214 of FIG. 6AQ. For example, shaded region 6217 of FIG. 6AP has moved with respect to shaded region 6217 of FIG. 6AQ. In some embodiments, the direction of movement 6218aq (e.g., tilting up) corresponds to the direction of the change in view (e.g., pan up). Additionally, shaded region 6217 overlaps with shaded region 6206, as depicted by darker shaded region 6224. Darker shaded region 6224 is a schematic representation that updated adjustable view 6214 is based on a portion of image data that is used for representation 622-4. Because movement 6218aq has resulted in changing the view (e.g., to the face of user 622 and / or not a view of drawing 618), device 600-2 no longer displays digital sun 6222 in adjustable view 6214.

[0288] At FIG. 6AR, adjustable view 6214 includes boundary indicator 6226. Boundary indicator 6226 indicates that a boundary has been reached. In some embodiments, the boundary is configured (e.g., by a user) to set a limit on what portion of field-of-view 620 (or the environment captured by camera 602) is provided for display. For example, user 622 can limit what portion is available to user 623. In some embodiments, the boundary is defined by physical limitations of camera 602 (e.g., image sensors and / or lenses) that provide field-of-view 620. At FIG. 6AR, shaded region 6217 has not reached the limits of field-of-view 620. As such, boundary indicator 6226 is based on a configurable setting that limits what portion of field-of-view 620 is provided for display. Turning briefly to FIG. 6AT, boundary indicator 6226 is displayed in response to a determination that the perspective provided in adjustable view 6214 has reached the edge of field-of-view 620.

[0289] At FIG. 6AR, boundary indicator 6226 is depicted with cross-hatching. In some embodiments, security boundary indicator 6226 is a visual effect (e.g., a blur and / or fade) applied to adjustable view 6214 (and / or adjustable view 6216). In some embodiments, boundary indicator 6226 is displayed along an edge of adjustable view 6214 (and / or 6216) to indicate the position of boundary. At FIG. 6AR, boundary indicator 6226 is displayed along the top and side edge to indicate that the user cannot see above and / or further to the side of boundary indicators 6226. While displaying interface 604-2 at FIG. 6AR, device 600-2 detects movement 6218ar (e.g., rotation and / or lowering). In response to detecting movement 6218ar, device 600-2 displays interface 604-2, as depicted in FIG. 6AS. In response to detecting movement 6218ar, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as depicted in FIG. 6AS.

[0290] At FIG. 6AS, interface 604-2 includes an updated adjustable view 6214, which includes the image of drawing 618. At FIG. 6AS, device 600-2 is in a similar position as device 600-2 was in FIG. 6AO. As such, adjustable view 6214 of FIG. 6AS includes the same perspective of the image of drawing 618 in adjustable view 6214 as the perspective of the image of drawing 618 in adjustable view 6214 in FIG. 6AO. Notably, device 600-2 displays digital sun 6222 in adjustable view 6214 of FIG. 6AS. The position of digital sun 6222 with respect to the house of drawing 618 in FIG. 6AS is similar to the position of digital sun 6222 with respect to the house of drawing 618 in FIG. 6AQ, except with slight differences based on the different view. As such, digital sun 6222 appears to be fixed in physical space, as if it were drawn next to drawing 618. Fixing the position of a digital mark in physical space facilitates better collaboration between the users, since a user can digitally draw or write in one view, move the device to see a different view, and then move the device back so as to re-display the digital drawings or writings and the context in which they were made.

[0291] For the sake of clarity, shaded regions 6217 and 6206 and field-of-view 620 have been have been omitted from FIGS. 6AS-6AU. In some embodiments, representations 622-1 and adjustable views 6214 and 6216 correspond to views associated with shaded regions 6217 and 6206 and field-of-view 620 of FIG. 6AO.

[0292] At FIG. 6AS, device 600-2 (and / or device 600-1) detects movement of drawing 618 and maintains display of the image of drawing 618 in adjustable view 6214. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to modify (e.g., zoom, skew, and / or rotate) image data so as to maintain display of the image of drawing 618 in adjustable view 6214. While displaying interface 604-2, device 600-2 (and / or device 600-1) detects horizontal movement 6230 of drawing 618. In response to detecting horizontal movement 6230 of drawing 618, device 600-2 displays interface 604-2, as depicted in FIG. 6AT. In some embodiments, in response to detecting horizontal movement 6230 of drawing 618, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as depicted in FIG. 6AT. In some embodiments, in response to device 600-1 detecting horizontal movement 6230 of drawing 618, device 600-2 displays (and / or device 600-1 causes device 600-2 to display) interface 602-4, as depicted in FIG. 6AT.

[0293] At FIG. 6AT, drawing 618 has been moved to the edge of desk 621, which is further away from (e.g., and to the side) of camera 602. Despite the change in position, interface 602-4 of FIG. 6AT includes image of drawing 618 in adjustable view 6214 that appears mostly unchanged from the image of drawing 618 in adjustable view 6214 of interface 602-4 of FIG. 6AS. For example, adjustable view 6214 provides a perspective that makes it appear that drawing 618 is still straight in front of camera 602, similar to the position of drawing 618 in FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct (e.g., by skewing and / or magnifying) the image of drawing 618 based on a new position with respect to camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it moves with respect to camera 602. In some embodiments, adjustable view 6214 of interface 604-2 of FIG. 6AT is provided without any change in position (e.g., location, orientation, and / or rotation) of camera 602.

[0294] At FIG. 6AT, device 600-2 displays boundary indicator 6226 in adjustable view 6214 (similar to adjustable view 6214 displayed by device 600-1 in adjustable view 6216). As discussed above with respect to FIG. 6AR, boundary indicator 6226 indicates that a limit of the field-of-view or physical space has been reached. At FIG. 6AT, device 600-2 displays boundary indicator 6226 in adjustable view 6214 to indicate that an edge of field-of-view 620 has been reached. Boundary indicator 6226 is along the right edge of adjustable view 6214 (and adjustable view 6216) indicating that views to the right of the current view exceed the field-of-view of camera 602.

[0295] At FIG. 6AT, digital sun 6222 maintains a similar respective position in relation to the house in the image of drawing 618 in adjustable view 6214 as the respective position of digital sun 6222 in relationship to the house in the image of drawing 618 in adjustable view 6214 of FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) displays digital sun 6222 overlaid on the image of drawing 618 that has been corrected based on the new position of drawing 618.

[0296] Returning briefly to FIG. 6AS, while displaying interface 602-4, device 600-2 (and / or device 600-1) detects rotation 6232 of drawing 618. In response to detecting rotation 6232 of drawing 618, device 600-2 displays interface 604-2, as depicted in FIG. 6AU. In some embodiments, in response to detecting rotation 6232 of drawing 618, device 600-2 causes device 600-1 to display interface 601-4, as depicted in FIG. 6AU. In some embodiments, in response to device 600-1 detecting rotation 6232 of drawing 618, device 600-2 displays (or device 600-1 causes device 600-2 to display) interface 602-4, as depicted in FIG. 6AU.

[0297] At FIG. 6AU, drawing 618 has been rotated with respect to edge of desk 621. Despite the change in position, interface 602-4 in FIG. 6AU includes image of drawing 618 in adjustable view 6214 that appears mostly unchanged from the image of drawing 618 in adjustable view 6214 of interface 604-2 in FIG. 6AS. That is, adjustable view 6214 of FIG. 6AU provides a perspective that makes it appear as if drawing 618 was not rotated, similar to the position of drawing 618 in FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct (e.g., by skewing and / or rotating) the image of drawing 618 based on the new position with respect to camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it rotates with respect to camera 602. In some embodiments, adjustable view 6214 of interface 604-2 of FIG. 6AU is provided without any change in position (e.g., location, orientation, and / or rotation) of camera 602. Adjustable view 6216 is updated in a similar manner as adjustable view 6214.

[0298] At FIG. 6AU, digital sun 6222 maintains a similar position in relation to the house in the image of drawing 618 in adjustable view 6214 as the position of digital sun 6222 in relationship to the house in the image of drawing 618 in adjustable view 6214 of FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) displays digital sun 6222 overlaid on the image of drawing 618 that has been corrected based on the rotation of drawing 618.

[0299] At FIG. 6AV, device 600-2 displays interface 604-2, which is similar to interface 604-2 of FIG. 6AU but having a different state (e.g., representation 622-2 of John and options menu 609 have been added to user interface 604-2). Device 600-4 is no longer being used in the live communication session. Additionally, device 600-2 has been moved from its position in FIG. 6AU to the same position device 600-2 had in FIG. 6AQ. As such, device 600-2 updates adjustable view 6214 of FIG. 6AV to include the same perspective as adjustable view 6214 of FIG. 6AQ. As illustrated, adjustable view 6214 includes a top-view perspective. Additionally, digital sun 6222 is displayed as having the same position of digital sun 6222 in relationship to the house in the image of drawing 618 in adjustable view 6214 of FIG. 6AQ.

[0300] At FIG. 6AV, while displaying digital sun 6222 in adjustable view 6214, device 600-2 detects movement 6218av (e.g., rotation and / or lifting). In response to detecting movement 6218av, device 600-2 displays interface 604-2, as depicted in FIG. 6AW. In response to detecting movement 6218aw, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as depicted in FIG. 6AW.

[0301] At FIG. 6AW, interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1) similar to adjustable view 6214 of FIG. 6AR. Notably, device does not update representation 622-2 in response to detecting movement 6218aw. Accordingly, in some embodiments, device 600-2 displays a dynamic representation that is updated based on the position of device 600-2 and a static representation that is not updated based on the position of device 600-2. Interface 604-2 also includes boundary indicator 6226 in adjustable view 6214, similar to boundary indicator 6226 of FIG. 6AR.

[0302] At FIG. 6AW, while displaying interface 604-2, device 600-2 detects movement 6218aw (e.g., rotation and / or lowering). In response to detecting movement 6218aw, device 600-2 displays interface 604-2, as depicted in FIG. 6AX. In response to detecting movement 6218aw, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as depicted in FIG. 6AX.

[0303] At FIG. 6AX, interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). Because adjustable view 6214 is substantially the same view provided by representation 622-2, shaded region 6206 overlaps shaded region 6217. Because movement 6218aq results in changing the view to the face of user 622 and / or not a view of drawing 618, device 600-2 no longer displays digital sun 6222 in adjustable view 6214. While displaying interface 604-2 at FIG. 6AX, device 600-2 (and / or device 600-1) detects a set of one or more inputs (e.g., similar to the inputs and / or hand gestures described in reference to FIGS. 6A-6AL) corresponding to a request to display a surface view. In some such embodiments, 616-1 of FIG. 6C, 616-2 of FIG. 6G, preview mode 674-1 of FIG. 6H, representation 676 of preview mode 674-2 in FIG. 6I, representation 676 of preview mode interface 674-3 in FIG. 6J, affordances 648-1, 648-2, 648-3 of FIGS. 6N-6Q are displayed at device 600-2 so as to allow device 600-2 to control the representation of the modified image of drawing 618 in the same manner as the detected inputs at device 600-1. In response to detecting the set of one or more inputs corresponding to a request to display a surface view, device 600-2 displays interface 604-2, as depicted in FIG. 6AY. Additionally or alternatively, in response to detecting the set of one or more inputs, device 600-1 displays interface 604-2, as depicted in FIG. 6AY. In some embodiments, device 600-1 detects the set of one or more inputs, as described in reference to FIGS. 6A-6AL. In some embodiments, device 600-2 detects the set of one or more inputs. In such embodiments, device 600-2 detects a selection of view affordance 6236 of options menu 609, which is similar to view affordance 607-2 of option menu 608 described in reference to FIG. 6F. In response, a view menu similar to view menu 616-2 as described with reference to FIG. 6G includes an affordance to request display of a surface view of a remote participant.

[0304] At FIG. 6AY, adjustable view 6214 includes a surface view, which is similar to representation 624-1 depicted in FIG. 6M, for example, and described above. As depicted in FIG. 6AY, adjustable view 6214 includes an image that is modified such that user 623 has a similar perspective looking down at the image of drawing 618 displayed on device 600-2 as the perspective user 622 has when looking down at drawing 618 in the physical environment, as described in greater detail with respect to FIGS. 6A-6AL. Notably, digital sun 6222 of FIG. 6AY is displayed as having the same position in relationship to the house in the image of drawing 618 in adjustable view 6214 as does digital sun 6222 of FIG. 6AQ.

[0305] FIG. 7 is a flow diagram illustrating a method for managing a live video communication session using a computer system, in accordance with some embodiments. Method 700 is performed at a computer system (e.g., 600-1, 600-2, 600-3, 600-4, 906a, 906b, 906c, 906d, 6100-1, 6100-2, 1100a, 1100b, 1100c, and / or 1100d) (e.g., a smartphone, a tablet, a laptop computer, and / or a desktop computer) (e.g., 100, 300, or 500) that is in communication with a display generation component (e.g., 601, 683, and / or 6101) (e.g., a display controller, a touch-sensitive display system, and / or a monitor), one or more cameras (e.g., 602, 682, and / or 6102) (e.g., an infrared camera, a depth camera, and / or a visible light camera), and one or more input devices (e.g., 601, 683, and / or 6103) (e.g., a touch-sensitive surface, a keyboard, a controller, and / or a mouse). Some operations in method 700 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.

[0306] As described below, method 700 provides an intuitive way for managing a live video communication session. The method reduces the cognitive burden on a user for managing a live video communication session, thereby creating a more efficient human-machine interface. For battery-operated computing devices, enabling a user to manage a live video communication session faster and more efficiently conserves power and increases the time between battery charges.

[0307] In method 700, computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays (702), via the display generation component, a live video communication interface (e.g., 604-1, 604-2, 6120, 6121, 6130, and / or 6131) for a live video communication session (e.g., an interface for an incoming and / or outgoing live audio / video communication session). In some embodiments, the live communication session is between at least the computer system (e.g., a first computer system) and a second computer system.

[0308] The live video communication interface includes a representation (e.g., 622-1, 622-2, 6124, and / or 6132) of at least a portion of a field-of-view (e.g., 620, 688, 6144, 6146, and / or 6148) of the one or more cameras (e.g., a first representation). In some embodiments, the first representation includes images of a physical environment (e.g., a scene and / or area of the physical environment that is within the field-of-view of the one or more cameras). In some embodiments, the representation includes a portion (e.g., a first cropped portion) of the field-of-view of the one or more cameras. In some embodiments, the representation includes a static image. In some embodiments, the representation includes series of images (e.g., a video). In some embodiments, the representation includes a live (e.g., real-time) video feed of the field-of-view (or a portion thereof) of the one or more cameras. In some embodiments, the field-of-view is based on physical characteristics (e.g., orientation, lens, focal length of the lens, and / or sensor size) of the one or more cameras. In some embodiments, the representation is displayed in a window (e.g., a first window). In some embodiments, the representation of at least the portion of the field-of-view includes an image of a first user (e.g., a face of a first user). In some embodiments, the representation of at least the portion of the field-of-view is provided by an application (e.g., 6110) providing the live video communication session. In some embodiments, the representation of at least the portion of the field-of-view is provided by an application (e.g., 6108) that is different from the application providing the live video communication session (e.g., 6110).

[0309] While displaying the live video communication interface, the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) detects (704), via the one or more input devices (e.g., 601, 683, and / or 6103), one or more user inputs including a user input (e.g., 612c, 612d, 614, 612g, 612i, 612j, 6112, 6118, 6128, and / or 6138) (e.g., a tap on a touch-sensitive surface, a keyboard input, a mouse input, a trackpad input, a gesture (e.g., a hand gesture), and / or an audio input (e.g., a voice command)) directed to a surface (e.g., 619) (e.g., a physical surface; a surface of a desk and / or a surface of an object (e.g., book, paper, tablet) resting on the desk; or a surface of a wall and / or a surface of an object (e.g., a whiteboard or blackboard) on a wall; or other surface (e.g., a freestanding whiteboard or blackboard)) in a scene (e.g., physical environment) that is in the field-of-view of the one or more cameras. In some embodiments, the user input corresponds to a request to display a view of the surface. In some embodiments, detecting user input via the one or more input devices includes obtaining image data of the field-of-view of the one or more cameras that includes a gesture (e.g., a hand gesture, eye gesture, or other body gesture). In some embodiments, the computer system determines, from the image data, that the gesture satisfies predetermined criteria.

[0310] In response to detecting the one or more user inputs, the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays, via the display generation component (e.g., 601, 683, and / or 6101), a representation (e.g., image and / or video) of the surface (e.g., 624-1, 624-2, 6140, and / or 6142) (e.g., a second representation). In some embodiments, the representation of the surface is obtained by digitally zooming and / or panning the field-of-view captured by the one or more cameras. In some embodiments, the representation of the surface is obtained by moving (e.g., translating and / or rotating) the one or more cameras. In some embodiments, the second representation is displayed in a window (e.g., a second window, the same window in which the first representation is displayed, or a different window than a window in which the first representation is displayed). In some embodiments, the second window is different from the first window. In some embodiments, the second window (e.g., 6140 and / or 6142) is provided by the application (e.g., 6110) providing the live video communication session (e.g., as shown in FIG. 6AJ). In some embodiments, the second window (e.g., 6114) is provided by an application (e.g., 6108) different from the application providing the live video communication session (e.g., as shown in FIG. 6AK). In some embodiments, the second representation includes a cropped portion (e.g., a second cropped portion) of the field-of-view of the one or more cameras. In some embodiments, the second representation is different from the first representation. In some embodiments, the second representation is different from the first representation because the second representation displays a portion (e.g., a second cropped portion) of the field-of-view that is different from a portion (e.g., the first cropped portion) that is displayed in the first representation (e.g., a panned view, a zoomed out view, and / or a zoomed in view). In some embodiments, the second representation includes images of a portion of the scene that is not included in the first representation and / or the first representation includes images of a portion of the scene that is not included in the second representation. In some embodiments, the surface is not displayed in the first representation.

[0311] The representation (e.g., 624-1, 624-2, 6140, and / or 6142) of the surface includes an image (e.g., photo, video, and / or live video feed) of the surface (e.g., 619) captured by the one or more cameras (e.g., 602, 682, and / or 6102) that is (or has been) modified (e.g., to correct distortion of the image of the surface) (e.g., adjusted, manipulated, corrected) based on a position (e.g., location and / or orientation) of the surface relative to the one or more cameras (sometimes referred to as the representation of the modified image of the surface). In some embodiments, the image of the surface displayed in the second representation is based on image data that is modified using image processing software (e.g., skewing, rotating, flipping, and / or otherwise manipulating image data captured by the one or more cameras). In some embodiments, the image of the surface displayed in the second representation is modified without physically adjusting the camera (e.g., without rotating the camera, without lifting the camera, without lowering the camera, without adjusting an angle of the camera, and / or without adjusting a physical component (e.g., lens and / or sensor) of the camera). In some embodiments, the image of the surface displayed in the second representation is modified such that the camera appears to be pointed at the surface (e.g., facing the surface, aimed at the surface, pointed along an axis that is normal to the surface). In some embodiments, the image of the surface displayed in the second representation is corrected such that the line of sight of the camera appears to be perpendicular to the surface. According to the claimed invention, an image of the scene displayed in the first representation is not modified based on the location of the surface relative to the one or more cameras. According to the claimed invention, the representation of the surface is concurrently displayed with the first representation (e.g., the first representation (e.g., of a user of the computer system) is maintained and an image of the surface is displayed in a separate window). In some embodiments, the image of the surface is automatically modified in real time (e.g., during the live video communication session). In some embodiments, the image of the surface is automatically modified (e.g., without user input) based on the position of the surface relative to the one or more first cameras. Displaying a representation of a surface including an image of the surface that is modified based on a position of the surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the surface despite its position relative to the camera without requiring further input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0312] The computer system (e.g., 600-1 and / or 600-2) receives, during the live video communication session, image data captured by a camera (e.g., 602) (e.g., a wide angle camera) of the one or more cameras. The computer system displays, via the display generation component, the representation of the at least a portion of the field-of-view (e.g., 622-1 and / or 622-2) (e.g., the first representation) based on the image data captured by the camera. The computer system displays, via the display generation component, the representation of the surface (e.g., 624-1 and / or 624-2) (e.g., the second representation) based on the image data captured by the camera (e.g., the representation of at least a portion of the field-of view of the one or more cameras and the representation of the surface are based on image data captured by a single (e.g. only one) camera of the one or more cameras). Displaying the representation of the at least a portion of the field-of-view and the representation of the surface captured from the same camera enhances the video communication session experience by displaying content captured by the same camera at different perspectives without requiring input from the user, which reduces the number of inputs (and / or devices) needed to perform an operation.

[0313] In some embodiments, the image of the surface is modified (e.g., by the computer system) by rotating the image of the surface relative to the representation of at least a portion of the field-of-view-of the one or more cameras (e.g., the image of the surface in 624-2 is rotated 180 degrees relative to representation 622-2). In some embodiments, the representation of the surface is rotated 180 degrees relative to the representation of at least a portion of the field-of-view of the one or more cameras. Rotating the image of the surface relative to the representation of at least a portion of the field-of-view of the one or more cameras enhances the video communication session experience as content associated with the surface can be viewed from a different perspective that other portions of the field-of-view without requiring input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0314] In some embodiments, the image of the surface is rotated based on a position (e.g., location and / or orientation) of the surface (e.g., 619) relative to a user (e.g., 622) (e.g., a position of a user) in the field-of-view of the one or more cameras. In some embodiments, a representation of the user is displayed at a first angle and the image of the surface is rotated to a second angle that is different from the first angle (e.g., even though the image of the user and the image of the surface are captured at the same camera angle). Rotating the image of the surface based on a position of the surface relative to a user in the field-of-view of the one or more cameras enhances the video communication session experience as content associated with the surface can be viewed from a perspective that is based on the position of the surface without requiring input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0315] In some embodiments, in accordance with a determination that the surface is in a first position (e.g., surface 619 is positioned in front of user 622 on desk 621 in FIG. 6A) (e.g., a predefined position) relative to a user in the field-of-view of the one or more cameras (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by at least 45 degrees relative to a representation of the user in the field-of-view of the one or more cameras (e.g., the image of surface 619 in representation 624-1 is rotated 180 degrees relative to representation 622-1 in FIG. 6M). In some embodiments, the image of the surface is rotated in the range of 160 degrees to 200 degrees (e.g., 180 degrees). In some embodiments, in accordance with a determination that the surface is in a first position relative to a user in the field-of-view of the one or more cameras (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by a first amount. In some embodiments, the first amount is in the range of 160 degrees to 200 degrees (e.g., 180 degrees). In some embodiments, in accordance with a determination that the surface is in a second position relative to a user in the field-of-view of the one or more cameras (e.g., to a side of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by a second amount. In some embodiments, the second amount is in the range of 45 degrees to 120 degrees (e.g., 90 degrees). Rotating the image of the surface by at least 45 degrees relative to a representation of the user captured in the field-of-view of the one or more cameras when the surface is in a first position relative to the user enhances the video communication session experience by adjusting an image to provide a more natural, intuitive image without requiring further input from the user, which provides improved visual feedback and performs an operation when a set of conditions has been met without requiring further user input.

[0316] In some embodiments, the representation of the at least a portion of the field-of-view includes a user and is concurrently displayed with the representation of the surface (e.g., representations 622-1 and 624-1 or representations 622-2 and 624-2 in FIG. 6M). In some embodiments, the representation of the at least a portion of the field-of-view and the representation of the surface are captured by the same camera (e.g., a single camera of the one or more cameras) and are displayed concurrently. In some embodiments, the representation of the at least a portion of the field-of-view and the representation of the surface are displayed in separate windows that are concurrently displayed. Including a user in the representation of the at least a portion of the field-of-view and concurrently displaying the representation with the representation of the surface enhances the video communication session experience by allowing a user to view a reaction of participant while the representation of the surface is displayed without requiring further input from the user, which provides improved visual feedback and performs an operation when a set of conditions has been met without requiring further user input.

[0317] In some embodiments, in response to detecting the one or more user inputs and prior to displaying the representation of the surface, the computer system displays a preview of image data for the field-of-view of the one or more cameras (e.g., as depicted in FIGS. 6H-6J) (e.g., in a preview mode of the live video communication interface), the preview including an image of the surface that is not modified based on the position of the surface relative to the one or more cameras (sometimes referred to as the representation of the unmodified image of the surface). In some embodiments, the preview of the field-of-view is displayed after displaying the representation of the image of the surface (e.g., in response to detecting user input corresponding to selection of the representation of the surface. Displaying a preview including an image of the surface that is not modified based on the position of the surface relative to the one or more cameras allows the user to quickly identify the surface within the preview as no distortion correction has been applied, which provides improved visual feedback.

[0318] In some embodiments, displaying the preview of image data for the field-of-view of the one or more cameras includes displaying a plurality of selectable options (e.g., 636-1 and / or 636-2 of FIG. 6I, or 636a-i of FIG. 6J) corresponding to respective portions of (e.g., surfaces within) the field-of-view of the one or more cameras. In some embodiments, the computer system detects an input (e.g., 612i or 612j) selecting one of the plurality of options corresponding to respective portions of the field-of-view of the one or more cameras. In response to detecting the input selecting one of the plurality of options corresponding to respective portions of the field-of-view of the one or more cameras and in accordance with a determination that the input selecting one of the plurality of options corresponding to respective portions of the field-of-view of the one or more cameras is directed to a first option corresponding to a first portion of the field-of-view of the one or more cameras, the computer system displays the representation of the surface based on the first portion of the field-of-view of the one or more cameras (e.g., selection of 636h in FIG. 6J causes display of the corresponding portion) (e.g., the computer system displays a modified version of an image of the first portion of the field-of-view, optionally with a first distortion correction). In response to detecting the input selecting one of the plurality of options corresponding to respective portions of the field-of-view of the one or more cameras and in accordance with a determination that the input selecting one of the plurality of options corresponding to respective portions of the field-of-view of the one or more cameras is directed to a second option corresponding to a second portion of the field-of-view of the one or more cameras, the computer system displays the representation of the surface based on the second portion of the field-of-view of the one or more cameras (e.g., selection of 636g in FIG. 6J causes display of the corresponding portion) (e.g., the computer system displays a modified version of an image of the second portion of the field-of-view, optionally with a second distortion correction that is different from the first distortion correction), wherein the second option is different from the first option. Displaying a plurality of selectable options corresponding to respective portions of the field-of-view of the one or more cameras in the preview of image data allows a user to identify portions of the field-of-view that are capable of being displayed as a representation in the video conference interface, which provides improved visual feedback.

[0319] In some embodiments, displaying the preview of image data for the field-of-view of the one or more cameras includes displaying a plurality of regions (e.g., distinct regions, non-overlapping regions, rectangular regions, square regions, and / or quadrants) of the preview (e.g., 636-1, 636-2 of FIG. 6I, and / or 636a-i of FIG. 6J) (e.g., the one or more regions may correspond to distinct portions of the image data for the field-of-view.). In some embodiments, the computer system detects a user input (e.g., 612i and / or 612j) corresponding to one or more regions of the plurality of regions. In response to detecting the user input corresponding to the one or more regions and in accordance with a determination that the user input corresponding to the one or more regions corresponds to a first region of the one or more regions, the computer system displays a representation of the first region in the live video communication interface (e.g., as described with reference to FIGS. 6I-6J) (e.g., with a distortion correction based on the first region). In response to detecting the user input corresponding to the one or more regions and in accordance with a determination that the user input corresponding to the one or more regions corresponds to a second region of the one or more regions, the computer system displays a representation of the second region as a representation in the live video communication interface (e.g., with a distortion correction based on the second region that is different from the distortion correction based on the first region). Displaying a representation of the first region or a representation of the second region in the live video communication interface enhances the video communication session experience by allowing a user to efficiently manage what is displayed in the live video communication interface, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0320] In some embodiments, the one or more user inputs include a gesture (e.g., 612d) (e.g., a body gesture, a hand gesture, a head gesture, an arm gesture, and / or an eye gesture) in the field-of-view of the one or more cameras (e.g., a gesture performed in the field-of-view of the one or more cameras that is directed to the physical position surface). Utilizing a gesture in the field-of-view of the one or more cameras as an input enhances the video communication session experience by allowing a user to control what is displayed without physically touching a device, which provides additional control options without cluttering the user interface.

[0321] In some embodiments, the computer system displays a surface-view option (e.g., 610) (e.g., icon, button, affordance, and / or user-interactive graphical interface object), wherein the one or more user inputs include an input (e.g., 612c and / or 612g) directed to the surface-view option (e.g., a tap input on a touch-sensitive surface, a click with a mouse while a cursor is over the surface-view option, or an air gesture while gaze is directed to the surface-view option). In some embodiments, the surface-view option is displayed in the representation of at least a portion of a field-of-view of the one or more cameras. Displaying a surface-view option enhances the video communication session experience by allowing a user to efficiently manage what is displayed in the live video communication interface, which provides additional control options without cluttering the user interface.

[0322] In some embodiments, the computer system detects a user input corresponding to selection of the surface-view option. In response to detecting the user input corresponding to selection of the surface-view option, the computer system displays a preview of image data for the field-of-view of the one or more cameras (e.g., as depicted in FIGS. 6H-6J) (e.g., in a preview mode of the live video communication interface), the preview including a plurality of portions of the field-of-view of the one or more cameras including the at least a portion of the field-of-view of the one or more cameras (e.g., 636-1, 636-2 of FIG. 6I, and / or 636a-i of FIG. 6J), wherein the preview includes a visual indication (e.g., text, a graphic, an icon, and / or a color) of an active portion of the field-of-view (e.g., 641-1 of FIG. 6I, and / or 640a-f of FIG. 6J) (e.g., the portion of the field-of-view that is being transmitted to and / or displayed by other participants of the live video communication session). In some embodiments, the visual indication indicates that a single portion (e.g., only one) portion of the plurality of portions of the field-of-view is active. In some embodiments, the visual indication indicates that two or more portions of the plurality of portions of the field-of-view are active. Displaying a preview of a plurality of portions of the field-of-view of the one or more cameras, where the preview includes a visual indication of an active portion of the field-of-view, enhances the video communication session experience by providing feedback to a user as to which portion of the field-of-view is active, which provides improved visual feedback.

[0323] In some embodiments, the computer system detects a user input corresponding to selection of the surface-view option (e.g., 612c, 612d, 614, 612g, 612i, and / or 612j). In response to detecting the user input corresponding to selection of the surface-view option, the computer system displays a preview (e.g., 674-2 and / or 674-3) of image data for the field-of-view of the one or more cameras (e.g., as described in FIGS. 6I-6J) (e.g., in a preview mode of the live video communication interface), the preview including a plurality of selectable visually distinguished portions overlaid on a representation of the field-of-view of the one or more cameras (e.g., as described in FIGS. 6I-6J). Displaying a preview including a plurality of selectable visually distinguished portions overlaid on a representation of the field-of-view of the one or more cameras, enhances the video communication session experience by providing feedback to a user as to which portions of the field-of-view are selectable for display as a representation during the video communication session, which provides improved visual feedback.

[0324] In some embodiments, the surface is a vertical surface (e.g., as depicted in FIG. 6J) (e.g., wall, easel, and / or whiteboard) in the scene (e.g., the surface is within a predetermined angle (e.g., 5 degrees, 10 degrees, or 20 degrees) of the direction of gravity). Displaying a representation of a vertical surface that includes an image of the vertical surface that is modified based on a position of the vertical surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the vertical surface despite its position relative to the camera without requiring further input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0325] In some embodiments, the surface is a horizontal surface (e.g., 619) (e.g., table, floor, and / or desk) in the scene (e.g., the surface is within a predetermined angle (e.g., 5 degrees, 10 degrees, or 20 degrees of a plane that is perpendicular to the direction of gravity). Displaying a representation of a horizontal surface that includes an image of the horizontal surface that is modified based on a position of the horizontal surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the horizontal surface despite its position relative to the camera without requiring further input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0326] In some embodiments, displaying the representation of the surface includes displaying a first view of the surface (e.g., 624-1 in FIG. 6N) (e.g., at a first angle of rotation and / or a first zoom level). In some embodiments, while displaying the first view of the surface, the computer system displays one or more shift-view options (e.g., 648-1 and / or 648-2) (e.g., buttons, icons, affordances, and / or user-interactive graphical user interface objects). The computer system detects a user input (e.g., 650a and / or 650b) directed to a respective shift-view option of the one or more shift-view options. In response to detecting the user input directed to the respective shift-view option, the computer system displays a second view of the surface (e.g., 624-1 in FIG. 6P and / or 624-1 in FIG. 6Q) (e.g., a second angle of rotation that is different from the first angle of rotation and / or a second zoom level that is different than the first zoom level) that is different from the first view of the surface (e.g., shifting the view of the surface from the first view to the second view). Providing a shift-view option to display the second view of the surface that is currently being displayed at the first view of the surface enhances the video communication session experience by allowing a user to view content associated with the surface at a different perspective, which provides additional control options without cluttering the user interface.

[0327] In some embodiments, displaying the first view of the surface includes displaying an image of the surface that is modified in a first manner (e.g., as depicted in FIG. 6N) (e.g., with a first distortion correction applied), and wherein displaying the second view of the surface includes displaying an image of the surface that is modified in a second manner (e.g., as depicted in FIG. 6P and / or FIG. 6Q) (e.g., with a second distortion correction applied) that is different from the first manner (e.g., the computer system changes (e.g., shifts) the distortion correction applied to the image of the surface based on the view (e.g., orientation and / or zoom) of the surface that is to be displayed). Displaying an image of the surface that is modified in a first manner and displaying the second view of the surface includes displaying an image of the surface that is modified in a second manner enhances the video communication session experience by allowing a user to automatically view content that is modified without requiring further input from the user, which provides improved visual feedback and reduces the number of inputs needed to perform an operation.

[0328] In some embodiments, the representation of the surface is displayed at a first zoom level (e.g., as depicted in FIG. 6N). In some embodiments, while displaying the representation of the surface at the first zoom level, the computer system detects a user input (e.g., 650b and / or 654) corresponding to a request to change a zoom level of the representation of the surface (e.g., selection of a zoom option (e.g., button, icon, affordance, and / or user-interactive user interface element). In response detecting the user input corresponding to a request to change a zoom level of the representation of the surface, the computer system displays the representation of the surface at a second zoom level that is different from the first zoom level (e.g., as depicted in FIG. 6Q and / or FIG. 6R) (e.g., zooming in or zooming out). Displaying the representation of the surface at a second zoom level that is different from the first zoom level when user input corresponding to a request to change a zoom level of the representation of the surface is detected enhances the video communication session experience by allowing a user to view content associated with the surface at a different level of granularity without further input, which provides improved visual feedback and additional control options without cluttering the user interface.

[0329] In some embodiments, while displaying the live video communication interface, the computer system displays (e.g., in a user interface (e.g., a menu, a dock region, a home screen, and / or a control center) that includes a plurality of selectable control options that, when selected, perform a function and / or set a parameter of the computer system, in the representation of at least a portion of the field-of-view of the one or more cameras, and / or in the live video communication interface) a selectable control option (e.g., 610, 6126, and / or 6136-1) (e.g., a button, icon, affordance, and / or user-interactive graphical user interface object) that, when selected, causes the representation of the surface to be displayed. In some embodiments, the one or more inputs include a user input corresponding to selection of the control option (e.g., 612c and / or 612g). In some embodiments, the computer system displays (e.g., in the live video communication interface and / or in a user interface of a different application) a second control option that, when selected, causes a representation of a user to be displayed in the live video communication session and causes the representation of the surface to cease being displayed. Displaying the control option that, when selected, displays the representation of the surface enhances the video communication session experience by allowing a user to modify what content is displayed, which provides additional control options without cluttering the user interface.

[0330] In some embodiments, the live video communication session is provided by a first application (e.g., 6110) (e.g., a video conferencing application and / or an application for providing an incoming and / or outgoing live audio / video communication session) operating at the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2). In some embodiments, the selectable control option (e.g., 610, 6126, 6136-1, and / or 6136-3) is associated with a second application (e.g., 6108) (e.g., a camera application and / or a presentation application) that is different from the first application.

[0331] In some embodiments, in response to detecting the one or more inputs, wherein the one or more inputs include the user input (e.g., 6128 and / or 6138) corresponding to selection of the control option (e.g., 6126 and / or 6136-3), the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays a user interface (e.g., 6140) of the second application (e.g., 6108) (e.g., a first user interface of the second application). Displaying a user interface of the second application in response to detecting the one or more inputs, wherein the one or more inputs include the user input corresponding to selection of the control option, provides access to the second application without having to navigate various menu options, which reduces the number of inputs needed to perform an operation. In some embodiments, displaying the user interface of the second application includes launching, activating, opening, and / or bringing to the foreground the second application. In some embodiments, displaying the user interface of the second application includes displaying the representation of the surface using the second application.

[0332] In some embodiments, prior to displaying the live video communication interface (e.g., 6121 and / or 6131) for the live video communication session (e.g., and before the first application (e.g., 6110) is launched), the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays a user interface (e.g., 6114 and / or 6116) of the second application (e.g., 6108) (e.g., a second user interface of the second application). Displaying a user interface of the second application prior to displaying the live video communication interface for the live video communication session, provides access to the second application without having to access the live video communication interface, which provides additional control options without cluttering the user interface. In some embodiments, the second application is launched before the first application is launched. In some embodiments, the first application is launched before the second application is launched.

[0333] In some embodiments, the live video communication session (e.g., 6120, 6121, 6130, and / or 6131) is provided using a third application (e.g., 6110) (e.g., a video conferencing application) operating at the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2). In some embodiments, the representation of the surface (e.g., 6116 and / or 6140) is provided by (e.g., displayed using a user interface of) a fourth application (e.g., 6108) that is different from the third application.

[0334] In some embodiments, the representation of the surface (e.g., 6116 and / or 6140) is displayed using a user interface (e.g., 6114) of the fourth application (e.g., 6108) (e.g., an application window of the fourth application) that is displayed in the live video communication session (e.g., 6120 and / or 6121) (e.g., the application window of the fourth application is displayed with the live video communication interface that is being displayed using the third application (e.g., 6110)). Displaying the representation of the surface using a user interface of the fourth application that is displayed in the live video communication session provides access to the fourth application, which provides additional control options without cluttering the user interface. In some embodiments, the user interface of the fourth application (e.g., the application window of the fourth application) is separate and distinct from the live video communication interface.

[0335] In some embodiments, the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays, via the display generation component (e.g., 601, 683, and / or 6101) a graphical element (e.g., 6108, 6108-1, 6126, and / or 6136-1) corresponding to the fourth application (e.g., a camera application associated with camera application icon 6108) (e.g., a selectable icon, button, affordance, and / or user-interactive graphical user interface object that, when selected, launches, opens, and / or brings to the foreground the fourth application), including displaying the graphical element in a region (e.g., 6104 and / or 6106) that includes (e.g., is configurable to display) a set of one or more graphical elements (e.g., 6110-1) corresponding to an application other than the fourth application (e.g., a set of application icons each corresponding to different applications). Displaying a graphical element corresponding to the fourth application in a region that includes a set of one or more graphical elements corresponding to an application other than the fourth application, provides controls for accessing the fourth application without having to navigate various menu options, which provides additional control options without cluttering the user interface. In some embodiments, the graphical element corresponding to the fourth application is displayed in, added to, and / or displayed adjacent to an application dock (e.g., 6104 and / or 6106) (e.g., a region of a display that includes a plurality of application icons for launching respective applications). In some embodiments, the set of one or more graphical elements includes a graphical element (e.g., 6110-1) that corresponds to the third application (e.g., video conferencing application associated with video conferencing application icon 6110) that provides the live video communication session. In some embodiments, in response to detecting the one or more user inputs (e.g., 6112 and / or 6118) (e.g., including an input on the graphical element corresponding to the fourth application), the computer system displays an animation of the graphical element corresponding to the fourth application, e.g., bouncing in the application dock.

[0336] In some embodiments, displaying the representation of the surface includes displaying, via the display generation component, an animation of a transition (e.g., a transition that gradually progresses through a plurality of intermediate states over time inclu...

Claims

1. A method, comprising: at a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field-of-view of the one or more cameras; while displaying the live video communication interface, detecting, via the one or more input devices, one or more user inputs including a user input corresponding to a request to display a view of a surface in a scene that is in a field-of-view of a first camera of the one or more cameras; in response to detecting the one or more user inputs, and while the first camera captures image data of the scene that is in the field-of-view of the first camera, concurrently displaying, via the display generation component: a representation of the surface, wherein the representation of the surface includes an image of the surface that is based on the image data being captured by the first camera of the one or more cameras, wherein the image of the surface is modified based on a position of the surface relative to the first camera of the one or more cameras; and a representation of at least a portion of the scene that is based on the image data being captured by the first camera of the one or more cameras that is not modified based on the position of the surface relative to the first camera of the one or more cameras; and wherein: the image of the surface is modified by rotating the image of the surface relative to the representation of at least a portion of the scene that is based on the image data being captured by the first camera of the one or more cameras that is not modified based on the position of the surface relative to the one or more cameras; and the image of the surface is rotated based on a position of the surface relative to a user in the field-of-view of the first camera of the one or more cameras.

2. The method of claim 1, further comprising: in response to detecting the one or more user inputs: prior to displaying the representation of the surface, displaying a preview of the image data being captured by the first camera of the one or more cameras, the preview including an image of the surface that is not modified based on the position of the surface relative to the one or more cameras.

3. The method of claim 2, wherein displaying the preview of the image data being captured by the first camera of the one or more cameras includes displaying a plurality of selectable options corresponding to respective portions of the field-of-view of the first camera of the one or more cameras, the method further comprising: detecting an input selecting one of the plurality of selectable options corresponding to respective portions of the field-of-view of the first camera of the one or more cameras; and in response to detecting the input selecting one of the plurality of selectable options corresponding to respective portions of the field-of-view of the first camera of the one or more cameras: in accordance with a determination that the input selecting one of the plurality of selectable options corresponding to respective portions of the field-of-view of the first camera of the one or more cameras is directed to a first option corresponding to a first portion of the field-of-view of the first camera of the one or more cameras, displaying the representation of the surface based on the first portion of the field-of-view of the first camera of the one or more cameras; and in accordance with a determination that the input selecting one of the plurality of selectable options corresponding to respective portions of the field-of-view of the first camera of the one or more cameras is directed to a second option corresponding to a second portion of the field-of-view of the first camera of the one or more cameras, displaying the representation of the surface based on the second portion of the field-of-view of the first camera of the one or more cameras, wherein the second option is different from the first option.

4. The method of any one of claims 1-3, wherein the one or more user inputs include a gesture in the field-of-view of the one or more cameras.

5. The method of any one of claims 1-4, wherein displaying the representation of the surface includes displaying a first view of the surface, the method further comprising: while displaying the first view of the surface, displaying one or more shift-view options; detecting a user input directed to a respective shift-view option of the one or more shift-view options; and in response to detecting the user input directed to the respective shift-view option, displaying a second view of the surface that is different from the first view of the surface.

6. The method of any one of claims 1-5, wherein the representation of the surface is displayed at a first zoom level, the method further comprising: while displaying the representation of the surface at the first zoom level, detecting a user input corresponding to a request to change a zoom level of the representation of the surface; and in response detecting the user input corresponding to a request to change a zoom level of the representation of the surface, displaying the representation of the surface at a second zoom level that is different from the first zoom level.

7. The method of any one of claims 1-6, further comprising: while displaying the live video communication interface, displaying a selectable control option that, when selected, causes the representation of the surface to be displayed, wherein the one or more user inputs include a user input corresponding to selection of the selectable control option.

8. The method of claim 7, wherein: the live video communication session is provided by a first application operating at the computer system; and the selectable control option is associated with a second application that is different from the first application.

9. The method of claim 8, further comprising: in response to detecting the one or more user inputs, wherein the one or more user inputs include the user input corresponding to selection of the selectable control option, displaying a user interface of the second application.

10. The method of any one of claims 1-9, wherein: the live video communication session is provided using a third application operating at the computer system; and the representation of the surface is provided by a fourth application that is different from the third application.

11. The method of any one of claims 1-10, wherein the computer system is in communication with a second computer system that is in communication with a second display generation component, the method further comprising: displaying the representation of at least a portion of the scene on the display generation component; and causing display of the representation of the surface on the second display generation component.

12. The method of claim 11, further comprising: in response to detecting a change in an orientation of the second computer system, updating the display of the representation of the surface that is displayed at the second display generation component from displaying a first view of the surface to displaying a second view of the surface that is different from the first view.

13. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions which, when executed by the one or more processors, cause the computer system to perform the method of any one of claims 1-12.

14. A computer system that is configured to communicate with a display generation component, one or more cameras, and one or more input devices, the computer system comprising: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions which, when executed by the one or more processors, cause the computer system to perform the method of any one of claims 1-12.