Wide angle video conference
The described methods and interfaces enhance the efficiency of live video communication sessions by detecting user inputs and modifying surface representations, addressing inefficiencies in existing systems and conserving energy.
Patent Information
- Application Number
- JP2025045614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-22
- Filing Date
- 2025-03-19
- Publication Date
- 2025-08-20
AI Technical Summary
Existing techniques for managing live video communication sessions are cumbersome and inefficient, often requiring multiple key presses or strokes, wasting time and energy, particularly in battery-operated devices.
Implementing methods and interfaces that allow for faster and more efficient management of live video communication sessions by detecting user inputs directed at a surface within the camera's field of view and displaying a modified representation of the surface based on its position relative to the cameras, reducing cognitive burden and conserving power.
The solution provides a more efficient human-machine interface, reducing the time required for user interactions and extending battery life in battery-operated devices.
Smart Images

Figure 2025121896000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. patent application Ser. No. 17 / 950,868 entitled "WIDE ANGLE VIDEO CONFERENCE," filed September 22, 2022, which claims priority to U.S. patent application Ser. No. 17 / 950,900 entitled "WIDE ANGLE VIDEO CONFERENCE," filed September 22, 2022, which claims priority to U.S. patent application Ser. No. 17 / 950,922 entitled "WIDE ANGLE VIDEO CONFERENCE," filed September 22, 2022, which claims priority to U.S. patent application Ser. No. 63 / 392,096 entitled "WIDE ANGLE VIDEO CONFERENCE," filed July 25, 2022, and U.S. patent application Ser. No. 63 / 392,096 entitled "WIDE ANGLE VIDEO CONFERENCE," filed June 30, 2022. This application claims priority to U.S. patent application Ser. No. 63 / 357,605 entitled "WIDE ANGLE VIDEO CONFERENCE," filed June 5, 2022, which claims priority to U.S. patent application Ser. No. 63 / 349,134 entitled "WIDE ANGLE VIDEO CONFERENCE," filed February 8, 2022, which claims priority to U.S. patent application Ser. No. 63 / 307,780 entitled "WIDE ANGLE VIDEO CONFERENCE," filed September 24, 2021, which claims priority to U.S. patent application Ser. No. 63 / 248,137 entitled "WIDE ANGLE VIDEO CONFERENCE," filed September 24, 2021, the contents of each of which are incorporated herein by reference in their entirety.
[0002] TECHNICAL FIELD The present disclosure relates generally to computer user interfaces, and more particularly to techniques for managing live video communication sessions and / or managing digital content. [Background technology]
[0003] The computer system may include hardware and / or software for displaying an interface for a live video communication session. Summary of the Invention
[0004] However, some techniques for managing live video communication sessions using electronic devices are generally cumbersome and inefficient. For example, some existing techniques use complex and time-consuming user interfaces that may involve multiple key presses or strokes. Existing techniques take longer than necessary, wasting the user's time and the device's energy. This latter consideration is particularly important in battery-operated devices.
[0005] The present technology thus provides electronic devices with faster, more efficient methods and interfaces for managing live video communication sessions and / or managing digital content. Such methods and interfaces optionally complement or replace other methods for managing live video communication sessions and / or managing digital content. Such methods and interfaces reduce the cognitive burden on users and create a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power and extend the time between battery charges.
[0006] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The method includes: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; detecting, while displaying the live video communication interface, one or more user inputs via the one or more input devices, including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and displaying, via the display generation component, a representation of the surface, the representation of the surface including an image of the surface captured by the one or more cameras modified based on the position of the surface relative to the one or more cameras.
[0007] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions to: display, via the display generation component, a live video communication interface for a live video communication session including a representation of at least a portion of a field of view of the one or more cameras; detect, while displaying the live video communication interface, via the one or more input devices, one or more user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and, in response to detecting the one or more user inputs, display, via the display generation component, a representation of the surface, the representation of the surface including an image of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
[0008] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions to: display, via the display generation component, a live video communication interface for a live video communication session including a representation of at least a portion of a field of view of the one or more cameras; detect, while displaying the live video communication interface, via the one or more input devices, one or more user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and, in response to detecting the one or more user inputs, display, via the display generation component, a representation of the surface, the representation of the surface including an image of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
[0009] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: display, via the display generation component, a live video communication interface for a live video communication session including a representation of at least a portion of a field of view of the one or more cameras; detect, while displaying the live video communication interface, one or more user inputs via the one or more input devices, including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and, in response to detecting the one or more user inputs, display, via the display generation component, a representation of the surface, the representation of the surface including an image of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
[0010] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system comprises: means for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; means for acquiring, via the one or more cameras, image data for the field of view of the one or more cameras while displaying the live video communication interface, the image data including a first gesture; means for displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the representation of the second portion of the scene including different visual content from the representation of the first portion of the scene, in response to acquiring the image data for the field of view of the one or more cameras and in accordance with a determination that the first gesture satisfies a first set of criteria; and means for continuing to display, via the display generation component, the representation of the first portion of the scene in accordance with a determination that the first gesture satisfies a second set of criteria different from the first set of criteria.
[0011] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; acquiring, while displaying the live video communication interface, via the one or more cameras, image data for the field of view of the one or more cameras, the image data including a first gesture; displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the representation of the second portion of the scene including visual content different from the representation of the first portion of the scene, in response to acquiring the image data for the field of view of the one or more cameras and in accordance with a determination that the first gesture satisfies a first set of criteria; and continuing to display, via the display generation component, the representation of the first portion of the scene in accordance with a determination that the first gesture satisfies a second set of criteria different from the first set of criteria.
[0012] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The method includes detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0013] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface being displayed by the one or more first cameras and one or more input devices of the first computer system. a first representation of the field of view of a first camera of the first computer system, a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in the first scene within the field of view of the one or more first cameras of the first computer system, a first representation of the field of view of one or more second cameras of the second computer system, and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in the second scene within the field of view of the one or more second cameras of the second computer system.
[0014] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface being displayed by the one or more first cameras and one or more input devices of the first computer system. a first representation of the field of view of a first camera of the first computer system, a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in the first scene within the field of view of the one or more first cameras of the first computer system, a first representation of the field of view of one or more second cameras of the second computer system, and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in the second scene within the field of view of the one or more second cameras of the second computer system.
[0015] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of the one or more first cameras of the first computer system; The computer system includes a second representation of a field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of the second computer system; and a second representation of a field of view of one or more second cameras of the second computer system, the second representation of a field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0016] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes means for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; and means for displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0017] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more first cameras, and one or more input devices. The one or more programs include instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0018] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The method includes detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0019] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface being displayed by the one or more first cameras and one or more input devices of the first computer system. a first representation of the field of view of a first camera of the first computer system, a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in the first scene within the field of view of the one or more first cameras of the first computer system, a first representation of the field of view of one or more second cameras of the second computer system, and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in the second scene within the field of view of the one or more second cameras of the second computer system.
[0020] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface being displayed by the one or more first cameras and one or more input devices of the first computer system. a first representation of the field of view of a first camera of the first computer system, a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in the first scene within the field of view of the one or more first cameras of the first computer system, a first representation of the field of view of one or more second cameras of the second computer system, and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in the second scene within the field of view of the one or more second cameras of the second computer system.
[0021] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of the one or more first cameras of the first computer system; The computer system includes a second representation of a field of view of one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of the second computer system; and a second representation of a field of view of one or more second cameras of the second computer system, the second representation of a field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0022] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes means for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; and means for displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0023] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more first cameras, and one or more input devices. The one or more programs include instructions for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants, and displaying, via a display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface including: a first representation of a field of view of one or more first cameras of a first computer system; a second representation of the field of view of one or more first cameras of the first computer system, the second representation of the field of view of one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of one or more second cameras of the second computer system, the second representation of the field of view of one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0024] According to some embodiments, a method is described that includes, at a first computer system in communication with a first display generating component and one or more sensors, displaying, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system while the first computer system is in a live video communication session with a second computer system, detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment, and, in response to detecting the change in position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment within a field of view of the one or more cameras of the second computer system, the second view of the physical environment within a field of view of the one or more cameras of the second computer system, the second view of the physical environment within a field of view of the one or more cameras of the second computer system.
[0025] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more sensors, the one or more programs including instructions to: display, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system while the first computer system is in a live video communication session with a second computer system; detect, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; and, in response to detecting the change in position of the first computer system, display, via the first display generating component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, that differs from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0026] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more sensors, the one or more programs including instructions to: display, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system while the first computer system is in a live video communication session with a second computer system; detect, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; and, in response to detecting the change in position of the first computer system, display, via the first display generating component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, that differs from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0027] According to some embodiments, a computer system configured to communicate with a first display generating component and one or more sensors is described, the computer system comprising one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: while the first computer system is engaged in a live video communication session with a second computer system, display, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detect, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; and, in response to detecting the change in position of the first computer system, display, via the first display generating component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the representation different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0028] According to some embodiments, a computer system configured to communicate with a first display generating component and one or more sensors is described, the computer system including means for displaying, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system while the first computer system is in a live video communication session with a second computer system, detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment, and, in response to detecting the change in position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the representation differing from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0029] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more sensors, the one or more programs including instructions for: displaying, via the first display generating component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system while the first computer system is engaged in a live video communication session with a second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; and, in response to detecting the change in position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the representation different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0030] According to some embodiments, a method is described that includes, in a computer system in communication with a display generation component, displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical mark and a physical background, where displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background within the field of view of the one or more cameras, acquiring data including a new physical mark in the physical environment while displaying the representation of the physical mark without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, and in response to acquiring the data representing the new physical mark in the physical environment, displaying the representation of the new physical mark without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras.
[0031] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for displaying, via the display generation component, representations of physical marks in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical marks and a physical background, displaying the representations of the physical marks includes displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, acquiring data including a new physical mark in the physical environment while displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, and in response to acquiring data representing the new physical mark in the physical environment, displaying the representations of the new physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras.
[0032] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, representations of physical marks in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical marks and a physical background, displaying the representations of the physical marks based on the view of the physical environment without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, acquiring data including a new physical mark in the physical environment while displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, and in response to acquiring the data representing the new physical mark in the physical environment, displaying the representations of the new physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras.
[0033] According to some embodiments, a computer system configured to communicate with a display generation component is described, the computer system comprising one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to, via the display generation component, display representations of physical marks in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical marks and a physical background, displaying the representations of the physical marks includes displaying the representations of the physical marks in the physical environment without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, acquire data including a new physical mark in the physical environment while displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, and in response to acquiring data representing the new physical mark in the physical environment, display the representations of the new physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras.
[0034] According to some embodiments, a computer system configured to communicate with a display generation component is described, the computer system including: means for displaying, via the display generation component, a representation of a physical mark in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical mark and a physical background, where displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background within the field of view of the one or more cameras; means for acquiring data including a new physical mark in the physical environment while displaying the representation of the physical mark without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras; and means for displaying the representation of the new physical mark without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras in response to acquiring the data representing the new physical mark in the physical environment.
[0035] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for, via the display generation component, displaying representations of physical marks in a physical environment based on a view of the physical environment within a field of view of one or more cameras, the view of the physical environment including the physical marks and a physical background, displaying the representations of the physical marks includes displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, acquiring data including a new physical mark in the physical environment while displaying the representations of the physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras, and in response to acquiring the data representing the new physical mark in the physical environment, displaying the representations of the new physical marks without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras.
[0036] According to some embodiments, a method is described that includes, in a computer system in communication with a display generation component and one or more cameras, displaying an electronic document via the display generation component, detecting, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, and in response to detecting handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, displaying digital text within the electronic document that corresponds to the handwriting within the field of view of the one or more cameras.
[0037] According to some embodiments, a non-transitory computer-readable storage medium is described. The exemplary non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions to: display, via the display generation component, an electronic document; detect, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system; and, in response to detecting handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, display digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
[0038] According to some embodiments, a temporary computer-readable storage medium is described. The exemplary temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions to: display, via the display generation component, an electronic document; detect, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system; and, in response to detecting handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, display digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
[0039] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described, the computer system comprising one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: display, via the display generation component, an electronic document; detect, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system; and, in response to detecting, within the field of view of the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, display digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
[0040] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described, the computer system including: means for displaying an electronic document via the display generation component; means for detecting, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system; and means for displaying digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras in response to detecting handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system.
[0041] According to some embodiments, a computer program product is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for: displaying, via the display generation component, an electronic document; detecting, via the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system; and, in response to detecting, within the field of view of the one or more cameras, handwriting including physical marks on a physical surface within the field of view of the one or more cameras and separate from the computer system, displaying digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
[0042] According to some embodiments, a method is described that is executed on a first computer system in communication with a display generation component, one or more cameras, and one or more input devices, the method including: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application to display a visual representation of a surface within a field of view of the one or more cameras; and, in response to detecting the one or more first user inputs and following a determination that a first set of one or more criteria is satisfied, simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication that a first region of the field of view of the one or more cameras is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0043] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions to: detect, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application to display a visual representation of a surface within the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs and following a determination that one or more first criteria are satisfied, simultaneously display, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication indicative of a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicative of a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0044] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a first computer system configured in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions to: detect, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application to display a visual representation of a surface within the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs and following a determination that one or more first criteria are satisfied, simultaneously display, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication indicative of a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicative of a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0045] According to some embodiments, a first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: detect, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs and following a determination that one or more first criteria are satisfied, simultaneously display, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication indicative of a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicative of a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0046] According to some embodiments, a first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described, the computer system including: means for detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; and means for simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication indicative of a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicative of a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0047] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs and following a determination that one or more first criteria are satisfied, simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras and a visual indication indicative of a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicative of a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
[0048] According to some embodiments, a method is described that includes, in a computer system in communication with a display generation component and one or more input devices, detecting, via the one or more input devices, a request to use a function on the computer system, and, in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function including a virtual demonstration of the function, wherein displaying the tutorial includes displaying the virtual demonstration having a first appearance in accordance with a determination that a property of the computer system has a first value, and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
[0049] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, the one or more programs including instructions for detecting, via the one or more input devices, a request to use a function on the computer system, and displaying, via the display generating component, a tutorial for using the function including a virtual demonstration of the function in response to detecting the request to use the function on the computer system, the tutorial including a virtual demonstration of the function, in accordance with determining that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance, and in accordance with determining that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance different from the first appearance.
[0050] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more input devices, the one or more programs including instructions for detecting, via the one or more input devices, a request to use a function on the computer system, and, in response to detecting the request to use the function on the computer system, displaying, via the display generating component, a tutorial for using the function including a virtual demonstration of the function, including displaying the virtual demonstration having a first appearance in accordance with a determination that a property of the computer system has a first value, and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
[0051] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for detecting, via the one or more input devices, a request to use a function on the computer system; and displaying, via the display generation component, a tutorial for using the function including a virtual demonstration of the function in response to detecting the request to use the function on the computer system, the tutorial including a virtual demonstration of the function, the tutorial including instructions for displaying the virtual demonstration having a first appearance in accordance with determining that a property of the computer system has a first value; and the virtual demonstration having a second appearance different from the first appearance in accordance with determining that the property of the computer system has a second value.
[0052] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described, the computer system comprising: means for detecting, via the one or more input devices, a request to use a feature on the computer system; and means for displaying, via the display generation component, a tutorial for using the feature including a virtual demonstration of the feature in response to detecting the request to use the feature on the computer system, the tutorial including a virtual demonstration of the feature, the means including: means for displaying the virtual demonstration having a first appearance in accordance with a determination that a property of the computer system has a first value; and means for displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
[0053] According to some embodiments, a computer program product is described, the computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs including instructions for detecting, via the one or more input devices, a request to use a function on the computer system, and, in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function including a virtual demonstration of the function, the tutorial including a virtual demonstration of the function, in accordance with determining that a property of the computer system has a first value, displaying the virtual demonstration having a first appearance, and in accordance with determining that the property of the computer system has a second value, displaying the virtual demonstration having a second appearance different from the first appearance.
[0054] Executable instructions to perform these functions are optionally contained in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions to perform these functions are optionally contained in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0055] This provides devices with faster, more efficient methods and interfaces for managing live video communication sessions, thereby increasing the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces may complement or replace other methods for managing live video communication sessions.
[0056] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout: [Brief explanation of the drawings]
[0057] [Figure 1A] FIG. 1 is a block diagram illustrating a portable multifunction device with a touch-sensitive display in accordance with some embodiments.
[0058] [Figure 1B] FIG. 2 is a block diagram illustrating exemplary components for event processing according to some embodiments.
[0059] [Figure 2] FIG. 1 illustrates a portable multifunction device with a touch screen in accordance with some embodiments.
[0060] [Figure 3] FIG. 1 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface in accordance with some embodiments.
[0061] [Figure 4A] 1 illustrates an exemplary user interface for a menu of applications on a portable multifunction device in accordance with some embodiments.
[0062] [Figure 4B] 1 illustrates an exemplary user interface for a multifunction device having a touch-sensitive surface that is separate from the display in accordance with some embodiments.
[0063] [Figure 5A] 1 illustrates a personal electronic device according to some embodiments.
[0064] [Figure 5B] FIG. 1 is a block diagram illustrating a personal electronic device according to some embodiments.
[0065] [Figure 5C] 1 illustrates an exemplary diagram of a communication session between electronic devices according to some embodiments.
[0066] [Figure 6A] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6B] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6C] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6D] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6E] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6F] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6G] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6H] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6I] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6J] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6K] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6L] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6M]1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6N] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6O] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6P] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6Q] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6R] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6S] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6T] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6U] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6V] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6W] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6X] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6Y] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6Z] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AA] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AB] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AC] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AD] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AE] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AF] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AG] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AH] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AI] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AJ] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AK] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AL] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AM] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AN] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AO] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AP] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AQ] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AR] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AS] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AT] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AU] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AV] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AW] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AX] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 6AY]1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments.
[0067] [Figure 7] 1 is a flow diagram illustrating a method for managing a live video communication session according to some embodiments.
[0068] [Figure 8] 1 is a flow diagram illustrating a method for managing a live video communication session according to some embodiments.
[0069] [Figure 9A] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9B] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9C] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9D] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9E] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9F] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9G] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9H] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9I]1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9J] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9K] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9L] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9M] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9N] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9O] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9P] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9Q] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9R] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9S] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 9T] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments.
[0070] [Figure 10]1 is a flow diagram illustrating a method for managing a live video communication session according to some embodiments.
[0071] [Figure 11A] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11B] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11C] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11D] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11E] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11F] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11G] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11H] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11I] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11J] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11K] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11L] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11M]1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11N] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11O] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 11P] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments.
[0072] [Figure 12] FIG. 1 is a flow diagram illustrating a method for managing digital content, according to some embodiments.
[0073] [Figure 13A] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13B] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13C] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13D] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13E] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13F] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13G] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13H]1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13I] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13J] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments. [Figure 13K] 1 illustrates an exemplary user interface for managing digital content, according to some embodiments.
[0074] [Figure 14] FIG. 1 is a flow diagram illustrating a method for managing digital content, according to some embodiments.
[0075] [Figure 15] 1 is a flow diagram illustrating a method for managing a live video communication session according to some embodiments.
[0076] [Figure 16A] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16B] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16C] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16D] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16E] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16F] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16G] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16H] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16I] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16J] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16K] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16L] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16M] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16N] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16O] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16P] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments. [Figure 16Q] 1 illustrates an exemplary user interface for managing a live video communication session according to some embodiments.
[0077] [Figure 17] 1 is a flow diagram illustrating a method for managing a live video communication session according to some embodiments.
[0078] [Figure 18A] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18B] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18C] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18D] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18E] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18F] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18G] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18H] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18I] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18J] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18K] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18L] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18M] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments. [Figure 18N] 1 illustrates an exemplary user interface for displaying a tutorial for a feature on a computer system, according to some embodiments.
[0079] [Figure 19] FIG. 1 is a flow diagram illustrating a method for displaying a tutorial for a feature on a computer system, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0080] The following description sets forth example methods, parameters, etc. However, it should be recognized that such description is not intended as a limitation on the scope of the present disclosure, but rather is provided as a description of example embodiments.
[0081] There is a need for electronic devices that provide efficient methods and interfaces for managing live video communication sessions and / or digital content. For example, there is a need for electronic devices that improve content sharing. Such techniques can reduce the cognitive burden on users sharing content during live video communication sessions and / or managing digital content in electronic documents, thereby increasing productivity. Furthermore, such techniques can reduce processor and battery power that would otherwise be wasted on redundant user input.
[0082] The following Figures 1A and 1B, 2, 3, 4A and 4B, and 5A-5C provide descriptions of example devices for managing live video communication sessions and / or implementing techniques for managing digital content. Figures 6A-6AY illustrate example user interfaces for managing live video communication sessions. Figures 7 and 8, and 15 are flow diagrams illustrating methods for managing live video communication sessions, according to some embodiments. The user interfaces in Figures 6A-6AY are used to explain processes described below, including the processes in Figures 7 and 8, and 15. Figures 9A-9T illustrate example user interfaces for managing live video communication. Figure 10 is a flow diagram illustrating a method for managing live video communication, according to some embodiments. The user interfaces in Figures 9A-9T are used to explain processes described below, including the process in Figure 10. Figures 11A-11P illustrate example user interfaces for managing digital content. Figure 12 is a flow diagram illustrating a method for managing digital content, according to some embodiments. The user interfaces in FIGS. 11A-11P are used to illustrate processes described below, including the process in FIG. 12. FIGS. 13A-13K illustrate exemplary user interfaces for managing digital content, according to some embodiments. FIG. 14 is a flow diagram illustrating a method for managing digital content, according to some embodiments. The user interfaces in FIGS. 13A-13K are used to explain processes described below, including the process in FIG. 14. FIGS. 16A-16O illustrate exemplary user interfaces for managing a live video communication session, according to some embodiments. FIG. 17 is a flow diagram illustrating a method for managing a live video communication session, according to some embodiments. The user interfaces in FIGS. 16A-16Q are used to explain processes described below, including the process in FIG. 17. FIGS. 18A-18N illustrate exemplary user interfaces for displaying tutorials for features on a computer system, according to some embodiments.19 is a flow diagram illustrating a method for displaying a tutorial for a feature on a computer system, according to some embodiments. The user interfaces in FIGS. 18A-18N are used to illustrate processes described below, including the process in FIG. 19.
[0083] The processes described below improve device usability and streamline the user-device interface (e.g., by assisting users in providing appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an action, providing additional control options without cluttering the user interface with additional displayed controls, performing an action without requiring further user input when a set of conditions is met, improving efficiency in managing digital content, improving collaboration between users during a live communication session, improving the live communication session experience, and / or additional techniques. These techniques also reduce power usage and improve device battery life by allowing users to use the device more quickly and efficiently.
[0084] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions for performing a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0085] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by these terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch can be referred to as a second touch, and similarly, a second touch can be referred to as a first touch, without departing from the scope of various described embodiments. In some embodiments, a first touch and a second touch are two separate references to the same touch. Although a first touch and a second touch are both touches, they are not the same touch.
[0086] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. In the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0087] The term "if" is interpreted, optionally, according to the context, to mean "when" or "upon," or "in response to determining" or "in response to detecting." Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" are interpreted, optionally, according to the context, to mean "upon determining" or "in response to determining," or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."
[0088] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Optionally, other portable electronic devices, such as a laptop computer or tablet computer having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad), are also used. It should also be understood that in some embodiments, the device is not a portable communication device, but rather a desktop computer having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad). In some embodiments, the electronic device is a computer system in communication (e.g., via wired communication, via wireless communication) with a display generation component. The display generation component is configured to provide a visual output, such as a display via a CRT display, a display via an LED display, or a display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes displaying content (e.g., video data rendered or decoded by display controller 156) by transmitting data (e.g., image data or video data) over a wired or wireless connection to an integrated or external display generation component to visually generate the content.
[0089] In the following discussion, electronic devices are described that include a display and a touch-sensitive surface. However, it should be understood that the electronic device optionally includes one or more other physical user-interface devices, such as a physical keyboard, a mouse, and / or a joystick.
[0090] The device typically supports a variety of applications such as one or more of a drawing application, a presentation application, a word processing application, a website creation application, a disc authoring application, a spreadsheet application, a gaming application, a telephone application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0091] Various applications running on the device optionally use at least one common physical user-interface device, such as a touch-sensitive surface. One or more features of the touch-sensitive surface and corresponding information displayed on the device are optionally adjusted and / or changed for each application and / or within individual applications. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally supports various applications with user interfaces that are intuitive and transparent to the user.
[0092] Attention now turns to embodiments of portable devices with touch-sensitive displays. FIG. 1A is a block diagram illustrating portable multifunction device 100 having touch-sensitive display system 112, according to some embodiments. Touch-sensitive display 112 may conveniently be referred to as a "touch screen" and may also be known or referred to as a "touch-sensitive display system." Device 100 includes memory 102 (optionally including one or more computer-readable storage media), memory controller 122, one or more processing units (CPUs) 120, peripherals interface 118, RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, input / output (I / O) subsystem 106, other input control devices 116, and external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 that detect the intensity of a contact on device 100 (e.g., a touch-sensitive surface, such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 that generate tactile output on device 100 (e.g., generate tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0093] As used herein and in the claims, the term “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or a proxy for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four distinct values and more typically includes hundreds (e.g., at least 256) distinct values. The intensity of a contact is optionally determined (or measured) using various techniques and various sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in the contact area detected on the touch-sensitive surface, the capacitance and / or change in the capacitance of the touch-sensitive surface proximate the contact, and / or the resistance and / or change in the capacitance of the touch-sensitive surface proximate the contact are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the surrogate measure of the force or pressure of the contact is used directly to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measure). In some implementations, the surrogate measure of the contact force or pressure is converted to an estimate of the force or pressure, and the estimate of the force or pressure is used to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using contact intensity as an attribute of user input allows users to access additional device functionality (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as knobs or buttons) that may not otherwise be accessible to users on devices of reduced size that have limited footprint for displaying affordances.
[0094] As used herein and in the claims, the term “tactile output” refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, that will be detected by a user with the user's sense of touch. For example, in a situation where a device or a component of a device is in contact with a touch-sensitive surface of a user (e.g., the fingers, palm, or other part of the user's hand), the tactile output produced by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a “downclick” or “upclick” of a physical actuator button. In some cases, a user feels a tactile sensation such as a “downclick” or “upclick” even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's action. As another example, movement of a touch-sensitive surface is optionally interpreted or perceived by a user as "roughness" of the touch-sensitive surface, even if there is no change in the smoothness of the touch-sensitive surface. While such user interpretation of touch depends on the user's personal sensory perception, there are many sensory perceptions of touch that are common to the majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "upclick," "downclick," "roughness"), unless otherwise specified, the generated tactile output corresponds to a physical displacement of the device, or a component of the device, that produces the described sensory perception for a typical (or average) user.
[0095] It should be understood that device 100 is only one example of a portable multifunction device, and that device 100 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of its components. The various components shown in Figure 1A are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.
[0096] Memory 102 optionally includes high-speed random access memory, and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.
[0097] Peripheral interface 118 may be used to couple input and output peripherals of the device to CPU 120 and memory 102. One or more processors 120 operate or execute various software programs (e.g., computer programs (e.g., including instructions)) and / or instruction sets stored in memory 102 to perform various functions and process data for device 100. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0098] RF (radio frequency) circuitry 108 transmits and receives RF signals, also called electromagnetic signals. RF circuitry 108 converts electrical signals to electromagnetic signals or electromagnetic signals to electrical signals and communicates with communication networks and other communication devices via electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. RF circuitry 108 optionally communicates via wireless communication with networks, such as the Internet, also called the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and with other devices. RF circuitry 108 optionally includes well-known circuitry for detecting near field communication (NFC) fields, such as by short-range radios. Wireless communication optionally includes, but is not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), Long Term Evolution (LTE), and other standards.evolution (LTE), near field communications (NFC), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol), The present invention may use any of a number of communication standards, protocols, and technologies, including the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (XMPP), the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), the Instant Messaging and Presence Service (IMPS), and / or the Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this application.
[0099] Audio circuit 110, speaker 111, and microphone 113 provide an audio interface between a user and device 100. Audio circuit 110 receives audio data from peripherals interface 118, converts the audio data into electrical signals, and transmits the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves audible to humans. Audio circuit 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuit 110 converts the electrical signals into audio data and transmits the audio data to peripherals interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to memory 102 and / or RF circuit 108 by peripherals interface 118. In some embodiments, audio circuit 110 also includes a headset jack (e.g., 212 in FIG. 2 ). The headset jack provides an interface between audio circuitry 110 and a detachable audio input / output peripheral, such as an output-only headphone or a headset with both an output (e.g., mono or binaural headphones) and an input (e.g., a microphone).
[0100] I / O subsystem 106 couples input / output peripherals on device 100, such as touchscreen 112 and other input control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, depth camera controller 169, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. One or more input controllers 160 receive / send electrical signals from / to other input control devices 116. Other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some embodiments, input controller(s) 160 are optionally coupled to any (or none) of a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 208 in FIG. 2 ) optionally include up / down buttons for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include push buttons (e.g., 206 in FIG. 2 ). In some embodiments, the electronic device is a computer system in communication with one or more input devices (e.g., via wireless communication over wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a trackpad as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system.In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).
[0101] A quick press of a push button optionally unlocks the touchscreen 112 or optionally initiates the process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application Serial No. 11 / 322,549, filed December 23, 2005, "Unlocking a Device by Performing Gestures on an Unlock Image," U.S. Patent No. 7,657,849, which is incorporated herein by reference in its entirety. A longer press of a push button (e.g., 206) optionally turns power on or off to the device 100. The functionality of one or more of the buttons is optionally customizable by the user. The touchscreen 112 is used to implement virtual or soft buttons and one or more soft keyboards.
[0102] Touch-sensitive display 112 provides an input and output interface between the device and a user. Display controller 156 receives and / or sends electrical signals to touchscreen 112. Touchscreen 112 displays visual output to the user. This visual output optionally includes graphics, text, icons, animation, and any combination thereof (collectively "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0103] Touchscreen 112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touchscreen 112 and display controller 156 (along with any associated modules and / or instruction sets in memory 102) detects contacts (and any movement or cessation of contact) on touchscreen 112 and translates the detected contacts into interactions with user interface objects (e.g., one or more softkeys, icons, web pages, or images) displayed on touchscreen 112. In an exemplary embodiment, the point of contact between touchscreen 112 and the user corresponds to the user's finger.
[0104] Touchscreen 112 optionally uses LCD (liquid crystal display), LPD (light emitting polymer display), or LED (light emitting diode) technology, although other display technologies are used in other embodiments. Touchscreen 112 and display controller 156 optionally use any of a number of now known or later developed touch sensing technologies to detect contact and any movement or disruption thereof, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more points of contact with touchscreen 112. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California.
[0105] The touch-sensitive display in some embodiments of touchscreen 112 is optionally similar to the multi-touch-sensing touchpad described in U.S. Patent Nos. 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman), and / or U.S. Patent Application Publication No. 2002 / 0015024 A1, each of which is incorporated by reference herein in its entirety. However, touchscreen 112 displays visual output from device 100, whereas touch-sensitive touchpads do not provide visual output.
[0106] The touch-sensitive display in some embodiments of touch screen 112 is described in the following applications: (1) U.S. patent application Ser. No. 11 / 381,313, filed May 2, 2006, entitled "Multipoint Touch Surface Controller," (2) U.S. patent application Ser. No. 10 / 840,862, filed May 6, 2004, entitled "Multipoint Touchscreen," (3) U.S. patent application Ser. No. 10 / 903,964, filed July 30, 2004, entitled "Gestures For Touch Sensitive Input Devices," (4) U.S. patent application Ser. No. 11 / 048,264, filed January 31, 2005, entitled "Gestures For Touch Sensitive Input Devices," and (5) U.S. patent application Ser. No. 11 / 038,590, filed January 18, 2005, entitled "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices." No. 11 / 228,758, filed September 16, 2005, entitled "Virtual Input Device Placement On A Touch Screen User Interface," (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, entitled "Operation Of A Computer With A Touch Screen Interface," (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, entitled "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, entitled "Multi-Functional Hand-Held Device," all of which are incorporated herein by reference in their entireties.
[0107] Touchscreen 112 optionally has a video resolution greater than 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. A user optionally contacts touchscreen 112 using any suitable object or accessory, such as a stylus, a finger, or the like. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates the coarse finger input into precise pointer / cursor positions or commands to perform the action desired by the user.
[0108] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. The touchpad is optionally a touch-sensitive surface separate from touchscreen 112 or an extension of the touch-sensitive surface formed by the touchscreen.
[0109] Device 100 also includes a power system 162 that provides power to the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, power failure detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with generating, managing, and distributing electrical power within a portable device.
[0110] Device 100 also optionally includes one or more optical sensors 164. FIG. 1A shows an optical sensor coupled to optical sensor controller 158 in I / O subsystem 106. Optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Optical sensor 164 receives light from the environment projected through one or more lenses and converts the light into data representing an image. Optical sensor 164 optionally works in conjunction with imaging module 143 (also called a camera module) to capture still images or video. In some embodiments, the optical sensor is located on the back side of device 100 opposite touchscreen display 112 on the front of the device, so that the touchscreen display can be used as a viewfinder for capturing still images and / or video. In some embodiments, the optical sensor is located on the front of the device so that an image of a user is optionally captured for video conferencing while the user views other video conference participants on the touchscreen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), so that a single optical sensor 164 is used for both video conferencing and capturing still images and / or video, along with a touchscreen display.
[0111] Device 100 also optionally includes one or more depth camera sensors 175. FIG. 1A shows a depth camera sensor coupled to depth camera controller 169 in I / O subsystem 106. Depth camera sensor 175 receives data from the environment and creates a three-dimensional model of an object (e.g., a face) in a scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with imaging module 143 (also referred to as a camera module), depth camera sensor 175 is optionally used to determine a depth map of different portions of an image captured by imaging module 143. In some embodiments, a depth camera sensor is located on the front of device 100 to obtain images of the user with depth information for videoconferences and to capture selfie images with depth map data while the user views other videoconference participants on a touchscreen display. In some embodiments, depth camera sensor 175 is located on the back of the device, or on the back and front of device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), so that the depth camera sensor 175 is used for both video conferencing and capturing still images and / or video, in conjunction with a touchscreen display.
[0112] In some embodiments, a depth map (e.g., a depth map image) contains information (e.g., values) about the distance of objects in a scene from a viewpoint (e.g., a camera, light sensor, depth camera sensor). In one embodiment of a depth map, each depth pixel defines a position on the Z axis of the viewpoint where its corresponding two-dimensional pixel is located. In some embodiments, a depth map is made up of pixels, each defined by a value (e.g., 0 to 255). For example, a value of "0" represents a pixel located furthest in a "3D" scene, and a value of "255" represents a pixel located closest to the viewpoint (e.g., a camera, light sensor, depth camera sensor) in the "3D" scene. In other embodiments, a depth map represents the distance between objects in a scene and the plane of the viewpoint. In some embodiments, a depth map contains information about the relative depth of various features of an object of interest as seen by a depth camera (e.g., the relative depth of the eyes, nose, mouth, and ears on a user's face). In some embodiments, the depth map contains information that allows the device to determine the contours of the target object in the z direction.
[0113] Device 100 also optionally includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled to intensity sensor controller 159 in I / O subsystem 106. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with or proximate to the touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the back of device 100, opposite touchscreen display 112, which is located on the front of device 100.
[0114] Device 100 also optionally includes one or more proximity sensors 166. Figure 1A shows proximity sensor 166 coupled to peripherals interface 118. Alternatively, proximity sensor 166 is optionally coupled to input controller 160 within I / O subsystem 106. Proximity sensor 166 optionally functions as described in U.S. patent application Ser. Nos. 11 / 241,839, "Proximity Detector In Handheld Device," 11 / 240,788, "Proximity Detector In Handheld Device," 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output," 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices," and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals," which are incorporated herein by reference in their entireties. In some embodiments, the proximity sensor turns off and disables touchscreen 112 when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).
[0115] Device 100 also optionally includes one or more tactile output generators 167. FIG. 1A shows tactile output generators coupled to haptic feedback controller 161 in I / O subsystem 106. Tactile output generator 167 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile output on the device). Contact intensity sensor 165 receives tactile feedback generation instructions from haptic feedback module 133 and generates a tactile output on device 100 that can be sensed by a user of device 100. In some embodiments, at least one tactile output generator is juxtaposed with or proximate to a touch-sensitive surface (e.g., touch-sensitive display system 112) and generates a tactile output, optionally by moving the touch-sensitive surface vertically (e.g., in / out of the surface of device 100) or horizontally (e.g., back and forth in the same plane as the surface of device 100). In some embodiments, at least one tactile output generator sensor is located on the back of device 100, opposite touchscreen display 112, which is located on the front of device 100.
[0116] Device 100 also optionally includes one or more accelerometers 168. FIG. 1A shows accelerometer 168 coupled to peripherals interface 118. Alternatively, accelerometer 168 is optionally coupled to input controller 160 in I / O subsystem 106. Accelerometer 168 optionally functions as described in U.S. Patent Application Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and U.S. Patent Application Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated herein by reference in their entireties. In some embodiments, information is displayed on the touchscreen display in portrait or landscape orientation based on an analysis of data received from the one or more accelerometers. In addition to accelerometer(s) 168, device 100 optionally includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for obtaining information about the location and orientation (e.g., vertical or horizontal) of device 100.
[0117] In some embodiments, software components stored in memory 102 include operating system 126, communication module (or instruction set) 128, touch / motion module (or instruction set) 130, graphics module (or instruction set) 132, text input module (or instruction set) 134, Global Positioning System (GPS) module (or instruction set) 135, and applications (or instruction set) 136. Additionally, in some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) stores device / global internal state 157, as shown in FIGS. 1A and 3. Device / global internal state 157 includes one or more of: active application state indicating which applications, if any, are currently active; display state indicating which applications, views, or other information occupy various regions of touchscreen display 112; sensor state including information obtained from the device's various sensors and input control devices 116; and location information regarding the device's location and / or orientation.
[0118] Operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware and software components.
[0119] Communications module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by RF circuitry 108 and / or external port 124. External port 124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, similar to, and / or compatible with the 30-pin connector used on iPod® (trademark of Apple Inc.) devices.
[0120] Contact / motion module 130, optionally in conjunction with display controller 156, detects contact with touchscreen 112 and other touch-sensing devices (e.g., a touchpad or physical click wheel). Contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact occurs (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is contact movement and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger-drag events), and determining whether the contact has ceased (e.g., detecting a finger-up event or an interruption of the contact). Contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact, as represented by the series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact. These actions are optionally applied to a single contact (e.g., a single finger contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, contact / motion module 130 and display controller 156 detect contacts on the touchpad.
[0121] In some embodiments, contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator, but can be adjusted without modifying the physical hardware of device 100). For example, the mouse “click” threshold of a trackpad or touchscreen display can be set to any of a wide range of predefined thresholds without modifying the trackpad or touchscreen display hardware. Additionally, in some implementations, a user of the device is provided with a software setting to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once via a system-level click “intensity” parameter).
[0122] Contact / motion module 130 optionally detects gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different movements, timing, and / or intensities of detected contacts). Thus, gestures are optionally detected by detecting particular contact patterns. For example, detecting a finger tap gesture includes detecting a finger down event, followed by detecting a finger up (lift off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger down event, followed by one or more finger drag events, followed by detecting a finger up (lift off) event.
[0123] Graphics module 132 includes various known software components that render and display graphics on touchscreen 112 or other display, including components that modify the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual properties) of the displayed graphics. As used herein, the term "graphic" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (such as user interface objects including soft keys), digital images, video, animation, etc.
[0124] In some embodiments, graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 132 receives one or more codes specifying the graphics to be displayed, including coordinate data and other graphic property data, as needed, from an application or the like, and then generates screen image data to output to display controller 156.
[0125] The tactile feedback module 133 includes various software components for generating instructions used by the tactile output generator(s) 167 to generate tactile outputs at one or more locations on the device 100 in response to a user's interaction with the device 100.
[0126] Text input module 134 is optionally a component of graphics module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).
[0127] The GPS module 135 determines the location of the device and provides this information for use within various applications (e.g., to the phone 138 for use in location-based dialing, to the camera 143 as picture / video metadata, and to applications that provide location-based services such as a weather widget, a local yellow pages widget, and a maps / navigation widget).
[0128] Application 136 optionally includes the following modules (or instruction sets), or a subset or superset thereof: • a contacts module 137 (sometimes called an address book or contact list); ●Telephone module 138, ●Videoconferencing module 139, ● an email client module 140; ● Instant messaging (IM) module 141, ●Training support module 142, camera module 143 for still images and / or video; ● Image management module 144; ●Video player module, ●Music player module, ● Browser module 147, ●Calendar module 148, • A widget module 149 optionally including one or more of a weather widget 149-1, a stock price widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, and a user-created widget 149-6; a widget creator module 150 for creating user-created widgets 149-6; ● Search module 151, A video and music player module 152 that integrates a video player module and a music player module; ● Memo module 153, Map module 154, and / or ●Online video module 155.
[0129] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice duplication.
[0130] Contacts module 137, along with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, is optionally used to manage an address book or contact list (e.g., stored in memory 102 or in the application internal state 192 of contacts module 137 in memory 370), including adding name(s) to the address book, deleting name(s) from the address book, associating phone number(s), email address(es), street address(es), or other information with names, associating images with names, categorizing and sorting names, providing phone numbers or email addresses to initiate and / or facilitate communication by phone 138, videoconferencing module 139, email 140, or IM 141, etc. used to manage the address book or contact list.
[0131] Telephone module 138, in conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, is optionally used to enter a series of characters corresponding to a telephone number, access one or more telephone numbers in contacts module 137, modify entered telephone numbers, dial individual telephone numbers, place calls, and disconnect and hang up when the call is complete. As previously mentioned, wireless communication optionally uses any of a number of communication standards, protocols, and technologies.
[0132] Videoconferencing module 139 includes executable instructions to cooperate with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, touch / motion module 130, graphics module 132, text input module 134, contacts module 137, and telephone module 138 to initiate, conduct, and end a videoconference between a user and one or more other participants according to the user's commands.
[0133] Email client module 140, in conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, contains executable instructions for composing, sending, receiving, and managing emails in response to user commands. In conjunction with image management module 144, email client module 140 greatly facilitates the creation and sending of emails with still or video images captured by camera module 143.
[0134] Instant messaging module 141 includes executable instructions, in conjunction with RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, to enter a series of characters corresponding to an instant message, modify previously entered characters, send individual instant messages (e.g., using Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for telephony-based instant messaging, or XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receive instant messages, and view received instant messages. In some embodiments, sent and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments, such as those supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephony-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0135] In conjunction with the RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, the training support module 142 includes executable instructions to create workouts (e.g., with time, distance, and / or calorie burn goals), communicate with training sensors (sports devices), receive training sensor data, calibrate sensors used to monitor workouts, select and play music for workouts, and display, store, and transmit workout data.
[0136] Camera module 143, in conjunction with touchscreen 112, display controller 156, optical sensor(s) 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, includes executable instructions to capture and store still images or video (including video streams) in memory 102, modify characteristics of still images or video, or delete still images or video from memory 102.
[0137] Image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still and / or video images in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and camera module 143.
[0138] Browser module 147, in conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, contains executable instructions for browsing the Internet according to user commands, including retrieving, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.
[0139] The calendar module 148 includes executable instructions to cooperate with the RF circuitry 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, the email client module 140, and the browser module 147 to create, display, modify, and store calendars and data associated with the calendars (e.g., calendar items, to-do lists, etc.) according to user instructions.
[0140] Widget module 149, in conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, optionally provides mini-applications (e.g., weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) downloaded and used by a user, or mini-applications created by a user (e.g., user-created widget 149-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! Widgets).
[0141] The widget creator module 150, in conjunction with the RF circuitry 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, and the browser module 147, is optionally used by a user to create a widget (e.g., turn a user-specified portion of a web page into a widget).
[0142] The search module 151 includes executable instructions for working in conjunction with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134 to search for text, music, sound, images, video, and / or other files in the memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with a user's commands.
[0143] Video and music player module 152 includes executable instructions that, in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, enable a user to download and play pre-recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing videos (e.g., on touchscreen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (a trademark of Apple Inc.).
[0144] The notes module 153 includes executable instructions for working with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134 to create and manage notes, to-do lists, and the like according to user commands.
[0145] Map module 154, in conjunction with RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, is used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data regarding businesses and other points of interest at or near a particular location, and other location-based data), optionally in accordance with user instructions.
[0146] Online video module 155, in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, contains instructions that enable a user to access, browse for, receive (e.g., by streaming and / or downloading), and play (e.g., on the touchscreen or on an external display connected via external port 124) particular online videos, send emails with links to particular online videos, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 141 is used to send links to particular online videos, rather than email client module 140. For additional description of online video applications, see U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of which are incorporated herein by reference in their entireties.
[0147] The above-identified modules and applications each correspond to sets of executable instructions that perform one or more of the functions and methods described herein (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as respective software programs (e.g., computer programs (e.g., including instructions)), procedures, or modules; thus, in various embodiments, various subsets of these modules are optionally combined or otherwise rearranged. For example, a video player module is optionally combined with a music player module into a single module (e.g., video and music player module 152 of FIG. 1A). In some embodiments, memory 102 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 102 optionally stores additional modules and data structures not described above.
[0148] In some embodiments, device 100 is a device in which operation of a predetermined set of functions on the device is performed solely via a touchscreen and / or touchpad. Using the touchscreen and / or touchpad as the primary input control device for operation of device 100 optionally reduces the number of physical input control devices (push buttons, dials, etc.) on device 100.
[0149] The set of predefined functions performed only through the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 to a main menu, home menu, or root menu from any user interface displayed on device 100. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touchpad.
[0150] 1B is a block diagram illustrating exemplary components for event processing, according to some embodiments. In some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) includes an event sorter 170 (e.g., within operating system 126) and a separate application 136-1 (e.g., any of applications 137-151, 155, 380-390 described above).
[0151] Event sorter 170 receives the event information and determines which application 136-1 to deliver the event information to and application view 191 for application 136-1. Event sorter 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192 that indicates the current application view(s) that are displayed on touch-sensitive display 112 when the application is active or running. In some embodiments, device / global internal state 157 is used by event sorter 170 to determine which application(s) are currently active, and application internal state 192 is used by event sorter 170 to determine which application(s) are currently active, and application internal state 192 is used by event sorter 170 to determine which application view(s) to deliver the event information to.
[0152] In some embodiments, application internal state 192 includes additional information such as one or more of resume information to be used when application 136-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 136-1, state cues that allow the user to return to a previous state or view of application 136-1, and redo / undo cues of previous actions taken by the user.
[0153] Event monitor 171 receives event information from peripherals interface 118. The event information includes information about sub-events (e.g., a user touch as part of a multi-touch gesture on touch-sensitive display 112). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, accelerometer(s) 168, and / or microphone 113 (via audio circuitry 110). Information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display 112 or a touch-sensitive surface.
[0154] In some embodiments, event monitor 171 sends requests to peripherals interface 118 at predetermined intervals. In response, peripherals interface 118 transmits event information. In other embodiments, peripherals interface 118 transmits event information only when there is a significant event (e.g., receipt of an input above a predetermined noise threshold and / or for more than a predetermined duration).
[0155] In some embodiments, the event sorter 170 also includes a hit view determination module 172 and / or an active event recognizer determination module 173 .
[0156] Hit view determination module 172 provides software procedures that determine where a sub-event occurred within one or more views when touch-sensitive display 112 displays more than one view. A view consists of the controls and other elements that a user can see on the display.
[0157] Another aspect of a user interface associated with an application is the set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of individual applications) in which touches are detected optionally correspond to programmatic levels within the application's programmatic or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events that are recognized as appropriate inputs is optionally determined based at least in part on the hit view of the initial touch that initiates the touch gesture.
[0158] Hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy that should process the sub-events. In most situations, the hit view is the lowest-level view in which an initiating sub-event occurs (e.g., the first sub-event in a series of sub-events that form an event or potential event). Once a hit view is identified by hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source as the touch or input source identified as the hit view.
[0159] Active event recognizer determination module 173 determines which view(s) in the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that contain the physical location of the sub-events are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to the area associated with one particular view, views higher in the hierarchy still remain actively participating views.
[0160] Event dispatcher module 174 dispatches event information to event recognizers (e.g., event recognizer 180). In embodiments that include active event recognizer determination module 173, event dispatcher module 174 delivers event information to the event recognizers determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores event information in an event queue, which is retrieved by individual event receivers 182.
[0161] In some embodiments, operating system 126 includes event sorter 170. Alternatively, application 136-1 includes event sorter 170. In still other embodiments, event sorter 170 is a stand-alone module or is part of another module stored in memory 102, such as contact / motion module 130.
[0162] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each containing instructions for processing touch events that occur within a separate view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, an individual application view 191 includes multiple event recognizers 180. In other embodiments, one or more of the event recognizers 180 are part of a separate module, such as a user interface kit or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, individual event handlers 190 include one or more of data updater 176, object updater 177, GUI updater 178, and / or event data 179 received from event sorter 170. Event handler 190 optionally utilizes or invokes data updater 176, object updater 177, or GUI updater 178 to update application internal state 192. Alternatively, one or more of the application views 191 include one or more respective event handlers 190. Also, in some embodiments, one or more of the data updater 176, the object updater 177, and the GUI updater 178 are included in individual application views 191.
[0163] A separate event recognizer 180 receives event information (e.g., event data 179) from event sorter 170 and identifies events from the event information. Event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event recognizer 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (optionally including sub-event delivery instructions).
[0164] The event receiver 182 receives event information from the event sorter 170. The event information includes information about a sub-event, e.g., a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves a movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also called the device's posture).
[0165] The event comparator 184 compares the event information with predefined event or sub-event definitions and determines the event or sub-event, or determines or updates the state of the event or sub-event, based on the comparison. In some embodiments, the event comparator 184 includes an event definition 186. The event definition 186 includes definitions of events (e.g., a predefined set of sub-events), such as Event 1 (187-1) and Event 2 (187-2). In some embodiments, sub-events within an event (187) include, for example, touch start, touch end, touch movement, touch cancellation, and multiple touches. In one example, the definition for Event 1 (187-1) is a double tap on a displayed object. The double tap includes, for example, a first touch on a displayed object relative to a predetermined phase (touch start), a first lift-off (touch end) relative to the predetermined phase, a second touch on a displayed object relative to the predetermined phase (touch start), and a second lift-off (touch end) relative to the predetermined phase. In another example, a definition of event 2 (187-2) is a drag on a displayed object. Drag includes, for example, a touch (or contact) on the displayed object to a predetermined stage, a movement of the touch across the touch-sensitive display 112, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event handlers 190.
[0166] In some embodiments, event definition 187 includes definitions of events for individual user interface objects. In some embodiments, event comparator 184 performs a hit test to determine which user interface objects are associated with the sub-event. For example, if a touch is detected on touch-sensitive display 112 in an application view in which three user interface objects are displayed on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a separate event handler 190, event comparator 184 uses the results of the hit test to determine which event handler 190 to activate. For example, event comparator 184 selects the event handler associated with the sub-event and object that triggers the hit test.
[0167] In some embodiments, the definition of an individual event 187 also includes a delay action that delays delivery of the event information until it is determined whether a set of sub-events corresponds to the event recognizer's event type.
[0168] If the individual event recognizer 180 determines that the sequence of sub-events does not match any of the events in the event definition 186, the individual event recognizer 180 enters an event disabled, event failed, or event finished state, after which it ignores the next sub-event of the touch-based gesture. In this situation, any other event recognizers that remain active for the hit view continue to track and process sub-events of the ongoing touch gesture.
[0169] In some embodiments, individual event recognizers 180 include metadata 183 with configurable properties, flags, and / or lists that indicate to actively participating event recognizers how the event delivery system should perform sub-event delivery. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact with each other or how event recognizers are allowed to interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how sub-events are delivered to various levels in the view or programmatic hierarchy.
[0170] In some embodiments, an individual event recognizer 180 activates an event handler 190 associated with an event when one or more specific sub-events of the event are recognized. In some embodiments, the individual event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is separate from sending (and postponing sending) sub-events to the individual hit view. In some embodiments, the event recognizer 180 pops a flag associated with the recognized event, and the event handler 190 associated with the flag catches the flag and performs a predetermined process.
[0171] In some embodiments, the event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with a set of sub-events or to an actively participating view. The event handler associated with the set of sub-events or the actively participating view receives the event information and performs a predetermined process.
[0172] In some embodiments, data updater 176 creates and updates data used by application 136-1. For example, data updater 176 updates phone numbers used by contacts module 137 or stores video files used by a video player module. In some embodiments, object updater 177 creates and updates objects used by application 136-1. For example, object updater 177 creates new user interface objects or updates the positions of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on the touch-sensitive display.
[0173] In some embodiments, event handler(s) 190 include or have access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of an individual application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0174] It should be understood that the foregoing description of event processing of a user's touch on a touch-sensitive display also applies to other forms of user input for operating multifunction device 100 using input devices, although not all of them are initiated on a touchscreen. For example, mouse movements and mouse button presses, contact movements such as tapping, dragging, scrolling on a touchpad, optionally coordinated with single or multiple keyboard presses or holds, pen stylus input, device movement, verbal commands, detected eye movements, biometric input, and / or any combination thereof, optionally utilize as inputs corresponding to sub-events that define the event to be recognized.
[0175] FIG. 2 illustrates portable multifunction device 100 having touchscreen 112, according to some embodiments. The touchscreen optionally displays one or more graphics within user interface (UI) 200. In this embodiment, as well as other embodiments described below, a user may select one or more of the graphics by performing a gesture on the graphics, for example, using one or more fingers 202 (not drawn to scale) or one or more styluses 203 (not drawn to scale). In some embodiments, selection of one or more graphics is performed when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward and / or downward), and / or rolling (right to left, left to right, upward and / or downward) of a finger in contact with device 100. In some implementations or situations, accidental contact with a graphic does not select the graphic, for example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon optionally does not select the corresponding application.
[0176] Device 100 also optionally includes one or more physical buttons, such as a "home" button or menu button 204. As previously mentioned, menu button 204 is optionally used to navigate to any application 136 within a set of applications running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touchscreen 112.
[0177] In some embodiments, device 100 includes touchscreen 112, menu button 204, pushbutton 206 for powering the device on / off and locking the device, volume control button(s) 208, subscriber identity module (SIM) card slot 210, headset jack 212, and external docking / charging port 124. Pushbutton 206 is optionally used to power the device on / off by pressing and holding the button down for a predetermined period of time, to lock the device by pressing and releasing the button before the predetermined time has elapsed, and / or to unlock the device or initiate the unlocking process. In alternative embodiments, device 100 also accepts verbal input via microphone 113 to activate or deactivate certain functions. Device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of a contact on touchscreen 112 and / or one or more tactile output generators 167 for generating a tactile output for a user of device 100.
[0178] FIG. 3 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or commercial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communication between system components. Device 300 includes input / output (I / O) interface 330, including display 340, which is typically a touchscreen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 that generates tactile output on device 300 (e.g., similar to tactile output generator(s) 167 described above with reference to FIG. 1A ), and sensors 359 (e.g., light, acceleration, proximity, touch-sensing, and / or contact intensity sensors similar to contact intensity sensor(s) 165 described above with reference to FIG. 1A ). Memory 370 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or a subset of, programs, modules, and data structures stored in memory 102 of portable multifunction device 100 (FIG. 1A). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores drawing module 380, presentation module 382, word processing module 384, website creation module 386, disk authoring module 388, and / or spreadsheet module 390, whereas memory 102 of portable multifunction device 100 (FIG. 1A) optionally does not store these modules.
[0179] Each of the above-identified elements of FIG. 3 is optionally stored in one or more of the memory devices mentioned above. Each of the above-identified modules corresponds to an instruction set that performs the functions described above. The above-identified modules or computer programs (e.g., including an instruction set or instructions) need not be implemented as separate software programs (e.g., computer programs (e.g., including instructions)), procedures, or modules; thus, in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. In some embodiments, memory 370 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 370 optionally stores additional modules and data structures not described above.
[0180] Attention is now optionally directed to user interface embodiments, for example, as implemented on portable multifunction device 100.
[0181] 4A shows an exemplary user interface for a menu of applications on portable multifunction device 100, according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof: signal strength indicator(s) 402 for wireless communication(s), such as cellular and Wi-Fi signals; ●Time 404, ●Bluetooth indicator 405, ● Battery status indicator 406, Tray 408 with icons of frequently used applications, such as: An icon 416 for the phone module 138, labeled "Phone," optionally including an indicator 414 of the number of missed calls or voicemail messages; icon 418 of the email client module 140, labeled "Mail," optionally including an indicator 410 of the number of unread emails; ○ An icon 420 for the browser module 147, labeled "Browser"; and ○ An icon 422 for the video and music player module 152, also called the iPod (trademark of Apple Inc.) module 152, labeled "iPod"; and ● Icons of other applications, such as: ○ Icon 424 of IM module 141, labeled "Messages"; ○ Icon 426 of the calendar module 148, labeled "Calendar" ○ Icon 428 of the image management module 144, labeled "Photos" ○ An icon 430 of the camera module 143, labeled "camera"; ○ Icon 432 of the online video module 155, labeled "Online Video"; Icon 434 of Stock Price Widget 149-2, labeled "Stock Price" ○ Icon 436 of the map module 154, labeled "Map"; Icon 438 of weather widget 149-1, labeled "Weather" ○ Icon 440 of alarm clock widget 149-4, labeled "Clock" ○ Icon 442 of Training Support Module 142, labeled "Training Support"; ○ An icon 444 of the Notes module 153 labeled "Notes," and A settings application or module icon 446 labeled "Settings" that provides access to settings for the device 100 and its various applications 136.
[0182] Note that the icon labels shown in FIG. 4A are merely exemplary. For example, icon 422 of video and music player module 152 is labeled "Music" or "Music Player." Other labels are optionally used for various application icons. In some embodiments, the label for an individual application icon includes the name of the application that corresponds to the individual application icon. In some embodiments, the label of a particular application icon is different from the name of the application that corresponds to that particular application icon.
[0183] 4B shows an example user interface on a device (e.g., device 300 of FIG. 3 ) that has touch-sensitive surface 451 (e.g., tablet or touchpad 355 of FIG. 3 ) that is separate from display 450 (e.g., touchscreen display 112). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) that detect the intensity of a contact on touch-sensitive surface 451, and / or one or more tactile output generators 357 that generate a tactile output for a user of device 300.
[0184] Although some of the following examples are given with reference to input on touchscreen display 112 (which combines a touch-sensitive surface and a display), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, as shown in FIG. 4B . In some embodiments, the touch-sensitive surface (e.g., 451 in FIG. 4B ) has a primary axis (e.g., 452 in FIG. 4B ) that corresponds to a primary axis (e.g., 453 in FIG. 4B ) on the display (e.g., 450). According to these embodiments, the device detects contact with touch-sensitive surface 451 (e.g., 460 and 462 in FIG. 4B ) at locations that correspond to respective locations on the display (e.g., in FIG. 4B , 460 corresponds to 468 and 462 corresponds to 470). In this way, user input (e.g., contacts 460 and 462 and their movement) detected by the device on the touch-sensitive surface (e.g., 451 in FIG. 4B ) is used by the device to operate a user interface on the display (e.g., 450 in FIG. 4B ) of the multifunction device when the touch-sensitive surface is separate from the display. It should be understood that similar methods are optionally used for the other user interfaces described herein.
[0185] Additionally, while the following examples are given primarily with reference to finger input (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs are replaced with input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact) followed by movement of a cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click (e.g., instead of detecting a contact and then ceasing contact detection) while the cursor is positioned over the location of the tap gesture. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice are optionally used simultaneously, or a mouse and finger contacts are optionally used simultaneously.
[0186] FIG. 5A shows an exemplary personal electronic device 500. Device 500 includes a main body 502. In some embodiments, device 500 can include some or all of the functionality described with respect to devices 100 and 300 (e.g., FIGS. 1A-4B ). In some embodiments, device 500 has a touch-sensitive display screen 504, hereafter touchscreen 504. Alternatively, or in addition to touchscreen 504, device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, touchscreen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors that detect the intensity of contact (e.g., touches) being applied. The one or more intensity sensors of touchscreen 504 (or the touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of device 500 can respond to touches based on their intensity, meaning that touches of different intensities can invoke different user interface actions on device 500.
[0187] For exemplary techniques for detecting and processing touch intensity, see, for example, related applications International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," published as International Publication No. WO / 2013 / 169849, and International Patent Application No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," published as International Publication No. WO / 2014 / 105276, each of which is incorporated herein by reference in its entirety.
[0188] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508, if included, may be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms, if included, may allow device 500 to be attached to, for example, hats, eyewear, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow device 500 to be worn by a user.
[0189] FIG. 5B illustrates an exemplary personal electronic device 500. In some embodiments, device 500 can include some or all of the components described with respect to FIGS. 1A, 1B, and 3. Device 500 has a bus 512 operably coupling an I / O section 514 to one or more computer processors 516 and memory 518. I / O section 514 can be connected to a display 504, which can have touch-sensing components 522 and, optionally, an intensity sensor 524 (e.g., a contact intensity sensor). Additionally, I / O section 514 can be connected to a communication unit 530 that receives application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 can include input mechanisms 506 and / or 508. Input mechanism 506 is optionally, for example, a rotatable input device or a depressible and rotatable input device. In some embodiments, input mechanism 508 is optionally a button.
[0190] In some embodiments, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which may be operably connected to I / O section 514.
[0191] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform, for example, the techniques described below, including processes 700, 800, 1000, 1200, 1400, 1500, 1700, and 1900 (FIGS. 7, 8, 10, 12, 14, 15, 17, and 19). A computer-readable storage medium may be any medium that can tangibly contain or store computer-executable instructions used by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transitory computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include, but is not limited to, magnetic, optical, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memory such as flash and solid-state drives. The personal electronic device 500 is not limited to the components and configuration of FIG. 5B and may include other or additional components in multiple configurations.
[0192] As used herein, the term "affordance" refers to a user-interactive graphical user interface object, optionally displayed on a display screen of device 100, 300, and / or 500 (FIGS. 1A, 3, and 5A-5C). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each, optionally, constitute an affordance.
[0193] As used herein, the term “focus selector” refers to an input element that indicates the current portion of the user interface with which the user is interacting. In some implementations involving a cursor or other location marker, the cursor acts as the “focus selector,” such that when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 of FIG. 3 or touch-sensitive surface 451 of FIG. 4B) while the cursor is positioned over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations involving a touchscreen display that allows direct interaction with user interface elements on the touchscreen display (e.g., touch-sensitive display system 112 of FIG. 1A or touchscreen 112 of FIG. 4A), a detected contact on the touchscreen acts as the “focus selector,” such that when input (e.g., a press input by contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one region of the user interface to another region of the user interface without a corresponding cursor movement or contact movement on the touchscreen display (e.g., by using the tab key or arrow keys to move focus from one button to another), and in these implementations, the focus selector moves to follow the movement of focus between various regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on a touchscreen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface through which the user intends to interact).For example, the position of a focus selector (e.g., cursor, touch, or selection box) over an individual button while a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen) indicates that the user intends to activate that individual button (and not other user interface elements shown on the device's display).
[0194] As used herein and in the claims, the term "characteristic intensity" of a contact refers to a characteristic of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on a plurality of intensity samples. The characteristic intensity is optionally based on a predetermined number of intensity samples, i.e., a set of intensity samples collected during a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detecting the contact, before detecting lift-off of the contact, before or after detecting the start of contact movement, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The characteristic intensity of the contact is optionally based on one or more of the maximum intensity of the contact, the mean intensity of the contact, the average intensity of the contact, the top 10 percentile intensity of the contact, half the maximum intensity of the contact, 90 percent of the maximum intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action is performed by the user. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact having a characteristic intensity that does not exceed the first threshold results in a first action, a contact having a characteristic intensity above the first intensity threshold but not above the second intensity threshold results in a second action, and a contact having a characteristic intensity above the second threshold results in a third action. In some embodiments, the comparison between the characteristic intensity and the one or more thresholds is not used to determine whether to perform the first action or the second action, but rather to determine whether to perform one or more actions (e.g., whether to perform an individual action or to refrain from performing an individual action).
[0195] 5C shows an exemplary diagram of a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500 and each share one or more data connections 510, such as an Internet connection, a Wi-Fi connection, a cellular connection, a short-range communication connection, and / or any other such data connection or network, with each other to facilitate real-time communication of audio and / or video data between the respective devices over a period of time. In some embodiments, the exemplary communication session may include a shared data session in which data is communicated from one or more of the electronic devices to other electronic devices to enable simultaneous output of individual content at the electronic devices. In some embodiments, the exemplary communication session may include a videoconferencing session in which audio and / or video data is communicated between devices 500A, 500B, and 500C so that users of the respective devices can engage in real-time communication using the electronic devices.
[0196] 5C, device 500A represents an electronic device associated with user A. Device 500A is in communication (via data connection 510) with devices 500B and 500C associated with users B and C, respectively. Device 500A includes a camera 501A used to capture video data of the communication session and a display 504A (e.g., a touchscreen) used to display content associated with the communication session. Device 500A also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.
[0197] Device 500A displays, via display 504A, communications UI 520A, a user interface that facilitates a communications session (e.g., a video conferencing session) between device 500B and device 500C. Communications UI 520A includes video feed 525-1A and video feed 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and communicated from device 500B to devices 500A and 500C during the communications session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and communicated from device 500C to devices 500A and 500B during the communications session.
[0198] Communication UI 520A includes camera preview 550A, which is a representation of video data captured at device 500A via camera 501A. Camera preview 550A represents to user A the expected video feed of user A that will be displayed on each of devices 500B and 500C.
[0199] The communications UI 520A includes one or more controls 555A that control one or more aspects of the communications session. For example, the controls 555A may include controls for muting the audio of the communications session, changing the camera view of the communications session (e.g., changing which camera is used to capture video of the communications session, adjusting the zoom value), ending the communications session, applying visual effects to the camera view of the communications session, and activating one or more modes associated with the communications session. In some embodiments, the one or more controls 555A are optionally displayed on the communications UI 520A. In some embodiments, the one or more controls 555A are displayed separately from the camera preview 550A. In some embodiments, the one or more controls 555A are displayed over at least a portion of the camera preview 550A.
[0200] 5C, device 500B represents an electronic device associated with user B, who is in communication with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B used to capture video data for the communication session and a display 504B (e.g., a touchscreen) used to display content associated with the communication session. Device 500B also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.
[0201] Device 500B displays, via touchscreen 504B, a communications UI 520B similar to communications UI 520A of device 500A. Communications UI 520B includes video feed 525-1B and video feed 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and communicated from device 500A to devices 500B and 500C during the communications session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and communicated from device 500C to devices 500A and 500B during the communications session. Communications UI 520B also includes camera preview 550B, which is a representation of video data captured at device 500B via camera 501B, and one or more controls 555B, similar to control 555A, for controlling one or more aspects of the communications session. Camera preview 550B represents to user B the expected video feed of user B as it would appear on each device 500A and 500C.
[0202] 5C, device 500C represents an electronic device associated with user C, who is in communication with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C used to capture video data of the communication session and a display 504C (e.g., a touchscreen) used to display content associated with the communication session. Device 500C also includes other components, such as a microphone (e.g., 113) for recording audio for the communication session and a speaker (e.g., 111) for outputting audio for the communication session.
[0203] Device 500C displays, via touchscreen 504C, a communication UI 520C similar to communication UI 520A of device 500A and communication UI 520B of device 500B. Communication UI 520C includes video feed 525-1C and video feed 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and communicated from device 500B to devices 500A and 500C during the communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and communicated from device 500A to devices 500B and 500C during the communication session. Communications UI 520C also includes camera preview 550C, which is a representation of video data captured at device 500C via camera 501C, and one or more controls 555C for controlling one or more aspects of the communications session, similar to controls 555A and 555B. Camera preview 550C represents to user C the expected video feed of user C that will be displayed on each of devices 500A and 500B.
[0204] While the diagram shown in FIG. 5C depicts a communication session between three electronic devices, a communication session can be established between two or more electronic devices, and the number of devices participating in a communication session can change as electronic devices join or leave the communication session. For example, if one of the electronic devices leaves the communication session, the audio and video data from the device that stopped participating in the communication session is no longer represented on the participating devices. For example, if device 500B stops participating in the communication session, data connection 510 between devices 500A and 500C and data connection 510 between devices 500C and 500B no longer exist. In addition, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, when a device joins a communication session, a connection is established between the joining device and the existing devices, and video and audio data is shared among all devices so that each device can output data communicated from the other devices.
[0205] The embodiment shown in Figure 5C represents a diagram of a communication session between multiple electronic devices, including the exemplary communication sessions shown in Figures 6A-6AY, 9A-9T, 11A-11P, 13A-13K, and 16A-16Q. In some embodiments, the communication sessions shown in Figures 6A-6AY, 9A-9T, 13A-13K, and 16A-16Q include two or more electronic devices, even if other electronic devices participating in the communication session are not shown in the figures.
[0206] Attention is now directed to embodiments of user interfaces (“UIs”) and related processes implemented on an electronic device such as portable multifunction device 100, device 300, or device 500.
[0207] 6A-6AY show exemplary user interfaces for managing a live video communication session, according to some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes in FIGS. 7 and 8 and 15.
[0208] 6A-6AY show exemplary user interfaces for managing a live video communication session from the perspective of different users (e.g., users participating in the live video communication session from different devices, different types of devices, devices with different applications installed, and / or devices with different operating system software).
[0209] 6A , device 600-1, in some embodiments, corresponds to user 622 (e.g., “John”), who is a participant in a live video communication session. Device 600-1 includes a display (e.g., a touch-sensitive display) 601 and a camera 602 (e.g., a front-facing camera) having a field of view 620. In some embodiments, camera 602 is configured to capture image data and / or depth data of the physical environment within field of view 620. Field of view 620 may be referred to herein as the available field of view, the full field of view, or the camera field of view. In some embodiments, camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600-1 includes multiple cameras. Thus, although device 600-1 is described herein as using camera 602 to capture image data during a live video communication session, it will be understood that device 600-1 may use multiple cameras to capture image data.
[0210] 6A , device 600-2 corresponds to user 623 (e.g., “Jane”), who, in some embodiments, is a participant in a live video communication session. Device 600-2 includes a display (e.g., a touch-sensitive display) 683 and a camera 682 (e.g., a front-facing camera) having a field of view 688. In some embodiments, camera 682 is configured to capture image data and / or depth data of the physical environment within field of view 688. Field of view 688 may be referred to herein as the available field of view, the full field of view, or the camera field of view. In some embodiments, camera 682 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600-2 includes multiple cameras. Thus, while this specification describes device 600-2 using camera 682 to capture image data during a live video communication session, it will be understood that device 600-2 may use multiple cameras to capture image data.
[0211] As shown, user 622 (“John”) is positioned (e.g., sitting) in front of desk 621 (and device 600-1) in environment 615. In some examples, user 622 is positioned in front of desk 621 such that user 622 is captured within field of view 620 of camera 602. In some embodiments, one or more objects proximate user 622 are positioned such that the objects are captured within field of view 620 of camera 602. In some embodiments, both user 622 and objects proximate user 622 are simultaneously captured within field of view 620. For example, as shown, drawing 618 is positioned on surface 619 in front of user 622 (relative to camera 602) such that both user 622 and drawing 618 are captured within field of view 620 of camera 602 and displayed within representation 622-1 (displayed by device 600-1) and representation 622-2 (displayed by device 600-2).
[0212] Similarly, user 623 ("Jane") is positioned (e.g., sitting) in front of desk 686 (and device 600-2) in environment 685. In some embodiments, user 623 is positioned in front of desk 686 so that user 623 is captured within field of view 688 of camera 682. As shown, user 623 is displayed in representation 623-1 (displayed by device 600-1) and representation 623-2 (displayed by device 600-2).
[0213] Generally, during operation, devices 600-1, 600-2 capture image data, which is exchanged between devices 600-1, 600-2 and used by devices 600-1, 600-2 to display various representations of content during the live video communication session. While each of devices 600-1, 600-2 is illustrated, the described embodiments are primarily directed to the user interface displayed on device 600-1 and / or user input detected by device 600-1. It should be understood that in some embodiments, electronic device 600-2 operates similarly to electronic device 600-1 during the live video communication session. In some embodiments, devices 600-1, 600-2 display user interfaces similar to those described below and / or perform similar operations.
[0214] As described in further detail below, in some examples, such representations include images modified during a live video communication session to provide an improved perspective of surfaces and / or objects within the fields of view (also referred to herein as “fields of view”) of the cameras of devices 600-1, 600-2. The images may be modified using any known image processing techniques, including, but not limited to, image rotation and / or distortion correction (e.g., image skew). Thus, image data may be captured from a camera having a particular location relative to a user, but the representation may provide a perspective that shows the user (and / or surfaces or objects in the user's environment) from a perspective different from that of the camera capturing the image data. The embodiments of FIGS. 6A-6AY disclose displaying elements and detecting inputs (including hand gestures) at device 600-1 that control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2). In some embodiments, device 600-2 displays similar elements and detects similar inputs (including hand gestures) at device 600-2 that control how image data captured by camera 602 is displayed (at either device 600-1 and / or device 600-2).
[0215] 6A , device 600-1 displays video conferencing interface 604-1 on display 601. Video conferencing interface 604-1 includes representation 622-1, which includes an image (e.g., a frame of a video stream) of a physical environment (e.g., a scene) within field of view 620 of camera 602. In some embodiments, the image of representation 622-1 includes the entire field of view 620. In other embodiments, the image of representation 622-1 includes a portion (e.g., a crop or subset) of the entire field of view 620. As shown, in some embodiments, the image of representation 622-1 includes user 622 and / or surface 619 proximate user 622 on which drawing 618 is located.
[0216] Videoconferencing interface 604-1 further includes representation 623-1, which includes an image of the physical environment within field of view 688 of camera 682. In some embodiments, the image of representation 623-1 includes the entire field of view 688. In other embodiments, the image of representation 623-1 includes a portion (e.g., a crop or subset) of the entire field of view 688. As shown, in some embodiments, the image of representation 623-1 includes user 623. As shown, representation 623-1 is displayed at a larger size than representation 622-1. In this manner, user 622 can better observe and / or interact with user 623 during the live communication session.
[0217] Device 600-2 displays videoconferencing interface 604-2 on display 683. Videoconferencing interface 604-2 includes representation 622-2, which includes an image of the physical environment within field of view 620 of camera 602. Videoconferencing interface 604-2 further includes representation 623-2, which includes an image of the physical environment within field of view 688 of camera 682. As shown, representation 622-2 is displayed at a larger size than representation 623-2. In this manner, user 623 can better observe and / or interact with user 622 during the live communication session.
[0218] In Figure 6A, device 600-1 is displaying interface 604-1. While displaying interface 604-1, device 600-1 detects input 612a (e.g., a swipe input) corresponding to a request to display a settings interface. In response to detecting input 612a, device 600-1 displays settings interface 606, as shown in Figure 6B. As shown, in some embodiments, settings interface 606 is overlaid on interface 604-1.
[0219] In some embodiments, settings interface 606 includes one or more affordances that control settings of device 600-1 (e.g., volume, display brightness, and / or Wi-Fi settings). For example, settings interface 606 includes view affordance 607-1, which, when selected, causes device 600-1 to display a view menu, as shown in FIG. 6B.
[0220] As shown in Figure 6B, while displaying settings interface 606, device 600-1 detects input 612b, which, in some embodiments, is a tap gesture on view affordance 607-1. In response to detecting input 612b, device 600-1 displays view menu 616-1, as shown in Figure 6C.
[0221] Generally, view menu 616-1 includes one or more affordances that can be used to manage (e.g., control) how representations are displayed during a live video communication session. By way of example, selection of a particular affordance may cause device 600-1 to display or cease displaying a representation in an interface (e.g., interface 604-1 or interface 604-2).
[0222] View menu 616-1 includes surface view affordance 610, which, for example, when selected, causes device 600-1 to display a representation including a modified image of the surface. In some embodiments, when surface view affordance 610 is selected, the user interface transitions directly to the user interface of FIG. 6M. Additionally or alternatively, FIGS. 6D-6L (described below) illustrate other user interfaces that may be displayed before the user interface of FIG. 6M, as well as other inputs for initiating the process of displaying the user interface shown in FIG. 6M. For example, while displaying view menu 616-1, device 600-1 detects input 612c corresponding to selection of surface view affordance 610. In some embodiments, input 612c is a touch input. In response to detecting input 612c, device 600-1 displays representation 624-1, as shown in FIG. 6M. In response to further detecting input 612c, device 600-2 displays representation 624-2. As described, in some embodiments, images are modified during a live video communication session to provide an image having a particular perspective. Thus, in some examples, representation 624-1 is provided by generating an image from image data captured by camera 602, modifying the image (or a portion of the image), and displaying representation 624-1 along with the modified image. In some embodiments, the image is modified using any known image processing technique, including, but not limited to, image rotation and / or distortion correction (e.g., image skew). An image of representation 624-2 is also provided in this manner in some embodiments.
[0223] In some embodiments, the image of representation 624-1 is modified to provide a desired perspective (e.g., a surface view). In some embodiments, the image of representation 624-1 is modified based on the position of surface 619 relative to camera 602. By way of example, device 600-1 may rotate the image of representation 624-1 a predetermined amount (e.g., 45 degrees, 90 degrees, or 180 degrees) to provide a more intuitive view of surface 619 in representation 624-1. As shown in FIG. 6M , for example, if camera 602 captures surface 619 from a perspective facing user 622, the image of representation 624-1 is rotated 180 degrees to provide a view of the image from the perspective of user 622. Thus, during a live video communication session, devices 600-1, 600-2 display surface 619 (and thus drawing 618) from the perspective of user 622 during the live communication session. An image of representation 624-2 is also provided in this manner in some examples.
[0224] In some embodiments, device 600-2 maintains display of representation 624-2 to ensure that user 623 maintains a view of user 622 while representation 622-2 includes the modified image of surface 619. As shown in FIG. 6M , maintaining display of representation 622-2 in this manner may include adjusting the size and / or position of representation 622-2 within interface 604-2. Optionally, in some embodiments, device 600-2 ceases displaying representation 622-2 and provides a larger-sized representation 624-2. Optionally, in some embodiments, device 600-1 ceases displaying representation 622-1 and provides a larger-sized representation 624-1.
[0225] Representations 624-1, 624-2 include an image of drawing 618 that has been modified with respect to the position (e.g., location and / or orientation) of drawing 618 relative to camera 602. For example, as shown in FIG. 6A , before modification, the image is shown in representations 622-1, 622-2 as having a particular orientation (e.g., upside down). As a result of modifying the image, the image of drawing 618 is rotated and / or tilted so that the perspective of representations 624-1, 624-2 appears to be from the perspective of user 622. In this way, the modified image of drawing 618 provides a perspective that is different from that of representations 624-1, 624-2 so as to give user 623 (and / or user 622) a more natural and direct view of drawing 618. Thus, drawing 618 may be more easily and intuitively viewed by user 623 during a live video communication session.
[0226] As described, a representation including a modified image of a surface is provided in response to selecting a surface image affordance (e.g., surface view affordance 610). In some examples, a representation including a modified view of a surface is provided in response to detecting other types of input.
[0227] 6D , in some examples, a representation including a modified image of the surface is provided in response to one or more gestures. As one example, device 600-1 may detect a gesture using camera 602 and, in response to detecting the gesture, determine whether the gesture satisfies a set of criteria (e.g., a gesture criteria set). In some embodiments, the criteria include a requirement that the gesture be a pointing gesture, and optionally a requirement that the pointing gesture have a particular orientation and / or be directed toward a surface and / or object. For example, with reference to FIG. 6D , device 600-1 detects gesture 612d and determines that gesture 612d is a pointing gesture directed toward drawing 618. In response, device 600-1 displays a representation including a modified image of surface 619, as described with reference to FIG. 6M .
[0228] In some embodiments, the set of criteria includes a requirement that the gesture be performed for at least a threshold amount of time. For example, with reference to FIG. 6E , in response to detecting a gesture, device 600-1 overlays representation 622-1 with graphical object 626 indicating that device 600-1 has detected that the user is currently performing a gesture, such as 612d. As shown, in some embodiments, device 600-1 expands representation 622-1 to help user 622 better see the detected gesture and / or graphical object 626.
[0229] In some embodiments, the graphical object 626 includes a timer 628 (e.g., a numeric timer, a ring that fills over time, and / or a bar that fills over time) that indicates the amount of time that the gesture 612d has been detected. In some embodiments, the timer 628 also (or alternatively) indicates a threshold amount of time that the gesture 612d continues to be provided to satisfy a set of criteria. In response to the gesture 612d meeting the threshold amount of time (e.g., 0.5 seconds, 2 seconds, and / or 5 seconds), the device 600-1 displays a representation 624-1 ( FIG. 6M ) that includes a modified image of the surface, as described.
[0230] In some examples, graphical object 626 indicates a type of gesture currently being detected by device 600-1. In some examples, graphical object 626 is an outline of a hand performing a gesture of the detected type and / or an image of the gesture of the detected type. Graphical object 626 may include, for example, a hand performing a pointing gesture in response to device 600-1 detecting that user 622 is performing a pointing gesture. Additionally or alternatively, graphical object 626 may optionally indicate a zoom level (e.g., the zoom level at which a representation of a second portion of the scene is or will be displayed).
[0231] In some examples, the representation with the modified image is provided in response to one or more voice inputs. For example, during a live communication session, device 600-1 receives a voice input, such as voice input 614 of FIG. 6D ("Look at my drawing."). In response, device 600-1 displays representation 624-1 (FIG. 6M) that includes the modified image of the surface, as described.
[0232] In some embodiments, audio input received by device 600-1 may include a reference to any surface and / or object recognizable by device 600-1, and in response, device 600-1 provides a representation that includes a modified image of the referenced object or surface. For example, device 600-1 may receive audio input that references a wall (e.g., the wall behind user 622). In response, device 600-1 provides a representation that includes a modified image of the wall.
[0233] In some embodiments, audio input can be used in combination with other types of input, such as gestures (e.g., gesture 612d). Thus, in some embodiments, device 600-1 displays a modified image of a surface (or object) in response to detecting both a gesture and audio input corresponding to a request to provide a modified image of the surface.
[0234] In some embodiments, the surface viewer affordances are provided in other ways. For example, referring to Figure 6F, video conferencing interface 604-1 includes options menu 608. Options menu 608 includes a set of affordances that can be used to control device 600-1 during a live video communication session, including viewer affordance 607-2.
[0235] While displaying options menu 608, device 600-1 detects input 612f corresponding to selection of view affordance 607-2. In response to detecting input 612f, device 600-1 displays view menu 616-2, as shown in FIG. 6G. View menu 616-2 can be used to control how representations are displayed during a live video communication session, as described with respect to FIG. 6C.
[0236] Although options menu 608 is shown throughout the figures as being permanently displayed on video conferencing interface 604-1, options menu 608 may be hidden and / or redisplayed at any time during a live video communication session with device 600-1. For example, options menu 608 may be displayed and / or removed from the display in response to one or more user inputs and / or the detection of a period of inactivity.
[0237] Although detecting an input directed at a surface has been described as causing device 600-1 to display a representation including a modified image of the surface (e.g., in response to detecting input 612c of FIG. 6C, device 600-1 displays representation 624-1 as shown in FIG. 6M), in some embodiments, detecting an input directed at a surface may cause device 600-1 to enter a preview mode (e.g., FIGS. 6H-6J), for example, before displaying representation 624-1.
[0238] 6H illustrates an example of device 600-1 operating in preview mode. In general, preview mode can be used to selectively provide portions or regions of an image of a representation to one or more other users during a live video communication session.
[0239] In some embodiments, before operating in preview mode, device 600-1 detects an input (e.g., input 612c) directed at surface viewer affordance 610. In response, device 600-1 initiates preview mode. While operating in preview mode, device 600-1 displays preview interface 647-1. Preview interface 647-1 includes left scroll affordance 634-2, right scroll affordance 634-1, and preview 636.
[0240] In some embodiments, selection of the left scroll affordance causes device 600-1 to change (e.g., replace) preview 636. For example, selection of left scroll affordance 634-2 or right scroll affordance 634-1 causes device 600-1 to cycle through various images (the user's image, the unmodified image of the surface, and / or the modified image of surface 619), allowing the user to select a particular viewpoint to be shared upon exiting preview mode, e.g., in response to detecting input directed at preview 636. Additionally or alternatively, these techniques may be used to cycle through and / or select particular surfaces (e.g., vertical and / or horizontal surfaces) and / or particular portions (e.g., crops or subsets) within the field of view.
[0241] As shown, in some embodiments, preview user interface 674-1 is displayed on device 600-1 but not on device 600-2. For example, device 600-2 displays videoconferencing interface 604-2 (including representation 622-2), and device 600-1 displays preview interface 674-1. Thus, preview user interface 674-1 allows user 622 to select a view before sharing the view with user 623.
[0242] FIG. 6I shows an example in which device 600-1 is operating in preview mode. As shown, while device 600-1 is operating in preview mode, device 600-1 displays preview interface 674-2. In some embodiments, preview interface 674-2 includes representation 676 having regions 636-1, 636-2. In some embodiments, representation 676 includes an image that is the same as or substantially similar to the image included in representation 622-1. Optionally, as shown, the size of representation 676 is larger than representation 622-1 in FIG. 6A . The position of representation 676 differs from the position of representation 622-1. Adjusting the size and / or position of the representation in preview interface 674-2 compared to the size and / or position of a representation containing a similar or identical image in videoconferencing interface 604-1 allows user 622 to better view the image before sharing it with user 623.
[0243] In some embodiments, region 636-1 and region 636-2 correspond to respective portions of representation 676. For example, as shown, region 636-1 corresponds to an upper portion of representation 676 (e.g., a portion including the upper body of user 622) and region 636-2 corresponds to a lower portion of representation 676 (e.g., a portion including the lower body of user 622 and / or drawing 618).
[0244] In some embodiments, region 636-1 and region 636-2 are displayed as separate regions (e.g., non-overlapping regions). In some embodiments, region 636-1 and region 636-2 overlap. Additionally or alternatively, one or more graphical objects 638-1 (e.g., lines, boxes, and / or dashes) may distinguish (e.g., visually distinguish) region 636-1 from region 636-2.
[0245] In some embodiments, preview interface 674-2 includes one or more graphical objects that indicate whether a region is active or inactive. In the example of FIG. 6I, preview interface 674-2 includes graphical objects 641 a, 641 b. The appearance (e.g., shape, size, and / or color) of graphical objects 641 a, 641 b indicates whether the respective region is active and / or inactive in some embodiments.
[0246] When active, a region is shared with one or more other users of the live video communication session. For example, referring to FIG. 6I, graphical user interface object 641 indicates that region 636-1 is active. As a result, image data corresponding to region 636-1 is displayed by device 600-2 in representation 622-2. In some examples, device 600-1 shares only the image data for the active region. In some embodiments, device 600-1 shares all the image data and instructs device 600-2 to display an image based only on the portion of the image data that corresponds to active region 636-1.
[0247] While displaying interface 674-2, device 600-1 detects input 612i at a location corresponding to region 636-2. Input 612i is a touch input in some embodiments. In response to detecting input 612i, device 600-1 activates region 636-2. As a result, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, region 636-1 remains active in response to input 612i (e.g., user 623 can see user 622, e.g., in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates region 636-1 in response to input 612i (e.g., user 623 can no longer see user 622, e.g., in representation 622-2).
[0248] Although the example of Figure 6I is described with respect to a preview mode having a representation including two regions 636-1, 636-2, in some embodiments, other numbers of regions can be used. For example, referring to Figure 6J, device 600-1 is operating in a preview mode in which preview interface 674-3 includes representation 676 including regions 636a-636i.
[0249] In some embodiments, multiple regions are active (and / or can be activated). For example, as shown, device 600-1 displays regions 636a-636i, of which regions 636a-f are active. Consequently, device 600-2 displays representation 622-2.
[0250] In some embodiments, device 600-1 rectifies images of surfaces having any type of orientation, including any angle relative to gravity (e.g., between 0 and 90 degrees). For example, device 600-1 can rectify images of surfaces when the surfaces are horizontal (e.g., surfaces that lie in a plane that is within 70-110 degrees of the direction of gravity). As another example, device 600-1 can rectify images of surfaces when the surfaces are vertical (e.g., surfaces that lie in a plane that is up to 30 degrees of the direction of gravity).
[0251] While displaying interface 674-3, electronic device 600-1 detects input 612j at a location corresponding to region 636h. In response to detecting input 612j, device 600-1 activates region 636-2. As a result, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, regions 636a-f remain active in response to input 612j (e.g., user 623 can see user 622, e.g., in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates regions 636a-f in response to input 612j (e.g., user 623 can no longer see user 622, e.g., in representation 622-2).
[0252] 6K-6L illustrate example animations that may be displayed by device 600-1 and / or device 600-2. As described in FIGS. 6A-6I, device 600-1 may display a representation that includes a modified image. In some embodiments, device 600-1 and / or device 600-2 display an animation that transitions between views and / or illustrates modifications to the image over time. The animation may include, for example, panning, rotating, and / or otherwise modifying the image to provide the modified image. Additionally or alternatively, the animation occurs in response to detecting input directed at the surface (e.g., selection of surface view affordance 610, gesture, and / or voice input).
[0253] 6K shows an example animation in which device 600-2 pans and rotates the image of representation 642a. During the animation, the image of representation 642a pans downward to view surface 619 from a more "overhead" perspective. The animation also includes rotating the image of representation 642a so that surface 619 is seen from the perspective of user 622. Although four frames of the animation are shown, the animation may include any number of frames. Optionally, in some embodiments, device 600-1 pans and rotates the image of a representation (e.g., representation 622-1).
[0254] 6L shows an example in which device 600-2 zooms in and rotates the image of representation 642a. During the animation, representation 642a is zoomed in until the desired zoom level is achieved. The animation also includes rotating representation 642a until the image of drawing 618 is directed toward the viewpoint of user 622, as described. Although four frames of animation are shown, the animation may include any number of frames. Optionally, in some embodiments, device 600-1 zooms in and rotates the image of a representation (e.g., representation 622-1).
[0255] 6N-6R show examples where the modified image of the surface is further modified during a live communication session.
[0256] FIG. 6N illustrates an example of a live communication session in which users provide various inputs. For example, while displaying interface 678, device 600-1 detects input 677 corresponding to a rotation of device 600-1. As shown in FIG. 6O, in response to detecting input 677, device 600-1 modifies interface 678 to compensate for the rotation (e.g., of camera 602). As shown in FIG. 6O, device 600-1 arranges representations 623-1 and 624-1 of interface 678 in a vertical configuration. In addition, representation 624-1 is rotated in accordance with the rotation of device 600-1 such that the viewpoint of representation 624-1 maintains the same orientation relative to user 622. In addition, the viewpoint of representation 624-2 maintains the same orientation relative to user 623.
[0257] 6N, in some examples, device 600-1 displays control affordances 648-1, 648-2 that modify the image of representation 624-1. Control affordances 648-1, 648-2 can be displayed in response to one or more inputs corresponding to, for example, selection of an affordance in options menu 608 (e.g., FIG. 6B).
[0258] As shown, in some embodiments, device 600-1 displays representation 624-1 including a modified image of a surface. Rotation affordance 648-1, when selected, causes device 600-1 to rotate the image of representation 624-1. For example, while displaying interface 678, device 600-1 detects input 650a corresponding to selection of rotation affordance 648-1. In response to input 650a, device 600-1 modifies the orientation of the image of representation 624-1 from a first orientation (shown in FIG. 6N) to a second orientation (shown in FIG. 6O). In some embodiments, the image of representation 624-1 is rotated a predetermined amount (e.g., 90 degrees).
[0259] When selected, zoom affordance 648-2 modifies the zoom level of the image of representation 624-1. For example, as shown in FIG. 6N, the image of representation 624-1 is displayed at a first zoom level (e.g., "1X"). While displaying zoom affordance 648-2, device 600-1 detects input 650b corresponding to selection of zoom affordance 648-2. In response to input 650b, device 600-1 modifies the zoom level of the image of representation 624-1 from the first zoom level (e.g., "1X") to a second zoom level (e.g., "2X"), as shown in FIG. 6Q.
[0260] Additionally or alternatively, in some embodiments, videoconferencing interface 604-1 includes an option to display a magnified view of at least a portion of the image of representation 624-1, as shown in FIG. 6R. For example, while displaying representation 624-1, device 600-1 may detect input 654 (e.g., a gesture directed at a surface and / or object) corresponding to a request to display a magnified view of a portion of the image of representation 624-1. In response to detecting input 654, device 600-1 displays magnified portion 652-1 at a greater zoom level than second portion 652-2 of representation 624-1. In some embodiments, the portion of the image of representation 624-1 that is magnified is determined based on the location of input 654. In some embodiments, in response to detecting input 650c (FIGS. 6R and 6Q), device 600-1 ceases displaying control affordances 648-1, 648-2.
[0261] 6S-6AC illustrate examples of a device modifying an image of a representation in response to user input. As described in more detail below, device 600-1 can modify an image of a representation (e.g., representation 622-1) in videoconferencing interface 604-1 in response to non-touch user input, including gestures and / or audio input, thereby improving how a user interacts with the device to manage and / or modify the representation during a live video communication session.
[0262] 6S-6T illustrate examples in which a device obscures at least a portion of an image of a representation in response to a gesture. As shown in FIG. 6S, device 600-1 detects gesture 656a corresponding to a request to modify at least a portion of an image of representation 622-1. In some examples, gesture 656a is a gesture in which user 622 points upward near user 622's mouth (e.g., a "shhh!" gesture). In response, as shown in FIG. 6T, device 600-1 replaces representation 622-1 with representation 622-1', which includes a modified image that includes modified portion 658-1 (e.g., the background of user 622's physical environment). In some examples, modifying portion 658-1 in this manner includes blurring, graying, or otherwise obscuring portion 658-1. In some examples, device 600-1 does not modify portion 658-2 in response to gesture 656a.
[0263] 6U-6V illustrate examples in which a device may magnify a portion of an image of a representation in response to detecting a gesture. As shown in FIG. 6U, in some embodiments, device 600-1 detects pointing gesture 656b corresponding to a request to magnify at least a portion of representation 622-1. As shown, pointing gesture 656b is directed toward object 660.
[0264] 6V, in response to pointing gesture 656b, device 600-1 replaces representation 622-1 with representation 622-1′ including the modified image by magnifying a portion of the image of representation 622-1 including object 660. In some embodiments, the magnification is based on the location of object 660 (e.g., relative to camera 602) and / or the size of object 660.
[0265] 6W-6X show examples in which a device magnifies a portion of a view of a representation in response to detecting a gesture. As shown in FIG. 6W, in some embodiments, device 600-1 detects framing gesture 656c corresponding to a request to magnify at least a portion of representation 622-1. As shown, framing gesture 656c is directed toward object 660 such that framing gesture 656c at least partially frames, surrounds, and / or outlines object 660.
[0266] 6X, in response to framing gesture 656c, device 600-1 modifies the image of representation 622-1 by magnifying a portion of the image of representation 622-1 that includes object 660. In some embodiments, the magnification is based on the location of object 660 (e.g., relative to camera 602) and / or the size of object 660. Additionally or alternatively, after magnifying the portion of the image of representation 622-1, device 600-1 may track the movement of framing gesture 656c. In response, device 600-1 may pan to different portions of the image.
[0267] 6Y-6Z illustrate examples of a device panning an image of a representation in response to detecting a gesture. As shown in FIG. 6Y, device 600-1 detects pointing gesture 656d, which corresponds to a request to pan the view of the image of representation 622-1 in a particular direction (e.g., pan horizontally). As shown, pointing gesture 656d is directed to the left of user 622.
[0268] As shown in FIG. 6Z, in response to pointing gesture 656d, device 600-1 replaces representation 622-1 with representation 622-1' that includes a modified image based on panning the image of representation 622-1 in the direction of pointing gesture 656d (e.g., to the left of user 622).
[0269] In some embodiments, as shown in FIG. 6Z, a portion of user 622 (e.g., user 622's right shoulder) may be excluded from the image of representation 622-1' due to the panning action, but in some embodiments, device 600-1 may adjust the zoom level of the image of representation 622-1' when panning to ensure that user 622 remains entirely within the image.
[0270] 6AA-6AB show examples of devices modifying the zoom level of a representation in response to detecting pinch and / or spread gestures. As shown in FIG. 6AA, in some embodiments, device 600-1 detects spread gesture 656e in which user 622 increases the distance between the thumb and index finger of user 622's right hand.
[0271] 6AB, in response to spread gesture 656e, device 600-1 replaces representation 622-1 with 622-1′ by magnifying a portion of the image of representation 622-1. In some embodiments, the magnification is based on the location of spread gesture 656e (e.g., relative to camera 602) and / or the magnitude of spread gesture 656e. In some embodiments, the portion of the image is magnified according to a predetermined zoom level.
[0272] 6AA, in some embodiments, in response to detecting spread gesture 656e, device 600-1 displays zoom indicator 662 that indicates the zoom level of the image of representation 622-1′. Once user 622 completes spread gesture 656e and device 600-1 enlarges a portion of representation 622-1′, device 600-1 updates the display of zoom indicator 662 to indicate the current zoom level of the image of representation 622-1′. In some embodiments, zoom indicator 662 is dynamically updated as user 622 performs gesture 656e.
[0273] Although this specification describes increasing the zoom level of an image in response to a spread gesture 656e, in some implementations, the zoom level of an image is decreased in response to a gesture (e.g., another type of gesture, such as a pinch gesture).
[0274] FIG. 6AC illustrates various gestures that can be used to modify the image of the representation. In some embodiments, for example, a user can indicate a zoom level using a gesture. By way of example, gesture 664 can be used to indicate that the zoom level of the image of the representation should be “1X,” and in response to detecting gesture 664, device 600-1 can modify the image of the representation to have the “1X” zoom level. Similarly, gesture 666 can be used to indicate that the zoom level of the image of the representation should be “2X,” and in response to detecting gesture 666, device 600-1 can modify the image of the representation to have the “2X” zoom level. Although two zoom levels (e.g., “1X” and “2X” zoom levels) are described with respect to FIG. 6AC, in some embodiments, device 600-1 can modify the image of the representation to other zoom levels (e.g., 0.5X, 3X, 5X, or 10X) using the same or different gestures. In some embodiments, device 600-1 can modify the image of the representation at three or more different zoom levels, hi some embodiments, the zoom levels are discrete or continuous.
[0275] As another example, user 622 can use a finger curling gesture to adjust the zoom level. For example, gesture 668 (e.g., a gesture in which the fingers of the user's hand are curled in a direction 668b away from the camera when the back of the hand 668a is oriented toward the camera) can be used to indicate that the zoom level of the image should be increased (e.g., zoomed in). Gesture 670 (e.g., a gesture in which the fingers of the user's hand are curled in a direction 670b toward the camera when the palm of the hand 668a is oriented toward the camera) can be used to indicate that the zoom level of the image should be decreased (e.g., zoomed out).
[0276] 6AD-6AE illustrate an example in which a user participates in a live video communication session using two devices.
[0277] As an example, as shown in Figure 6AD, user 623 is using additional device 600-3 during a live video communication session. In some embodiments, devices 600-2 and 600-3 simultaneously display representations that include images with different views. For example, device 600-2 displays representation 624-2 while device 600-3 displays representation 622-2.
[0278] In some embodiments, device 600-2 is positioned in front of user 622 on desk 686 to correspond to the position of surface 619 relative to user 623. User 623 can therefore see representation 624-2 (including an image of surface 619) in the same way that user 622 sees surface 619 in the physical environment.
[0279] 6AE, during a live communication session, user 623 can modify the image displayed in representation 624-2 by moving device 600-2. In response to user 623 changing the orientation of device 600-2, device 600-2, for example, modifies the image of representation 624-2 in a manner corresponding to the change in orientation of device 600-2. For example, in response to user 623 tilting device 600-2, device 600-2 pans upward to view other portions of surface 619. In this manner, user 623 can change the orientation of device 600-2 (in any direction) to view different portions of surface 619 that are not displayed when device 600-2 is in a different orientation.
[0280] FIGS. 6AF-6AL illustrate embodiments for accessing the various user interfaces shown and described with reference to FIGS. 6A-6AE. In the embodiments shown in FIGS. 6AF-6AL, the interfaces are shown using laptops (e.g., John's device 6100-1 and / or Jane's device 6100-2). It should be understood that the embodiments shown in FIGS. 6AF-6AL may be implemented using different devices, such as tablets (e.g., John's tablet 600-1 and / or Jane's device 600-2). Similarly, the embodiments shown in FIGS. 6A-6AE may be implemented using different devices, such as John's device 6100-1 and / or Jane's device 6100-2. Accordingly, various operations or functions described above with respect to FIGS. 6A-6AE will not be repeated below for the sake of brevity. For example, the applications, interfaces (e.g., 604-1 and / or 604-2), and display elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, and / or 624-2) described with respect to Figures 6A-6AE are similar to the applications, interfaces (e.g., 6121 and / or 6131), and display elements (e.g., 6124, 6132, 6122, 6134, 6116, 6140, and / or 6142) described with respect to Figures 6AF-6AL. Accordingly, details of these applications, interfaces, and displayed elements may not be repeated below for the sake of brevity.
[0281] FIG. 6AF shows John's device 6100-1, which includes a display 6101, one or more cameras 6102, and a keyboard 6103 (which, in some embodiments, includes a trackpad). John's device 6100-1 displays a home screen via display 6101, which includes a camera application icon 6108 and a video conferencing application icon 6110. Camera application icon 6108 corresponds to a camera application running on John's device 6100-1 that can be used to access camera 6102. Video conferencing application icon 6110 corresponds to a video conferencing application running on John's device 6100-1 that can be used to initiate and / or join a live video communication session (e.g., a video call and / or a video chat) similar to those described above with reference to FIGS. 6A-6AE. John's device 6100-1 also displays a dock 6104 that includes various application icons, including a subset of the icons displayed in dynamic area 6106. The icons displayed in dynamic area 6106 represent applications that are active (e.g., launched, open, and / or in use) on John's device 6100-1. In Figure 6AF, neither the camera application nor the video conferencing application is currently active. Therefore, no icons representing the camera application or the video conferencing application are displayed in dynamic area 6106, and John's device 6100-1 is not participating in a live video communication session.
[0282] In Figure 6AF, John's device 6100-1 detects input (e.g., cursor input caused by clicking a mouse, tapping a trackpad, and / or other such input) indicated by cursor 6112 selecting camera application icon 6108. In response, John's device 6100-1 launches the camera application and displays camera application window 6114, as shown in Figure 6AG. In the embodiment shown in Figure 6AG, the camera application is used to access camera 6102 to generate a surface view 6116, for example, similar to representation 624-1 shown in Figure 6M and described above. In some embodiments, the camera application can have different modes (e.g., user-selectable modes), such as, for example, an extended field-of-view mode (providing an extended field of view of camera 6102) and a surface view mode (providing the surface view illustrated in Figure 6AG). 6AG. Additionally, because John's laptop has launched the camera application, camera application icon 6108-1 is displayed in dynamic area 6106 of dock 6104 to indicate that the camera application is active. In some embodiments, the application icon (e.g., 6108-1) is displayed with an animation effect (e.g., bouncing) when it is added to the dynamic area of the dock.
[0283] 6AG, John's device 6100-1 detects input 6118 selecting videoconferencing application icon 6110. In response, John's device 6100-1 launches the videoconferencing application, displays videoconferencing application icon 6110-1 in dynamic region 6106, and displays videoconferencing application window 6120, as shown in FIG. 6AH. Videoconferencing application window 6120 includes videoconferencing interface 6121, which is similar to interface 604-1, and includes Jane's video feed 6122 (similar to representation 623-1) and John's video feed 6124 (similar to representation 622-1). In some embodiments, John's device 6100-1 displays videoconferencing application window 6120 with videoconferencing interface 6121 after detecting one or more additional inputs after input 6118. For example, such inputs may be inputs to initiate a video call with Jane's laptop or inputs to accept a request to join a video call with Jane's laptop.
[0284] 6AH , John's device 6100-1 displays videoconferencing application window 6120 partially overlapping camera application window 6114. In some embodiments, John's device 6100-1 can move camera application window 6114 to the front or foreground (e.g., partially overlapping videoconferencing application window 6120) in response to selecting camera application icon 6108, selecting icon 6108-1, and / or detecting input on camera application window 6114. Similarly, videoconferencing application window 6120 can move to the front or foreground (e.g., partially overlapping camera application window 6114) in response to selecting videoconferencing application icon 6110, selecting icon 6110-1, and / or detecting input on videoconferencing application window 6120.
[0285] 6AH, John's device 6100-1 is shown participating in a live video communication session with Jane's device 6100-2. Accordingly, Jane's device 6100-2 is shown displaying video conferencing application window 6130 similar to video conferencing application window 6120 on John's device 6100-1. Video conferencing application window 6130 includes video conferencing interface 6131 similar to interface 604-2, and includes John's video feed 6132 (similar to representation 622-2) and Jane's video feed 6134 (similar to representation 623-2).
[0286] 6AH, a video conferencing application is being used to access camera 6102 to generate video feed 6124 and video feed 6132. Video feeds 6124 and 6132 therefore represent views of image data that are acquired using camera 6102 and modified (e.g., enlarged and / or cropped) by the video conferencing application to generate the images (e.g., video) shown in video feed 6124 and video feed 6132. In some embodiments, the camera application and the video conferencing application may use different cameras to provide their respective video feeds.
[0287] Video conferencing application window 6120 includes menu option 6126 that can be selected to display different options for sharing content in the live video communication session. In FIG. 6AH, John's device 6100-1 detects input 6128 selecting menu option 6126 and, in response, displays share menu 6136, as shown in FIG. 6AI. Share menu 6136 includes options 6136-1, 6136-2, and 6136-3. Share option 6136-1 is an option that can be selected to share content from a camera application. Share option 6136-2 is an option that can be selected to share content from John's device 6100-1's desktop. Share option 6136-3 is an option that can be selected to share content from a presentation application. In response to detecting input 6138 on share option 6136-1, John's device 6100-1 begins sharing content from the camera application, as shown in FIGS. 6AJ and 6AK.
[0288] In Figure 6AJ, John's device 6100-1 updates video conferencing interface 6121 to include surface view 6140, which is shared with Jane's device 6100-2 in the live video communication session. In the embodiment shown in Figure 6AJ, John's device 6100-1 shares a video feed generated using a camera application (shown as surface view 6116 in camera application window 6114) and displays a representation of the video feed as surface view 6140 in video conferencing application window 6120. In addition, John's laptop enhances the display of surface view 6140 in video conferencing interface 6121 (e.g., by displaying the surface view at a larger size than the other video feeds) and reduces the display size of Jane's video feed 6122. In Figure 6AJ, John's device 6100-1 displays surface view 6140 simultaneously with John's video feed 6124 and Jane's video feed 6122 in video conferencing application window 6120. In some embodiments, the display of John's video feed 6124 and / or Jane's video feed 6122 in the videoconferencing application window 6120 is optional.
[0289] Jane's device 6100-2 updates video conference interface 6131 to show surface video feed 6142, which is the surface view (from a camera application) shared by John's device 6100-1. As shown in FIG. 6AJ, Jane's device 6100-2 adds surface video feed 6142 to video conference interface 6131, showing the surface video feed simultaneously with Jane's video feed 6134 and John's video feed 6132, optionally resized to accommodate the addition of surface video feed 6142. In some embodiments, Jane's device 6100-2 replaces John's video feed 6132 and / or Jane's video feed 6134 with surface video feed 6142.
[0290] 6AK shows an alternative embodiment illustrating sharing content from a camera application in response to detecting input 6138 on share option 6136-1. In FIG. 6AK, John's laptop displays camera application window 6114 along with surface view 6116 (optionally minimizing or hiding videoconferencing application window 6120). John's device 6100-1 also displays John's video feed 6115 (similar to John's video feed 6124) and Jane's video feed 6117 (similar to Jane's video feed 6122), indicating that John's laptop is sharing surface view 6116 with Jane's device 6100-2 in a live video communication session (e.g., a video chat provided by a videoconferencing application). In some embodiments, the display of John's video feed 6115 and / or Jane's video feed 6117 is optional. Similar to the embodiment shown in FIG. 6AJ, Jane's device 6100-2 shows a surface video feed 6142, which is a surface view (from a camera application) shared by John's device 6100-1.
[0291] Figure 6AL shows a schematic diagram illustrating the field of view of camera 6102 and the portions of the field of view used for video conferencing and camera applications for the embodiment shown in Figures 6AF-6AK. For example, in Figure 6AL, a profile view of John's laptop 6100 is shown in John's physical environment. Dashed lines 6145-1 and dotted lines 6147-2 represent the outer dimensions of the field of view of camera 6102, which in some embodiments is a wide-angle camera. The collective fields of view of cameras 6102 are indicated by shaded regions 6144, 6146, and 6148. The portions of the camera fields of view used for camera applications (e.g., for surface view 6116) are indicated by dotted lines 6147-1 and 6147-2 and shaded regions 6146 and 6148. In other words, surface view 6116 (and surface view 6140) is generated by a camera application using the portions of the camera's field of view represented by shaded regions 6146 and 6148 between dotted lines 6147-1 and 6147-2. The portions of the camera's field of view used for a video conferencing application (e.g., John's video feed 6124) are indicated by dashed lines 6145-1 and 6145-2 and shaded regions 6144 and 6146. In other words, John's video feed 6124 is generated by a video conferencing application using the portions of the camera's field of view represented by shaded regions 6144 and 6146 between dashed lines 6145-1 and 6145-2. Shaded region 6146 represents the overlap of the portions of the camera's field of view used to generate the video feeds for each camera and the video conferencing application.
[0292] 6AM-6AY illustrate embodiments for controlling and / or interacting with the various user interfaces and views shown and described with reference to FIGS. 6A-6AL. In the embodiments shown in FIGS. 6AM-6AY, the interfaces are shown using a tablet (e.g., John's tablet 600-1 and / or Jane's device 600-2) and a computer (e.g., Jane's computer 600-4). The embodiments shown in FIGS. 6AM-6AY are, optionally, implemented using a different device, such as a laptop (e.g., John's device 6100-1 and / or Jane's device 6100-2). Similarly, the embodiments shown in FIGS. 6A-6AL are, optionally, implemented using a different device, such as Jane's computer 6100-2. Accordingly, the various operations or functions described above with respect to FIGS. 6A-6AL will not be repeated below for the sake of brevity.
[0293] Additionally, the applications, interfaces (e.g., 604-1, 604-2, 6121, and / or 6131), and fields of view (e.g., 620, 688, 6145-1, and 6147-2) provided by one or more cameras (e.g., 602, 682, and / or 6102) described with respect to Figures 6A-6AL are similar to the applications, interfaces (e.g., 604-4), and fields of view (e.g., 620) provided by a camera (e.g., 602) described with respect to Figures 6AM-6AY. Accordingly, details of these applications, interfaces, and fields of view may not be repeated below for the sake of brevity. Additionally, options and requests (e.g., inputs and / or hand gestures) detected by device 600-1 to control views associated with display elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, 624-2, 6121, and / or 6131) described with respect to Figures 6A-6AL are optionally detected by device 600-2 and / or device 600-4 to control views associated with display elements (e.g., 622-1, 622-4, 623-1, 623-4, 6214, and / or 6216) described with respect to Figures 6AM-6AY (e.g., user 623 optionally provides input that causes device 600-1 and / or device 600-2 to provide representation 624-1 including a modified image of the surface). Additionally, devices 600-1 and 600-2 in Figures 6AM-6AY are described and illustrated as being in landscape orientation. In some embodiments, device 600-1 and / or device 600-2, like device 600-1 in Figure 6O, are in portrait orientation. Accordingly, details of these options and requests detected by device 600-2 may not be repeated below for the sake of brevity.
[0294] 6AM-6AJ show and describe exemplary user interfaces for controlling views of a physical environment. The user interfaces in these figures are used to describe processes described below, including the process in FIG. 15. In FIG. 6AM, device 600-1 and device 600-4 display interfaces 604-1 and 604-4, respectively. Interface 604-1 includes representation 622-1, and interface 604-4 includes representation 622-4. Representations 622-1 and 622-4 include images of image data from a portion of field of view 620, specifically shaded region 6206. As shown, representations 622-1 and 622-4 include images of the head and upper body of user 622 and do not include an image of drawing 618 on desk 621. Interfaces 604-1 and 604-4 include representations 623-1 and 623-4, respectively, that include an image of user 223 within field of view 6204 of camera 6202. Interfaces 604-1 and 604-4 further include an options menu 609 (similar to options menu 608 described with respect to Figures 6A-6AE, including Figures 6F-6G, for controlling image data captured by 602 and / or captured by camera 6202) that allows devices 600-1 and 600-4 to manage how image data is displayed.
[0295] 6AN, user 623 brings device 600-2 near device 600-4 during a live video communication session. As shown, in response to detecting device 600-2 (e.g., via wireless communication), device 600-4 displays additional notification 6210a. Similarly, in response to detecting device 600-4, device 600-2 displays additional notification 6210b via display 683 (e.g., a touch-sensitive display). In some embodiments, devices 600-2 and 600-4 use specific device criteria to trigger the display of additional notifications 6210a and 6210b. In some embodiments, the specific device criteria include criteria for a specific position (e.g., location, orientation, and / or angle) of device 600-2 that, when met, triggers the display of additional notifications 6210a and / or 6210b. In such embodiments, the particular position (e.g., location, orientation, and / or angle) of device 600-2 includes criteria that device 600-2 is at a particular angle or within a certain range of angles (e.g., an angle or range of angles indicative of the device being horizontal and / or lying flat on desk 686) and / or that display 683 is facing up (e.g., as opposed to facing down toward desk 686). In some embodiments, the particular device criteria include criteria that device 600-2 is near device 600-4 (e.g., within a threshold distance of device 600-4). In some embodiments, device 600-2 communicates wirelessly with device 600-4 to communicate the location and / or proximity of device 600-2 (e.g., using location data and / or short-range wireless communication such as Bluetooth and / or NFC). In some embodiments, the specific device criteria include that device 600-2 and device 600-4 are associated with the same user (eg, used by and / or logged in by the same user).In some embodiments, the particular device criteria include that device 600-2 has a particular state (e.g., unlocked and / or display powered on as opposed to locked and / or display powered off).
[0296] 6AN, connection notifications 6210a-6210b include an indication to include device 600-4 in the live video communication session. For example, additional notifications 6210a-6210b include an indication to add a representation including an image of field of view 620 captured by camera 602 for display on device 600-2. In some embodiments, additional notifications 6210a-6210b include an indication to add a representation including an image of field of view 6204 captured by camera 6202 for display on device 600-1.
[0297] 6AN, add notifications 6210a and 6210b include acceptance affordances 6212a and 6212b that, when selected, add (e.g., connect) device 600-2 to the live video communication session. Notifications 6210a and 6210b include rejection affordances 6213a and 6213b that, when selected, dismiss notifications 6210a and 6210b, respectively, without adding device 600-2 to the live video communication session. While displaying acceptance affordance 6212b, device 600-2 detects input 6250an (e.g., a tap, mouse click, or other selection input) directed at acceptance affordance 6212b. In response to detecting input 6250an, device 600-2 displays interface 604-2, as shown in FIG. 6AO.
[0298] In Figure 6AO, interface 604-2 is similar to interface 604-2 described herein (e.g., with reference to Figures 6A-6AE) and video conferencing interface 6131 described herein (e.g., with reference to Figures 6AH-6AK), but has different states. For example, interface 604-2 of Figure 6AO does not include representations 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and options menu 609. In some embodiments, interface 604-2 of Figure 6AO includes one or more of representations 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and / or options menu 609.
[0299] 6AO, interface 604-2 includes adjustable view 6214 of the video feed captured by camera 602 (similar to John's video feed 6132 and representation 622-2, but including a different portion of field of view 620). Adjustable view 6214 is associated with a portion of field of view 620 corresponding to shaded region 6217. In some embodiments, interface 604-2 of FIG. 6AO includes representations 622-4 and 623-4 and / or options menu 609. In some embodiments, representations 622-4 and 623-4 and / or options menu 609 are moved from interface 604-4 to interface 604-2 in response to input detected at device 600-2 and / or device 600-4 so as to be displayed simultaneously with adjustable view 6214. In such embodiments, display 6201 functions as a secondary display (e.g., an extended display) for display 604-1, and / or vice versa.
[0300] In Figure 6AO, in response to detecting input 6250an of Figure 6AN, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown in Figure 6AO. Interface 604-1 of Figure 6AO is similar to interface 604-1 of Figure 6AN, but has a different state (e.g., representations 623-1 and 622-1 are smaller in size and in different positions). Interface 604-1 includes adjustable view 6216 similar to adjustable view 6214 displayed on device 600-2 (e.g., adjustable view 6216 is associated with a portion of field of view 620 corresponding to shaded region 6217). Adjustable view 6216 is updated to include an image similar to adjustable view 6214 when an input described herein (e.g., movement of device 600-2) is detected by device 600-2. Displaying adjustable view 6216 allows user 624 to see which portion of field of view 620 user 622 is currently looking at, as user 623 optionally controls which view within field of view 620 is displayed, as described in more detail below.
[0301] In Figure 6AO, while displaying interface 604-2, device 600-2 detects movement 6218ao of device 600-2. In response to detecting movement 6218ao, device 600-2 displays interface 602-4 of Figure 6AP. Additionally, in response to detecting movement 6218ao, device 600-2 causes device 600-1 to display interface 604-1 of Figure 6AP.
[0302] In FIG. 6AP, interface 602-4 includes an updated adjustable view 6214. The adjustable view 6214 in FIG. 6AP is a different view within field of view 620 compared to the adjustable view 6214 in FIG. 6AO. For example, the shaded region 6217 in FIG. 6AP has moved relative to the shaded region 6217 in FIG. 6AO. Note that camera 602 has not moved. In some embodiments, the movement 6218ao of device 600-2 corresponds to (e.g., is proportional to) the amount of change in adjustable view 6214. For example, in some embodiments, the angular magnitude through which device 600-2 rotates (e.g., relative to gravity) corresponds to the amount of change in adjustable view 6214 (e.g., the amount by which image data is panned to include the new angle of view). In some embodiments, the direction (e.g., tilting downward and / or rotating downward) of the movement (e.g., movement 6218ao) of device 600-2 corresponds to the direction of change in adjustable view 6214 (e.g., panning downward). In some embodiments, the acceleration and / or velocity of the movement (e.g., movement 6218) corresponds to the rate at which the adjustable view 6214 changes. In some embodiments, device 600-2 (and / or device 600-1) displays a gradual transition (e.g., a series of views) from the adjustable view 6214 of FIG. 6AO to the adjustable view 6214 of FIG. 6AP. Additionally or alternatively, as shown in FIG. 6AP, device 600-2 is positioned flat on desk 686. In some embodiments, in response to detecting a particular position or a position within a predetermined range of positions (e.g., horizontal and / or display up), device 600-2 displays the adjustable view 6214 of FIG. 6AP. As shown, movement 6218ao in FIG. 6AO does not cause device 600-2 to update representations 622-4 and 623-4 (and / or representations 623-1 and 622-1 on device 600-1) in FIG. 6AP.
[0303] 6AP, the image of drawing 618 in adjustable view 6214 is at a perspective that is different from the perspective of the image of drawing 618 in adjustable view 6214 in FIG. 6AO. For example, adjustable view 6214 in FIG. 6AP includes a top view perspective, while adjustable view 6214 in FIG. 6AO includes a perspective that includes a combination of a side view and a top view. In some embodiments, the image of the drawing included in adjustable view 6214 in FIG. 6AP is based on image data that has been corrected (e.g., skewed and / or scaled) using similar techniques described with reference to FIGS. 6A-6AL. In some embodiments, the image of the drawing included in adjustable view 6214 in FIG. 6AO is based on image data that has not been corrected (e.g., skewed and / or scaled) and / or has been corrected differently (e.g., to a lesser extent) than the image of drawing 618 in adjustable view 6214 in FIG. 6AP (e.g., less skew and / or less scaled compared to the amount of skew and / or scale applied in FIG. 6AP). Providing a top-down perspective makes it easier to collaborate and share content because it gives the user 622 a view of the drawing that is similar to the view the user 623 would have if the user 623 were sitting across from the user 623 looking down at the surface 619 of the desk 621.
[0304] 6AP, adjustable view 6216 of interface 604-1 has been similarly updated. In some embodiments, the image in adjustable view 6216 and / or adjustable view 6214 is modified based on the position of surface 619 relative to camera 602, as described with reference to FIGS. 6A-6AL. In such embodiments, device 600-1 and / or device 600-2 rotates the image in adjustable view 6214 by an amount (e.g., 45 degrees, 90 degrees, or 180 degrees) so that the image in drawing 618 can be viewed more intuitively in adjustable view 6216 and / or adjustable view 6214 (e.g., the image in drawing 618 shows the house right-side up, rather than upside down).
[0305] In Figure 6AP, user 623 applies numeric marks to adjustable view 6214 using stylist 6220. For example, while displaying adjustable view 6214 of Figure 6AP, device 600-2 detects input corresponding to a request to add digital marks to adjustable view 6214 (e.g., using stylist 6220). In response to detecting input corresponding to a request to add digital marks to adjustable view 6214, device 600-2 displays interface 604-2, as shown in Figure 6AO. Additionally or alternatively, in response to detecting input corresponding to a request to add digital marks to adjustable view 6214, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown in Figure 6AQ.
[0306] 6AQ, interface 602-4 includes digital sun 6222 within adjustable view 6214, and interface 602-1 includes digital sun 6223 within adjustable view 6214. Displaying the digital sun on both devices enables users 623 and 622 to collaborate throughout the video communication session. Additionally, as shown, digital sun 6222 has a position relative to the image of drawing 618. As described in more detail below, digital sun 6222 maintains its position relative to the image of drawing 618 even if device 600-1 detects further movement and / or drawing 618 moves on surface 619. In some embodiments, device 600-2 stores data corresponding to a relationship between a digital mark (e.g., digital sun 6223) and an object (e.g., a house) detected in the image data to determine where (and / or where) digital sun 6222 should be displayed. In some embodiments, device 600-2 stores data corresponding to a relationship between a digital mark (e.g., digital sun 6223) and the position of device 600-2 to determine where (and / or if) digital sun 6222 should be displayed. In some embodiments, device 600-2 detects digital marks applied to other views within field of view 620. For example, digital marks can be applied within an image of the user's head, such as the image of the user's head in adjustable view 6214 of FIG. 6AR.
[0307] 6AQ, interface 604-2 includes control affordance 648-1 (similar to control affordance 648-1 in FIG. 6N) that modifies the image in adjustable view 6214. Rotation affordance 648-1, when selected, causes device 600-1 (and / or device 600-2) to rotate the image in adjustable view 6214, similar to how control affordance 648-1 modifies the image in representation 624-1 in FIG. 6N.
[0308] In Figure 6AQ, in some embodiments, interface 604-2 includes a zoom affordance similar to zoom affordance 648-2 of Figure 6N. In such embodiments, the zoom affordance modifies the image in adjustable view 6214 in a manner similar to how zoom affordance 648-2 modifies the image of representation 624-1 of Figure 6N. Control affordances 648-1, 648-2 can be displayed in response to one or more inputs corresponding, for example, to selection of an affordance in options menu 609 (e.g., Figure 6AM).
[0309] 6AQ, in some embodiments, the digital sun 6222 is projected onto the physical surface of the drawing 618, similar to how the markup 956 described in FIGS. 9K-9N is projected onto the surface 908b. In such embodiments, an electronic device (e.g., a projector and / or a light projector) is used to project an image and / or rendering of the digital sun 6222 within the physical environment 915. For example, the electronic device may use the techniques described with respect to FIGS. 9K-9N to display a projection of the digital sun next to the drawing 618 based on the relative location of the digital sun 6222 with respect to the drawing 618.
[0310] In Figure 6AQ, device 600-2 detects movement 6218aq (e.g., rotation and / or lifting) while displaying digital sun 6222 within adjustable view 6214. In response to detecting movement 6218aq, device 600-2 displays interface 604-2, as shown in Figure 6AR. In response to detecting movement 6218aq, device 600-1 displays (and / or causes device 600-2 to display) interface 604-1, as shown in Figure 6AR.
[0311] In FIG. 6AR, interface 604-2 includes updated adjustable view 6214 (which corresponds to updated adjustable view 6216 in interface 606-1). Adjustable view 6214 in FIG. 6AR is a different view within field of view 620 compared to adjustable view 6214 in FIG. 6AQ. For example, shaded region 6217 in FIG. 6AP has moved relative to shaded region 6217 in FIG. 6AQ. In some embodiments, the direction of movement 6218aq (e.g., tilting upward) corresponds to the direction of the change of view (e.g., panning upward). Furthermore, shaded region 6217 overlaps with shaded region 6206, as indicated by darker shaded region 6224. Darker shaded region 6224 is a schematic representation in which updated adjustable view 6214 is based on a portion of the image data used for representation 622-4. Because movement 6218aq resulted in a change of view (e.g., a change to the face of user 622 and / or a change other than the view of drawing 618), device 600-2 no longer displays digital sun 6222 in adjustable view 6214.
[0312] In FIG. 6AR, adjustable view 6214 includes boundary indicator 6226. Boundary indicator 6226 indicates that a boundary has been reached. In some embodiments, the boundary is configured (e.g., by a user) to set limits on which portions of field of view 620 (or the environment captured by camera 602) are provided for display. For example, user 622 can limit which portions are available to user 623. In some embodiments, the boundary is defined by the physical limits (e.g., image sensor and / or lens) of camera 602 providing field of view 620. In FIG. 6AR, shaded region 6217 does not reach the limits of field of view 620. Thus, boundary indicator 6226 is based on a configurable setting that limits which portions of field of view 620 are provided for display. Referring briefly to FIG. 6AT, boundary indicator 6226 is displayed in response to a determination that the viewpoint provided in adjustable view 6214 has reached the edge of field of view 620.
[0313] In FIG. 6AR, boundary indicator 6226 is shown with cross-hatching. In some embodiments, security boundary indicator 6226 is a visual effect (e.g., blur and / or fade) applied to adjustable view 6214 (and / or adjustable view 6216). In some embodiments, boundary indicator 6226 is displayed along the edge of adjustable view 6214 (and / or 6216) to indicate the position of the boundary. In FIG. 6AR, boundary indicator 6226 is displayed along the top and side edges to indicate that the user cannot see above and / or to the sides of boundary indicator 6226. While displaying interface 604-2 in FIG. 6AR, device 600-2 detects movement 6218ar (e.g., rotation and / or lowering). In response to detecting movement 6218ar, device 600-2 displays interface 604-2, as shown in FIG. 6AS. In response to detecting movement 6218ar, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as shown in FIG. 6AS.
[0314] In FIG. 6AS, interface 604-2 includes an updated adjustable view 6214 that includes an image of drawing 618. In FIG. 6AS, device 600-2 is in a position similar to that in FIG. 6AO. Thus, adjustable view 6214 in FIG. 6AS includes a perspective of the image of drawing 618 in adjustable view 6214 that is the same as the perspective of the image of drawing 618 in adjustable view 6214 in FIG. 6AO. In particular, device 600-2 displays digital sun 6222 in adjustable view 6214 in FIG. 6AS. The position of digital sun 6222 relative to the house in drawing 618 in FIG. 6AS is similar to the position of digital sun 6222 relative to the house in drawing 618 in FIG. 6AQ, except for minor differences based on the different views. Thus, digital sun 6222 appears fixed in physical space, as if it were drawn next to drawing 618. Fixing the position of digital marks in physical space facilitates better collaboration between users, as they can digitally draw or write in one view, move the device to see a different view, and then move the device back to re-display the digital drawing or writing and the context in which it was made.
[0315] For clarity, shaded regions 6217 and 6206 and field of view 620 have been omitted from Figures 6AS-6AU. In some embodiments, representation 622-1 and adjustable views 6214 and 6216 correspond to views associated with shaded regions 6217 and 6206 and field of view 620 in Figure 6AO.
[0316] 6AS, device 600-2 (and / or device 600-1) detects movement of drawing 618 and maintains the display of an image of drawing 618 within adjustable view 6214. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to modify (e.g., zoom, skew, and / or rotate) the image data to maintain the display of the image of drawing 618 within adjustable view 6214. While displaying interface 604-2, device 600-2 (and / or device 600-1) detects horizontal movement 6230 of drawing 618. In response to detecting horizontal movement 6230 of drawing 618, device 600-2 displays interface 604-2, as shown in FIG. 6AT. In some embodiments, in response to detecting horizontal movement 6230 of drawing 618, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as shown in Figure 6AT. In some embodiments, in response to device 600-1 detecting horizontal movement 6230 of drawing 618, device 600-2 displays (and / or device 600-1 causes device 600-2 to display) interface 602-4, as shown in Figure 6AT.
[0317] In FIG. 6AT, drawing 618 has been moved to the edge of desk 621, further away (e.g., laterally) from camera 602. Despite the change in position, interface 602-4 in FIG. 6AT includes an image of drawing 618 in adjustable view 6214 that appears largely unchanged from the image of drawing 618 in adjustable view 6214 of interface 602-4 in FIG. 6AS. For example, adjustable view 6214 provides a perspective from which drawing 618 still appears directly in front of camera 602, similar to the position of drawing 618 in FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct the image of drawing 618 (e.g., by skew and / or magnification) based on its new position relative to camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it moves relative to camera 602. In some embodiments, the adjustable view 6214 of the interface 604-2 of FIG. 6AT is provided without any change in the position (eg, location, orientation, and / or rotation) of the camera 602.
[0318] In Figure 6AT, device 600-2 displays boundary indicator 6226 within adjustable view 6214 (similar to adjustable view 6214 displayed in adjustable view 6216 by device 600-1). As discussed above with respect to Figure 6AR, boundary indicator 6226 indicates that the limits of the field of view or physical space have been reached. In Figure 6AT, device 600-2 displays boundary indicator 6226 within adjustable view 6214 to indicate that the edge of field of view 620 has been reached. Boundary indicator 6226 is along the right edge of adjustable view 6214 (and adjustable view 6216), indicating that the view to the right of the current view exceeds the field of view of camera 602.
[0319] In Figure 6AT, digital sun 6222 maintains an individual position relative to the house in the image of drawing 618 in adjustable view 6214, similar to the individual position of digital sun 6222 relative to the house in the image of drawing 618 in adjustable view 6214 in Figure 6AS. In some embodiments, device 600-2 (and / or device 600-1) displays digital sun 6222 overlaid on the image of drawing 618 corrected based on the new position of drawing 618.
[0320] Returning briefly to Figure 6AS, while displaying interface 602-4, device 600-2 (and / or device 600-1) detects rotation 6232 of drawing 618. In response to detecting rotation 6232 of drawing 618, device 600-2 displays interface 604-2, as shown in Figure 6AU. In some embodiments, in response to detecting rotation 6232 of drawing 618, device 600-2 causes device 600-1 to display interface 601-4, as shown in Figure 6AU. In some embodiments, in response to device 600-1 detecting rotation 6232 of drawing 618, device 600-2 displays (or device 600-1 causes device 600-2 to display) interface 602-4, as shown in Figure 6AU.
[0321] In FIG. 6AU, drawing 618 has been rotated relative to the edge of desk 621. Despite the change in position, interface 602-4 of FIG. 6AU includes an image of drawing 618 in adjustable view 6214 that appears largely unchanged from the image of drawing 618 in adjustable view 6214 of interface 604-2 of FIG. 6AS. That is, adjustable view 6214 of FIG. 6AU provides a perspective from which drawing 618 appears as if it were not rotated, similar to the position of drawing 618 in FIG. 6AS. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct the image of drawing 618 based on its new position relative to camera 602 (e.g., by skew and / or rotate). In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it rotates relative to camera 602. In some embodiments, adjustable view 6214 of interface 604-2 of FIG. 6AU is provided without any change in the position (e.g., location, orientation, and / or rotation) of camera 602. Adjustable view 6216 is updated in a similar manner to adjustable view 6214.
[0322] In Figure 6AU, digital sun 6222 maintains a position relative to the house in the image of drawing 618 in adjustable view 6214 similar to the position of digital sun 6222 relative to the house in the image of drawing 618 in adjustable view 6214 in Figure 6AS. In some embodiments, device 600-2 (and / or device 600-1) displays digital sun 6222 overlaid on the image of drawing 618 that has been corrected based on the rotation of drawing 618.
[0323] In FIG. 6AV, device 600-2 displays interface 604-2 that is similar to interface 604-2 of FIG. 6AU, but with a different state (e.g., representation 622-2 of John and options menu 609 have been added to user interface 604-2). Device 600-2 is no longer being used in a live communication session. Furthermore, device 600-2 has been moved from its position in FIG. 6AU to the same position that device 600-2 had in FIG. 6AQ. Accordingly, device 600-2 updates adjustable view 6214 of FIG. 6AV to include the same perspective as adjustable view 6214 of FIG. 6AQ. As shown, adjustable view 6214 includes a top perspective view. Additionally, digital sun 6222 appears to have the same position relative to the house as digital sun 6222 in the image of drawing 618 in adjustable view 6214 of FIG. 6AQ.
[0324] In Figure 6AV, device 600-2 detects movement 6218av (e.g., rotation and / or lifting) while displaying digital sun 6222 within adjustable view 6214. In response to detecting movement 6218av, device 600-2 displays interface 604-2, as shown in Figure 6AW. In response to detecting movement 6218aw, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown in Figure 6AW.
[0325] In FIG. 6AW, interface 604-2 includes an updated adjustable view 6214 similar to adjustable view 6214 of FIG. 6AR (which corresponds to updated adjustable view 6216 in interface 606-1). Notably, the device does not update representation 622-2 in response to detecting movement 6218aw. Thus, in some embodiments, device 600-2 displays a dynamic representation that updates based on the position of device 600-2 and a static representation that does not update based on the position of device 600-2. Interface 604-2 also includes a boundary indicator 6226 within adjustable view 6214 similar to boundary indicator 6226 of FIG. 6AR.
[0326] In Figure 6AW, device 600-2 detects movement 6218aw (e.g., a rotation and / or a downward movement) while displaying interface 604-2. In response to detecting movement 6218aw, device 600-2 displays interface 604-2, as shown in Figure 6AX. In response to detecting movement 6218aw, device 600-1 displays (and / or causes device 600-2 to display) interface 604-1, as shown in Figure 6AX.
[0327] In Figure 6AX, interface 604-2 includes updated adjustable view 6214 (corresponding to updated adjustable view 6216 in interface 606-1). Because adjustable view 6214 is substantially the same as the view provided by representation 622-2, shaded region 6206 overlaps shaded region 6217. Because movement 6218aq changes the view to the face of user 622 and / or does not change the view of drawing 618, device 600-2 no longer displays digital sun 6222 in adjustable view 6214. While displaying interface 604-2 in Figure 6AX, device 600-2 (and / or device 600-1) detects a set of one or more inputs (e.g., similar to the inputs and / or hand gestures described with reference to Figures 6A-6AL) corresponding to a request to display a surface view. In some such embodiments, 616-1 of FIG. 6C , 616-2 of FIG. 6G , preview mode 674-1 of FIG. 6H , representation 676 of preview mode interface 674-2 of FIG. 6I , representation 676 of preview mode interface 674-3 of FIG. 6J , and affordances 648-1, 648-2, and 648-3 of FIG. 6N-6Q are displayed on device 600-2 to enable device 600-2 to control the presentation of the modified image of drawing 618 in the same manner as inputs detected at device 600-1. In response to detecting a set of one or more inputs corresponding to a request to display a surface view, device 600-2 displays interface 604-2 as shown in FIG. 6AY. Additionally or alternatively, in response to detecting a set of one or more inputs, device 600-1 displays interface 604-2 as shown in FIG. 6AY. In some embodiments, device 600-1 detects a set of one or more inputs as described with reference to FIGS. 6A-6AL. In some embodiments, device 600-2 detects a set of one or more inputs, in such embodiments, device 600-2 detects a selection of viewer affordance 6236 of options menu 609, which is similar to viewer affordance 607-2 of options menu 608 described with reference to FIG.Accordingly, a view menu similar to view menu 616-2 described with reference to FIG. 6G includes an affordance that requests the display of a surface view of the remote participant.
[0328] In Figure 6AY, adjustable view 6214 includes a surface view similar to representation 624-1 shown in Figure 6M and described above, for example. As shown in Figure 6AY, adjustable view 6214 includes a modified image so that user 623 has a perspective looking down on the image of drawing 618 displayed on device 600-2 similar to the perspective that user 622 has when looking down on drawing 618 in a physical environment, as described in more detail with respect to Figures 6A-6AL. Notably, digital sun 6222 in Figure 6AY appears to have the same position relative to the house in the image of drawing 618 in adjustable view 6214 as digital sun 6222 in Figure 6AQ.
[0329] 7 is a flow diagram illustrating a method for managing a live video communication session using a computer system, according to some embodiments. Method 700 involves display generation components (e.g., 601, 683, and / or 6101) (e.g., a display controller, a touch-sensitive display system, and / or a monitor), one or more cameras (e.g., 602, 682, and / or 6102) (e.g., an infrared camera, a depth camera, and / or a visible light camera), and one or more input devices (e.g., 601, 683, and / or 6103) (e.g., a touch-sensitive surface, a keyboard, a controller, a 7. The method 700 is performed in a computer system (e.g., 600-1, 600-2, 600-3, 600-4, 906a, 906b, 906c, 906d, 6100-1, 6100-2, 1100a, 1100b, 1100c, and / or 1100d) (e.g., a smartphone, a tablet, a laptop computer, and / or a desktop computer) (e.g., 100, 300, or 500) in communication with a keyboard, a mouse, a display, a display device, a display controller, and / or a mouse. Some operations of method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0330] As described below, method 700 provides an intuitive way to manage a live video communication session. The method reduces the cognitive load on a user managing a live video communication session, thereby creating a more efficient human-machine interface. For battery-operated computing devices, allowing a user to manage a live video communication session more quickly and efficiently conserves power, saving energy and extending the time between battery charges.
[0331] In method 700, a computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2), via a display generation component, displays (702) a live video communication interface (e.g., 604-1, 604-2, 6120, 6121, 6130, and / or 6131) for a live video communication session (e.g., an interface for an incoming and / or outgoing live audio / video communication session). In some embodiments, the live communication session is between at least a computer system (e.g., a first computer system) and a second computer system.
[0332] The live video communication interface includes a representation (e.g., 622-1, 622-2, 6124, and / or 6132) (e.g., a first representation) of at least a portion of a field of view (e.g., 620, 688, 6144, 6146, and / or 6148) of one or more cameras. In some embodiments, the first representation includes an image of the physical environment (e.g., a scene and / or area of the physical environment that is within the field of view of one or more cameras). In some embodiments, the representation includes a portion (e.g., a first crop) of the field of view of one or more cameras. In some embodiments, the representation includes a still image. In some embodiments, the representation includes a series of images (e.g., a video). In some embodiments, the representation includes a live (e.g., real-time) video feed of the field of view (or portion thereof) of one or more cameras. In some embodiments, the field of view is based on physical characteristics of the one or more cameras (e.g., orientation, lens, focal length of the lens, and / or sensor size). In some embodiments, the representation is displayed in a window (e.g., a first window). In some embodiments, the representation of at least a portion of the field of view includes an image of the first user (e.g., the first user's face). In some embodiments, the representation of at least a portion of the field of view is provided by an application (e.g., 6110) that provides the live video communication session. In some embodiments, the representation of at least a portion of the field of view is provided by an application (e.g., 6108) that is different from the application that provides the live video communication session (e.g., 6110).
[0333] While displaying the live video communication interface, the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) may receive user input (e.g., taps on a touch-sensitive surface, keyboard input, mouse input, trackpad input, gestures (e.g., hand gestures), and / or audio) directed via one or more input devices (e.g., 601, 603, and / or 6103) to a surface (e.g., 619) within a scene (e.g., physical environment) within the field of view of one or more cameras. An input (e.g., a voice command) (e.g., 612c, 612d, 614, 612g, 612i, 612j, 6112, 6118, 6128, and / or 6138) (e.g., one or more user inputs including a physical surface, a surface of a desk and / or an object resting on the desk (e.g., a book, paper, tablet), a surface of a wall and / or an object on the wall (e.g., a whiteboard or blackboard), or other surface (e.g., a freestanding whiteboard or blackboard)) is detected (704). In some embodiments, the user input corresponds to a request to display a view of a surface. In some embodiments, detecting the user input via one or more input devices includes acquiring image data of a field of view of one or more cameras including a gesture (e.g., a hand gesture, an eye gesture, or other body gesture). In some embodiments, the computer system determines from the image data that the gesture meets predetermined criteria.
[0334] In response to detecting one or more user inputs, a computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays, via a display generation component (e.g., 601, 683, and / or 6101), a representation (e.g., image and / or video) (e.g., second representation) of a surface (e.g., 624-1, 624-2, 6140, and / or 6142). In some embodiments, the representation of the surface is obtained by digitally zooming and / or panning a field of view captured by one or more cameras. In some embodiments, the representation of the surface is obtained by moving (e.g., translating and / or rotating) one or more cameras. In some embodiments, the second representation is displayed in a window (e.g., a second window, the same window in which the first representation is displayed, or a window different from the window in which the first representation is displayed). In some embodiments, the second window is different from the first window. In some embodiments, the second window (e.g., 6140 and / or 6142) is provided by an application (e.g., 6110) providing the live video communication session (e.g., as shown in FIG. 6AJ). In some embodiments, the second window (e.g., 6114) is provided by an application (e.g., 6108) that is different from the application providing the live video communication session (e.g., as shown in FIG. 6AK). In some embodiments, the second representation includes a crop (e.g., the second crop) of the field of view of one or more cameras. In some embodiments, the second representation is different from the first representation. In some embodiments, the second representation is different from the first representation because the second representation displays a different portion of the field of view (e.g., the second crop) than the portion (e.g., the first crop) displayed in the first representation (e.g., a panned view, a zoomed-out view, and / or a zoomed-in view). In some embodiments, the second representation includes an image of a portion of the scene not included in the first representation, and / or the first representation includes an image of a portion of the scene not included in the second representation, hi some embodiments, the surface is not displayed in the first representation.
[0335] The representation of the surface (e.g., 624-1, 624-2, 6140, and / or 6142) includes images (e.g., photographs, videos, and / or live video feeds) of the surface (e.g., 619) captured by one or more cameras (e.g., 602, 682, and / or 6102) that have been rectified (or corrected) (e.g., adjusted, manipulated, corrected) based on the position (e.g., location and / or orientation) of the surface relative to the one or more cameras (e.g., to correct distortions in the image of the surface) (sometimes referred to as a representation of the rectified image of the surface). In some embodiments, the image of the surface displayed in the second representation is based on image data that has been rectified (e.g., skewed, rotated, flipped, and / or otherwise manipulated) using image processing software (e.g., to skew, rotate, flip, and / or otherwise manipulate the image data captured by the one or more cameras). In some embodiments, the image of the surface displayed in the second representation is modified without physically adjusting the camera (e.g., without rotating the camera, lifting the camera, lowering the camera, adjusting the camera angle, and / or adjusting physical components of the camera (e.g., lens and / or sensor)). In some embodiments, the image of the surface displayed in the second representation is modified so that the camera appears to be pointed at the surface (e.g., facing the surface, pointed at the surface, pointed along an axis perpendicular to the surface). In some embodiments, the image of the surface displayed in the second representation is corrected so that the camera's line of sight appears perpendicular to the surface. In some embodiments, the image of the scene displayed in the first representation is not modified based on the location of the surface relative to one or more cameras. In some embodiments, the representation of the surface is displayed simultaneously with the first representation (e.g., the first representation (e.g., of a user of a computer system) is maintained and the image of the surface is displayed in a separate window). In some embodiments, the image of the surface is automatically modified in real time (e.g., during a live video communication session). In some embodiments, the image of the surface is automatically rectified (eg, without user input) based on the position of the surface relative to the one or more first cameras.Displaying a representation of a surface that includes an image of the surface that is modified based on the position of the surface relative to one or more cameras enhances the video communication session experience by providing a clearer view of the surface regardless of its position relative to the camera without requiring further input from the user, provides improved visual feedback, and reduces the number of inputs required to perform actions.
[0336] In some embodiments, a computer system (e.g., 600-1 and / or 600-2) receives image data captured by a camera (e.g., 602) (e.g., a wide-angle camera) of one or more cameras during a live video communication session. The computer system, via a display generation component, displays a representation (e.g., 622-1 and / or 622-2) (e.g., a first representation) of at least a portion of a field of view based on the image data captured by the camera. The computer system, via a display generation component, displays a representation (e.g., 624-1 and / or 624-2) (e.g., a second representation) of a surface based on the image data captured by the camera (e.g., the representation of at least a portion of a view of the one or more cameras and the representation of the surface are based on image data captured by a single (e.g., only one) camera of the one or more cameras). Displaying the representation of at least a portion of a field of view and the representation of the surface captured from the same camera enhances the video communication session experience and reduces the number of inputs (and / or devices) required to perform operations by displaying content captured by the same camera from different perspectives without requiring input from a user.
[0337] In some embodiments, the image of the surface is rectified (e.g., by a computer system) by rotating the image of the surface relative to the representation of at least a portion of the field of view of one or more cameras (e.g., the image of the surface in 624-2 is rotated 180 degrees relative to representation 622-2). In some embodiments, the representation of the surface is rotated 180 degrees relative to the representation of at least a portion of the field of view of one or more cameras. Rotating the image of the surface relative to the representation of at least a portion of the field of view of one or more cameras enhances the video communication session experience, provides improved visual feedback, and reduces the number of inputs required to perform actions, because content associated with the surface can be viewed from a different perspective than other portions of the field of view without requiring input from the user.
[0338] In some embodiments, the image of the surface is rotated based on the position (e.g., location and / or orientation) of the surface (e.g., 619) relative to the user (e.g., 622) (e.g., the user's position) within the field of view of one or more cameras. In some embodiments, a representation of the user is displayed at a first angle and the image of the surface is rotated to a second angle different from the first angle (e.g., even though the image of the user and the image of the surface are captured at the same camera angle). Rotating the image of the surface based on the position of the surface relative to the user within the field of view of one or more cameras enhances the video communication session experience, provides improved visual feedback, and reduces the number of inputs required to perform actions because content associated with the surface can be viewed from a perspective based on the position of the surface without requiring input from the user.
[0339] In some embodiments, pursuant to a determination that the surface is in a first position (e.g., surface 619 is positioned in front of user 622 on desk 621 of FIG. 6A ) (e.g., a predefined position) (e.g., in front of the user, between the user and one or more cameras, and / or in a substantially horizontal plane) relative to the user within the field of view of one or more cameras, the image of the surface is rotated at least 45 degrees relative to the representation of the user within the field of view of the one or more cameras (e.g., the image of surface 619 in representation 622-1 is rotated 180 degrees relative to representation 624-1 of FIG. 6M ). In some embodiments, the image of the surface is rotated in a range of 160 degrees to 200 degrees (e.g., 180 degrees). In some embodiments, pursuant to a determination that the surface is in a first position (e.g., in front of the user, between the user and one or more cameras, and / or in a substantially horizontal plane) relative to the user within the field of view of one or more cameras, the image of the surface is rotated a first amount. In some embodiments, the first amount is within a range of 160 degrees to 200 degrees (e.g., 180 degrees). In some embodiments, pursuant to a determination that the surface is in a second position relative to the user within the field of view of the one or more cameras (e.g., to the side of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by the second amount. In some embodiments, the second amount is within a range of 45 degrees to 120 degrees (e.g., 90 degrees). Rotating the image of the surface by at least 45 degrees relative to a representation of the user captured within the field of view of the one or more cameras when the surface is in the first position relative to the user enhances the video communication session experience by adjusting the image to provide a more natural and intuitive image without requiring further input from the user, provides improved visual feedback, and performs an action when a set of conditions is met without requiring further user input.
[0340] In some embodiments, the representation of at least a portion of the field of view includes the user and is displayed simultaneously with a representation of the surface (e.g., representations 622-1 and 624-1 or representations 622-2 and 624-2 in FIG. 6M). In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are captured by the same camera (e.g., a single camera of the one or more cameras) and displayed simultaneously. In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are displayed in separate windows that are displayed simultaneously. Including the user in the representation of at least a portion of the field of view and displaying the representation simultaneously with the representation of the surface enhances the video communication session experience by allowing the user to see participants' reactions while the representation of the surface is displayed without requiring further input from the user, provides improved visual feedback, and performs actions when a set of conditions is met without requiring further user input.
[0341] In some embodiments, in response to detecting one or more user inputs, and before displaying the representation of the surface, the computer system displays a preview of image data for the field of view of one or more cameras (e.g., in a preview mode of a live video communication interface, as shown in FIGS. 6H-6J), where the preview includes an image of the surface that has not been rectified based on the position of the surface relative to the one or more cameras (sometimes referred to as a representation of an unrectified image of the surface). In some embodiments, the preview of the field of view is displayed after displaying a representation of the image of the surface (e.g., in response to detecting user input corresponding to a selection of the representation of the surface). Displaying a preview that includes an image of the surface that has not been rectified based on the position of the surface relative to the one or more cameras allows a user to quickly identify the surface in the preview as one that has not had distortion correction applied, and provides improved visual feedback.
[0342] In some embodiments, displaying a preview of the image data for the field of view of the one or more cameras includes displaying a plurality of selectable options (e.g., 636-1 and / or 636-2 of FIG. 6I or 636a-i of FIG. 6J) corresponding to respective portions of (e.g., surfaces within) the field of view of the one or more cameras. In some embodiments, the computer system detects an input (e.g., 612i or 612j) selecting one of the plurality of options corresponding to the respective portions of the field of view of the one or more cameras. In response to detecting the input selecting one of the plurality of options corresponding to the respective portions of the field of view of the one or more cameras, and in accordance with a determination that the input selecting one of the plurality of options corresponding to the respective portions of the field of view of the one or more cameras is directed to a first option corresponding to a first portion of the field of view of the one or more cameras, the computer system displays a representation of the surface based on the first portion of the field of view of the one or more cameras (e.g., selecting 636h of FIG. 6J displays the corresponding portion) (e.g., the computer system optionally displays a modified version of the image of the first portion of the field of view with a first distortion correction). In response to detecting an input selecting one of a plurality of options corresponding to a respective portion of the field of view of the one or more cameras, and in accordance with a determination that the input selecting one of the plurality of options corresponding to a respective portion of the field of view of the one or more cameras is directed to a second option corresponding to a second portion of the field of view of the one or more cameras, the computer system displays a representation of the surface based on the second portion of the field of view of the one or more cameras (e.g., selecting 636g in FIG. 6J displays the corresponding portion) (e.g., the computer system optionally displays a modified version of the image of the second portion of the field of view with a second distortion correction that is different from the first distortion correction), the second option being different from the first option. Displaying a plurality of selectable options corresponding to a respective portion of the field of view of the one or more cameras in the preview of the image data enables a user to identify portions of the field of view that may be displayed as a representation in the videoconferencing interface, providing improved visual feedback.
[0343] In some embodiments, displaying a preview of image data for the field of view of the one or more cameras includes displaying multiple regions (e.g., distinct regions, non-overlapping regions, rectangular regions, square regions, and / or quadrants) of the preview (e.g., 636-1, 636-2 of FIG. 6I and / or 636a-i of FIG. 6J) (e.g., one or more regions may correspond to distinct portions of the image data for the field of view). In some embodiments, the computer system detects user input (e.g., 612i and / or 612j) corresponding to one or more regions of the multiple regions. In response to detecting the user input corresponding to the one or more regions, and following a determination that the user input corresponding to the one or more regions corresponds to a first region of the one or more regions, the computer system displays a representation of the first region within the live video communication interface (e.g., with distortion correction based on the first region) (e.g., as described with reference to FIGS. 6I-6J). In response to detecting user input corresponding to the one or more regions, and in accordance with determining that the user input corresponding to the one or more regions corresponds to a second region of the one or more regions, the computer system displays a representation of the second region (e.g., with distortion correction based on the second region that is different from the distortion correction based on the first region) as a representation within the live video communication interface. Displaying a representation of the first region or a representation of the second region within the live video communication interface enhances the video communication session experience by allowing a user to efficiently manage what is displayed within the live video communication interface, provides improved visual feedback, and reduces the number of inputs required to perform actions.
[0344] In some embodiments, the one or more user inputs include gestures (e.g., 612d) (e.g., body gestures, hand gestures, head gestures, arm gestures, and / or eye gestures) within the field of view of one or more cameras (e.g., gestures performed within the field of view of one or more cameras pointed at the physical position surface). Utilizing gestures within the field of view of one or more cameras as inputs enhances the video communication session experience by allowing a user to control what is displayed without physically touching the device, providing additional control options without cluttering the user interface.
[0345] In some embodiments, the computer system displays a s...
Claims
1. 1. A method comprising: A computer system in communication with a display generating component, one or more cameras, and one or more input devices, comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the live video communication interface, the user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and displaying, via the display generation component, a representation of the surface in response to detecting the one or more user inputs, wherein the representation of the surface includes an image of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
2. receiving image data captured by a camera of the one or more cameras during the live video communication session; displaying, via the display generation component, the representation of the at least a portion of the field of view based on the image data captured by the camera; and The method of claim 1 , further comprising: displaying, via the display generation component, the representation of the surface based on the image data captured by the camera.
3. The method of claim 1 or 2, wherein the image of the surface is rectified by rotating the image of the surface relative to the representation of at least a portion of the field of view of the one or more cameras.
4. The method of claim 3 , wherein the image of the surface is rotated based on the position of the surface relative to the user within the field of view of the one or more cameras.
5. 4. The method of claim 3, wherein, in accordance with a determination that the surface is in a first position relative to a user within the field of view of the one or more cameras, the image of the surface is rotated at least 45 degrees relative to a representation of the user within the field of view of the one or more cameras.
6. The method of claim 1 , wherein the representation of the at least part of the field of view includes the user and is displayed simultaneously with the representation of the surface.
7. In response to detecting the one or more user inputs, 7. The method of claim 1, further comprising displaying a preview of image data for the field of view of the one or more cameras before displaying the representation of the surface, the preview including an image of the surface that is not modified based on the position of the surface relative to the one or more cameras.
8. wherein displaying the preview of image data for the field of view of the one or more cameras includes displaying a plurality of selectable options corresponding to respective portions of the field of view of the one or more cameras, the method comprising: detecting an input selecting one of the plurality of options corresponding to a respective portion of the field of view of the one or more cameras; In response to detecting the input, selecting one of the plurality of selectable options corresponding to a respective portion of the field of view of the one or more cameras; displaying the representation of the surface based on the first portion of the field of view of the one or more cameras in accordance with a determination that the input selecting one of the plurality of selectable options corresponding to a respective portion of the field of view of the one or more cameras is directed to a first option corresponding to the first portion of the field of view of the one or more cameras; 8. The method of claim 7, further comprising: displaying the representation of the surface based on a second portion of the field of view of the one or more cameras in accordance with a determination that the input selecting one of the plurality of selectable options corresponding to a respective portion of the field of view of the one or more cameras is directed toward a second option corresponding to a second portion of the field of view of the one or more cameras, wherein the second option is different from the first option.
9. Displaying the preview of image data for the field of view of the one or more cameras includes displaying a plurality of regions of the preview, the method comprising: detecting a user input corresponding to one or more of the plurality of regions; In response to detecting the user input corresponding to the one or more regions, pursuant to determining that the user input corresponding to the one or more regions corresponds to a first region of the one or more regions, displaying a representation of the first region in the live video communication interface; 10. The method of claim 7, further comprising: according to a determination that the user input corresponding to the one or more regions corresponds to a second region of the one or more regions, displaying a representation of the second region as a representation within the live video communication interface.
10. The method of claim 1 , wherein the one or more user inputs comprise gestures within the field of view of the one or more cameras.
11. The method of claim 1 , further comprising displaying a surface view option, wherein the one or more user inputs include inputs directed to the surface view option.
12. detecting a user input corresponding to a selection of the surface view option; 12. The method of claim 11 , further comprising: in response to detecting the user input corresponding to selection of the surface view option, displaying a preview of image data for the field of view of the one or more cameras, the preview including multiple portions of the field of view of the one or more cameras including the at least a portion of the field of view of the one or more cameras, and the preview including a visual indication of an active portion of the field of view.
13. detecting a user input corresponding to a selection of the surface view option; 12. The method of claim 11 , further comprising: in response to detecting the user input corresponding to a selection of the surface view option, displaying a preview of image data for the field of view of the one or more cameras, the preview including a plurality of selectable visually distinct portions overlaid on a representation of the field of view of the one or more cameras.
14. The method of claim 1 , wherein the surface is a vertical surface in the scene.
15. The method of claim 1 , wherein the surface is a horizontal surface in the scene.
16. Displaying the representation of the surface includes displaying a first view of the surface, the method comprising: displaying one or more shift view options while displaying the first view of the surface; and detecting a user input directed to a respective shift view option of the one or more shift view options; 16. The method of claim 1, further comprising: in response to detecting the user input directed to the respective shift view option, displaying a second view of the surface that differs from the first view of the surface.
17. 17. The method of claim 16, wherein displaying the first view of the surface comprises displaying an image of the surface modified in a first manner, and wherein displaying the second view of the surface comprises displaying an image of the surface modified in a second manner different from the first manner.
18. The representation of the surface is displayed at a first zoom level, and the method comprises: While displaying the representation of the surface at the first zoom level, detecting a user input corresponding to a request to change the zoom level of the representation of the surface; 18. The method of claim 1, further comprising: in response to detecting the user input corresponding to a request to change a zoom level of the representation of the surface, displaying the representation of the surface at a second zoom level different from the first zoom level.
19. while displaying the live video communication interface, displaying a selectable control option that, when selected, causes the representation of the surface to be displayed; The method of claim 1 , wherein the one or more user inputs include a user input corresponding to a selection of the selectable control option.
20. the live video communication session is provided by a first application running on the computer system; 20. The method of claim 19, wherein the selectable control option is associated with a second application that is different from the first application.
21. 21. The method of claim 20, further comprising: displaying a user interface of the second application in response to detecting the one or more user inputs, including the user input corresponding to a selection of the selectable control option.
22. 21. The method of claim 20, further comprising displaying a user interface of the second application before displaying the live video communication interface for the live video communication session.
23. the live video communication session is provided using a third application running on the computer system; 23. The method of any one of claims 1 to 22, wherein the representation of the surface is provided by a fourth application different from the third application.
24. 24. The method of claim 23, wherein the representation of the surface is displayed using a user interface of the fourth application that is displayed during the live video communication session.
25. 25. The method of claim 23 or 24, further comprising displaying, via the display generation component, a graphical element corresponding to the fourth application, comprising displaying the graphical element in a region including a set of one or more graphical elements corresponding to applications other than the fourth application.
26. 26. The method of claim 1, wherein displaying the representation of the surface comprises displaying, via the display generation component, an animation of a transition from the representation of at least a portion of the field of view of the one or more cameras to the representation of the surface.
27. The computer system is in communication with a second computer system in communication with a second display generation component, and the method includes: displaying the representation of at least a portion of the field of view of the one or more cameras on the display generation component; 27. The method of claim 1, further comprising: displaying the representation of the surface on the second display generating component.
28. 28. The method of claim 27, further comprising, in response to detecting a change in orientation of the second computer system, updating the display of the representation of the surface displayed on the second display generation component from displaying a first view of the surface to displaying a second view of the surface that differs from the first view.
29. 29. The method of claim 1, wherein displaying the representation of the surface comprises displaying an animation of a transition from the display of the representation of the at least part of the field of view of the one or more cameras to the display of the representation of the surface, the animation comprising panning a view of the field of view of the one or more cameras and rotating the view of the field of view of the one or more cameras.
30. 30. The method of claim 1, wherein displaying the representation of the surface comprises displaying an animation of a transition from the display of the representation of the at least part of the field of view of the one or more cameras to the display of the representation of the surface, the animation comprising zooming a view of the field of view of the one or more cameras and rotating the view of the field of view of the one or more cameras.
31. 31. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 1 to 30.
32. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors, said one or more programs including instructions for performing the method of any one of claims 1 to 30.
33. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: A computer system comprising means for carrying out the method of any one of claims 1 to 30.
34. 31. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 1 to 30.
35. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the live video communication interface, the user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; a display generating component configured to display, in response to detecting the one or more user inputs, a representation of the surface, the representation of the surface including an image of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras;
36. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the live video communication interface, the user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; and instructions for displaying, via the display generation component, a representation of the surface in response to detecting the one or more user inputs, the representation of the surface including images of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
37. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: means for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; means for detecting one or more user inputs via the one or more input devices while displaying the live video communication interface, the user inputs including user inputs directed at surfaces in a scene within the field of view of the one or more cameras; means for displaying, via the display generation component, a representation of the surface in response to detecting the one or more user inputs, wherein the representation of the surface includes images of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
38. 1. A computer program product storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of at least a portion of a field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the live video communication interface, the user inputs including user inputs directed at a surface in a scene within the field of view of the one or more cameras; responsive to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, the representation of the surface including images of the surface captured by the one or more cameras modified based on a position of the surface relative to the one or more cameras.
39. 1. A method comprising: A computer system in communication with a display generating component, one or more cameras, and one or more input devices, comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; acquiring, while displaying the live video communication interface, image data via the one or more cameras about the field of view of the one or more cameras, the image data including a first gesture; In response to acquiring the image data for the field of view of the one or more cameras, displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the second portion comprising different visual content than the representation of the first portion of the scene, in accordance with a determination that the first gesture satisfies a first set of criteria; and continuing to display, via the display generation component, the representation of the first portion of the scene according to a determination that the first gesture satisfies a second set of criteria different from the first set of criteria.
40. 40. The method of claim 39, wherein the representation of the first portion of the scene is displayed simultaneously with the representation of the second portion of the scene.
41. In response to acquiring the image data for the field of view of the one or more cameras, 41. The method of claim 39 or 40, further comprising: displaying, via the display generation component, a representation of a third portion of the scene within the field of view of the one or more cameras in accordance with a determination that the first gesture satisfies a third set of criteria different from the first set of criteria and the second set of criteria, wherein the representation of the third portion of the scene includes different visual content than the representation of the first portion of the scene and different visual content than the representation of the second portion of the scene.
42. acquiring image data including a user's hand movement while displaying the representation of the second portion of the scene; In response to acquiring image data including the movement of the hand of the user, 42. The method of claim 39, further comprising: displaying a representation of a fourth portion of the scene that is different from the second portion of the scene and that includes the user's hand, wherein displaying comprises tracking movement of the user's hand from the second portion of the scene to the fourth portion of the scene.
43. acquiring image data including a third gesture; In response to acquiring the image data including the third gesture, 43. The method of claim 39, further comprising: according to a determination that the third gesture satisfies a zoom criterion, changing a zoom level of the respective representation of the portion of the scene from a first zoom level to a second zoom level different from the first zoom level.
44. 44. The method of claim 43, wherein the third gesture comprises a pointing gesture, and wherein changing the zoom level comprises zooming in on an area of the scene that corresponds to the pointing gesture.
45. 44. The method of claim 43, wherein the separate representation displayed at the first zoom level is centered on a first position in the scene, and the separate representation displayed at the second zoom level is centered on the first position in the scene.
46. Varying the zoom level of the individual representations may include: changing the zoom level of a first portion of the individual representation from the first zoom level to the second zoom level; and displaying a second portion of the individual representation at the first zoom level, the second portion being different from the first portion.
47. In response to acquiring the image data for the field of view of the one or more cameras, 47. The method of any one of claims 39 to 46, further comprising: displaying a first graphical indication that a gesture has been detected in accordance with the determination that the first gesture satisfies the first set of criteria.
48. Displaying the first graphical indication includes: displaying the first graphical indication having a first appearance in accordance with a determination that the first gesture comprises a first type of gesture; and displaying the first graphical indication having a second appearance different from the first appearance in accordance with a determination that the first gesture comprises a second type of gesture.
49. In response to acquiring the image data for the field of view of the one or more cameras, 49. The method of claim 39, further comprising: displaying a second graphical object indicating progress toward satisfying a threshold amount of time in accordance with the determination that the first gesture satisfies a fourth set of criteria.
50. 50. The method of claim 49, wherein the first set of criteria includes criteria that are met if the first gesture is maintained for the threshold amount of time.
51. 50. The method of claim 49, wherein the second graphical object is a timer.
52. 50. The method of claim 49, wherein the second graphical object comprises an outline of a representation of a gesture.
53. 50. The method of claim 49, wherein the second graphical object indicates a zoom level.
54. detecting an audio input before displaying the representation of the second portion of the scene; 54. The method of any one of claims 39 to 53, wherein the first set of criteria includes criteria based on the audio input.
55. the first gesture includes a pointing gesture; the representation of the first portion of the scene is displayed at a first zoom level; Displaying the representation of the second portion comprises:
55. The method of any one of claims 39 to 54, comprising, in accordance with a determination that the pointing gesture is directed at an object within the scene, displaying a representation of the object at a second zoom level different from the first zoom level.
56. the first gesture includes a framing gesture; the representation of the first portion of the scene is displayed at a first zoom level; Displaying the representation of the second portion comprises:
56. The method of any one of claims 39 to 55, comprising, in accordance with a determination that the framing gesture is directed at an object within the scene, displaying a representation of the object at a second zoom level different from the first zoom level.
57. the first gesture includes a pointing gesture; Displaying the representation of the second portion comprises: panning image data in the first direction of the pointing gesture according to a determination that the pointing gesture is in a first direction; 57. The method of any one of claims 39 to 56, comprising: panning image data in the second direction of the pointing gesture in accordance with a determination that the pointing gesture is in a second direction.
58. displaying the representation of the first portion of the scene includes displaying a user representation; 58. The method of claim 57, wherein displaying the representation of the second portion includes maintaining a display of the representation of the user.
59. the first gesture comprises a hand gesture; Displaying the representation of the first portion of the scene includes displaying the representation of the first portion of the scene at a first zoom level; 59. The method of any one of claims 39 to 58, wherein displaying the representation of the second portion of the scene comprises displaying the representation of the second portion of the scene at a second zoom level different from the first zoom level.
60. 60. The method of claim 59, wherein the hand gesture for displaying the representation of the second portion of the scene at the second zoom level comprises a hand pose with two fingers lifted corresponding to an amount of zoom.
61. 60. The method of claim 59, wherein the hand gesture to display the representation of the second portion of the scene at the second zoom level comprises a hand movement corresponding to an amount of zoom.
62. the representation of the first portion of the scene includes a representation of a first area of the scene and a representation of a second area of the scene; Displaying the representation of the second portion of the scene includes: maintaining the appearance of the representation of the first area of the scene; and and modifying the appearance of the representation of the second area of the scene.
63. 63. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 39 to 62.
64. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: one or more processors; and a memory storing one or more programs configured to be executed by said one or more processors, said one or more programs including instructions for performing the method of any one of claims 39 to 62.
65. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising:
63. A computer system comprising means for carrying out the method of any one of claims 39 to 62.
66. 63. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising instructions for performing the method of any one of claims 39 to 62.
67. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; while displaying the live video communication interface, acquiring, via the one or more cameras, image data about the field of view of the one or more cameras, the image data including a first gesture; In response to acquiring the image data for the field of view of the one or more cameras, displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the second portion comprising different visual content than the representation of the first portion of the scene, in accordance with a determination that the first gesture satisfies a first set of criteria; and continuing to display, via the display generation component, the representation of the first portion of the scene according to a determination that the first gesture satisfies a second set of criteria that is different from the first set of criteria.
68. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; while displaying the live video communication interface, acquiring, via the one or more cameras, image data about the field of view of the one or more cameras, the image data including a first gesture; In response to acquiring the image data for the field of view of the one or more cameras, displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the second portion comprising different visual content than the representation of the first portion of the scene, in accordance with a determination that the first gesture satisfies a first set of criteria; and instructions for continuing to display, via the display generation component, the representation of the first portion of the scene according to a determination that the first gesture satisfies a second set of criteria different from the first set of criteria.
69. 1. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: means for displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; means for acquiring, via the one or more cameras while displaying the live video communication interface, image data about the field of view of the one or more cameras, the image data including a first gesture; In response to acquiring the image data for the field of view of the one or more cameras, displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the second portion comprising different visual content than the representation of the first portion of the scene, in accordance with a determination that the first gesture satisfies a first set of criteria; means for continuing to display, via the display generation component, the representation of the first portion of the scene according to a determination that the first gesture satisfies a second set of criteria that is different from the first set of criteria.
70. 1. A computer program product storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: displaying, via the display generation component, a live video communication interface for a live video communication session, the live video communication interface including a representation of a first portion of a scene within a field of view captured by the one or more cameras; while displaying the live video communication interface, acquiring, via the one or more cameras, image data about the field of view of the one or more cameras, the image data including a first gesture; In response to acquiring the image data for the field of view of the one or more cameras, displaying, via the display generation component, a representation of a second portion of the scene within the field of view of the one or more cameras, the second portion comprising different visual content than the representation of the first portion of the scene, in accordance with a determination that the first gesture satisfies a first set of criteria; pursuant to a determination that the first gesture satisfies a second set of criteria different from the first set of criteria, continue to display, via the display generation component, the representation of the first portion of the scene.
71. 1. A method comprising: A first computer system in communication with a display generating component, one or more first cameras, and one or more input devices, Detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; and displaying, via the display generation component, a live video communication interface for the live video communication session in response to detecting the set of one or more user inputs, the live video communication interface comprising: a first representation of a field of view of the one or more first cameras of the first computer system; and a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene that are within the field of view of the one or more second cameras of the second computer system.
72. receiving image data captured by a first camera of the one or more first cameras during the live video communication session; Displaying the live video communication interface for the live video communication session includes: displaying, via the display generation component, the first representation of the field of view of the one or more first cameras of the first computer system based on the image data captured by the first cameras; and displaying, via the display generation component, the second representation of the field of view of the one or more first cameras of the first computer system based on the image data captured by the first camera.
73. Displaying the live video communication interface for the live video communication session includes: displaying the first representation of the field of view of the one or more first cameras of the first computer system within a predetermined distance from the second representation of the field of view of the one or more first cameras of the first computer system; and displaying the first representation of the field of view of the one or more second cameras of the second computer system within the predetermined distance from the second representation of the field of view of the one or more second cameras of the second computer system.
74. 74. The method of any one of claims 71 to 73, wherein displaying the live video communication interface for the live video communication session includes displaying the second representation of the field of view of the one or more first cameras of the first computer system superimposed on the second representation of the field of view of the one or more second cameras of the second computer system.
75. Displaying the live video communication interface for the live video communication session includes: displaying the second representation of the field of view of the one or more first cameras of the first computer system in a first visually defined area of the live video communication interface; and displaying the second representation of the field of view of the one or more second cameras of the second computer system in a second visually defined area of the live video communication interface, wherein the first visually defined area does not overlap with the second visually defined area.
76. the second representation of the field of view of the one or more first cameras of the first computer system is based on image data captured by the one or more first cameras of the first computer system corrected with a first distortion correction to change the perspective from which the image data captured by the one or more first cameras of the first computer system appears to be captured; and / or 76. The method of any one of claims 71 to 75, wherein the second representation of the field of view of the one or more second cameras of the second computer system is based on image data captured by the one or more second cameras of the second computer system that has been corrected by a second distortion correction to change the perspective from which the image data captured by the one or more second cameras of the second computer system appears to be captured.
77. the second representation of the field of view of the one or more first cameras of the first computer system is based on image data captured by the one or more first cameras of the first computer system corrected with a first distortion correction; 77. The method of any one of claims 71 to 76, wherein the second representation of the field of view of the one or more second cameras of the second computer system is based on image data captured by the one or more second cameras of the second computer system that has been corrected using a second distortion correction that is different from the first distortion correction.
78. the second representation of the field of view of the one or more first cameras of the first computer system is based on image data captured by the one or more first cameras of the first computer system rotated relative to the position of the surface within the first scene; and / or 78. The method of any one of claims 71 to 77, wherein the second representation of the field of view of the one or more second cameras of the second computer system is based on image data captured by the one or more second cameras of the second computer system rotated relative to the position of the surface within the second scene.
79. the second representation of the field of view of the one or more first cameras of the first computer system is based on image data captured by the one or more first cameras of the first computer system rotated by a first amount relative to a position of the surface within the first scene; 79. The method of claim 78, wherein the second representation of the field of view of the one or more second cameras of the second computer system is based on image data captured by the one or more second cameras of the second computer system rotated by a second amount relative to the position of the surface in the second scene, the first amount being different from the second amount.
80. displaying the live video communication interface displaying a graphical object in the live video communication interface; to the live video communication interface via the display generation component; the second representation of the field of view of the one or more first cameras of the first computer system on the graphical object; and and simultaneously displaying the second representation of the field of view of the one or more second cameras of the second computer system on the graphical object.
81. detecting a first user input via the one or more input devices while simultaneously displaying the second representation of the field of view of the one or more first cameras of the first computer system on the graphical object and the second representation of the field of view of the one or more second cameras of the second computer system on the graphical object; In response to detecting the first user input, changing a zoom level of the second representation of the field of view of the one or more first cameras of the first computer system in accordance with a determination that the first user input corresponds to the second representation of the field of view of the one or more first cameras of the first computer system; 81. The method of claim 80, further comprising: altering a zoom level of the second representation of the field of view of the one or more second cameras of the second computer system in accordance with a determination that the first user input corresponds to the second representation of the field of view of the one or more second cameras of the second computer system.
82. 81. The method of claim 80, wherein the graphical object is based on an image of a physical object in the first scene or the second scene.
83. detecting a second user input via the one or more input devices while simultaneously displaying the second representation of the field of view of the one or more first cameras of the first computer system on the graphical object and the second representation of the field of view of the one or more second cameras of the second computer system on the graphical object; In response to detecting the second user input, moving the first representation of the field of view of the one or more first cameras of the first computer system from a first position on the graphical object to a second position on the graphical object; moving the second representation of the field of view of the one or more first cameras of the first computer system from a third position on the graphical object to a fourth position on the graphical object; moving the first representation of the field of view of the one or more second cameras of the second computer system from a fifth position on the graphical object to a sixth position on the graphical object; 81. The method of claim 80, further comprising: moving the second representation of the field of view of the one or more second cameras of the second computer system from a seventh position on the graphical object to an eighth position on the graphical object.
84. moving the first representation of the field of view of the one or more first cameras of the first computer system from a first position on the graphical object to a second position on the graphical object includes displaying an animation of the first representation of the field of view of the one or more first cameras of the first computer system moving from the first position on the graphical object to the second position on the graphical object; moving the second representation of the field of view of the one or more first cameras of the first computer system from a third position on the graphical object to a fourth position on the graphical object includes displaying an animation of the second representation of the field of view of the one or more first cameras of the first computer system moving from the third position on the graphical object to the fourth position on the graphical object; moving the first representation of the field of view of the one or more second cameras of the second computer system from a fifth position on the graphical object to a sixth position on the graphical object includes displaying an animation of the first representation of the field of view of the one or more second cameras of the second computer system moving from the fifth position on the graphical object to the sixth position on the graphical object; 84. The method of claim 83, wherein moving the second representation of the field of view of the one or more second cameras of the second computer system from a seventh position on the graphical object to an eighth position on the graphical object comprises displaying an animation of the second representation of the field of view of the one or more second cameras of the second computer system moving from the seventh position on the graphical object to the eighth position on the graphical object.
85. displaying the live video communication interface displaying the first representation of the field of view of the one or more first cameras at a smaller size than the second representation of the field of view of the one or more first cameras; and displaying the first representation of the field of view of the one or more second cameras at a smaller size than the second representation of the field of view of the one or more second cameras.
86. while simultaneously displaying the second representation of the field of view of the one or more first cameras of the first computer system on the graphical object and the second representation of the field of view of the one or more second cameras of the second computer system on the graphical object; displaying the first representation of the field of view of the one or more first cameras at an orientation based on a position of the second representation of the field of view of the one or more first cameras on the graphical object; 86. The method of claim 85, further comprising: displaying the first representation of the field of view of the one or more second cameras at an orientation based on a position of the second representation of the field of view of the one or more second cameras on the graphical object.
87. the second representation of the field of view of the one or more first cameras of the first computer system includes a representation of a drawing on the surface within the first scene; and / or 87. The method of any one of claims 71 to 86, wherein the second representation of the field of view of the one or more second cameras of the second computer system comprises a representation of a drawing on the surface within the second scene.
88. the second representation of the field of view of the one or more first cameras of the first computer system includes a representation of physical objects on the surface in the first scene; and / or 88. The method of any one of claims 71 to 87, wherein the second representation of the field of view of the one or more second cameras of the second computer system includes a representation of physical objects on the surface in the second scene.
89. detecting a third user input via the one or more input devices while displaying the live video communication interface; 89. The method of any one of claims 71 to 88, further comprising: in response to detecting the third user input, displaying visual markup content within the second representation of the field of view of the one or more second cameras of the second computer system in accordance with the third user input.
90. The visual markup content is displayed on a representation of an object within the second representation of the field of view of the one or more second cameras of the second computer system, and the method further comprises: while displaying the visual markup content on the representation of the object within the second representation of the field of view of the one or more second cameras of the second computer system; receiving an indication of movement of the object in the second representation of the field of view of the one or more second cameras of the second computer system; in response to receiving the indication of movement of the object within the second representation of the field of view of the one or more second cameras of the second computer system, moving the representation of the object within the second representation of the field of view of the one or more second cameras of the second computer system according to the movement of the object; 90. The method of claim 89, further comprising: moving the visual markup content within the second representation of the field of view of the one or more second cameras of the second computer system in accordance with the movement of the object, wherein moving the visual markup content comprises maintaining a position of the visual markup content relative to the representation of the object.
91. the visual markup content is displayed on a representation of a page within the second representation of the field of view of the one or more second cameras of the second computer system, the method comprising: receiving an indication that the page has been turned; 90. The method of claim 89, further comprising: ceasing display of the visual markup content in response to receiving the indication that the page has been turned.
92. receiving an indication that the page will be redisplayed after ceasing to display the visual markup content; and 92. The method of claim 91 , further comprising: in response to receiving an indication that the page will be redisplayed, displaying the visual markup content on the representation of the page within the second representation of the field of view of the one or more second cameras of the second computer system.
93. while displaying the visual markup content within the second representation of the field of view of the one or more second cameras of the second computer system; receiving an indication of a request to modify the visual markup content within the live video communication session detected by the second computer system; and 90. The method of claim 89, further comprising: in response to receiving the indication of the request to modify the visual markup content in the live video communication session detected by the second computer system, modifying the visual markup content in the second representation of the field of view of the one or more second cameras of the second computer system in accordance with the request to modify the visual markup content.
94. 90. The method of claim 89, further comprising, after displaying the visual markup content within the second representation of the field of view of the one or more second cameras of the second computer system, fading out the display of the visual markup content over time.
95. detecting, via the one or more input devices, a voice input comprising a query while displaying the live video communication interface including the representation of the surface in the first scene; In response to detecting the voice input, 95. The method of any one of claims 71 to 94, further comprising: outputting a response to the query based on visual content within the second representation of the field of view of the one or more first cameras of the first computer system and / or the second representation of the field of view of the one or more second cameras of the second computer system.
96. while displaying the live video communication interface; detecting that the second representation of the field of view of the one or more first cameras of the first computer system includes a representation of a third computer system in the first scene in communication with a third display generation component; in response to detecting that the second representation of the field of view of the one or more first cameras of the first computer system includes the representation of the third computer system in the first scene in communication with the third display generation component; 96. The method of any one of claims 71 to 95, further comprising: displaying, on the live video communication interface, visual content corresponding to display data received from the third computing system that corresponds to visual content displayed on the third display generation component.
97. 97. The method of any one of claims 71 to 96, further comprising displaying content included in the live video communication session on a physical object.
98. 98. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more first cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 71 to 97.
99. 1. A computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 71 to 97.
100. 1. A computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices, comprising:
98. A computer system comprising means for carrying out the method of any one of claims 71 to 97.
101. 98. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more first cameras, and one or more input devices, the one or more programs comprising instructions for performing the method of any one of claims 71 to 97.
102. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more first cameras, and one or more input devices, the one or more programs comprising: Detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; responsive to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for the live video communication session, the live video communication interface comprising: a first representation of a field of view of the one or more first cameras of the first computer system; and a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene that are within the field of view of the one or more second cameras of the second computer system.
103. a first computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices; one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: Detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; responsive to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for the live video communication session, the live video communication interface comprising: a first representation of a field of view of the one or more first cameras of the first computer system; and a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is within the field of view of the one or more second cameras of the second computer system.
104. a first computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices; means for detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; means for displaying, via the display generation component, a live video communication interface for a live video communication session in response to detecting the set of one or more user inputs, the live video communication interface comprising: a first representation of a field of view of the one or more first cameras of the first computer system; and a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is within the field of view of the one or more second cameras of the second computer system.
105. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generation component, one or more first cameras, and one or more input devices, the one or more programs comprising: Detecting a set of one or more user inputs corresponding to a request to display a user interface for a live video communication session including a plurality of participants; responsive to detecting the set of one or more user inputs, displaying, via the display generation component, a live video communication interface for the live video communication session, the live video communication interface comprising: a first representation of a field of view of the one or more first cameras of the first computer system; and a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene that is within the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene that is within the field of view of the one or more second cameras of the second computer system.
106. 1. A method comprising: a first computer system in communication with a first display generating component and one or more sensors; While the first computer system is in a live video communication session with a second computer system, displaying, via the first display generation component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
107. While the first computer system is in the live video communication session with the second computer system, Detecting handwriting from the image data, the handwriting including physical marks on a physical surface that is within a field of view of the one or more cameras of the second computer system and that is separate from the second computer system; 107. The method of claim 106, further comprising: in response to detecting the handwriting that is within the field of view of the one or more cameras of the second computer system and that includes a physical mark on the physical surface that is separate from the second computer system, displaying digital text that corresponds to the handwriting that is within the field of view of the one or more cameras of the second computer system.
108. Displaying the representation of the second view of the physical environment within the field of view of the one or more cameras of the second computer system includes: in response to determining that the change in the position of the first computer system includes a first amount of change in angle of the first computer system, the second view of the physical environment differs from the first view of the physical environment by a first amount of angle; 108. The method of claim 106 or 107, wherein, in accordance with a determination that the change in the position of the first computer system includes a second amount of change in angle of the first computer system that is different from the first amount of change in angle of the first computer system, the second view of the physical environment differs from the first view of the physical environment by a second angular amount that is different from the first angular amount.
109. Displaying the representation of the second view of the physical environment within the field of view of the one or more cameras of the second computer system includes: In accordance with determining that the change in the position of the first computer system includes a first direction of change in position of the first computer system, the second view of the physical environment is in a first direction in the physical environment from the first view of the physical environment; 109. The method of any one of claims 106 to 108, comprising: in accordance with a determination that the change in the position of the first computer system comprises a second direction that is different from the first direction of change in position of the first computer system, wherein the second direction of change in position of the first computer system comprises a second direction that is different from the first direction of change in position of the first computer system, the second view of the physical environment is in a second direction of the physical environment from the first view of the physical environment, and the second direction of the physical environment is different from the first direction of the physical environment.
110. The changing of the position of the first computer system includes changing an angle of the first computer system, and displaying the representation of the second view of the physical environment within the field of view of the one or more cameras of the second computer system includes:
110. The method of any one of claims 106 to 109, comprising displaying a gradual transition from the representation of the first view of the physical environment to the representation of the second view of the physical environment based on the change in angle of the first computer system.
111. the representation of the first view includes a representation of a user's face within the field of view of the one or more cameras of the second computer system; 111. The method of any one of claims 106 to 110, wherein the representation of the second view includes a representation of a physical mark within the field of view of the one or more cameras of the second computer system.
112. detecting user input corresponding to a digital mark via one or more input devices while displaying the representation of the physical mark; 112. The method of claim 111, further comprising: in response to detecting the user input, displaying a representation of the digital mark simultaneously with the representation of the physical mark.
113. 113. The method of claim 112, wherein the representation of the digital mark is displayed via the first display generation component.
114. 114. The method of claim 113, further comprising, in response to detecting the digital mark, causing the second computer system to display a representation of the digital mark.
115. 115. The method of claim 114, wherein the representation of the digital mark is displayed on the representation of the physical mark at the second computer system.
116. 116. A method according to claim 114 or 115, wherein the representation of the digital mark is displayed on a physical object in the physical environment of the second computer system.
117. While the first computer system is in the live video communication session with the second computer system, 117. The method of any one of claims 106 to 116, further comprising displaying, via the first display generation component, a representation of a third view of the physical environment within the field of view of the one or more cameras of the second computer system, the third view including a face of a user within the field of view of the one or more cameras of the second computer system, the representation of the face of the user being displayed simultaneously with the representation of the second view of the physical environment.
118. displaying the representation of the first view of the physical environment includes displaying the representation of the first view of the physical environment based on the image data captured by a first camera of the one or more cameras of the second computer system; 118. The method of any one of claims 106 to 117, wherein displaying the representation of the second view of the physical environment comprises displaying the representation of the second view of the physical environment based on the image data captured by the first camera of the one or more cameras of the second computer system.
119. 119. The method of any one of claims 106 to 118, wherein displaying the representation of the second view of the physical environment within the field of view of the one or more cameras of the second computer system is performed in accordance with a determination that permission to modify the view of the physical environment displayed at the first computer system has been provided to the first computer system.
120. detecting, via the one or more sensors, a respective change in position of the first computer system while displaying a representation of the third view of the physical environment; In response to detecting the individual change in the position of the first computer system, displaying, via the first display generation component, a representation of the distinct view of the physical environment within the field of view of the one or more cameras of the second computer system in accordance with determining that the distinct change in the position of the first computer system corresponds to a distinct view within a defined portion of the physical environment; 120. The method of any one of claims 106 to 119, further comprising: ceasing to display the representation of the individual view of the physical environment within the field of view of the one or more cameras of the second computer system in accordance with a determination that the individual change in the position of the first computer system corresponds to an individual view that is not within the defined portion of the physical environment.
121. In response to detecting the individual change in the position of the first computer system, and displaying, via the first display generation component, an obscured representation of the portion of the physical environment that is not within the defined portion of the physical environment in accordance with the determination that the individual change in the position of the first computer system corresponds to the view that is not within the defined portion of the physical environment.
122. The second view of the physical environment includes physical objects within the physical environment, and the method further comprises: acquiring image data including movement of the physical object within the physical environment while displaying the representation of the second view of the physical environment; In response to acquiring image data including the movement of the physical object, 122. The method of any one of claims 106 to 121, further comprising: displaying a representation of a fourth view of the physical environment, different from the second view and including the physical object.
123. The first computer system is in communication with a second display generation component, and the method includes:
123. The method of any one of claims 106 to 122, further comprising displaying via the second display generation component a representation of the user within the field of view of the one or more cameras of the second computer system, wherein the representation of the user is displayed simultaneously with the representation of the second view of the physical environment displayed via the first display generation component.
124. While the first computer system is in the live video communication session with the second computer system, 124. The method of any one of claims 106 to 123, further comprising: displaying an affordance in accordance with a determination that a third computer system satisfies a first set of criteria, wherein selection of the affordance causes the representation of the second view to be displayed on the third computer system, and wherein the first set of criteria includes a location criterion that the third computer system is within a threshold distance of the first computer system.
125. 125. The method of claim 124, wherein the first set of criteria includes a second set of criteria that are different from the location criteria and that are based on characteristics of the third computer system.
126. 126. The method of claim 125, wherein the second set of criteria includes orientation criteria that are satisfied when the third computer system is in a predetermined orientation.
127. 126. The method of claim 125, wherein the second set of criteria includes user account criteria that are met when the first computer system and the third computer system are associated with the same user account.
128. 128. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more sensors, the one or more programs including instructions for performing the method of any one of claims 106 to 127.
129. a computer system configured to communicate with a first display generating component and one or more sensors, one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 106 to 127.
130. a computer system configured to communicate with a first display generating component and one or more sensors, 128. A computer system comprising means for carrying out the method of any one of claims 106 to 127.
131. 128. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a first display generating component and one or more sensors, the one or more programs comprising instructions for performing the method of any one of claims 106 to 127.
132. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a first display generating component and one or more sensors, the one or more programs comprising: While the first computer system is in a live video communication session with a second computer system, displaying, via the first display generation component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; a first display generating component configured to generate a representation of the physical environment within the field of view of the one or more cameras of the second computer system, the first display generating component being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system, in response to detecting the change in the position of the first computer system;
133. a first computer system configured to communicate with a first display generating component and one or more sensors, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: While the first computer system is in a live video communication session with a second computer system, displaying, via the first display generation component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; a first computer system including instructions for, in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
134. a first computer system configured to communicate with a first display generating component and one or more sensors, While the first computer system is engaged in a live video communication session with a second computer system, displaying, via the first display generation component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; a first computer system, comprising: means for, in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
135. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system in communication with a first display generating component and one or more sensors, the one or more programs comprising: While the first computer system is in a live video communication session with a second computer system, displaying, via the first display generation component, a representation of a first view of a physical environment within a field of view of one or more cameras of the second computer system; detecting, via the one or more sensors, a change in position of the first computer system while displaying the representation of the first view of the physical environment; responsive to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
136. 1. A method comprising: A computer system in communication with a display generation component, via the display generation component, displaying a representation of a physical mark within a physical environment, a view of the physical environment within a field of view of one or more cameras, the view of the physical environment includes the physical mark and a physical background; displaying the representation of the physical mark based on a view of the physical environment, the displaying including displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is within the field of view of the one or more cameras; acquiring data including a new physical mark in the physical environment while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras; and in response to acquiring data representing the new physical mark within the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras.
137. 137. The method of claim 136, wherein the portion of the physical background is adjacent to and / or at least partially surrounds the physical mark.
138. 138. The method of claim 136 or 137, wherein the portion of the physical background is at least partially surrounded by the physical mark.
139. 139. The method of any one of claims 136 to 138, further comprising displaying a representation of a user's hand within the field of view of the one or more cameras without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras, wherein the user's hand is in the foreground of the one or more elements of the portion of the physical background within the field of view of the one or more cameras.
140. 140. The method of any one of claims 136 to 139, further comprising displaying a representation of a marking implement without displaying the one or more elements of the portion of the physical background that is within the field of view of the one or more cameras.
141. simultaneously displaying the representation of the physical mark with a first emphasis relative to a representation of the one or more elements of the portion of the physical background prior to displaying the representation of the physical mark without displaying one or more elements of the portion of the physical background within the field of view of the one or more cameras; detecting a user input corresponding to a request to modify the representation of the one or more elements of the portion of the physical background while simultaneously displaying the representation of the physical mark and the representation of the one or more elements of the portion of the physical background; 141. A method according to any one of claims 136 to 140, further comprising, in response to detecting the user input corresponding to the request to modify the representation of the one or more elements of the portion of the physical background, displaying the representation of the physical mark with a second degree of emphasis relative to the representation of the one or more elements of the portion of the physical background, the second degree of emphasis being higher than the first degree of emphasis.
142. 142. The method of claim 141, wherein detecting the user input corresponding to the request to modify the representation of the one or more elements of the portion of the physical background comprises detecting user input directed at a control including a set of highlighting options for the representation of the one or more elements of the portion of the physical background.
143. 143. The method of claim 141 or 142, wherein the user input corresponding to the request to modify the representation of the one or more elements of the portion of the physical background comprises detecting user input directed at a selectable user interface object.
144. the physical mark in the physical environment is a first physical mark; the first physical mark is within the field of view of one or more cameras of the computer system; The method comprises:
144. A method according to any one of claims 136 to 143, further comprising displaying, via the display generation component, a representation of a second physical mark within the physical environment based on a view of the physical environment within the field of view of one or more cameras of an external computer system, wherein the representation of the second physical mark is displayed simultaneously with the representation of the first physical mark.
145. The representation of the first physical mark is a first representation of the first physical mark and is displayed in a first portion of a user interface, and the method further comprises: detecting a first set of one or more user inputs including inputs directed to a first selectable user interface object while displaying the first representation of the first physical mark in the first portion of the user interface; 145. The method of claim 144, further comprising: in response to detecting the first set of one or more user inputs, displaying a second representation of the first physical mark in a second portion of the user interface different from the first portion of the user interface.
146. the representation of the second physical mark is a first representation of the second physical mark and is displayed in a third portion of the user interface, and the method further comprises: detecting a second set of one or more user inputs corresponding to a request to display a second representation of the second physical mark in a fourth portion of the user interface different from the third portion of the user interface; 146. The method of claim 144 or 145, further comprising: displaying the second representation of the second physical mark in the fourth portion of the user interface in response to detecting the second set of one or more user inputs corresponding to the request to display the second representation of the second physical mark in the fourth portion of the user interface.
147. Detecting a request to display a digital mark corresponding to the third physical mark; 147. The method of any one of claims 136 to 146, further comprising: in response to detecting the request to display the digital mark, displaying the digital mark corresponding to the third physical mark.
148. detecting, while displaying the digital mark, a request to modify the digital mark corresponding to the third physical mark; 148. The method of claim 147, further comprising: in response to detecting the request to modify the digital mark corresponding to the third physical mark, displaying a new digital mark that differs from the representation of the digital mark corresponding to the third physical mark.
149. 149. A method according to any one of claims 136 to 148, wherein displaying the representation of the physical mark is based on image data captured by a first camera having a field of view that includes the user's face and the physical mark.
150. further comprising displaying a representation of the face of the user based on the image data captured by the first camera; the field of view of the first camera includes the face of the user and the physical background of the user; 150. The method of claim 149, wherein displaying the representation of the user's face includes displaying the representation of the user's face together with a representation of the user's physical background, wherein the user's face is in the foreground of the one or more elements of the portion of the physical background that is within the field of view of the one or more cameras.
151. 149. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for performing the method of any one of claims 136 to 148.
152. a computer system configured to communicate with a display generation component, one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 136 to 148.
153. a computer system configured to communicate with a display generation component, Means for carrying out the method of any one of claims 136 to 148 A computer system comprising:
154. 149. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for performing the method of any one of claims 136 to 148.
155. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising: via the display generation component, displaying a representation of a physical mark within a physical environment, a view of the physical environment within a field of view of one or more cameras, the view of the physical environment includes the physical mark and a physical background; displaying the representation of the physical mark based on a view of the physical environment, the displaying including displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is within the field of view of the one or more cameras; acquiring data including a new physical mark in the physical environment while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras; a non-transitory computer-readable storage medium comprising instructions for, in response to acquiring data representing the new physical mark within the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras.
156. a computer system configured to communicate with a display generation component, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: via the display generation component, displaying a representation of a physical mark within a physical environment, a view of the physical environment within a field of view of one or more cameras, the view of the physical environment includes the physical mark and a physical background; displaying the representation of the physical mark based on a view of the physical environment, the displaying including displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is within the field of view of the one or more cameras; acquiring data including a new physical mark within the physical environment while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras; a computer system including instructions for, in response to acquiring data representing the new physical mark within the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras.
157. a computer system configured to communicate with a display generation component, via the display generation component, displaying a representation of a physical mark within a physical environment, a view of the physical environment within a field of view of one or more cameras, the view of the physical environment includes the physical mark and a physical background; means for displaying based on a view of the physical environment, wherein displaying the representation of the physical mark includes displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is within the field of view of the one or more cameras; means for acquiring data including a new physical mark within the physical environment while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras; means for displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras in response to acquiring data representing the new physical mark within the physical environment.
158. 1. A computer program product storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising: via the display generation component, displaying a representation of a physical mark within a physical environment, a view of the physical environment within a field of view of one or more cameras, the view of the physical environment includes the physical mark and a physical background; displaying the representation of the physical mark based on a view of the physical environment, the displaying including displaying the representation of the physical mark without displaying one or more elements of a portion of the physical background that is within the field of view of the one or more cameras; acquiring data including a new physical mark within the physical environment while displaying the representation of the physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras; responsive to acquiring data representing the new physical mark within the physical environment, displaying a representation of the new physical mark without displaying the one or more elements of the portion of the physical background within the field of view of the one or more cameras.
159. 1. A method comprising:
1. A computer system in communication with a display generation component and one or more cameras, comprising: displaying the electronic document via the display generation component; detecting, via the one or more cameras, handwriting including physical marks on a physical surface within a field of view of the one or more cameras and separate from the computer system; and in response to detecting the handwriting including a physical mark on the physical surface within the field of view of the one or more cameras and separate from the computer system, displaying digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
160. acquiring data representing new handwriting including a first new physical mark on the physical surface within the field of view of the one or more cameras while displaying the digital text; 160. The method of claim 159, further comprising: in response to obtaining data representing the new handwriting, displaying new digital text corresponding to the new handwriting.
161. 161. The method of claim 160, wherein obtaining data representing the new handwriting comprises detecting the new physical mark while the new physical mark is being applied to the physical surface.
162. 161. The method of claim 160, wherein acquiring data representing the new handwriting comprises detecting the new physical mark when the physical surface including the new physical mark is brought within the field of view of the one or more cameras.
163. acquiring data representing new handwriting including a second new physical mark on the physical surface within the field of view of the one or more cameras while displaying the digital text; 163. The method of any one of claims 159 to 162, further comprising: in response to obtaining data representing the new handwriting, displaying updated digital text corresponding to the new handwriting.
164. 164. The method of claim 163, wherein displaying the updated digital text includes modifying the digital text corresponding to the handwriting.
165. 164. The method of claim 163, wherein displaying the updated digital text includes ceasing to display a portion of the digital text.
166. Displaying the updated digital text includes: displaying new digital text corresponding to the one or more new written characters in accordance with a determination that the second new physical mark satisfies a first criterion; and and ceasing to display a portion of the digital text corresponding to one or more previously written characters in accordance with a determination that the second new physical mark satisfies a second criterion.
167. detecting an input corresponding to a request to display digital text corresponding to each physical mark in the electronic document while displaying a separate handwritten representation including each physical mark on the physical surface; 167. The method of any one of claims 159 to 166, further comprising: in response to detecting the input corresponding to a request to display digital text corresponding to each physical mark, displaying digital text corresponding to each physical mark within the electronic document.
168. Detecting user input directed to a selectable user interface object; in response to detecting a user input directed at the selectable user interface object; pursuant to a determination that the user input directed at the selectable user interface object enables adding digital text based on a live camera feed, displaying digital text corresponding to individual handwriting including each physical mark while the respective physical mark is being applied to the physical surface; 168. The method of any one of claims 159 to 167, further comprising: ceasing to display the digital text corresponding to the individual handwriting including the respective physical marks while the respective physical marks are being applied to the physical surface in accordance with a determination that the user input directed at the user interface selectable object does not enable adding digital text based on the live camera feed.
169. Further comprising displaying the handwritten representation including the physical mark via the display generation component.
169. The method of any one of claims 159 to 168.
170. 170. The method of any one of claims 159 to 169, further comprising displaying, via the display generation component, graphical elements overlaid on individual representations of physical marks corresponding to respective digital text of the electronic document.
171. 171. The method of any one of claims 159 to 170, wherein detecting handwriting is based on image data captured by a first camera having a field of view that includes a user's face and the physical surface.
172. 172. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 159 to 171.
173. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for performing the method of any one of claims 159 to 171.
174. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising:
172. A computer system comprising means for carrying out the method of any one of claims 159 to 171.
175. 172. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 159 to 171.
176. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs comprising: displaying the electronic document via the display generation component; detecting, via the one or more cameras, handwriting including physical marks on a physical surface within a field of view of the one or more cameras and separate from the computer system; a non-transitory computer-readable storage medium comprising instructions for, in response to detecting handwriting including a physical mark on the physical surface within the field of view of the one or more cameras and separate from the computer system, displaying digital text within the electronic document corresponding to the handwriting within the field of view of the one or more cameras.
177. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: displaying the electronic document via the display generation component; detecting, via the one or more cameras, handwriting including physical marks on a physical surface within a field of view of the one or more cameras and separate from the computer system; responsive to detecting handwriting that includes a physical mark on the physical surface that is within the field of view of the one or more cameras and that is separate from the computer system, displaying digital text in the electronic document that corresponds to the handwriting that is within the field of view of the one or more cameras.
178. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising: means for displaying an electronic document via said display generation component; means for detecting, via said one or more cameras, handwriting including physical marks on a physical surface within the field of view of said one or more cameras and separate from said computer system; means for displaying digital text in the electronic document corresponding to the handwriting within the field of view of the one or more cameras in response to detecting the handwriting including a physical mark on the physical surface that is within the field of view of the one or more cameras and separate from the computer system.
179. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising: displaying the electronic document via the display generation component; detecting, via the one or more cameras, handwriting including physical marks on a physical surface within a field of view of the one or more cameras and separate from the computer system; responsive to detecting handwriting that includes a physical mark on the physical surface that is within the field of view of the one or more cameras and that is separate from the computer system, displaying digital text in the electronic document that corresponds to the handwriting that is within the field of view of the one or more cameras.
180. 1. A method comprising: A first computer system in communication with a display generating component, one or more cameras, and one or more input devices, detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within a field of view of the one or more cameras; In response to detecting the one or more first user inputs, pursuant to a determination that a first set of one or more criteria are satisfied, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and and simultaneously displaying a visual indication indicating a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
181. 181. The method of claim 180, wherein the visual representation of the first portion of the field of view of the one or more cameras and the visual indication of the first area of the field of view are displayed simultaneously while the first computer system does not share the second portion of the field of view of the one or more cameras with the second computer system.
182. 182. The method of claim 180 or 181, wherein the second portion of the field of view of the one or more cameras includes an image of a surface positioned between the one or more cameras and a user within the field of view of the one or more cameras.
183. 183. The method of any one of claims 180 to 182, wherein the surface comprises a vertical surface.
184. 184. A method according to any one of claims 180 to 183, wherein the view of the surface presented by the second computer system includes an image of the surface that has been modified based on the position of the surface relative to the one or more cameras.
185. 185. The method of any one of claims 180 to 184, wherein the first portion of the field of view of the one or more cameras includes an image of a user within the field of view of the one or more cameras.
186. upon detecting a change in the position of the one or more cameras, via the display generation component: a visual representation of a third portion of the field of view of the one or more cameras; and 186. The method of any one of claims 180 to 185, further comprising simultaneously displaying the visual indication, wherein the visual indication indicates a second region of the field of view of the one or more cameras that is a subset of the third portion of the field of view of the one or more cameras, and the second region indicates a fourth portion of the field of view of the one or more cameras that is presented as a view of the surface by the second computer system.
187. detecting one or more second user inputs via the one or more input devices while the one or more cameras are substantially stationary and while displaying the visual representation and the visual indication of the first portion of the field of view of the one or more cameras; in response to detecting the one or more second user inputs, via the display generation component, while the one or more cameras remain substantially stationary; the visual representation of the first portion of the field of view; and 187. The method of any one of claims 180 to 186, further comprising simultaneously displaying the visual indication and a third region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the third region indicating a fifth portion of the field of view different from the second portion, which is presented as a view of the surface by the second computer system.
188. detecting, while displaying the visual representation of the first portion of the field of view of the one or more cameras and the visual indication, user input via the one or more input devices directed at a control including a set of options for the visual indication; in response to detecting the user input directed to the control; 188. The method of any one of claims 180 to 187, further comprising: displaying the visual indication to show a fourth region of the field of view of the one or more cameras including a sixth portion of the field of view different from the second portion, as presented as a view of the surface by the second computer system.
189. in response to detecting the user input directed to the control; maintaining a position of a first portion of a boundary of a portion of the field of view presented as a view of the surface by the second computer system; 189. The method of claim 188, further comprising: modifying a position of a second portion of the boundary of the portion of the field of view presented as a view of the surface by the second computer system.
190. 190. The method of claim 189, wherein the first portion of the visual indication corresponds to the top edge of the second portion of the field of view presented as the view of the surface by the second computer system.
191. 191. The method of any one of claims 180 to 190, wherein the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras presented as the view of the surface by the second computer system are based on image data captured by a first camera.
192. detecting one or more third user inputs via the one or more input devices corresponding to a request to display the user interface of the application displaying a visual representation of a surface within the field of view of the one or more cameras; in response to detecting the one or more third user inputs; pursuant to a determination that the one or more first set of criteria are satisfied, via the display generation component: a visual representation of a seventh portion of the field of view of the one or more cameras; and 192. The method of any one of claims 180 to 191, further comprising simultaneously displaying a visual indication indicating a fifth region of the field of view of the one or more cameras that is a subset of the seventh portion of the field of view of the one or more cameras, the fifth region indicating an eighth portion of the field of view of the one or more cameras that is presented as a view of the surface by a third computer system different from the second computer system.
193. 193. The method of claim 192, wherein visual characteristics of the visual indication are user-configurable, and wherein the first computer system displays the visual indication indicating the fifth region as having visual characteristics based on visual characteristics of the visual indication used during recent use of the one or more cameras to present as a view of the surface by a remote computer system.
194. detecting, while displaying the visual representation of the first portion of the field of view of the one or more cameras and the visual indication, one or more fourth user inputs via the one or more input devices corresponding to a request to modify a visual characteristic of the visual indication; in response to detecting the one or more fourth user inputs; displaying the visual indication to show a sixth region of the field of view of the one or more cameras including a ninth portion different from the second portion of the field of view presented as a view of the surface by the second computer system; detecting one or more user inputs corresponding to a request to share a view of the surface while displaying the visual indication that the sixth region of the field of view of the one or more cameras, including the ninth portion of the field of view, is presented as a view of the surface by the second computer system; 194. The method of any one of claims 180 to 193, further comprising: in response to detecting the one or more user inputs corresponding to a request to share a view of the surface, sharing the ninth portion of the field of view for presentation by the second computer system.
195. In response to detecting the one or more first user inputs, in response to a determination that one or more second sets of criteria different from the one or more first sets of criteria are satisfied; 195. The method of any one of claims 180 to 194, further comprising displaying the second portion of the field of view as a view of the surface presented by the second computer system, the second portion of the field of view including an image of the surface modified based on the position of the surface relative to the one or more cameras.
196. 196. The method of any one of claims 180 to 195, further comprising, while providing the second portion of the field of view as a view of the surface for presentation by the second computer system, displaying, via the display generation component, controls for modifying the portion of the field of view of the one or more cameras that is presented as a view of the surface by the second computer system.
197. displaying, via the display generation component, the control for modifying the portion of the field of view of the one or more cameras that is presented as the view of the surface by the second computer system in accordance with a determination that focus is directed to an area corresponding to the view of the surface; 200. The method of claim 196, further comprising: ceasing to display the controls that modify the portion of the field of view of the one or more cameras that is presented as the view of the surface by the second computer system in accordance with a determination that the focus is not directed to the area corresponding to the view of the surface.
198. the second portion of the field of view includes a first boundary; detecting one or more fifth user inputs directed at the control to modify a portion of the field of view of the one or more cameras to be presented as a view of the surface by the second computer system; in response to detecting the one or more fifth user inputs; maintaining a position of the first boundary of the second portion of the field of view; 198. The method of claim 196 or 197, further comprising modifying the amount of the portion of the field of view that is included in the second portion of the field of view.
199. detecting one or more sixth user inputs via the one or more input devices while the one or more cameras are substantially stationary and while displaying the visual representation and the visual indication of the first portion of the field of view of the one or more cameras; in response to detecting the one or more sixth user inputs and while the one or more cameras remain substantially stationary, via the display generation component: a visual representation of an eleventh portion of the field of view of the one or more cameras that is different from the first portion of the field of view of the one or more cameras; and 202. The method of any one of claims 180 to 198, further comprising simultaneously displaying the visual indication indicating a seventh region of the field of view of the one or more cameras that is a subset of the eleventh portion of the field of view of the one or more cameras, the seventh region indicating a twelfth portion of the field of view that is different from the second portion and that is presented as a view of the surface by the second computer system.
200. Displaying the visual indication comprises: displaying the visual indication having a first appearance in accordance with a determination that a set of one or more alignment criteria is satisfied, including alignment criteria based on an alignment between a current region of the field of view of the one or more cameras indicated by the visual indication and a designated portion of the field of view of the one or more cameras; and displaying the visual indication having a second appearance different from the first appearance in accordance with a determination that the alignment criterion is not satisfied.
201. 201. The method of any one of claims 180 to 200, further comprising simultaneously displaying a visual representation of a thirteenth portion of the field of view of the one or more cameras and the visual indication while the visual indication indicates an eighth region of the field of view of the one or more cameras, and a target area indication indicating a first designated region of the field of view of the one or more cameras, the first designated region indicating a determined portion of the field of view of the one or more cameras based on the position of the surface within the field of view of the one or more cameras.
202. 202. The method of claim 201, wherein the target area indication is stationary relative to the surface.
203. 203. The method of claim 202, wherein the area of interest indication is selected based on an edge of the surface.
204. 204. The method of claim 202 or 203, wherein the area of interest indication is selected based on a position of a person within the field of view of the one or more cameras.
205. A method according to any one of claims 201 to 204, further comprising displaying the area of interest indication via the display generation component after detecting a change in the position of the one or more cameras, the area of interest indication indicating a second designated area of the field of view of the one or more cameras, the second designated area indicating a second determined portion of the field of view of the one or more cameras based on the position of the surface within the field of view of the one or more cameras after the change in position of the one or more cameras.
206. 206. The method of any one of claims 180 to 205, further comprising displaying a surface view representation of the surface in a ninth region of the field of view of the one or more cameras indicated by the visual indication, presented as a view of the surface by a second computer system simultaneously with the visual representation of the field of view of the one or more cameras and the visual indication, the surface view representation including an image of the surface captured by the one or more cameras modified based on the position of the surface relative to the one or more cameras to correct for the viewpoint of the surface.
207. 207. The method of claim 206, wherein displaying the surface view representation comprises displaying the surface view representation within a visual representation of a portion of the field of view of the one or more cameras that includes a person.
208. detecting a change in the field of view of the one or more cameras indicated by the visual indication after displaying the surface view representation of the surface in the ninth region of the field of view of the one or more cameras indicated by the visual indication; 208. The method of claim 207, further comprising: displaying the surface view representation in response to detecting the change in the field of view of the one or more cameras indicated by the visual indication, wherein the surface view representation includes the surface within the ninth region of the field of view of the one or more cameras indicated by the visual indication after the change in the field of view of the one or more cameras indicated by the visual indication.
209. 209. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 180 to 208.
210. a first computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices; one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 180 to 208.
211. a first computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices; 209. A first computer system comprising means for performing the method of any one of claims 180 to 208.
212. 209. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 180 to 208.
213. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; In response to detecting the one or more first user inputs, pursuant to a determination that a first set of one or more criteria are satisfied, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and and a visual indication indicating a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
214. a first computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices; one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; wherein the one or more programs are detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; In response to detecting the one or more first user inputs, pursuant to a determination that a first set of one or more criteria are satisfied, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and and a visual indication indicating a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
215. a first computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices; means for detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; In response to detecting the one or more first user inputs, pursuant to a determination that a first set of one or more criteria are satisfied, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and and a visual indication indicating a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
216. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising: detecting, via the one or more input devices, one or more first user inputs corresponding to a request to display a user interface of an application for displaying a visual representation of a surface within the field of view of the one or more cameras; In response to detecting the one or more first user inputs, pursuant to a determination that a first set of one or more criteria are satisfied, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and and a visual indication indicating a first region of the field of view of the one or more cameras that is a subset of the first portion of the field of view of the one or more cameras, the first region indicating a second portion of the field of view of the one or more cameras that is presented as a view of the surface by a second computer system.
217. 1. A method comprising: A computer system in communication with a display generation component and one or more input devices, comprising: detecting a request to use a function on the computer system via the one or more input devices; and in response to detecting the request to use the feature on the computer system, displaying, via the display generation component, a tutorial for using the feature, the tutorial including a virtual demonstration of the feature, wherein displaying the tutorial comprises: displaying the virtual demonstration having a first appearance in accordance with determining that the property of the computer system has a first value; and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
218. 218. The method of claim 217, wherein the virtual demonstration has an appearance based on what type of device is being used to provide access to the functionality.
219. 219. The method of claim 217 or 218, wherein the virtual demonstration has an appearance based on which model of device is being used to provide access to the functionality.
220. 220. The method of any one of claims 217 to 219, wherein the virtual demonstration has an appearance based on whether the computer system is coupled to an external device to provide access to the functionality.
221. 221. The method of claim 220, wherein, in accordance with a determination that the computer system is coupled to an external device, displaying the tutorial includes displaying a graphical representation of the external device in a selected orientation of a plurality of possible orientations.
222. 222. The method of any one of claims 217 to 221, wherein the virtual demonstration has an appearance based on a system language of the computer system.
223. 223. The method of any one of claims 217 to 222, wherein the virtual demonstration has an appearance based on a color associated with the computer system.
224. 224. The method of any one of claims 217 to 223, wherein displaying the tutorial comprises displaying a graphical indication of the range of field of view of one or more cameras in a simulated representation of a physical environment.
225. 225. The method of any one of claims 217 to 224, wherein displaying the tutorial comprises displaying a graphical representation of an input area and a graphical representation of an output area.
226. 226. The method of any one of claims 217 to 225, wherein displaying the tutorial comprises displaying a graphical representation of an input.
227. 227. The method of claim 226, wherein displaying the tutorial includes displaying a graphical representation of a first output of the input and a graphical representation of a second output of the input.
228. displaying the graphical representation of the first output includes displaying the graphical representation of the first output on a graphical representation of a physical surface; 228. The method of claim 226 or 227, wherein displaying the graphical representation of the second output comprises displaying the graphical representation of the second output on a graphical representation of the computer system.
229. displaying the graphical representation of the input includes displaying a graphical representation of a writing instrument making a mark; 229. The method of any one of claims 226 to 228, wherein displaying the tutorial includes displaying movement of the graphical representation of the writing instrument after display of the graphical representation of the input is completed.
230. Viewing the tutorial Displaying a graphical representation of a physical object from a first viewpoint at a first time; 230. The method of any one of claims 217 to 229, comprising: displaying the graphical representation of the physical object from a second perspective at a second time, the second perspective being different from the first perspective and the second time being different from the first time.
231. Viewing the tutorial displaying a graphical representation of the computer system from a first perspective at a first time; and displaying the graphical representation of the computer system from a second perspective at a second time, the second perspective being different from the first perspective and the second time being different from the first time.
232. Viewing the tutorial displaying a first virtual demonstration of the functionality; and displaying a second virtual demonstration of the feature after displaying the first virtual demonstration of the feature.
233. detecting a second request to use the function on the computer system; in response to detecting the second request to use the function on the computer system; displaying the tutorial for using the feature, including the virtual demonstration of the feature, in accordance with a determination that a set of criteria is met; and 233. The method of any one of claims 217 to 232, further comprising: ceasing to display the tutorial for using the feature, including the virtual demonstration of the feature, in accordance with a determination that the set of criteria is not met.
234. 234. The method of claim 233, wherein the set of criteria includes criteria that are met if the feature is used a threshold amount of times.
235. displaying selectable continuation options after detecting the request to use the feature on the computer system; and detecting a selection of the selectable continue option; 235. The method of any one of claims 217 to 234, further comprising: in response to detecting selection of the selectable continue option, executing a process on the computer system that uses the feature.
236. displaying selectable information options after detecting the request to use the feature on the computer system; and detecting a selection of the selectable information option; 236. The method of any one of claims 217 to 235, further comprising: in response to detecting a selection of the selectable information option, displaying a user interface that provides information for using the function on the computer system.
237. 236. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs including instructions for performing the method of any one of claims 217 to 235.
238. 1. A computer system configured to communicate with a display generation component and one or more input devices, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 217 to 235.
239. 1. A computer system configured to communicate with a display generation component and one or more input devices, comprising:
236. A computer system comprising means for carrying out the method of any one of claims 217 to 235.
240. 236. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising instructions for performing the method of any one of claims 217 to 235.
241. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising: detecting a request to use a function on the computer system via the one or more input devices; and instructions for displaying, via the display generation component, a tutorial for using the feature, the tutorial including a virtual demonstration of the feature, in response to detecting the request to use the feature on the computer system, wherein displaying the tutorial includes: displaying the virtual demonstration having a first appearance in accordance with determining that the property of the computer system has a first value; and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
242. 1. A computer system configured to communicate with a display generation component and one or more input devices, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: detecting a request to use a function on the computer system via the one or more input devices; and instructions for displaying, via the display generation component, a tutorial for using the feature, the tutorial including a virtual demonstration of the feature, in response to detecting the request to use the feature on the computer system, wherein displaying the tutorial includes: displaying the virtual demonstration having a first appearance in accordance with determining that the property of the computer system has a first value; and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
243. 1. A computer system configured to communicate with a display generation component and one or more input devices, comprising: means for detecting, via said one or more input devices, a request to use a function on said computer system; and means for displaying, via the display generation component, a tutorial for using the feature, the tutorial including a virtual demonstration of the feature, in response to detecting the request to use the feature on the computer system, the means for displaying comprising: means for displaying the virtual demonstration having a first appearance in accordance with a determination that the property of the computer system has a first value; means for displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
244. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, the one or more programs comprising: detecting a request to use a function on the computer system via the one or more input devices; and instructions for displaying, via the display generation component, a tutorial for using the feature, the tutorial including a virtual demonstration of the feature, in response to detecting the request to use the feature on the computer system, wherein displaying the tutorial includes: displaying the virtual demonstration having a first appearance in accordance with determining that the property of the computer system has a first value; and displaying the virtual demonstration having a second appearance different from the first appearance in accordance with a determination that the property of the computer system has a second value.
Citation Information
Patent Citations
Image input device and image transmitter using it
JP1997233384A
User interface for capturing and managing visual medium
JP2021040300A
Mobile communication terminal having touch screen and method of controlling display thereof
US20090046075A1
Systems and methods for configuring a HUB-centric virtual / augmented reality environment
US20210333864A1
Designated view within a multi-view composited webcam signal
US20220046186A1