Wide angle video conferencing
By communicating with display generation components, cameras and input devices in a computer system, a faster and more efficient real-time video communication interface is realized, solving the complex and time-consuming problems of user interfaces in the prior art, improving efficiency and saving device energy.
Patent Information
- Application Number
- CN202510448790.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-22
- Filing Date
- 2022-09-23
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art When managing real-time video communication sessions, the user interface is complex and time-consuming, resulting in wasting user time and device energy.
By communicating with display generation components, cameras, and input devices in a computer system, a faster and more efficient real-time video communication interface includes displaying a partial representation of the camera's field of view and modifying the image of the surface in response to user input.
Reduces user cognitive burden, improves the efficiency of the human-computer interface, and saves power in battery-powered devices, extends battery charging intervals.
Smart Images

Figure CN120075385A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention titled "Wide Angle Video Conference" with the application date of September 23, 2022, application number 202280037932.6.
[0002] Cross - Reference to Related Applications
[0003] This application claims priority to U.S. Patent Application No. 17 / 950,868, titled "WIDE ANGLE VIDEO CONFERENCE", filed on September 22, 2022; and claims priority to U.S. Patent Application No. 17 / 950,900, titled "WIDE ANGLE VIDEO CONFERENCE", filed on September 22, 2022; and claims priority to U.S. Patent Application No. 17 / 950,922, titled "WIDE ANGLE VIDEO CONFERENCE", filed on September 22, 2022; and claims priority to U.S. Provisional Patent Application No. 63 / 392,096, titled "WIDE ANGLE VIDEO CONFERENCE", filed on July 25, 2022; and claims priority to U.S. Provisional Patent Application No. 63 / 357,605, titled "WIDE ANGLE VIDEO CONFERENCE", filed on June 30, 2022; and claims priority to U.S. Provisional Patent Application No. 63 / 349,134, titled "WIDE ANGLE VIDEO CONFERENCE", filed on June 5, 2022; and claims priority to U.S. Provisional Patent Application No. 63 / 307,780, titled "WIDE ANGLE VIDEO CONFERENCE", filed on February 8, 2022; and claims priority to U.S. Provisional Patent Application No. 63 / 248,137, titled "WIDE ANGLE VIDEO CONFERENCE", filed on September 24, 2021. The content of each of these patent applications is hereby incorporated by reference in its entirety. Technical Field
[0004] The present disclosure generally relates to computer user interfaces, and more particularly, to techniques for managing real-time video communication sessions and / or managing digital content. Background Art
[0005] A computer system may include hardware and / or software for displaying an interface for a real-time video communication session. Summary of the Invention
[0006] However, some techniques for using an electronic device to manage a real-time video communication session are generally cumbersome and inefficient. For example, some prior art uses complex and time-consuming user interfaces, which may include multiple button presses or keystrokes. The prior art takes more time than necessary, which results in wasted user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0007] Accordingly, the present technology provides faster and more efficient methods and interfaces for an electronic device to manage a real-time video communication session and / or manage digital content. Such methods and interfaces optionally supplement or replace other methods for managing a real-time video communication session and / or managing digital content. Such methods and interfaces reduce the cognitive burden imposed on the user and result in a more effective human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.
[0008] According to some embodiments, a method is described that is performed at a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The method includes: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting, via the one or more input devices while the real-time video communication interface is being displayed, one or more user inputs, the one or more user inputs including a user input that points to a surface in a scene within the field of view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras that is modified based on the position of the surface relative to the one or more cameras.
[0009] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting, while displaying the real-time video communication interface, one or more user inputs via the one or more input devices, the one or more user inputs including a user input pointing to a surface in a scene within the field of view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras modified based on the position of the surface relative to the one or more cameras.
[0010] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting, while displaying the real-time video communication interface, one or more user inputs via the one or more input devices, the one or more user inputs including a user input pointing to a surface in a scene within the field of view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras modified based on the position of the surface relative to the one or more cameras.
[0011] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting, while displaying the real-time video communication interface, one or more user inputs via the one or more input devices, the one or more user inputs including a user input pointing to a surface in a scene within the field of view of the one or more cameras; and in response to detecting the one or more user inputs, displaying, via the display generation component, a representation of the surface, wherein the representation of the surface includes an image of the surface captured by the one or more cameras modified based on the position of the surface relative to the one or more cameras.
[0012] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: means for displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of a first portion of a scene in the field of view captured by the one or more cameras; and means for obtaining, while displaying the real-time video communication interface, image data of the field of view of the one or more cameras via the one or more cameras, the image data including a first gesture; and in response to obtaining the image data of the field of view of the one or more cameras, means for: displaying, via the display generation component, a representation of a second portion of a scene in the field of view of the one or more cameras according to a determination that the first gesture meets a first set of criteria, the representation of the second portion of the scene including visual content different from the representation of the first portion of the scene; and continuing to display, via the display generation component, the representation of the first portion of the scene according to a determination that the first gesture meets a second set of criteria different from the first set of criteria.
[0013] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a real-time video communication interface for a real-time video communication session, the real-time video communication interface including a representation of a first portion of a scene in a field of view captured by the one or more cameras; and while displaying the real-time video communication interface, obtaining, via the one or more cameras, image data of a field of view of the one or more cameras, the image data including a first gesture; and in response to obtaining the image data of the field of view of the one or more cameras: displaying, via the display generation component, a representation of a second portion of the scene in the field of view of the one or more cameras based on determining that the first gesture meets a first set of criteria, the representation of the second portion of the scene including visual content different from the representation of the first portion of the scene; and continuing to display, via the display generation component, the representation of the first portion of the scene based on determining that the first gesture meets a second set of criteria different from the first set of criteria.
[0014] According to some embodiments, a method is described for execution at a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The method includes: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including a plurality of participants; and in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of a field of view of the one or more first cameras of the first computer system; a second representation of a field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of a field of view of one or more second cameras of a second computer system; and a second representation of a field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0015] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0016] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0017] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0018] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: means for detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; means for, in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0019] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request for a user interface for displaying a real-time video communication session including multiple participants; and in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0020] According to some embodiments, a method is described that is performed at a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The method includes: detecting a set of one or more user inputs corresponding to a request for a user interface for displaying a real-time video communication session including multiple participants; and in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene in the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene in the field of view of the one or more second cameras of the second computer system.
[0021] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0022] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more first cameras, and one or more input devices, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0023] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0024] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: means for detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; means for, in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0025] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request for a user interface displaying a real-time video communication session including a plurality of participants; in response to detecting the set of one or more user inputs, displaying, via the display generation component, a real-time video communication interface for the real-time video communication session, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of a surface in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of the one or more second cameras of the second computer system; and a second representation of the field of view of the one or more second cameras of the second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of a surface in a second scene within the field of view of the one or more second cameras of the second computer system.
[0026] According to some embodiments, a method is described. The method includes: at a first computer system in communication with a first display generation component and one or more sensors: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment within the field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment within the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment within the field of view of the one or more cameras of the second computer system.
[0027] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a first display generating component and one or more sensors, the one or more programs including instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generating component, a representation of a first view of a physical environment in a field of view of one or more cameras of the second computer system; when displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0028] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a first display generating component and one or more sensors, the one or more programs including instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generating component, a representation of a first view of a physical environment in a field of view of one or more cameras of the second computer system; when displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0029] According to some embodiments, a computer system configured to communicate with a first display generation component and one or more sensors is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment in a field of view of one or more cameras of the second computer system; when displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0030] According to some embodiments, a computer system configured to communicate with a first display generation component and one or more sensors is described. The computer system includes means for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generation component, a representation of a first view of a physical environment in a field of view of one or more cameras of the second computer system; when displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generation component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0031] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a first display generating component and one or more sensors, the one or more programs including instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generating component, a representation of a first view of a physical environment in a field of view of one or more cameras of the second computer system; when displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0032] According to some embodiments, a method is described. The method includes: at a computer system in communication with a display generating component: displaying, via the display generating component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; when displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras, obtaining data including a new physical marker in the physical environment; and in response to obtaining the data representing the new physical marker in the physical environment, displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras.
[0033] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; obtaining data including a new physical marker in the physical environment while displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras; and in response to obtaining data representing the new physical marker in the physical environment, displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras.
[0034] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; obtaining data including a new physical marker in the physical environment while displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras; and in response to obtaining data representing the new physical marker in the physical environment, displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras.
[0035] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; obtaining data including a new physical marker in the physical environment when displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras; and in response to obtaining data representing a new physical marker in the physical environment, displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras.
[0036] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: means for displaying, via the display generation component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; means for obtaining data including a new physical marker in the physical environment when displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras; and means for displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras in response to obtaining data representing a new physical marker in the physical environment.
[0037] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs including instructions for: displaying, via the display generation component, a representation of a physical marker in a physical environment based on a view of the physical environment in a field of view of one or more cameras, wherein: the view of the physical environment includes the physical marker and a physical background, and displaying the representation of the physical marker includes displaying the representation of the physical marker without displaying one or more elements of a portion of the physical background in the field of view of the one or more cameras; obtaining, when displaying the representation of the physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras, data including a new physical marker in the physical environment; and in response to obtaining data representing the new physical marker in the physical environment, displaying a representation of the new physical marker without displaying the one or more elements of the portion of the physical background in the field of view of the one or more cameras.
[0038] According to some embodiments, a method is described. The method includes, at a computer system that communicates with a display generation component and one or more cameras: displaying an electronic document via the display generation component; detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and is separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0039] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more cameras, the one or more programs including instructions for: displaying an electronic document via the display generation component; detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and is separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0040] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more cameras. The one or more programs include instructions for: displaying an electronic document via the display generation component; detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and is separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0041] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: displaying an electronic document via the display generation component; detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and is separate from the computer system; and in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and is separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0042] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes: means for displaying an electronic document via the display generation component; means for detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and is separate from the computer system; and means for, in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and is separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0043] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more cameras. The one or more programs include instructions for: displaying an electronic document via the display generation component; detecting, via the one or more cameras, handwriting of a physical marker on a physical surface that is included in the field of view of the one or more cameras and separate from the computer system; and in response to detecting the handwriting of the physical marker on the physical surface that is included in the field of view of the one or more cameras and separate from the computer system, displaying digital text corresponding to the handwriting in the field of view of the one or more cameras in the electronic document.
[0044] According to some embodiments, a method is described for execution at a first computer system that communicates with a display generation component, one or more cameras, and one or more input devices. The method includes: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application for displaying a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more first user inputs: based on determining that a first set of one or more criteria is satisfied, simultaneously displaying via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that is to be presented by a second computer system as a view of the surface.
[0045] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application, the user interface for displaying a visual representation of a surface in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: based on determining that a first set of one or more criteria is met, simultaneously displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion to be presented by a second computer system as a view of the surface.
[0046] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application, the user interface for displaying a visual representation of a surface in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: based on determining that a first set of one or more criteria is met, simultaneously displaying, via the display generation component: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion to be presented by a second computer system as a view of the surface.
[0047] According to some embodiments, a first computer system is described that is configured to communicate with a display generation component, one or more cameras, and one or more input devices. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application, the user interface for displaying a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, based on determining that a first set of one or more criteria is met: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented by a second computer system as a view of the surface.
[0048] According to some embodiments, a first computer system is described that is configured to communicate with a display generation component, one or more cameras, and one or more input devices. The computer system includes: means for detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application, the user interface for displaying a visual representation of a surface in the field of view of the one or more cameras; and means for, in response to detecting the one or more first user inputs, performing the following operations: simultaneously displaying, via the display generation component, based on determining that a first set of one or more criteria is met: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented by a second computer system as a view of the surface.
[0049] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface of a display application, the user interface for displaying a visual representation of a surface in a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, based on determining that a first set of one or more criteria is met: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion to be presented by a second computer system as a view of the surface.
[0050] According to some embodiments, a method is described. The method includes: at a computer system communicating with a display generation component and one or more input devices: detecting, via the one or more input devices, a request to use a function on the computer system; and in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function including a virtual demonstration of the function, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0051] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicating with a display generation component and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a request to use a function on the computer system; and in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function including a virtual demonstration of the function, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0052] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices. The one or more programs include instructions for: detecting, via the one or more input devices, a request to use a function on the computer system; and in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function that includes a virtual demonstration of the function, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0053] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: detecting, via the one or more input devices, a request to use a function on the computer system; and in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function that includes a virtual demonstration of the function, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0054] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system includes: means for detecting, via the one or more input devices, a request to use a function on the computer system; and means for displaying, in response to detecting the request to use the function on the computer system, via the display generation component, a tutorial for using the function that includes a virtual demonstration of the function, including: means for displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and means for displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0055] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more input devices. The one or more programs include instructions for: detecting, via the one or more input devices, a request to use a function on the computer system; and in response to detecting the request to use the function on the computer system, displaying, via the display generation component, a tutorial for using the function that includes a virtual demonstration of the function, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that the attribute of the computer system has a second value.
[0056] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are optionally included in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0057] Thus, a faster and more efficient method and interface for managing real-time video communication sessions are provided for a device, thereby enhancing the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces can supplement or replace other methods for managing real-time video communication sessions. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals indicate corresponding parts in all the drawings.
[0059] Figure 1A is a block diagram showing a portable multifunctional device having a touch-sensitive display in accordance with some embodiments.
[0060] Figure 1B is a block diagram showing exemplary components for event handling in accordance with some embodiments.
[0061] Figure 2 shows a portable multifunctional device having a touch screen in accordance with some embodiments.
[0062] Figure 3 is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface in accordance with some embodiments.
[0063] Figure 4A shows an exemplary user interface for a menu of an application on a portable multifunctional device in accordance with some embodiments.
[0064] Figure 4B Shows an exemplary user interface for a multifunctional device having a touch-sensitive surface separate from a display, according to some embodiments.
[0065] Figure 5A Shows a personal electronic device according to some embodiments.
[0066] Figure 5B Is a block diagram showing a personal electronic device according to some embodiments.
[0067] Figure 5C Shows an exemplary diagram of a communication session between electronic devices according to some embodiments.
[0068] Figures 6A to 6AY Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.
[0069] Figure 7 Shows a flowchart illustrating a method for managing a real-time video communication session according to some embodiments.
[0070] Figure 8 Shows a flowchart illustrating a method for managing a real-time video communication session according to some embodiments.
[0071] Figures 9A to 9T Shows an exemplary user interface for managing a real-time video communication session according to some embodiments.
[0072] Figure 10 Shows a flowchart illustrating a method for managing a real-time video communication session according to some embodiments.
[0073] Figures 11A to 11P Shows an exemplary user interface for managing digital content according to some embodiments.
[0074] Figure 12 Is a flowchart illustrating a method for managing digital content according to some embodiments.
[0075] Figures 13A to 13K Shows an exemplary user interface for managing digital content according to some embodiments.
[0076] Figure 14 Is a flowchart illustrating a method for managing digital content according to some embodiments.
[0077] Figure 15 Shows a flowchart illustrating a method for managing a real-time video communication session according to some embodiments.
[0078] Figures 16A to 16QShows an exemplary user interface for managing a real-time video communication session according to some embodiments.
[0079] Figure 17 Is a flowchart showing a method for managing a real-time video communication session according to some embodiments.
[0080] Figures 18A to 18N Shows an exemplary user interface for displaying a tutorial on functions on a computer system according to some embodiments.
[0081] Figure 19 Is a flowchart showing a method for displaying a tutorial on functions on a computer system according to some embodiments. Detailed Description
[0082] The following description sets forth exemplary methods, parameters, etc. However, it should be recognized that such description is not intended to limit the scope of the present disclosure, but rather is provided as a description of exemplary embodiments.
[0083] There is a need for electronic devices that provide effective methods and interfaces for managing real-time video communication sessions and / or managing digital content. For example, there is a need for electronic devices to improve content sharing. Such technologies can reduce the cognitive burden on users sharing content during a real-time video communication session and / or managing digital content in an electronic document, thereby increasing productivity. In addition, such technologies can reduce processor power and battery power otherwise wasted on redundant user input.
[0084] Below, Figures 1A to 1B 、 Figure 2 、 Figure 3 、 Figures 4A to 4B and Figures 5A to 5C Provide a description of exemplary devices for performing techniques for managing real-time video communication sessions and / or managing digital content. Figures 6A to 6AY Shows an exemplary user interface for managing a real-time video communication session. Figures 7 to 8 and Figure 15 Is a flowchart showing a method for managing a real-time video communication session according to some embodiments. Figures 6A to 6AY The user interface in Figures 7 to 8 and Figure 15 is used to show the processes described below including the processes in Figures 9A to 9T Shows an exemplary user interface for managing real-time video communication. Figure 10 Is a flowchart showing a method for managing real-time video communication according to some embodiments. Figures 9A to 9T The user interface in Figure 10 is used to show the processes described below including the processes in Figures 11A to 11P Shows an exemplary user interface for managing digital content.Figure 12 is a flowchart showing a method of managing digital content according to some embodiments. Figures 11A to 11P The user interface in is used to show including Figure 12 the processes described below of the process in. Figures 13A to 13K An exemplary user interface for managing digital content according to some embodiments is shown. Figure 14 is a flowchart showing a method of managing digital content according to some embodiments. Figures 13A to 13K The user interface in is used to show including Figure 14 the processes described below of the process in. Figures 16A to 16O An exemplary user interface for managing a real-time video communication session according to some embodiments is shown. Figure 17 is a flowchart showing a method of managing a real-time video communication session according to some embodiments. Figures 16A to 16Q The user interface in is used to show including Figure 17 the processes described below of the process in. Figures 18A to 18N An exemplary user interface for displaying a tutorial of functions on a computer system according to some embodiments is shown. Figure 19 is a flowchart showing a method of displaying a tutorial of functions on a computer system according to some embodiments. Figures 18A to 18N The user interface in is used to show including Figure 19 the processes described below of the process in.
[0085] The processes described below enhance the operability of the device and make the user-device interface more effective through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including by providing the user with improved visual feedback, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving the efficiency of managing digital content, improving collaboration between users in a real-time communication session, improving the real-time communication session experience, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more effectively.
[0086] In addition, in a method where one or more of the steps described herein depend on one or more conditions being met, it should be understood that the method can be repeated in multiple iterations such that, during the repetition, all the conditions that determine the steps in the method are met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and a second step (if the condition is not met), one of ordinary skill in the art will know to repeat the stated steps until both the condition being met and the condition not being met (in no particular order) occur. Thus, a method described as having one or more steps that depend on one or more conditions being met can be rewritten as a method that repeats until each condition described in the method is met. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium contains instructions for performing the conditional operations based on the satisfaction of the corresponding one or more conditions and thus be able to determine whether the possible conditions have been met without explicitly repeating the steps of the method until all the conditions that determine the steps in the method are met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all the conditional steps have been performed.
[0087] Although the following description uses the terms "first", "second", etc. to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch could be named a second touch and similarly a second touch could be named a first touch without departing from the scope of the various described embodiments. In some embodiments, the first touch and the second touch are two separate references to the same touch. In some embodiments, both the first touch and the second touch are touches, but they are not the same touch.
[0088] The terms used in the description of the various described embodiments herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0089] Depending on the context, the term "if" is optionally interpreted to mean "when", "upon", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined that..." or "if [stated condition or event] is detected" is optionally interpreted to mean "when it is determined that..." or "in response to determining that..." or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]".
[0090] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions such as PDA and / or music player functions. Exemplary embodiments of the portable multifunctional device include, but are not limited to, devices from Apple Inc. (Cupertino, California), devices, iPod devices, and devices. Other portable electronic devices, such as a laptop or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad), are optionally used. It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). In some embodiments, the electronic device is a computer system that communicates (e.g., via wireless communication, via wired communication) with a display generation component. The display generation component is configured to provide a visual output, such as a display via a CRT monitor, a display via an LED monitor, or a display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes displaying content (e.g., video data rendered or decoded by a display controller 156) by transmitting data (e.g., image data or video data) to an integrated or external display generation component via a wired or wireless connection to visually generate the content.
[0091] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick.
[0092] The device generally supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, gaming applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, and / or digital video player applications.
[0093] The various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or within the respective applications. Thus, a common physical architecture of the device (such as a touch-sensitive surface) optionally supports a variety of applications with a user interface that is intuitive and clear to the user.
[0094] Attention is now turned to an embodiment of a portable device having a touch-sensitive display. Figure 1A FIG. is a block diagram of a portable multifunctional device 100 having a touch-sensitive display system 112 in accordance with some embodiments. The touch-sensitive display 112 is sometimes called a "touch screen" for convenience and is sometimes referred to as or called a "touch-sensitive display system". The device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. The device 100 optionally includes one or more optical sensors 164. The device 100 optionally includes one or more contact intensity sensors 165 for detecting the intensity of a contact on the device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100). The device 100 optionally includes one or more tactile output generators 167 for generating tactile output on the device 100 (e.g., generating tactile output on a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100 or the touchpad 355 of the device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0095] As used in this specification and the claims, the “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on the touch-sensitive surface, or to a surrogate for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using a variety of methods and a variety of sensors or combinations of sensors. For example, one or more force sensors beneath or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in size of the contact area detected on the touch-sensitive surface, the capacitance and / or change in capacitance of the touch-sensitive surface near the contact, and / or the resistance and / or change in resistance of the touch-sensitive surface near the contact are optionally used as surrogates for the force or pressure of a contact on the touch-sensitive surface. In some embodiments, the surrogate measurements of contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measurements). In some embodiments, the surrogate measurements of contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of user input allows a user to access additional device functions that would otherwise be inaccessible to the user on a smaller device with limited footprint, the smaller device being used to (e.g., on a touch-sensitive display) display affordances and / or receive user input (e.g., via a touch-sensitive display, a touch-sensitive surface, or physical / mechanical controls such as knobs or buttons).
[0096] As used in this specification and the claims, the term "haptic output" refers to a physical displacement of the device relative to a previous portion of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, detected by a user using the user's sense of touch. For example, in the case of contact between the device or a component of the device and a surface of the user that is sensitive to touch (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a sense of touch that corresponds to a perceived change in the physical characteristics of the device or the component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or a touchpad) is optionally interpreted by the user as a "press click" or "release click" of a physical actuation button. In some cases, the user will feel a sense of touch, such as a "press click" or "release click", even when the physical actuation button associated with the touch-sensitive surface that is physically depressed (e.g., displaced) by the user's movement does not move. As another example, even when there is no change in the smoothness of the touch-sensitive surface, movement of the touch-sensitive surface is optionally interpreted or sensed by the user as "roughness" of the touch-sensitive surface. Although such interpretations of touch by the user will be limited by the user's individual sensory perception, many sensory perceptions of touch are common to most users. Thus, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated haptic output corresponds to a physical displacement of the device or a component thereof that would generate the described sensory perception of a typical (or average) user.
[0097] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0098] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.
[0099] The peripheral device interface 118 can be used to couple input and output peripheral devices of the device to the CPU 120 and the memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in the memory 102 to perform various functions of the device 100 and process data. In some embodiments, the peripheral device interface 118, the CPU 120, and the memory controller 122 are optionally implemented on a single chip such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0100] The RF (Radio Frequency) circuit 108 receives and transmits RF signals, which are also referred to as electromagnetic signals. The RF circuit 108 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with the communication network and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 108 optionally communicates with the network and other devices via wireless communication, and these networks are such as the Internet (also known as the World Wide Web (WWW)), an intranet, and / or a wireless network (such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN)). The RF circuit 108 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as via a short-range communication radio component. The wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution-Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long-Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), WiMAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the date of submission of this document.
[0101] The audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuitry 110 receives audio data from the peripheral interface 118, converts the audio data into an electrical signal, and transmits the electrical signal to the speaker 111. The speaker 111 converts the electrical signal into sound waves audible to humans. The audio circuitry 110 also receives the electrical signal converted from sound waves by the microphone 113. The audio circuitry 110 converts the electrical signal into audio data and transmits the audio data to the peripheral interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuitry 108 by the peripheral interface 118. In some embodiments, the audio circuitry 110 also includes an earphone jack (e.g., Figure 2 212 in
[0102] The I / O subsystem 106 couples input / output peripheral devices on the device 100, such as the touch screen 112 and other input control devices 116, to the peripheral interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / transmit electrical signals to the other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some embodiments, the input controller 160 is optionally coupled to (or not coupled to) any of the following: a keyboard, an infrared port, a USB port, and a pointing device such as a mouse. One or more buttons (e.g., Figure 2 208 in Figure 2206). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication, via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some embodiments, an air gesture is detected when the user does not touch an input element that is part of the device (or independently of an input element that is part of the device) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0103] Quickly pressing the depress button optionally disengages the lock of the touch screen 112 or optionally starts the process of unlocking the device using gestures on the touch screen, as described in U.S. Patent Application No. 11 / 322,549, filed Dec. 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image" (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. Long pressing the depress button (e.g., 206) optionally powers on or powers off the device 100. The functions of the one or more buttons are optionally user-customizable. The touch screen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0104] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from / to the touch screen 112. The touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0105] The touch screen 112 has a touch-sensitive surface, sensor, or group of sensors that accepts input from the user based on haptic and / or tactile contact. The touch screen 112 and the display controller 156 (along with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of that contact) on the touch screen 112 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.
[0106] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, but uses other display technologies in other embodiments. The touch screen 112 and the display controller 156 optionally use any of a variety of touch sensing technologies now known or later to be developed, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112 to detect contact and any movement or interruption thereof, the variety of touch sensing technologies including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as the technology used in and iPod used by.
[0107] The touch-sensitive display in some embodiments of the touch screen 112 optionally resembles the multi-touch sensitive touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, the touch screen 112 displays visual output from the device 100, while the touch-sensitive touchpad does not provide visual output.
[0108] Touch-sensitive displays in some embodiments of the touch screen 112 are described in the following applications: (1) U.S. Patent Application 11 / 381,313, "Multipoint Touch Surface Controller," filed May 2, 2006; (2) U.S. Patent Application 10 / 840,862, "Multipoint Touchscreen," filed May 6, 2004; (3) U.S. Patent Application 10 / 903,964, "Gestures For Touch Sensitive Input Devices," filed Jul. 30, 2004; (4) U.S. Patent Application 11 / 048,264, "Gestures For Touch Sensitive Input Devices," filed Jan. 31, 2005; (5) U.S. Patent Application 11 / 038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices," filed Jan. 18, 2005; (6) U.S. Patent Application 11 / 228,758, "Virtual Input Device Placement On A Touch Screen User Interface," filed Sep. 16, 2005; (7) U.S. Patent Application 11 / 228,700, "Operation Of A Computer With A Touch Screen Interface," filed Sep. 16, 2005; (8) U.S. Patent Application 11 / 228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," filed Sep. 16, 2005; and (9) U.S. Patent Application 11 / 367,749, "Multi-Functional Hand-Held Device," filed Mar. 3, 2006. All of these applications are hereby incorporated by reference in their entirety.
[0109] The touch screen 112 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user optionally uses any suitable object or attachment such as a stylus, finger, etc. to contact the touch screen 112. In some embodiments, the user interface is designed to work primarily through finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts finger-based rough input into precise pointer / cursor positions or commands for performing the actions desired by the user.
[0110] In some embodiments, in addition to the touch screen, the device 100 optionally further includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device, which, unlike the touch screen, does not display a visual output. The touchpad is optionally a touch-sensitive surface separate from the touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.
[0111] The device 100 further includes a power system 162 for powering various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0112] The device 100 optionally further includes one or more optical sensors 164. Figure 1AAn optical sensor coupled to the optical sensor controller 158 in the I / O subsystem 106 is shown. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In combination with the imaging module 143 (also referred to as the camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite the touch screen display 112 on the front of the device, such that the touch screen display can be used as a viewfinder for still image and / or video image capture. In some embodiments, the optical sensor is located on the front of the device such that an image of the user can optionally be captured for a video conference while the user views other video conference participants on the touch screen display. In some embodiments, the location of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that a single optical sensor 164 can be used with the touch screen display for both video conferencing and still image and / or video image capture.
[0113] The device 100 optionally further includes one or more depth camera sensors 175. Figure 1A A depth camera sensor coupled to the depth camera controller 169 in the I / O subsystem 106 is shown. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of an object (e.g., a face) within a scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in combination with the imaging module 143 (also referred to as the camera module), the depth camera sensor 175 is optionally used to determine depth maps of different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located on the front of the device 100 such that an image of the user with depth information can optionally be captured for a video conference while the user views other video conference participants on the touch screen display, and a selfie with depth map data can be captured. In some embodiments, the depth camera sensor 175 is located on the rear of the device, or on both the rear and the front of the device 100. In some embodiments, the location of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that the depth camera sensor 175 can be used with the touch screen display for both video conferencing and still image and / or video image capture.
[0114] In some embodiments, a depth map (e.g., a depth map image) contains information (e.g., values) related to the distance of objects in a scene from a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines a location in the Z-axis of the viewpoint where its corresponding two-dimensional pixel is located. In some embodiments, a depth map is composed of pixels, where each pixel is defined by a value (e.g., from 0 to 255). For example, a value of "0" represents a pixel that is the farthest from the viewpoint (e.g., a camera, an optical sensor, a depth camera sensor) in a "three-dimensional" scene, and a value of "255" represents a pixel that is the closest to the viewpoint in the "three-dimensional" scene. In other embodiments, a depth map represents the distance between an object in a scene and a plane of the viewpoint. In some embodiments, a depth map includes information about the relative depth of various features of an object of interest in the field of view of a depth camera (e.g., the relative depth of the eyes, nose, mouth, ears of a user's face). In some embodiments, a depth map includes information that enables a device to determine the profile of an object of interest in the z-direction.
[0115] Device 100 optionally further includes one or more contact intensity sensors 165. Figure 1A Shown is a contact intensity sensor coupled to an intensity sensor controller 159 in I / O subsystem 106. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-mechanical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the rear of device 100, opposite to the touchscreen display 112 located on the front of device 100.
[0116] Device 100 optionally further includes one or more proximity sensors 166. Figure 1AA proximity sensor 166 is shown coupled to the peripheral device interface 118. Alternatively, the proximity sensor 166 is optionally coupled to an input controller 160 in the I / O subsystem 106. The proximity sensor 166 optionally operates as described in the following U.S. patent applications: No. 11 / 241,839, titled "Proximity Detector In Handheld Device"; No. 11 / 240,788, titled "Proximity Detector In Handheld Device"; No. 11 / 620,702, titled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; No. 11 / 586,862, titled "Automated Response To And Sensing Of User Activity In Portable Devices"; and No. 11 / 638,251, titled "Methods And Systems For Automatic Configuration Of Peripherals", which are hereby incorporated by reference in their entirety. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touch screen 112.
[0117] Device 100 optionally further includes one or more haptic output generators 167. Figure 1AShows a haptic output generator coupled to a haptic feedback controller 161 in the I / O subsystem 106. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting an electrical signal into a haptic output on the device). The contact intensity sensor 165 receives haptic feedback generation instructions from the haptic feedback module 133 and generates a haptic output on the device 100 that can be sensed by a user of the device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., the touch-sensitive display system 112), and optionally generates a haptic output by moving the touch-sensitive surface vertically (e.g., into / out of the surface of the device 100) or laterally (e.g., backward and forward in the same plane as the surface of the device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.
[0118] The device 100 optionally further includes one or more accelerometers 168. Figure 1A Shows an accelerometer 168 coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to an input controller 160 in the I / O subsystem 106. The accelerometer 168 optionally operates as described in the following U.S. Patent Publications: U.S. Patent Publication No. 20050190059, titled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication No. 20060017692, titled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", both of which are incorporated herein by reference in their entireties. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on an analysis of data received from one or more accelerometers. The device 100 optionally further includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver in addition to the accelerometer 168 for obtaining information about the location and orientation (e.g., portrait or landscape) of the device 100.
[0119] In some embodiments, the software components stored in the memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and application programs (or instruction sets) 136. Additionally, in some embodiments, the memory 102 ( Figure 1A ) or 370 ( Figure 3 ) stores a device / global internal state 157, as shown in Figure 1A and Figure 3 . The device / global internal state 157 includes one or more of the following: an active application state, which indicates which applications (if any) are currently active; a display state, indicating what applications, views, or other information occupy the respective regions of the touchscreen display 112; a sensor state, including information obtained from the various sensors and input control devices 116 of the device; and location information related to the location and / or orientation of the device.
[0120] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between the various hardware components and software components.
[0121] The communication module 128 facilitates communication with other devices through one or more external ports 124, and also includes various software components for processing data received by the RF circuit 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices, or indirectly coupled through a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is the same as or similar to and / or compatible with the 30-pin connector used on (a trademark of Apple Inc.) devices and / or a multi-pin (e.g., 30-pin) connector.
[0122] The contact / motion module 130 optionally detects contact with the touch screen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or a physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has ceased (e.g., detecting a finger lift event or contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of the contact point being represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., “multi-touch” / multiple finger contact). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.
[0123] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by the user (e.g., determining whether the user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and can be adjusted without changing the physical hardware of the device 100). For example, the mouse “click” threshold of a touchpad or touch screen can be set to any one of a wide range of predefined thresholds without changing the touchpad or touch screen display hardware. Additionally, in some implementations, software settings are provided to the user of the device for adjusting one or more of the intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using a system-level click on an “intensity” parameter).
[0124] The touch / motion module 130 optionally detects gesture inputs made by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of the detected contacts). Thus, gestures are optionally detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift (lift-off) event.
[0125] The graphics module 132 includes various known software components for presenting and displaying graphics on the touch screen 112 or other display, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual attributes). As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0126] In some embodiments, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives from an application, etc., one or more codes for specifying the graphic to be displayed, and also receives coordinate data and other graphic attribute data as necessary, and then generates screen image data for output to the display controller 156.
[0127] The haptic feedback module 133 includes various software components for generating instructions that are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.
[0128] The text input module 134, which is optionally a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).
[0129] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., provided to the phone 138 for location-based dialing; provided to the camera 143 as picture / video metadata; and provided to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).
[0130] The application 136 optionally includes the following modules (or instruction sets), or subsets or supersets thereof:
[0131] · A contacts module 137 (sometimes referred to as an address book or contacts list);
[0132] · A phone module 138;
[0133] · A video conferencing module 139;
[0134] · An email client module 140;
[0135] · An instant messaging (IM) module 141;
[0136] · A fitness support module 142;
[0137] · A camera module 143 for still images and / or video images;
[0138] · An image management module 144;
[0139] · A video player module;
[0140] · A music player module;
[0141] · A browser module 147;
[0142] · A calendar module 148;
[0143] · A widget module 149, which optionally includes one or more of the following: a weather widget 149-1, a stock market widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6;
[0144] · A widget creator module 150 for forming user-created widgets 149-6;
[0145] · A search module 151;
[0146] · A video and music player module 152, which combines the video player module and the music player module;
[0147] · A notes module 153;
[0148] · A maps module 154; and / or
[0149] · An online video module 155.
[0150] Examples of other application programs 136 that are optionally stored in the memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice reproduction.
[0151] In conjunction with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the contacts module 137 is optionally used to manage an address book or contact list (e.g., in the application internal state 192 of the contacts module 137 stored in the memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating a phone number, email address, physical address, or other information with a name; associating an image with a name; categorizing and classifying names; providing a phone number or email address to initiate and / or facilitate communication via the phone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0152] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the phone module 138 is optionally used to input a character sequence corresponding to a phone number, access one or more phone numbers in the contacts module 137, modify an entered phone number, dial the corresponding phone number, conduct a session, and disconnect or hang up when the session is complete. As described above, wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies.
[0153] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contacts module 137, and phone module 138, the video conferencing module 139 includes executable instructions to initiate, conduct, and terminate a video conference between the user and one or more other participants according to user instructions.
[0154] In conjunction with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the email client module 140 includes executable instructions to create, send, receive, and manage emails in response to user instructions. In conjunction with the image management module 144, the email client module 140 makes it very easy to create and send emails with static images or video images captured by the camera module 143.
[0155] In combination with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the instant messaging module 141 includes executable instructions for: inputting a character sequence corresponding to an instant message, modifying a previously input character, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing the received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0156] In combination with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, the fitness support module 142 includes executable instructions for creating a fitness (e.g., having time, distance, and / or calorie burn goals); communicating with a fitness sensor (exercise device); receiving fitness sensor data; calibrating the sensors for monitoring fitness; selecting and playing music for the fitness; and displaying, storing, and transmitting fitness data.
[0157] In combination with the touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, the camera module 143 includes executable instructions for: capturing a still image or video (including a video stream) and storing them in the memory 102, modifying the characteristics of a still image or video, or deleting a still image or video from the memory 102.
[0158] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or video images.
[0159] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134, the browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, and linking to attachments and other files of web pages.
[0160] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, the email client module 140, and the browser module 147, the calendar module 148 includes executable instructions for creating, displaying, modifying, and storing calendars and data associated with the calendars (e.g., calendar entries, to-do items, etc.) according to user instructions.
[0161] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, and the browser module 147, the widget module 149 is a mini-application (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) optionally downloaded and used by the user or a mini-application created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (HyperText Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).
[0162] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, and the browser module 147, the widget creator module 150 is optionally used by the user to create widgets (e.g., transforming a user-specified portion of a web page into a widget).
[0163] In combination with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sounds, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.
[0164] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch screen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0165] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, note module 153 includes executable instructions for creating and managing notes, to-do items, etc. in accordance with user instructions.
[0166] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and data associated with the maps (e.g., driving directions, data related to stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.
[0167] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, e-mail client module 140, and browser module 147, online video module 155 includes instructions for performing the following operations: allowing a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on the touch screen or on an external display connected via external port 124), send an e-mail with a link to a particular online video, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, the instant message module 141 is used instead of the e-mail client module 140 to send a link to a particular online video. Other descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," filed on June 20, 2007, and U.S. Patent Application No. 11 / 968,067, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," filed on December 31, 2007, the contents of both of which are hereby incorporated by reference in their entirety.
[0168] Each of the above modules and applications corresponds to a set of executable instructions for performing one or more of the above functions and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, so various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, the video player module is optionally combined with the music player module into a single module (e.g., Figure 1A the video and music player module 152 in ). In some embodiments, memory 102 optionally stores a subgroup of the above modules and data structures. In addition, memory 102 optionally stores additional modules and data structures not described above.
[0169] In some embodiments, device 100 is a device in which the operation of a predefined set of functions on the device is performed exclusively via a touchscreen and / or a touchpad. By using the touchscreen and / or the touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on device 100 is optionally reduced.
[0170] The predefined set of functions performed exclusively via the touchscreen and / or the touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by the user, navigates device 100 from any user interface displayed on device 100 to a main menu, a main desktop menu, or a root menu. In such embodiments, the touchpad is used to implement a "menu button". In some other embodiments, the menu button is a physical push button or other physical input control device rather than the touchpad.
[0171] Figure 1B is a block diagram showing exemplary components for event handling according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 ) includes an event classifier 170 (e.g., in operating system 126) and corresponding application 136-1 (e.g., any one of the foregoing applications 137 to 151, 155, 380 to 390).
[0172] Event classifier 170 receives event information and determines the application 136-1 to which the event information is to be delivered and the application view 191 of application 136-1. Event classifier 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, application 136-1 includes an application internal state 192 that indicates one or more current application views displayed on the touch-sensitive display 112 when the application is active or executing. In some embodiments, the device / global internal state 157 is used by event classifier 170 to determine which application(s) is / are currently active, and the application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0173] In some embodiments, the application internal state 192 includes additional information such as one or more of the following: recovery information to be used when application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by application 136-1, a state queue for enabling the user to return to a previous state or view of application 136-1, and a repeat / undo queue of previous actions taken by the user.
[0174] The event monitor 171 receives event information from the peripheral device interface 118. The event information includes information about sub-events (e.g., a user touch on the touch-sensitive display 112 as part of a multi-touch gesture). The peripheral device interface 118 transmits information that it receives from the I / O subsystem 106 or sensors such as the proximity sensor 166, one or more accelerometers 168, and / or the microphone 113 (via the audio circuitry 110). The information that the peripheral device interface 118 receives from the I / O subsystem 106 includes information from the touch-sensitive display 112 or a touch-sensitive surface.
[0175] In some embodiments, the event monitor 171 sends requests to the peripheral device interface 118 at predetermined intervals. In response, the peripheral device interface 118 transmits event information. In other embodiments, the peripheral device interface 118 transmits event information only when there is a significant event (e.g., a received input that is above a predetermined noise threshold and / or a received input that exceeds a predetermined duration).
[0176] In some embodiments, the event classifier 170 further includes a hit view determination module 172 and / or an active event recognizer determination module 173.
[0177] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where within one or more of the views a sub-event has occurred. A view is composed of controls and other elements that a user can see on the display.
[0178] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events that are recognized as correct inputs is optionally determined at least in part based on the hit view of the initial touch that begins the touch-based gesture.
[0179] The hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest view in the hierarchical structure that should handle the sub-event. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or a potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view generally receives all sub-events related to the same touch or input source for which it was identified as the hit view.
[0180] The active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, the active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, the active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views and, thus, determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with a particular view, higher views in the hierarchy will still remain as actively participating views.
[0181] The event dispatcher module 174 distributes event information to event recognizers (e.g., event recognizer 180). In embodiments that include the active event recognizer determination module 173, the event dispatcher module 174 delivers the event information to the event recognizer determined by the active event recognizer determination module 173. In some embodiments, the event dispatcher module 174 stores the event information in an event queue, which is retrieved by the corresponding event receiver 182.
[0182] In some embodiments, the operating system 126 includes the event classifier 170. Alternatively, the application 136-1 includes the event classifier 170. In yet another embodiment, the event classifier 170 is a stand-alone module or part of another module (such as the contact / motion module 130) stored in the memory 102.
[0183] In some embodiments, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the user interface of the application. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other embodiments, one or more of the event recognizers 180 are part of an independent module that is a higher-level object such as a user interface toolkit or from which application 136-1 inherits methods and other properties. In some embodiments, the corresponding event handlers 190 include one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. The event handlers 190 optionally utilize or invoke the data updater 176, the object updater 177, or the GUI updater 178 to update the internal state 192 of the application. Alternatively, one or more of the application views 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of the data updater 176, the object updater 177, and the GUI updater 178 are included within the corresponding application view 191.
[0184] The corresponding event recognizer 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies an event based on the event information. The event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event recognizer 180 also includes at least a subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0185] The event receiver 182 receives event information from the event classifier 170. The event information includes information about sub-events such as a touch or a touch movement. Depending on the sub-event, the event information also includes additional information such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device pose).
[0186] The event comparator 184 compares the event information with predefined event or sub - event definitions and, based on the comparison, determines an event or sub - event or determines or updates the state of an event or sub - event. In some embodiments, the event comparator 184 includes an event definition 186. The event definition 186 contains definitions of events (e.g., a predefined sequence of sub - events), such as Event 1 (187 - 1), Event 2 (187 - 2), and others. In some embodiments, the sub - events in an event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi - touch. In one example, the definition of Event 1 (187 - 1) is a double - tap on a displayed object. For example, a double - tap includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift - off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on the displayed object, and a second lift - off (touch end) of a predetermined duration. In another example, the definition of Event 2 (187 - 2) is a drag on a displayed object. For example, a drag includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch - sensitive display 112, and lift - off of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0187] In some embodiments, the event definition 187 includes a definition of an event for a corresponding user interface object. In some embodiments, the event comparator 184 performs a hit test to determine which user interface object is associated with a sub - event. For example, in an application view that displays three user interface objects on the touch - sensitive display 112, when a touch is detected on the touch - sensitive display 112, the event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub - event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, the event comparator 184 selects the event handler associated with the sub - event and the object that triggered the hit test.
[0188] In some embodiments, the definition of a corresponding event (187) also includes a delay action that delays the delivery of event information until it has been determined that the sub - event sequence does or does not correspond to the event type of an event recognizer.
[0189] When the corresponding event recognizer 180 determines that the sub - event sequence does not match any event in the event definition 186, the corresponding event recognizer 180 enters an event - impossible, event - failed, or event - ended state, after which subsequent sub - events of the touch - based gesture are ignored. In such a case, other event recognizers (if any) for which the hit view remains active continue to track and process the sub - events of the ongoing touch - based gesture.
[0190] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists that indicate how the event delivery system should perform sub - event delivery to the event recognizers actively participating. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists that indicate how event recognizers interact with each other or can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists that indicate whether sub - events are delivered to different levels in the view or the programmatic hierarchy.
[0191] In some embodiments, when one or more specific sub - events of an event are recognized, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferring sending) sub - events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a token associated with the recognized event, and the event handler 190 associated with that token retrieves the token and executes a predefined process.
[0192] In some embodiments, the event delivery instruction 188 includes a sub - event delivery instruction that delivers event information about the sub - event without activating the event handler. Instead, the sub - event delivery instruction delivers the event information to the event handler associated with the sub - event sequence or to the actively participating view. The event handler associated with the sub - event sequence or with the actively participating view receives the event information and executes a predetermined process.
[0193] In some embodiments, the data updater 176 creates and updates data used in the application 136-1. For example, the data updater 176 updates the phone numbers used in the contact module 137 or stores video files used in the video player module. In some embodiments, the object updater 177 creates and updates objects used in the application 136-1. For example, the object updater 177 creates new user interface objects or updates parts of user interface objects. The GUI updater 178 updates the GUI. For example, the GUI updater 178 prepares display information and sends the display information to the graphics module 132 for display on the touch-sensitive display.
[0194] In some embodiments, the event handler 190 includes the data updater 176, the object updater 177, and the GUI updater 178, or has access to the data updater, the object updater, and the GUI updater. In some embodiments, the data updater 176, the object updater 177, and the GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0195] It should be understood that the above discussion of event handling for user touches on the touch-sensitive display also applies to other forms of user input for operating the multifunctional device 100 using an input device, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses optionally in cooperation with single or multiple keyboard presses or holds; contact movement on a touchpad, such as tapping, dragging, scrolling, etc.; stylus input; movement of the device; verbal instructions; detected eye movement; biometric input; and / or any combination thereof are optionally used as inputs corresponding to sub-events that define the events to be discriminated.
[0196] Figure 2FIG. 0 shows a portable multifunctional device 100 having a touch screen 112 according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this and other embodiments described below, a user can select one or more of these graphics by making gestures on the graphics using, for example, one or more fingers 202 (not drawn to scale in the figures) or one or more styli 203 (not drawn to scale in the figures). In some embodiments, selection of one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more flicks (from left to right, right to left, up, and / or down), and / or rolling of a finger that has made contact with the device 100 (from right to left, left to right, up, and / or down). In some implementations or in some cases, inadvertently contacting a graphic does not select the graphic. For example, when the gesture corresponding to selection is a tap, a flick gesture that sweeps over an application icon optionally does not select the corresponding application.
[0197] Device 100 optionally further includes one or more physical buttons, such as a "home" or menu button 204. As previously described, menu button 204 is optionally used to navigate to any of a set of applications 136 optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touch screen 112.
[0198] In some embodiments, device 100 includes a touch screen 112, a menu button 204, a depressible button 206 for powering the device on / off and for locking the device, one or more volume adjustment buttons 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. Depressible button 206 is optionally used to power the device on / off by depressing the button and holding the button in the depressed state for a predefined time interval; to lock the device by depressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also accepts voice input for activating or deactivating certain functions via a microphone 113. Device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of contact on touch screen 112, and / or one or more haptic output generators 167 for generating haptic output for a user of device 100.
[0199] Figure 3FIG. is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, in accordance with some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a gaming system, or a control device (e.g., a home or industrial controller). Device 300 generally includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects system components and controls the communication between them. Device 300 includes an input / output (I / O) interface 330 having a display 340, which is typically a touch screen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350, a touchpad 355, a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to haptic output generator 167 described above with reference to Figure 1A ), sensors 359 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors (similar to contact intensity sensor 165 described above with reference to Figure 1A ). Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU 310. In some embodiments, memory 370 stores programs, modules, and data structures similar to or a subset of those stored in memory 102 of portable multifunctional device 100 ( Figure 1A ). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunctional device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while memory 102 of portable multifunctional device 100 ( Figure 1A ) optionally does not store these modules.
[0200] Figure 3Each of the above elements in [element] is optionally stored in one or more of the memory devices of the previously mentioned memory device. Each of the above modules corresponds to a set of instructions for performing the above functions. The above modules or computer programs (e.g., sets of instructions or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, the memory 370 optionally stores a subgroup of the above modules and data structures. Additionally, the memory 370 optionally stores additional modules and data structures not described above.
[0201] Attention is now turned to an embodiment of a user interface optionally implemented on, for example, the portable multifunctional device 100.
[0202] Figure 4A An exemplary user interface of an application menu on the portable multifunctional device 100 according to some embodiments is shown. A similar user interface is optionally implemented on the device 300. In some embodiments, the user interface 400 includes the following elements or a subset or superset thereof:
[0203] · A signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0204] · Time 404;
[0205] · A Bluetooth indicator 405;
[0206] · A battery status indicator 406;
[0207] · A tray 408 with icons for common applications, such icons as:
[0208] ο An icon 416 marked "Phone" for the phone module 138, which icon 416 optionally includes an indicator 414 of the number of missed calls or voicemails;
[0209] ο An icon 418 marked "Mail" for the email client module 140, which icon 418 optionally includes an indicator 410 of the number of unread emails;
[0210] ο An icon 420 marked "Browser" for the browser module 147; and
[0211] ο An icon 422 marked "iPod" for the video and music player module 152 (also referred to as the iPod (trademark of Apple Inc.) module 152); and
[0212] · Icons for other applications, such as:
[0213] The icon 424 of the IM module 141 labeled "Message";
[0214] The icon 426 of the calendar module 148 labeled "Calendar";
[0215] The icon 428 of the image management module 144 labeled "Photo";
[0216] The icon 430 of the camera module 143 labeled "Camera";
[0217] The icon 432 of the online video module 155 labeled "Online Video";
[0218] The icon 434 of the stock market widget 149-2 labeled "Stock Market";
[0219] The icon 436 of the map module 154 labeled "Map";
[0220] The icon 438 of the weather widget 149-1 labeled "Weather";
[0221] The icon 440 of the alarm clock widget 149-4 labeled "Clock";
[0222] The icon 442 of the fitness support module 142 labeled "Fitness Support";
[0223] The icon 444 of the note module 153 labeled "Note"; and
[0224] The icon 446 of the settings application or module labeled "Settings", which provides access to the settings of the device 100 and its various applications 136.
[0225] It should be noted that Figure 4A The icon labels shown in Figure 4A are merely exemplary. For example, the icon 422 of the video and music player module 152 is labeled "Music" or "Music Player". Other labels may be optionally used for the various application icons. In some embodiments, the label of the corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.
[0226] Figure 4B A device (e.g., Figure 3 is shown having a touch-sensitive surface 451 (e.g., Figure 3 a tablet or touchpad 355) separate from the display 450 (e.g., a touchscreen display 112).Exemplary user interface on device 300). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting the intensity of a contact on the touch-sensitive surface 451 and / or one or more tactile output generators 357 for generating tactile output for a user of the device 300.
[0227] Although some examples below will be given with reference to input on a touch screen display 112 (where the touch-sensitive surface and the display are combined), in some embodiments, the device detects input on a touch-sensitive surface separate from the display, as Figure 4B shown. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451 in ) has a major axis (e.g., Figure 4B 453 in ) corresponding to the major axis on the display (e.g., Figure 4B 452 in ). According to these embodiments, the device detects contact (e.g., Figure 4B 460 and 462 in ) with the touch-sensitive surface 451 at a location corresponding to a respective location on the display (e.g., in Figure 4B 460 corresponds to 468 and 462 corresponds to 470). Thus, when the touch-sensitive surface (e.g., Figure 4B 451 in ) is separate from the display of the multifunctional device (e.g., Figure 4B 450 in ), user input (e.g., contacts 460 and 462 and their movement) detected by the device on the touch-sensitive surface is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0228] Additionally, although the examples below are mainly given with reference to finger input (e.g., finger contact, single-finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click when the cursor is above the location of the tap gesture (e.g., instead of detecting a contact, followed by stopping detection of the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or a mouse and a finger contact are optionally used simultaneously.
[0229] Figure 5AAn exemplary personal electronic device 500 is shown. The device 500 includes a body 502. In some embodiments, the device 500 may include some or all of the features described with respect to devices 100 and 300 (e.g., Figures 1A to 4B ). In some embodiments, the device 500 has a touch-sensitive display screen 504 hereinafter referred to as a touch screen 504. As an alternative or addition to the touch screen 504, the device 500 has a display and a touch-sensitive surface. As in the case of devices 100 and 300, in some embodiments, the touch screen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 504 (or the touch-sensitive surface) may provide output data representative of the intensity of the touch. The user interface of the device 500 may respond to the touch based on the intensity of the touch, meaning that touches of different intensities may invoke different user interface operations on the device 500.
[0230] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application", published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships", published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.
[0231] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) can be in physical form. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) may allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watchbands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear device 500.
[0232] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include some or all of the components referred to in Figure 1A , Figure 1B and Figure 3 . Device 500 has a bus 512 that operatively couples the I / O section 514 to one or more computer processors 516 and a memory 518. The I / O section 514 may be connected to a display 504, which may have a touch-sensitive component 522 and optionally a force sensor 524 (e.g., a contact force sensor). Additionally, the I / O section 514 may be connected to a communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 is optionally a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 508 is optionally a button.
[0233] In some examples, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.
[0234] The memory 518 of personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform the techniques described below, including processes 700, 800, 1000, 1200, 1400, 1500, 1700, and 1900 ( Figures 7 to 8 , Figure 10 , Figure 12 , Figure 14 , Figure 15 , Figure 17 andFigure 19 ). Computer-readable storage media can be any medium that can tangibly contain or store computer-executable instructions for use by or in conjunction with instruction execution systems, devices, and apparatuses. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. Personal electronic device 500 is not limited to Figure 5B components and configurations, but may include other components or additional components in a variety of configurations.
[0235] As used herein, the term "affordance" refers to an indication that is optionally displayed on a device 100, 300, and / or 500 ( Figure 1A , Figure 3 and Figures 5A to 5C ) on a display screen of a computer program product. For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) optionally each constitute an affordance.
[0236] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface that a user is interacting with. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 3 Touchpad 355 or Figure 4B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 451 in the display, the particular user interface element is adjusted according to the detected input. In the case that a touch-screen display (e.g., Figure 1A A touch-sensitive display system 112 or Figure 4AIn some specific implementations of the touch screen 112), the detected contact on the touch screen serves as a "focus selector", such that when an input (e.g., a press input made by the contact) is detected at the position of a specific user interface element (e.g., a button, a window, a slider, or other user interface element) on the touch screen display, the specific user interface element is adjusted according to the detected input. In some specific implementations, the focus moves from one area of the user interface to another area of the user interface without a corresponding movement of the cursor or a movement of the contact on the touch screen display (e.g., moving the focus from one button to another button by using the tab key or arrow keys); in these specific implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is generally a user interface element (or a contact on the touch screen display) that is controlled by the user to deliver the interaction with the user interface expected by the user (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or a touch screen), the position of the focus selector (e.g., a cursor, a contact, or a selection box) above the corresponding button will indicate that the user expects to activate the corresponding button (rather than other user interface elements shown on the device display).
[0237] As used in the specification and claims, the term "feature strength" of a contact refers to a feature of the contact based on one or more intensities of the contact. In some embodiments, the feature strength is based on a plurality of intensity samples. The feature strength is optionally based on a predefined number of intensity samples or a set of intensity samples collected during a predefined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after detecting the contact, before detecting the contact lift-off, before or after detecting the contact starts to move, before detecting the contact ends, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The feature strength of the contact is optionally based on one or more of the following: the maximum value of the intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the value at the top 10% of the intensity of the contact, the half-maximum value of the intensity of the contact, the 90% maximum value of the intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the feature strength (e.g., when the feature strength is the average value of the intensity of the contact over time). In some embodiments, the feature strength is compared with a set of one or more intensity thresholds to determine whether the user has performed an operation. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact whose feature strength does not exceed the first threshold results in a first operation, a contact whose feature strength exceeds the first intensity threshold but does not exceed the second intensity threshold results in a second operation, and a contact whose feature strength exceeds the second threshold results in a third operation. In some embodiments, the comparison between the feature strength and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to forgo performing the corresponding operation) rather than for determining whether to perform a first operation or a second operation.
[0238] Figure 5C FIG. depicts an exemplary diagram of a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500, and each device shares one or more data connections 510 (such as an Internet connection, a Wi-Fi connection, a cellular connection, a short-range communication connection, and / or any other such data connection or network) with each other to facilitate real-time communication of audio data and / or video data between the corresponding devices for a period of time. In some embodiments, the exemplary communication session may include a shared data session, whereby data is transferred from one or more of the electronic devices to other electronic devices to enable simultaneous output of the corresponding content at the electronic devices. In some embodiments, the exemplary communication session may include a video conferencing session, whereby audio data and / or video data is transferred between devices 500A, 500B, and 500C such that users of the corresponding devices can use the electronic devices for real-time communication.
[0239] In Figure 5C this example, device 500A represents an electronic device associated with user A. Device 500A communicates with devices 500B and 500C (via data connection 510), and devices 500B and 500C are associated with users B and C, respectively. Device 500A includes a camera 501A for capturing video data of a communication session, and a display 504A (e.g., a touch screen) for displaying content associated with the communication session. Device 500A also includes other components, such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0240] Device 500A displays a communication UI 520A via display 504A, which is a user interface for facilitating a communication session (e.g., a video conference session) between device 500B and device 500C. Communication UI 520A includes video feeds 525-1A and 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during the communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during the communication session.
[0241] Communication UI 520A includes a camera preview 550A, which is a representation of video data captured at device 500A via camera 501A. Camera preview 550A shows user A the expected video feeds of user A as displayed at the respective devices 500B and 500C.
[0242] Communication UI 520A includes one or more controls 555A for controlling one or more aspects of the communication session. For example, controls 555A may include controls for muting the audio of the communication session, changing the camera view of the communication session (e.g., changing the camera used to capture the communication session video, adjusting the zoom value), terminating the communication session, applying visual effects to the camera view of the communication session, activating one or more modes associated with the communication session. In some embodiments, one or more controls 555A are optionally displayed in communication UI 520A. In some embodiments, one or more controls 555A are displayed separately from camera preview 550A. In some embodiments, one or more controls 555A are displayed as covering at least a portion of camera preview 550A.
[0243] In Figure 5CIn this case, device 500B represents an electronic device associated with user B, and user B communicates with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B for capturing video data of a communication session, and a display 504B (e.g., a touch screen) for displaying content associated with the communication session. Device 500B also includes other components, such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0244] Device 500B displays a communication UI 520B similar to communication UI 520A of device 500A via touch screen 504B. Communication UI 520B includes video feeds 525-1B and 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during the communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during the communication session. Communication UI 520B also includes: a camera preview 550B, which is a representation of video data captured at device 500B via camera 501B; and one or more controls 555B similar to controls 555A, which are used to control one or more aspects of the communication session. Camera preview 550B represents to user B the expected video feeds of user B displayed at corresponding devices 500A and 500C.
[0245] In Figure 5C this case, device 500C represents an electronic device associated with user C, and user C communicates with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C for capturing video data of a communication session, and a display 504C (e.g., a touch screen) for displaying content associated with the communication session. Device 500C also includes other components, such as a microphone (e.g., 113) for recording the audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0246] Device 500C displays a communication UI 520C similar to communication UI 520A of device 500A and communication UI 520B of device 500B via a touch screen 504C. The communication UI 520C includes a video feed 525-1C and a video feed 525-2C. The video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during a communication session. The video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during a communication session. The communication UI 520C further includes: a camera preview 550C, which is a representation of video data captured at device 500C via camera 501C; and one or more controls 555C similar to controls 555A and 555B, which are used to control one or more aspects of the communication session. The camera preview 550C represents to user C the expected video feed of user C displayed at the respective devices 500A and 500B.
[0247] Although Figure 5C the diagrams depicted in show a communication session between three electronic devices, the communication session can be established between two or more electronic devices, and the number of devices participating in the communication session can change as electronic devices join or leave the communication session. For example, if one of the electronic devices leaves the communication session, the audio data and video data from the device that stops participating in the communication session are no longer represented on the participating devices. For example, if device 500B stops participating in the communication session, there is no data connection 510 between devices 500A and 500C, and there is no data connection 510 between devices 500C and 500B. Additionally, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and video data and audio data are shared among all devices such that each device can output the data transmitted from the other devices.
[0248] Figure 5C The diagrams depicted in show a communication session between multiple electronic devices, including Figures 6A to 6AY 、 Figures 9A to 9T 、 Figures 11A to 11P 、 Figures 13A to 13K and Figures 16A to 16Q the exemplary communication sessions depicted in. In some embodiments, Figures 6A to 6AY 、 Figures 9A to 9T 、 Figures 13A to 13K and Figures 16A to 16QThe communication session depicted includes two or more electronic devices, even if other electronic devices participating in the communication session are not depicted in the figures.
[0249] Attention is now turned to an implementation of a user interface (“UI”) and associated processes implemented on an electronic device such as portable multifunctional device 100, device 300, or device 500.
[0250] Figures 6A to 6AY Exemplary user interfaces for managing a real-time video communication session are shown in accordance with some embodiments. The user interfaces in these figures are used to illustrate processes described below that include Figures 7 to 8 and Figure 15 the processes in.
[0251] Figures 6A to 6AY Exemplary user interfaces for managing a real-time video communication session are shown from the perspective of different users (e.g., users participating in a real-time video communication session from different devices, different types of devices, devices with different applications installed, and / or devices with different operating system software).
[0252] Referring Figure 6A , device 600-1 corresponds to user 622 (e.g., “John”), who, in some embodiments, is a participant in a real-time video communication session. Device 600-1 includes a display (e.g., a touch-sensitive display) 601 and a camera 602 (e.g., a front camera) having a field of view 620. In some embodiments, camera 602 is configured to capture image data and / or depth data of the physical environment within field of view 620. Field of view 620 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view. In some embodiments, camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens having a relatively short focal length and a wide field of view). In some embodiments, device 600-1 may include multiple cameras. Thus, while device 600-1 is described herein as using camera 602 to capture image data during a real-time video communication session, it should be understood that device 600-1 may use multiple cameras to capture image data.
[0253] Referring Figure 6A, Device 600-2 corresponds to user 623 (e.g., "Jane"), who is, in some embodiments, a participant in a real-time video communication session. Device 600-2 includes a display (e.g., a touch-sensitive display) 683 and a camera 682 (e.g., a front camera) with a field of view 688. In some embodiments, camera 682 is configured to capture image data and / or depth data of the physical environment within field of view 688. Field of view 688 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view. In some embodiments, camera 682 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600-2 may include multiple cameras. Thus, while device 600-2 is described herein as using camera 682 to capture image data during a real-time video communication session, it should be understood that device 600-2 may use multiple cameras to capture image data.
[0254] As shown, user 622 ("John") is positioned (e.g., sitting) in front of desk 621 (and device 600-1) in environment 615. In some examples, user 622 is positioned in front of desk 621 such that user 622 is captured within the field of view 620 of camera 602. In some embodiments, one or more objects proximate to user 622 are positioned such that these objects are captured within the field of view 620 of camera 602. In some embodiments, both user 622 and the objects proximate to user 622 are captured within the field of view 620 simultaneously. For example, as shown, drawing 618 is positioned on surface 619 in front of user 622 (relative to camera 602) such that both user 622 and drawing 618 are captured in the field of view 620 of camera 602 and are displayed in representation 622-1 (displayed by device 600-1) and representation 622-2 (displayed by device 600-2).
[0255] Similarly, user 623 ("Jane") is positioned (e.g., sitting) in front of desk 686 (and device 600-2) in environment 685. In some examples, user 623 is positioned in front of desk 686 such that user 623 is captured within the field of view 688 of camera 682. As shown, user 623 is displayed in representation 623-1 (displayed by device 600-1) and representation 623-2 (displayed by device 600-2).
[0256] Generally, during operation, devices 600-1, 600-2 capture image data, which is then exchanged between devices 600-1, 600-2 and used by devices 600-1, 600-2 to display various representations of content during a real-time video communication session. Although each of devices 600-1, 600-2 is shown, the examples described are primarily directed to the user interface displayed on device 600-1 and / or user input detected by that device. It should be understood that in some examples, electronic device 600-2 operates in a manner similar to electronic device 600-1 during a real-time video communication session. In some examples, devices 600-1, 600-2 display similar user interfaces and / or cause operations similar to those described below to be performed.
[0257] As will be described in further detail below, in some examples, such representations include images that have been modified during a real-time video communication session to provide an improved view of surfaces and / or objects within the field of view of the cameras of devices 600-1, 600-2 (also referred to herein as the "field of view"). Any known image processing techniques can be used to modify the image, including but not limited to image rotation and / or distortion correction (e.g., image skew). Thus, although image data may be captured from a camera having a particular position relative to the user, the representation can provide a view of the user (and / or surfaces or objects in the user's environment) from a perspective different from the perspective from which the image data was captured by the camera. Figures 6A to 6AY Embodiments disclose displaying elements and detecting inputs (including hand gestures) at device 600-1 to control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2). In some embodiments, device 600-2 displays similar elements and detects similar inputs (including hand gestures) at device 600-2 to control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2).
[0258] Referring Figure 6A , device 600-1 displays a video conferencing interface 604-1 on display 601. Video conferencing interface 604-1 includes a representation 622-1, which in turn includes an image (e.g., a frame of a video stream) of the physical environment (e.g., a scene) within the field of view 620 of camera 602. In some examples, the image of representation 622-1 includes the entire field of view 620. In other examples, the image of representation 622-1 includes a portion (e.g., a cropped portion or subset) of the entire field of view 620. As shown, in some examples, the image of representation 622-1 includes user 622 and / or surface 619 proximate user 622 on which drawing 618 is located.
[0259] The video conferencing interface 604-1 also includes a representation 623-1, which in turn includes an image of the physical environment within the field of view 688 of the camera 682. In some examples, the image of the representation 623-1 includes the entire field of view 688. In other examples, the image of the representation 623-1 includes a portion (e.g., a cropped portion or subset) of the entire field of view 688. As shown, in some examples, the image of the representation 623-1 includes the user 623. As shown, the representation 623-1 is displayed with a greater magnitude than the representation 622-1. In this way, the user 622 can better observe and / or interact with the user 623 during a real-time communication session.
[0260] The device 600-2 displays the video conferencing interface 604-2 on the display 683. The video conferencing interface 604-2 includes a representation 622-2, which in turn includes an image of the physical environment within the field of view 620 of the camera 602. The video conferencing interface 604-2 also includes a representation 623-2, which in turn includes an image of the physical environment within the field of view 688 of the camera 682. As shown, the representation 622-2 is displayed with a greater magnitude than the representation 623-2. In this way, the user 623 can better observe and / or interact with the user 622 during a real-time communication session.
[0261] At Figure 6A the device 600-1 displays the interface 604-1. When displaying the interface 604-1, the device 600-1 detects an input 612a (e.g., a swipe input) corresponding to a request for a display settings interface. In response to detecting the input 612a, the device 600-1 displays the settings interface 606, as Figure 6B depicted. As shown, in some embodiments, the settings interface 606 is overlaid on the interface 604-1.
[0262] In some embodiments, the settings interface 606 includes one or more enabling representations for controlling the settings of the device 600-1 (e.g., volume, brightness of the display, and / or Wi-Fi settings). For example, the settings interface 606 includes a view enabling representation 607-1, which when selected causes the device 600-1 to display a view menu, as Figure 6B shown.
[0263] As Figure 6B shown, when displaying the settings interface 606, the device 600-1 detects the input 612b. In some embodiments, the input 612b is a tap gesture on the view enabling representation 607-1. In response to detecting the input 612b, the device 600-1 displays the view menu 616-1, as Figure 6C shown.
[0264] Generally, the view menu 616-1 includes one or more affordance representations that can be used to manage (e.g., control) the way in which representations are displayed during a real-time video communication session. For example, selection of a particular affordance representation may cause the device 600-1 to display or stop displaying a representation in an interface (e.g., interface 604-1 or interface 604-2).
[0265] The view menu 616-1 includes, for example, a surface view affordance representation 610 that, when selected, causes the device 600-1 to display a representation including a modified image of the surface. In some embodiments, when the surface view affordance representation 610 is selected, the user interface directly transitions to Figure 6M the user interface. Additionally or alternatively, Figures 6D to 6L (described below) shows other user interfaces that may be displayed prior to the user interface in Figure 6M and other inputs for initiating the process of displaying the user interface as shown in Figure 6M . For example, when the view menu 616-1 is displayed, the device 600-1 detects an input 612c corresponding to the selection of the surface view affordance representation 610. In some examples, the input 612c is a touch input. In response to detecting the input 612c, the device 600-1 displays a representation 624-1, as shown in Figure 6M . Further in response to detecting the input 612c, the device 600-2 displays a representation 624-2. As described, in some embodiments, the image is modified during the real-time video communication session to provide an image with a particular perspective. Thus, in some examples, the representation 624-1 is provided by generating an image from the image data captured by the camera 602, modifying the image (or a portion of the image), and displaying the representation 624-1 with the modified image. In some embodiments, any known image processing techniques are used to modify the image, including but not limited to image rotation and / or distortion correction (e.g., image skew). In some embodiments, the image of the representation 624-2 is also provided in this manner.
[0266] In some embodiments, the image of the representation 624-1 is modified to provide a desired perspective (e.g., surface view). In some embodiments, the image of the representation 624-1 is modified based on the position of the surface 619 relative to the camera 602. For example, the device 600-1 may rotate the image of the representation 624-1 by a predetermined amount (e.g., 45 degrees, 90 degrees, or 180 degrees) such that the surface 619 can be more intuitively viewed in the representation 624-1. As shown in Figure 6MAs shown, for example, where camera 602 captures surface 619 from the perspective facing user 622, the image representing 624-1 is rotated 180 degrees to provide a perspective of the image from the perspective of user 622. Thus, during a real-time video communication session, devices 600-1, 600-2 display surface 619 (and by extension drawing 618) from the perspective of user 622 during the real-time communication session. In some examples, the image representing 624-2 is also provided in this manner.
[0267] In some embodiments, to ensure that user 623 maintains the view of user 622 when the representation 624-2 includes a modified image of surface 619, device 600-2 maintains the display of representation 622-2. As Figure 6M shown, maintaining the display of representation 622-2 in this manner may include adjusting the size and / or position of representation 622-2 in interface 604-2. Optionally, in some embodiments, device 600-2 stops displaying representation 622-2 to provide a representation 624-2 of a larger size. Optionally, in some embodiments, device 600-1 stops displaying representation 622-1 to provide a representation 624-1 of a larger size.
[0268] Representations 624-1, 624-2 include images of drawing 618 that are modified relative to the position (e.g., location and / or orientation) of drawing 618 relative to camera 602. For example, as Figure 6A depicted, prior to modification, the image is shown as having a particular orientation (e.g., inverted) in representations 622-1, 622-1. As a result of modifying the image, the image of drawing 618 is rotated and / or skewed such that the perspective of representations 624-1, 624-2 appears to be from the perspective of user 622. In this manner, the modified image of drawing 618 provides a perspective different from the perspective of representations 624-1, 624-2 in order to give user 623 (and / or user 622) a more natural and direct view of drawing 618. Thus, during a real-time video communication session, user 623 can more easily and intuitively view drawing 618.
[0269] As described, a representation including a modified image of a surface is provided in response to a selection of a surface image enabling representation (e.g., surface view enabling representation 610). In some examples, a representation including a modified view of a surface is provided in response to detecting other types of input.
[0270] Referring to Figure 6D, in some examples, a representation of a modified image including a surface is provided in response to one or more gestures. For example, device 600-1 may use camera 602 to detect a gesture, and in response to detecting the gesture, determine whether the gesture meets a set of criteria (e.g., a set of gesture criteria). In some embodiments, the criteria includes the requirement that the gesture is a pointing gesture, and optionally the requirement that the pointing gesture has a specific orientation and / or points to a surface and / or object. For example, referring to Figure 6D , device 600-1 detects gesture 612d and determines that gesture 612d is a pointing gesture pointing to drawing 618. In response, device 600-1 displays a representation of a modified image including surface 619, as shown in reference Figure 6M described.
[0271] In some embodiments, the set of criteria includes the requirement that the gesture be performed for at least a threshold amount of time. For example, referring to Figure 6E , in response to detecting the gesture, device 600-1 overlays graphical object 626 on representation 622-1, indicating that device 600-1 has detected that the user is currently performing a gesture, such as 612d. As shown, in some embodiments, device 600-1 magnifies representation 622-1 to help user 622 better view the detected gesture and / or graphical object 626.
[0272] In some embodiments, graphical object 626 includes a timer 628 (e.g., a digital timer, a ring that fills over time, and / or a bar that fills over time) indicating the amount of time that gesture 612d has been detected. In some embodiments, timer 628 also (or alternatively) indicates the threshold amount of time that gesture 612d will continue to be provided to meet the set of criteria. In response to gesture 612 meeting the threshold amount of time (e.g., 0.5 seconds, 2 seconds, and / or 5 seconds), device 600-1 displays a representation 624-1 of a modified image including a surface ( Figure 6M ), as described.
[0273] In some examples, graphical object 626 indicates the type of gesture currently detected by device 600-1. In some examples, graphical object 626 is a silhouette of a hand performing the detected gesture type and / or an image of the detected gesture type. Graphical object 626 may, for example, include a hand performing a pointing gesture in response to device 600-1 detecting that user 622 is performing a pointing gesture. Additionally or alternatively, graphical object 626 may optionally indicate a zoom level (e.g., the zoom level at which a representation of a second portion of the scene is being displayed or will be displayed).
[0274] In some examples, a representation with a modified image is provided in response to one or more voice inputs. For example, during a real-time communication session, device 600-1 receives a voice input, such as Figure 6DThe voice input 614 in (“View my drawing”). In response, device 600-1 displays a representation 624-1 of a modified image that includes the surface, as described. Figure 6M )
[0275] In some examples, the voice input received by device 600-1 can include a reference to any surface and / or object that can be recognized by device 600-1, and in response, device 600-1 provides a representation of a modified image that includes the referenced object or surface. For example, device 600-1 can receive a voice input that references a wall (e.g., the wall behind user 622). In response, device 600-1 provides a representation of a modified image that includes the wall.
[0276] In some embodiments, voice input can be used in combination with other types of input, such as gestures (e.g., gesture 612d). Thus, in some embodiments, device 600-1 displays a modified image of a surface (or object) in response to detecting both a gesture and a voice input that correspond to a request to provide a modified image of the surface.
[0277] In some embodiments, a surface view affordance representation is provided in other ways. Refer to Figure 6F , for example, video conferencing interface 604-1 includes an options menu 608. Options menu 608 includes a set of affordance representations that can be used to control device 600-1 during a real-time video communication session, including view affordance representation 607-2.
[0278] When displaying options menu 608, device 600-1 detects an input 612f that corresponds to a selection of view affordance representation 607-2. In response to detecting input 612f, device 600-1 displays view menu 616-2, as shown in Figure 6G . View menu 616-2 can be used to control the way in which a representation is displayed during a real-time video communication session, as described with respect to Figure 6C above.
[0279] Although options menu 608 is shown in all of the figures as being persistently displayed in video conferencing interface 604-1, options menu 608 can be hidden and / or redisplayed by device 600-1 at any point during a real-time video communication session. For example, in response to detecting one or more inputs from the user and / or an inactivity period, options menu 608 can be displayed on and / or removed from the display.
[0280] Although detecting an input that points to a surface has been described as causing device 600-1 to display a representation of a modified image that includes the surface (e.g., in response to detecting input 612c, device 600-1 displays representation 624-1, as shown in Figure 6C ),Figure 6M as shown), but in some embodiments, detecting an input directed to the surface may cause device 600-1 to enter a preview mode (e.g., Figures 6H to 6J ) before displaying, for example, representation 624-1.
[0281] Figure 6H An example of device 600-1 operating in preview mode is shown. Generally, preview mode can be used to selectively provide a portion or region of an image of a representation to one or more other users during a real-time video communication session.
[0282] In some embodiments, before operating in preview mode, device 600-1 detects an input (e.g., input 612c) that points to the surface view enabling representation 610. In response, device 600-1 initiates preview mode. When operating in preview mode, device 600-1 displays preview interface 674-1. Preview interface 647-1 includes left scroll enabling representation 634-2, right scroll enabling representation 634-1, and preview 636.
[0283] In some embodiments, selection of the left scroll enabling representation causes device 600-1 to change (e.g., replace) preview 636. For example, selection of left scroll enabling representation 634-2 or right scroll enabling representation 634-1 causes device 600-1 to cycle through various images (the user's image, an unmodified image of the surface, and / or a modified image of surface 619) such that the user can select a particular perspective to share when exiting preview mode, e.g., in response to detecting an input directed to preview 636. Additionally or alternatively, these techniques can be used to cycle through and / or select a particular surface (e.g., a vertical and / or horizontal surface) and / or a particular portion (e.g., a cropped portion or subset) in the field of view.
[0284] As shown, in some embodiments, preview user interface 674-1 is displayed at device 600-1 and not at device 600-2. For example, device 600-2 displays video conferencing interface 604-2 (including representation 622-2), while device 600-1 displays preview interface 674-1. Thus, preview user interface 674-1 allows user 622 to select a view before sharing it with user 623.
[0285] Figure 6IShows an example of device 600-1 operating in preview mode. As depicted, when device 600-1 is operating in preview mode, device 600-1 displays preview interface 674-2. In some embodiments, preview interface 674-2 includes representation 676 having regions 636-1, 636-2. In some embodiments, representation 676 includes an image that is the same as or substantially similar to the image included in representation 622-1. Optionally, as shown, the size of representation 676 is greater than Figure 6A that of representation 622-1. The position of representation 676 is different from the position of representation 622-1. Adjusting the size and / or position of the representation in preview interface 674-2 allows user 622 to better view the image before sharing it with user 623 as compared to the size and / or position of a representation including a similar or identical image in video conferencing interface 604-1.
[0286] In some embodiments, regions 636-1 and 636-2 correspond to respective portions of representation 676. For example, as shown, region 636-1 corresponds to the upper portion of representation 676 (e.g., the portion including the upper body of user 622), and region 636-2 corresponds to the lower portion of representation 676 (e.g., the portion including the lower body of user 622 and / or drawing 618).
[0287] In some embodiments, regions 636-1 and 636-2 are shown as different regions (e.g., non-overlapping regions). In some embodiments, regions 636-1 and 636-2 overlap. Additionally or alternatively, one or more graphical objects 638-1 (e.g., lines, boxes, and / or dashed lines) may distinguish (e.g., visually distinguish) region 636-1 from region 636-2.
[0288] In some embodiments, preview interface 674-2 includes one or more graphical objects to indicate whether a region is active or inactive. In Figure 6I the example, preview interface 674-2 includes graphical objects 641a, 641b. In some embodiments, the appearance (e.g., shape, size, and / or color) of graphical objects 641a, 641b indicates whether the corresponding region is active or inactive.
[0289] When active, the region is shared with one or more other users of the real-time video communication session. For example, referring to Figure 6I, the graphical user interface object 641 indicates that the area 636-1 is active. Accordingly, the image data corresponding to the area 636-1 is displayed by the device 600-2 in the representation 622-2. In some examples, the device 600-1 shares only the image data of the active area. In some embodiments, the device 600-1 shares all the image data and instructs the device 600-2 to display the image only based on a portion of the image data corresponding to the valid area 636-1.
[0290] When the display interface 674-2 is presented, the device 600-1 detects the input 612i at the location corresponding to the area 636-2. In some embodiments, the input 612i is a touch input. In response to detecting the input 612i, the device 600-1 activates the area 636-2. Accordingly, the device 600-2 displays a representation of the modified image including the surface 619, such as the representation 624-2. In some embodiments, the area 636-1 remains active in response to the input 612i (e.g., the user 623 can see the user 622 in the representation 622-2, for example). Optionally, in some embodiments, the device 600-1 deactivates the area 636-1 in response to the input 612i (e.g., the user 623 can no longer see the user 622 in the representation 622-2, for example).
[0291] Although examples have been described with respect to a preview mode having a representation including two areas 636-1, 636-2 Figure 6I , in some embodiments, other numbers of areas may be used. For example, referring to Figure 6J , the device 600-1 is operating in a preview mode where the preview interface 674-3 includes a representation 676 that includes areas 636a-636i.
[0292] In some embodiments, multiple areas are active (and / or can be activated). For example, as shown, the device 600-1 displays areas 636a-636i, where areas 636a-f are active. Accordingly, the device 600-2 displays the representation 622-2.
[0293] In some embodiments, the device 600-1 modifies the image of a surface having any type of orientation (including any angle relative to gravity (e.g., between zero and ninety degrees)). For example, when the surface is a horizontal surface (e.g., a surface in a plane within the range of 70 degrees to 110 degrees from the direction of gravity), the device 600-1 can modify the image of the surface. As another example, when the surface is a vertical surface (e.g., a surface in a plane up to 30 degrees from the direction of gravity), the device 600-1 can modify the image of the surface.
[0294] When displaying interface 674-3, device 600-1 detects input 612j at a location corresponding to region 636h. In response to detecting input 612j, device 600-1 activates region 636-2. Accordingly, device 600-2 displays a representation of a modified image that includes surface 619, such as representation 624-2. In some embodiments, regions 636a-f remain active in response to input 612j (e.g., user 623 can see user 622 in representation 622-2, for example). Optionally, in some embodiments, device 600-1 deactivates regions 636a-f in response to input 612j (e.g., user 623 can no longer see user 622 in representation 622-2, for example).
[0295] Figures 6K to 6L An exemplary animation that can be displayed by device 600-1 and / or device 600-2 is shown. As Figures 6A to 6I discussed, device 600-1 can display a representation that includes a modified image. In some embodiments, device 600-1 and / or device 600-2 display an animation that transitions between views and / or shows modifications to an image over time. The animation can include, for example, translating, rotating, and / or otherwise modifying the image to provide a modified image. Additionally or alternatively, the animation occurs in response to detecting an input directed at the surface (e.g., selection of a surface view enabling representation 610, a gesture, and / or a voice input).
[0296] Figure 6K An exemplary animation is shown in which device 600-2 translates and rotates the image of representation 642a. During the animation, the image of representation 642a is translated downward to view surface 619 from a more "top-down" perspective. The animation also includes rotating the image of representation 642a such that surface 619 is viewed from the perspective of user 622. Although four frames of the animation are shown, the animation can include any number of frames. Optionally, in some embodiments, device 600-1 translates and rotates the image of a representation (e.g., representation 622-1).
[0297] Figure 6L An example of device 600-2 magnifying and rotating the image of representation 642a is shown. During the animation, the image of representation 642a is magnified until a desired zoom level is obtained. The animation also includes rotating the image of representation 642a until the image orientation of drawing 618 is in the perspective of user 622, as described. Although four frames of the animation are shown, the animation can include any number of frames. Optionally, in some embodiments, device 600-1 magnifies and rotates the image of a representation (e.g., representation 622-1).
[0298] Figures 6N to 6R An example of a modified image that further modifies the surface during a real-time communication session is shown.
[0299] Figure 6N illustrates an example of a real-time communication session in which a user provides various inputs. For example, when displaying interface 678, device 600-1 detects input 677 corresponding to the rotation of device 600-1. As Figure 6O depicted, in response to detecting input 677, device 600-1 modifies interface 678 to compensate for the rotation (e.g., of camera 602). As Figure 6O shown, device 600-1 arranges representations 623-1 and 624-1 of interface 678 in a vertical configuration. Additionally, representation 624-1 is rotated according to the rotation of device 600-1 such that the viewing angle of representation 624-1 remains in the same orientation relative to user 622. Additionally, the viewing angle of representation 624-2 remains in the same orientation relative to user 623.
[0300] Further reference Figure 6N , in some examples, device 600-1 displays control affordance representations 648-1, 648-2 to modify the image of representation 624-1. Control affordance representations 648-1, 648-2 may be displayed in response to one or more inputs corresponding to, for example, the selection of an affordance representation of option menu 608 (e.g., Figure 6B ).
[0301] As shown, in some embodiments, device 600-1 displays a representation 624-1 that includes a modified image of a surface. The rotation affordance representation 648-1, when selected, causes device 600-1 to rotate the image of representation 624-1. For example, when displaying interface 678, device 600-1 detects input 650a corresponding to the selection of rotation affordance representation 648-1. In response to input 650a, device 600-1 modifies the orientation of the image of representation 624-1 from a first orientation ( Figure 6N shown) to a second orientation ( Figure 6O shown). In some embodiments, the image of representation 624-1 rotates a pre-determined amount (e.g., 90 degrees).
[0302] The zoom affordance representation 648-2, when selected, modifies the zoom level of the image of representation 624-1. For example, as Figure 6N depicted, the image of representation 624-1 is displayed at a first zoom level (e.g., “1X”). When the zoom affordance representation 648-2 is displayed, device 600-1 detects input 650b corresponding to the selection of zoom affordance representation 648-2. In response to input 650b, device 600-1 modifies the zoom level of the image of representation 624-1 from a first zoom level (e.g., “1X”) to a second zoom level (e.g., “2X”), as Figure 6Q shown.
[0303] Additionally or alternatively, in some embodiments, the video conferencing interface 604-1 includes an option to display a magnified view of at least a portion of the image representing 624-1, as Figure 6R shown. For example, when displaying the representation 624-1, the device 600-1 may detect an input 654 (e.g., a gesture toward a surface and / or object) corresponding to a request for a magnified view of a portion of the image representing 624-1. In response to detecting the input 654, the device 600-1 displays the magnified portion 652-1 at a zoom level greater than that of the second portion 652-2 of the representation 624-1. In some embodiments, the portion of the image representing 624-1 to be magnified is determined based on the location of the input 654. In some embodiments, in response to detecting the input 650c( Figure 6R and Figure 6Q ), the device 600-1 stops displaying the control enabling representations 648-1, 648-2.
[0304] Figures 6S to 6AC An example of a device modifying an image of a representation in response to user input is shown. As described in more detail below, the device 600-1 may modify the image of a representation (e.g., the representation 622-1) in the video conferencing interface 604-1 in response to non-touch user input, including gestures and / or audio input, thereby improving the way in which a user interacts with the device to manage and / or modify the representation during a real-time video communication session.
[0305] Figures 6S to 6T An example of a device obscuring at least a portion of an image of a representation in response to a gesture is shown. As Figure 6S shown, the device 600-1 detects a gesture 656a corresponding to a request to modify at least a portion of the image of the representation 622-1. In some examples, the gesture 656a is a gesture by the user 622 in an upward direction near the user 622's mouth (e.g., a "shh" gesture). As Figure 6T shown, in response, the device 600-1 replaces the representation 622-1 with a representation 622-1' including a modified image that includes a modified portion 658-1 (e.g., the background of the physical environment of the user 622). In some examples, modifying the portion 658-1 in this way includes blurring, graying, or otherwise obscuring the portion 658-1. In some examples, the device 600-1 does not modify the portion 658-2 in response to the gesture 656a.
[0306] Figures 6U to 6V An example of a device magnifying a portion of an image of a representation in response to detecting a gesture is shown. As Figure 6UAs shown, in some embodiments, device 600-1 detects a pointing gesture 656b corresponding to a request for at least a portion of magnified representation 622-1. As shown, pointing gesture 656b points to object 660.
[0307] As Figure 6V depicted, in response to pointing gesture 656b, device 600-1 replaces representation 622-1 with a representation 622-1' that includes a modified image by magnifying a portion of the image of representation 622-1 that includes object 660. In some embodiments, the magnification is based on the position of object 660 (e.g., relative to camera 602) and / or the size of object 660.
[0308] Figures 6W to 6X An example of a device magnifying a portion of a view of a representation in response to detecting a gesture is shown. As Figure 6W shown, in some embodiments, device 600-1 detects a framing gesture 656c corresponding to a request for at least a portion of magnified representation 622-1. As shown, since framing gesture 656c at least partially frames, encloses, and / or outlines object 660, framing gesture 656c points to object 660.
[0309] As Figure 6X depicted, in response to framing gesture 656c, device 600-1 modifies the image of representation 622-1 by magnifying a portion of the image of representation 622-1 that includes object 660. In some embodiments, the magnification is based on the position of object 660 (e.g., relative to camera 602) and / or the size of object 660. Additionally or alternatively, after magnifying a portion of the image of representation 622-1, device 600-1 may track the movement of framing gesture 656c. In response, device 600-1 may pan to a different portion of the image.
[0310] Figures 6Y to 6Z An example of a device panning an image of a representation in response to detecting a gesture is shown. As Figure 6Y shown, device 600-1 detects a pointing gesture 656d corresponding to a request to pan the view of the image of representation 622-1 in a particular direction (e.g., horizontally). As shown, pointing gesture 656d points to the left of user 622.
[0311] As Figure 6Z shown, in response to pointing gesture 656d, device 600-1 replaces representation 622-1 with a representation 622-1' that includes a modified image based on panning the image of representation 622-1 in the direction of pointing gesture 656d (e.g., to the left of user 622).
[0312] Although in some embodiments, asFigure 6Z As shown, due to the panning operation, a portion of the user 622 (e.g., the right shoulder of the user 622) may be excluded from the image representing 622-1'. However, in some embodiments, the device 600-1 may adjust the zoom level of the image representing 622-1' when panning to ensure that the user 622 remains entirely within the image.
[0313] Figures 6AA to 6AB An example of a device modifying the zoom level of a representation in response to detecting a pinch and / or spread gesture is shown. As Figure 6AA shown, in some embodiments, the device 600-1 detects a spread gesture 656e, where the user 622 increases the distance between the thumb and index finger of the user 622's right hand.
[0314] As Figure 6AB depicted, in response to the spread gesture 656e, the device 600-1 replaces the representation 622-1 with 622-1' by magnifying a portion of the image representing 622-1. In some embodiments, the magnification is based on the position of the spread gesture 656e (e.g., relative to the camera 602) and / or the magnitude of the spread gesture 656e. In some embodiments, a portion of the image is magnified according to a pre-determined zoom level.
[0315] Referring Figure 6AA , in some embodiments, in response to detecting the spread gesture 656e, the device 600-1 displays a zoom indicator 662 indicating the zoom level of the image representing 622-1'. Once the user 622 has completed the spread gesture 656e and the device 600-1 has magnified the portion representing 622-1', the device 600-1 updates the display of the zoom indicator 662 to indicate the current zoom level of the image representing 622-1'. In some embodiments, the zoom indicator 662 is dynamically updated as the user 622 performs the gesture 656e.
[0316] Although described herein with respect to increasing the zoom level of an image in response to the spread gesture 656e, in some examples, the zoom level of the image is decreased in response to a gesture (e.g., other types of gestures such as a pinch gesture).
[0317] Figure 6ACShows various gestures that can be used to modify a represented image. In some embodiments, for example, a user can use gestures to indicate a zoom level. For example, gesture 664 can be used to indicate that the zoom level of the represented image should be "1X", and in response to detecting gesture 664, device 600-1 can modify the represented image to have a "1X" zoom level. Similarly, gesture 666 can be used to indicate that the zoom level of the represented image should be "2X", and in response to detecting gesture 666, device 600-1 can modify the represented image to have a "2X" zoom level. Although two zoom levels (e.g., "1X" and "2X" zoom levels) are described for Figure 6AC device 600-1, in some embodiments, device 600-1 can use the same gesture or different gestures to modify the represented image to other zoom levels (e.g., 0.5X, 3X, 5X, or 10X). In some embodiments, device 600-1 can modify the represented image to three or more different zoom levels. In some embodiments, the zoom levels are discrete or continuous.
[0318] As another example, a gesture in which user 622 curls their fingers can be used to adjust the zoom level. For example, gesture 668 (e.g., a gesture in which the fingers of the user's hand curl in the direction 668b away from the camera (e.g., when the back of hand 668a is oriented towards the camera)) can be used to indicate that the zoom level of the image should be increased (e.g., zoomed in). Gesture 670 (e.g., a gesture in which the fingers of the user's hand curl in the direction 670b towards the camera (e.g., when the palm of hand 668a is oriented towards the camera)) can be used to indicate that the zoom level of the image should be decreased (e.g., zoomed out).
[0319] Figures 6AD to 6AE Shows an example of a user using two devices to participate in a real-time video communication session.
[0320] For example, as Figure 6AD shown, user 623 is using additional device 600-3 during a real-time video communication session. In some embodiments, devices 600-2 and 600-3 simultaneously display representations including images with different views. For example, when device 600-3 displays representation 622-2, device 600-2 displays representation 624-2.
[0321] In some embodiments, device 600-2 is positioned on desk 686 in front of user 623 in a manner corresponding to the position of surface 619 relative to user 622. Thus, user 623 can view representation 624-2 (including an image of surface 619) in a manner similar to how user 622 views surface 619 in the physical environment.
[0322] As Figure 6AEAs shown, during a real-time communication session, user 623 can modify the image displayed in representation 624-2 via mobile device 600-2. In response to user 623 changing the orientation of device 600-2, device 600-2 modifies the image of representation 624-2, for example, in a manner corresponding to the change in the orientation of device 600-2. For example, in response to user 623 tilting device 600-2, device 600-2 translates upward to display other portions of surface 619. In this way, user 623 can change the orientation of device 600-2 (in any direction) to view various portions of surface 619 that are not otherwise displayed when device 600-2 is in different orientations.
[0323] Figures 6AF to 6AL Illustrates an implementation for accessing reference Figures 6A to 6AE Illustrates and describes various user interface implementations. In the Figures 6AF to 6AL implementation depicted, a laptop computer (e.g., John's device 6100-1 and / or Jane's device 6100-2) is used to illustrate these interfaces. It should be understood that Figures 6AF to 6AL the implementation shown in can be implemented using different devices, such as a tablet computer (e.g., John's tablet 600-1 and / or Jane's device 600-2). Similarly, Figures 6A to 6AE the implementation shown in can be implemented using different devices (such as John's device 6100-1 and / or Jane's device 6100-2). Therefore, for the sake of brevity, the various operations or features described above with respect to Figures 6A to 6AE are not repeated below. For example, with respect to Figures 6A to 6AE the applications, interfaces (e.g., 604-1 and / or 604-2), and display elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, and / or 624-2) discussed in are similar to the applications, interfaces (e.g., 6121 and / or 6131), and display elements (e.g., 6124, 6132, 6122, 6134, 6116, 6140, and / or 6142) discussed in Figures 6AF to 6AL . Therefore, for the sake of brevity, the details of these applications, interfaces, and display elements may not be repeated below.
[0324] Figure 6AFDepicts John's device 6100-1, which includes a display 6101, one or more cameras 6102, and a keyboard 6103 (in some embodiments, the keyboard includes a touchpad). John's device 6100-1 displays a home screen via the display 6101 that includes a camera application icon 6108 and a video conferencing application icon 6110. The camera application icon 6108 corresponds to a camera application operating on John's device 6100-1, which can be used to access the camera 6102. The video conferencing application icon 6110 corresponds to a video conferencing application operating on John's device 6100-1, which can be used to initiate and / or participate in a real-time video communication session similar to that discussed above with reference to Figures 6A to 6AE The real-time video communication sessions (e.g., video calls and / or video chats). John's device 6100-1 also displays a task bar 6104, which includes various application icons, including a subset of the icons displayed in the dynamic area 6106. The icons displayed in the dynamic area 6106 represent applications that are active (e.g., started, opened, and / or used) on John's device 6100-1. In Figure 6AF , neither the camera application nor the video conferencing application is currently active. Accordingly, the icons representing the camera application or the video conferencing application are not displayed in the dynamic area 6106, and John's device 6100-1 is not participating in a real-time video communication session.
[0325] In Figure 6AF , John's device 6100-1 detects an input (e.g., cursor input caused by clicking a mouse, tapping on a touchpad, and / or other such input) selecting the camera application icon 6108 as indicated by the cursor 6112. In response, John's device 6100-1 launches the camera application and displays a camera application window 6114, as shown in Figure 6AG . In the embodiment depicted in Figure 6AG , the camera application is being used to access the camera 6102 to generate a surface view 6116, which is similar to, for example, the representation 624-1 depicted in Figure 6M and described above. In some embodiments, the camera application may have different modes (e.g., user-selectable modes), such as, for example, an extended field of view mode (which provides an extended field of view of the camera 6102) and a surface view mode (which provides the surface view shown in Figure 6AG ). Accordingly, the surface view 6116 represents an image obtained using the camera 6102 and modified (e.g., zoomed, rotated, cropped, and / or skewed) by the camera application to produce Figure 6AGA view of the image data of the surface view 6116 shown. Additionally, since John's laptop has launched the camera application, the camera application icon 6108-1 is displayed in the dynamic area 6106 of the taskbar 6104, indicating that the camera application is active. In some embodiments, when application icons (e.g., 6108-1) are added to the dynamic area of the taskbar, these application icons are displayed with an animated effect (e.g., bouncing).
[0326] In Figure 6AG , John's device 6100-1 detects an input 6118 to select the video conferencing application icon 6110. In response, John's device 6100-1 launches the video conferencing application, displays the video conferencing application icon 6110-1 in the dynamic area 6106, and displays the video conferencing application window 6120, as Figure 6AH shown. The video conferencing application window 6120 includes a video conferencing interface 6121 similar to the interface 604-1, and includes Jane's video feed 6122 (similar to the representation 623-1) and John's video feed 6124 (similar to the representation 622-1). In some embodiments, John's device 6100-1 displays the video conferencing application window 6120 with the video conferencing interface 6121 after one or more additional inputs after detecting the input 6118. For example, such an input can be an input to initiate a video call with Jane's laptop or to accept a request to participate in a video call with Jane's laptop.
[0327] In Figure 6AH , John's device 6100-1 displays the video conferencing application window 6120 partially covering the camera application window 6114. In some embodiments, in response to detecting a selection of the camera application icon 6108, a selection of the icon 6108-1, and / or an input on the camera application window 6114, John's device 6100-1 can bring the camera application window 6114 to the front or foreground (e.g., partially covering the video conferencing application window 6120). Similarly, in response to detecting a selection of the video conferencing application icon 6110, a selection of the icon 6110-1, and / or an input on the video conferencing application window 6120, the video conferencing application window 6120 can be brought to the front or foreground (e.g., partially covering the camera application window 6114).
[0328] In Figure 6AHIn [the figure], John's device 6100-1 is shown as participating in a real-time video communication session with Jane's device 6100-2. Thus, Jane's device 6100-2 is depicted as displaying a video conferencing application window 6130, which is similar to the video conferencing application window 6120 on John's device 6100-1. The video conferencing application window 6130 includes a video conferencing interface 6131 similar to interface 604-2, and includes John's video feed 6132 (similar to representation 622-2) and Jane's video feed 6134 (similar to representation 623-2).
[0329] In Figure 6AH the depicted embodiment, the video conferencing application is used to access camera 6102 to generate video feeds 6124 and 6132. Thus, video feeds 6124 and 6132 represent views of the image data that is obtained using camera 6102 and modified (e.g., zoomed in and / or cropped) by the video conferencing application to produce the images (e.g., videos) shown in video feeds 6124 and 6132. In some embodiments, the camera application and the video conferencing application may use different cameras to provide the respective video feeds.
[0330] The video conferencing application window 6120 includes menu option 6126, which can be selected to display different options for sharing content in a real-time video communication session. In Figure 6AH [the figure], John's device 6100-1 detects an input 6128 selecting menu option 6126, and in response, displays a sharing menu 6136, as shown in Figure 6AI [the figure]. The sharing menu 6136 includes sharing options 6136-1, 6136-2, and 6136-3. Sharing option 6136-1 is an option that can be selected to share content from the camera application. Sharing option 6136-2 is an option that can be selected to share content from the desktop of John's device 6100-1. Sharing option 6136-3 is an option that can be selected to share content from a presentation application. In response to detecting an input 6138 on sharing option 6136-1, John's device 6100-1 begins sharing content from the camera application, as shown in Figure 6AJ and Figure 6AK [the figure].
[0331] In Figure 6AJ [the figure], John's device 6100-1 updates the video conferencing interface 6121 to include a surface view 6140, which is shared with Jane's device 6100-2 in the real-time video communication session. In Figure 6AJIn the depicted embodiment, John's device 6100-1 shares the video feed generated using the camera application (shown as surface view 6116 in the camera application window 6114), and displays a representation of the video feed as surface view 6140 in the video conferencing application window 6120. Additionally, John's laptop emphasizes the display of surface view 6140 in the video conferencing interface 6121 (e.g., by displaying the surface view in a larger size than other video feeds) and reduces the display size of Jane's video feed 6122. In Figure 6AJ John's device 6100-1 simultaneously displays surface view 6140 along with John's video feed 6124 and Jane's video feed 6122 in the video conferencing application window 6120. In some embodiments, the display of John's video feed 6124 and / or Jane's video feed 6122 in the video conferencing application window 6120 is optional.
[0332] Jane's device 6100-2 updates the video conferencing interface 6131 to show surface video feed 6142, which is the surface view (from the camera application) shared by John's device 6100-1. As Figure 6AJ shown, Jane's device 6100-2 adds surface video feed 6142 to the video conferencing interface 6131 to simultaneously show the surface video feed along with Jane's video feed 6134 and John's video feed 6132, optionally with the size of John's video feed adjusted to accommodate the addition of surface video feed 6142. In some embodiments, Jane's device 6100-2 replaces John's video feed 6132 and / or Jane's video feed 6134 with surface video feed 6142.
[0333] Figure 6AK An alternative embodiment is shown depicting sharing content from the camera application in response to detecting an input 6138 on the sharing option 6136-1. In Figure 6AKIn this case, John's laptop displays a camera application window 6114 with a surface view 6116 (optionally minimizing or hiding the video conferencing application window 6120). John's device 6100-1 also displays John's video feed 6115 (similar to John's video feed 6124) and Jane's video feed 6117 (similar to Jane's video feed 6122), thereby indicating that John's laptop is sharing the surface view 6116 with Jane's device 6100-2 in a real-time video communication session (e.g., a video chat provided by a video conferencing application). In some embodiments, the display of John's video feed 6115 and / or Jane's video feed 6117 is optional. Similar to Figure 6AJ the embodiment shown in, Jane's device 6100-2 shows a surface video feed 6142 that is the surface view (from the camera application) shared by John's device 6100-1.
[0334] Figure 6AL shows a schematic diagram representing the field of view of the camera 6102 and a portion of the field of view of the video conferencing application and the camera application for the Figures 6AF to 6AK embodiment depicted in. For example, in Figure 6AL this case, an outline view of John's laptop 6100 is shown in John's physical environment. The dashed lines 6145-1 and the dotted lines 6147-2 represent the outer dimensions of the field of view of the camera 6102, which in some embodiments is a wide-angle camera. The collective field of view of the camera 6102 is indicated by the shaded regions 6144, 6146, and 6148. The portion of the camera field of view that is being used for the camera application (e.g., for the surface view 6116) is indicated by the dotted lines 6147-1 and 6147-2 and the shaded regions 6146 and 6148. In other words, the surface view 6116 (and the surface view 6140) is generated by the camera application using the portion of the camera field of view represented by the shaded regions 6146 and 6148 between the dotted lines 6147-1 and 6147-2. The portion of the camera field of view that is being used for the video conferencing application (e.g., for John's video feed 6124) is indicated by the dashed lines 6145-1 and 6145-2 and the shaded regions 6144 and 6146. In other words, John's video feed 6124 is generated by the video conferencing application using the portion of the camera field of view represented by the shaded regions 6144 and 6146 between the dashed lines 6145-1 and 6145-2. The shaded region 6146 represents the overlap of the portion of the camera field of view that is being used to generate the video feeds for the respective camera and video conferencing applications.
[0335] Figures 6AM to 6AY shows a reference for controlling Figures 6A to 6ALEmbodiments of various user interfaces and views shown and described and / or interacting with them. In Figures 6AM to 6AY In the embodiments depicted, these interfaces are shown using a tablet computer (e.g., John's tablet computer 600-1 and / or Jane's device 600-2) and a computer (e.g., Jane's computer 600-4). Figures 6AM to 6AY The embodiments shown in Figures 6A to 6AL are optionally implemented using different devices such as laptop computers (e.g., John's device 6100-1 and / or Jane's device 6100-2). Similarly, Figures 6A to 6AL the embodiments shown in
[0336] are optionally implemented using different devices such as Jane's computer 6100-2. Therefore, for the sake of brevity, the various operations or features described above with respect to Figures 6A to 6AL are not repeated below. Figures 6AM to 6AY In addition, the applications, interfaces (e.g., 604-1, 604-2, 6121, and / or 6131), and fields of view (e.g., 620, 688, 6145-1, and 6147-2) provided by one or more cameras (e.g., 602, 682, and / or 6102) with respect to Figures 6A to 6AL are similar to the applications, interfaces (e.g., 604-4), and fields of view (e.g., 620) provided by a camera (e.g., 602) with respect to Figures 6AM to 6AY Therefore, for the sake of brevity, the details of these applications, interfaces, and fields of view may not be repeated below. In addition, the controls detected by device 600-1 with respect to Figures 6AM to 6AY the options and requests (e.g., inputs and / or hand gestures) for the views associated with the elements of the display (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, 624-2, 6121, and / or 6131) discussed with respect to Figure 6O are optionally detected by device 600-2 and / or device 600-4 to control the views (e.g., user 623 optionally provides an input to cause device 600-1 and / or device 600-2 to provide a representation 624-1 of a modified image including a surface) associated with the elements of the display (e.g., 622-1, 622-4, 623-1, 623-4, 6214, and / or 6216) discussed with respect to
[0337] Figures 6AM to 6AJAn exemplary user interface for controlling views of a physical environment is shown and described. The user interfaces in these figures are used to illustrate processes including the processes Figure 15 described below. At Figure 6AM , devices 600-1 and 600-4 respectively display interfaces 604-1 and 604-4. Interface 604-1 includes representation 622-1, and interface 604-4 includes representation 622-4. Representations 622-1 and 622-4 include images of image data from a portion of the field of view 620 (specifically, shaded region 6206). As shown, representations 622-1 and 622-4 include images of the head and upper torso of user 622 and do not include an image of drawing 618 on desk 621. Interfaces 604-1 and 604-4 respectively include representations 623-1 and 623-4, which include images of user 223 in the field of view 6204 of camera 6202. Interfaces 604-1 and 604-4 also include option menus 609 (similar to option menu 608 discussed with respect to Figures 6A to 6AE for controlling image data captured by 602 and / or captured by camera 6202, including Figures 6F to 6G ), which allow devices 600-1 and 600-4 to manage how the image data is displayed.
[0338] At Figure 6ANAt location, user 623 brings device 600-2 near device 600-4 during a real-time video communication session. As depicted, in response to detecting device 600-2 (e.g., via wireless communication), device 600-4 displays an add notification 6210a. Similarly, in response to detecting device 600-4, device 600-2 displays an add notification 6210b via display 683 (e.g., a touch-sensitive display). In some embodiments, devices 600-2 and 600-4 use specific device criteria to trigger the display of add notifications 6210a and 6210b. In some embodiments, the specific device criteria include criteria for a specific location (e.g., positioning, orientation, and / or angle) of device 600-2 that, when met, trigger the display of add notifications 6210a and / or 6210b. In such embodiments, the specific location (e.g., positioning, orientation, and / or angle) of device 600-2 includes criteria where device 600-2 has a specific angle or is within an angle range (e.g., an angle or angle range indicating that the device is horizontal and / or flat on desk 686) and / or the display 683 faces up (e.g., as opposed to facing down towards desk 686). In some embodiments, the specific device criteria include criteria that device 600-2 is near device 600-4 (e.g., within a threshold distance of device 600-4). In some embodiments, devices 600-2 and 600-4 communicate wirelessly to convey the location and / or proximity of device 600-2 (e.g., using location data and / or short-range wireless communication such as Bluetooth and / or NFC). In some embodiments, the specific device criteria include criteria that devices 600-2 and 600-4 are associated with the same user (e.g., are being used by the same user and / or logged in to the same user). In some embodiments, the specific device criteria include criteria that device 600-2 has a specific state (e.g., unlocked and / or display powered on, as opposed to locked and / or display powered off).
[0339] At Figure 6AN location, the connection notifications 6210a-6210b include an indication of including device 600-4 in the real-time video communication session. For example, the add notifications 6210a-6210b include an indication of adding a representation (for display on device 600-2) that includes an image of the field of view 620 captured by camera 602. In some embodiments, the add notifications 6210a-6210b include an indication of adding a representation (for display on device 600-1) that includes an image of the field of view 6204 captured by camera 6202.
[0340] At Figure 6ANAt this point, adding notifications 6210a and 6210b includes acceptance indicia 6212a and 6212b, which when selected add (e.g., connect) device 600-2 to a real-time video communication session. Notifications 6210a and 6210b include rejection indicia 6213a and 6213b, which when selected clear notifications 6210a and 6210b respectively without adding device 600-2 to the real-time video communication session. When displaying acceptance indicium 6212b, device 600-2 detects an input 6250an (e.g., a tap, a mouse click, or other selection input) directed at acceptance indicium 6212b. In response to detecting input 6250an, device 600-2 displays interface 604-2, as Figure 6AO depicted in.
[0341] At Figure 6AO this point, interface 604-2 is similar to interface 604-2 described herein (e.g., reference Figures 6A to 6AE ), and video conferencing interface 6131 as described herein (e.g., reference Figures 6AH to 6AK ), but has a different state. For example, Figure 6AO the interface 604-2 does not include indicia 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and option menu 609. In some embodiments, Figure 6AO the interface 604-2 includes one or more of indicia 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and / or option menu 609.
[0342] At Figure 6AO this point, interface 604-2 includes an adjustable view 6214 of the video feed captured by camera 602 (similar to John's video feed 6132 and indicia 622-2, but including a different portion of the field of view 620). Adjustable view 6214 is associated with a portion of the field of view 620 corresponding to shaded region 6217. In some embodiments, Figure 6AO the interface 604-2 includes indicia 622-4 and 623-4 and / or option menu 609. In some embodiments, in response to an input detected at device 600-2 and / or device 600-4, indicia 622-4 and 623-4 and / or option menu 609 are moved from interface 604-4 to interface 604-2 for simultaneous display with adjustable view 6214. In such embodiments, display 6201 acts as an auxiliary display (e.g., an extended display) to display 604-1 and / or vice versa.
[0343] At Figure 6AO this point, in response to an input detected at Figure 6ANAt input 6250an is detected, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as Figure 6AO depicted in. Figure 6AO The interface 604-1 of Figure 6AN is similar to the interface 604-1 of
[0344] but has different states (e.g., indicating that 623-1 and 622-1 are smaller in size and in different positions). Interface 604-1 includes an adjustable view 6216, which is similar to the adjustable view 6214 displayed at device 600-2 (e.g., the adjustable view 6216 is associated with a portion of the field of view 620 corresponding to the shaded region 6217). When the input described herein (e.g., movement of device 600-2) is detected by device 600-2, the adjustable view 6216 is updated to include an image similar to the adjustable view 6214. Displaying the adjustable view 6216 allows user 622 to see what portion of the field of view 620 user 624 is currently viewing, because, as described in more detail below, user 623 optionally controls what view within the display field of view 620 is shown.
[0344] At Figure 6AO while displaying interface 604-2, device 600-2 detects movement 6218ao of device 600-2. In response to detecting movement 6218ao, device 600-2 displays Figure 6AP interface 602-4. Additionally, in response to detecting movement 6218ao, device 600-2 causes device 600-1 to display Figure 6AP interface 604-1.
[0345] At Figure 6AP interface 602-4 includes an updated adjustable view 6214. Compared to the adjustable view 6214 of Figure 6AO the adjustable view 6214 of Figure 6AP is a different view within the field of view 620. For example, Figure 6AP the shaded region 6217 of Figure 6AOThe shaded region 6217 moves. Notably, the camera 602 does not move. In some embodiments, the movement 6218ao of the device 600-2 corresponds to the amount of change of the adjustable view 6214 (e.g., is proportional thereto). For example, in some embodiments, the magnitude of the angle by which the device 600-2 rotates (e.g., relative to gravity) corresponds to the amount of change of the adjustable view 6214 (e.g., the image data is translated to include the amount of the new angle of the view). In some embodiments, the direction of the movement of the device 600-2 (e.g., movement 6218ao) (e.g., tilting downwards and / or rotating downwards) corresponds to the direction of change of the adjustable view 6214 (e.g., translating downwards). In some embodiments, the acceleration and / or speed of the movement (e.g., movement 6218) corresponds to the speed of change of the adjustable view 6214. In some embodiments, the device 600-2 (and / or the device 600-1) displays a gradual transition (e.g., a series of views) of the adjustable view 6214 from Figure 6AO to the adjustable view 6214 in Figure 6AP . Additionally or alternatively, as depicted in Figure 6AP , the device 600-2 lies flat on the desk 686. In some embodiments, in response to detecting a position within a specific position or a predetermined range of positions (e.g., horizontal and / or upward display), the device 600-2 displays Figure 6AP of the adjustable view 6214. As depicted, Figure 6AO the movement 6218ao does not cause the device 600-2 to update Figure 6AP the representations 622-4 and 623-4 (and / or the representations 623-1 and 622-1 on the device 600-1).
[0346] At Figure 6AP , the image of the drawing 618 in the adjustable view 6214 is at a different perspective from the image of the drawing 618 in the adjustable view 6214 of Figure 6AO . For example, Figure 6AP the adjustable view 6214 of Figure 6AO includes a top view perspective, while the adjustable view 6214 of Figure 6AP includes a perspective that includes a combination of a side view and a top view. In some embodiments, the image of the drawing included in the adjustable view 6214 of Figures 6A to 6AL is based on image data that has been modified (e.g., skewed and / or magnified) using similar techniques described with reference to Figure 6AO . In some embodiments, the image of the drawing included in the adjustable view 6214 of Figure 6APThe image of the drawing 618 in the adjustable view 6214 is modified (e.g., to a lesser extent) in a different way (e.g., skewed and / or magnified less as compared to the amount of skew and / or magnification applied in Figure 6AP ). Providing a top-down view perspective provides greater convenience in terms of collaborating and sharing content because this gives user 623 a view of the drawing that would be similar to the view that user 623 would have if user 623 were sitting across from user 622 looking down at the surface 619 of the desk 621.
[0347] At Figure 6AP , the adjustable view 6216 of the interface 604-1 has also been updated in a similar manner. In some embodiments, the images of the adjustable view 6216 and / or the adjustable view 6214 are modified based on the position of the surface 619 relative to the camera 602, as described with reference to Figures 6A to 6AL . In such embodiments, the device 600-1 and / or the device 600-2 rotates the image of the adjustable view 6214 by a certain amount (e.g., 45 degrees, 90 degrees, or 180 degrees) such that the image of the drawing 618 can be viewed more intuitively in the adjustable view 6216 and / or the adjustable view 6214 (e.g., displaying the image of the drawing 618 such that the house is right-side up as opposed to upside down).
[0348] At Figure 6AP , user 623 applies a digital mark to the adjustable view 6214 using a stylus 6220. For example, when the adjustable view 6214 of Figure 6AP is displayed, the device 600-2 detects an input corresponding to a request to add a digital mark to the adjustable view 6214 (e.g., using the stylus 6220). In response to detecting an input corresponding to a request to add a digital mark to the adjustable view 6214, the device 600-2 displays the interface 604-2, as depicted in Figure 6AO . Additionally or alternatively, in response to detecting an input corresponding to a request to add a digital mark to the adjustable view 6214, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) the interface 604-1, as depicted in Figure 6AQ .
[0349] At Figure 6AQAt this point, interface 602-4 includes digital sun 6222 in adjustable view 6214, and interface 602-1 includes digital sun 6223 in adjustable view 6214. Displaying the digital sun at both devices allows users 623 and 622 to collaborate above the video communication session. Additionally, as depicted, digital sun 6222 has a position relative to the image of drawing 618. As described in more detail below, even if device 600-1 detects further movement and / or if drawing 618 moves on surface 619, digital sun 6222 maintains its position relative to the image of drawing 618. In some embodiments, device 600-2 stores data corresponding to the relationship between a digital marker detected in the image data (e.g., digital sun 6223) and an object (e.g., a house) in order to determine where (and / or whether) to display digital sun 6222. In some embodiments, device 600-2 stores data corresponding to the relationship between a digital marker (e.g., digital sun 6223) and the position of device 600-2 in order to determine where (and / or whether) to display digital sun 6222. In some embodiments, device 600-2 detects digital markers applied to other views in the field of view 620. For example, a digital marker may be applied to an image of the user's head, such as Figure 6AR the image of the head of user 622 in adjustable view 6214.
[0350] At Figure 6AQ this point, interface 604-2 includes control affordance representation 648-1 (similar to Figure 6N control affordance representation 648-1 in Figure 6N ) to modify the image in adjustable view 6214. Rotation affordance representation 648-1, when selected, causes device 600-1 (and / or device 600-2) to rotate the image of adjustable view 6214, similar to how control affordance representation 648-1 modifies the image of representation 624-1 in
[0351] At Figure 6AQ this point, in some embodiments, interface 604-2 includes a zoom affordance representation similar to zoom affordance representation 648-2 in Figure 6N . In such embodiments, the zoom affordance representation modifies the image in adjustable view 6214, similar to how zoom affordance representation 648-2 modifies the image of representation 624-1 in Figure 6N . Control affordance representations 648-1, 648-2 may be displayed in response to one or more inputs corresponding to, for example, the selection of an affordance representation of option menu 609 (e.g., Figure 6AM ).
[0352] At Figure 6AQAt, in some embodiments, a digital sun 6222 is projected onto the physical surface of the drawing 618, similar to how the marker 956 is projected onto the surface 908b, which is described in Figures 9K to 9N . In such embodiments, an electronic device (e.g., a projector and / or a light-emitting projector) is used to project an image and / or rendering of the digital sun 6222 within the physical environment 915. For example, the electronic device may use the techniques described with respect to Figures 9K to 9N to cause the projection of the digital sun to be displayed next to the drawing 618 based on the relative position of the digital sun 6222 with respect to the drawing 618.
[0353] At Figure 6AQ , when the digital sun 6222 is displayed in the adjustable view 6214, the device 600-2 detects a movement 6218aq (e.g., a rotation and / or a lift). In response to detecting the movement 6218aq, the device 600-2 displays the interface 604-2, as depicted in Figure 6AR . In response to detecting the movement 6218aq, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) the interface 604-1, as depicted in Figure 6AR .
[0354] At Figure 6AR , the interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in the interface 606-1). Compared to the adjustable view 6214 of Figure 6AQ , the adjustable view 6214 of Figure 6AR is a different view within the field of view 620. For example, Figure 6AP 's shadow region 6217 has moved relative to Figure 6AQ 's shadow region 6217. In some embodiments, the direction of the movement 6218aq (e.g., tilting up) corresponds to the direction of the change in the view (e.g., panning up). Additionally, the shadow region 6217 overlaps with the shadow region 6206, as depicted by the darker shadow region 6224. The darker shadow region 6224 is a schematic representation based on a portion of the image data used to represent 622-4. Because the movement 6218aq has caused a change in the view (e.g., a change to a view of the face of the user 622 and / or a view of something other than the drawing 618), the device 600-2 no longer displays the digital sun 6222 in the adjustable view 6214.
[0355] At Figure 6ARAt location, the adjustable view 6214 includes a boundary indicator 6226. The boundary indicator 6226 indicates that a boundary has been reached. In some embodiments, the boundary is configured to (e.g., by the user) set a limit on what portion of the field of view 620 (or the environment captured by the camera 602) is to be provided for display. For example, the user 622 may limit what portion is available to the user 623. In some embodiments, the boundary is defined by the physical limitations (e.g., image sensor and / or lens) of the camera 602 that provides the field of view 620. At Figure 6AR location, the shaded area 6217 has not reached the limit of the field of view 620. Thus, the boundary indicator 6226 is based on a configurable setting that limits what portion of the field of view 620 is provided for display. Temporarily turning to Figure 6AT , the boundary indicator 6226 is displayed in response to determining that the viewing angle provided in the adjustable view 6214 has reached the edge of the field of view 620.
[0356] At Figure 6AR location, the boundary indicator 6226 is depicted with cross-hatching. In some embodiments, the safety boundary indicator 6226 is a visual effect (e.g., blur and / or fade) applied to the adjustable view 6214 (and / or adjustable view 6216). In some embodiments, the boundary indicator 6226 is displayed along the edge of the adjustable view 6214 (and / or 6216) to indicate the location of the boundary. At Figure 6AR location, the boundary indicator 6226 is displayed along the top edge and side edges to indicate that the user cannot see above the boundary indicator 6226 and / or further to the side of the boundary indicator. When at Figure 6AR location, the device 600-2 detects movement 6218ar (e.g., rotation and / or lowering). In response to detecting the movement 6218ar, the device 600-2 displays the interface 604-2 as depicted in Figure 6AS . In response to detecting the movement 6218ar, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) the interface 604-2 as depicted in Figure 6AS .
[0357] At Figure 6AS location, the interface 604-2 includes an updated adjustable view 6214 that includes an image of the drawing 618. At Figure 6AS location, the device 600-2 is in a position similar to that of the device 600-2 in Figure 6AO . Thus, Figure 6AS the adjustable view 6214 includes an image of the drawing 618 in the adjustable view 6214 having the same viewing angle as the image of the drawing 618 in the adjustable view 6214 in Figure 6AO . Notably, the device 600-2 is atFigure 6AS The digital sun 6222 is shown in the adjustable view 6214. Figure 6AS The position of the digital sun 6222 in relation to the house in drawing 618 is similar to Figure 6AQ the position of the digital sun 6222 in relation to the house in drawing 618, except for minor differences based on different views. In this way, the digital sun 6222 appears fixed in physical space as if it were drawn next to drawing 618. Fixing the position of the digital markings in physical space facilitates better collaboration among users, as users can digitally draw or write in the view, move the device to see different views, and then move the device back to redisplay the digital drawing or text and the context in which they were created.
[0358] For clarity, the shaded regions 6217 and 6206 and the field of view 620 have been omitted from Figures 6AS to 6AU In some embodiments, the representations 622-1 and the adjustable views 6214 and 6216 correspond to views associated with Figure 6AO the shaded regions 6217 and 6206 and the field of view 620.
[0359] At Figure 6AS , the device 600-2 (and / or device 600-1) detects the movement of drawing 618 and maintains the display of the image of drawing 618 in the adjustable view 6214. In some embodiments, the device 600-2 (and / or device 600-1) uses image correction software to modify (e.g., scale, skew, and / or rotate) the image data in order to maintain the display of the image of drawing 618 in the adjustable view 6214. When displaying interface 604-2, the device 600-2 (and / or device 600-1) detects the horizontal movement 6230 of drawing 618. In response to detecting the horizontal movement 6230 of drawing 618, the device 600-2 displays interface 604-2 as depicted in Figure 6AT . In some embodiments, in response to detecting the horizontal movement 6230 of drawing 618, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) interface 604-2 as depicted in Figure 6AT . In some embodiments, in response to the device 600-1 detecting the horizontal movement 6230 of drawing 618, the device 600-2 displays (and / or the device 600-1 causes the device 600-2 to display) interface 602-4 as depicted in Figure 6AT .
[0360] At Figure 6AT , drawing 618 has been moved to the edge of the desk 621 that is further away from the camera 602 (e.g., and closer to the side of the camera). Despite the change in position, Figure 6ATThe interface 602-4 includes an image of the drawing 618 in the adjustable view 6214, which looks almost unchanged compared to the image of the drawing 618 in the adjustable view 6214 of the interface 602-4. For example, the adjustable view 6214 provides a perspective that makes it look like the drawing 618 is still directly in front of the camera 602, similar to the position of the drawing 618 in Figure 6AS . In some embodiments, the device 600-2 (and / or the device 600-1) uses image correction software to correct (e.g., by skewing and / or magnifying) the image of the drawing 618 based on the new position relative to the camera 602. In some embodiments, the device 600-2 (and / or the device 600-1) uses object detection software to track the drawing 618 as it moves relative to the camera 602. In some embodiments, the adjustable view 6214 of the interface 604-2 is provided without any change in the position (e.g., positioning, orientation, and / or rotation) of the camera 602. Figure 6AS At Figure 6AT , the device 600-2 displays a boundary indicator 6226 in the adjustable view 6216 (similar to the adjustable view 6214 displayed by the device 600-1 in the adjustable view 6214). As discussed above with respect to
[0361] , the boundary indicator 6226 indicates that the limit of the field of view or physical space has been reached. At Figure 6AT , the device 600-2 displays a boundary indicator 6226 in the adjustable view 6214 to indicate that the edge of the field of view 620 has been reached. The boundary indicator 6226 is along the right edge of the adjustable view 6214 (and the adjustable view 6216), indicating that the view to the right of the current view is outside the field of view of the camera 602. Figure 6AR As Figure 6AT discussed above, the boundary indicator 6226 indicates that the limit of the field of view or physical space has been reached. At
[0362] At Figure 6AT , the digital sun 6222 maintains a corresponding position relative to the house in the image of the drawing 618 in the adjustable view 6214 that is similar to the corresponding position of the digital sun 6222 relative to the house in the image of the drawing 618 in the adjustable view 6214 of Figure 6AS . In some embodiments, the device 600-2 (and / or the device 600-1) displays the digital sun 6222 overlaid on the image of the drawing 618 that has been corrected based on the new position of the drawing 618.
[0363] Returning temporarily to Figure 6AS , when displaying the interface 602-4, the device 600-2 (and / or the device 600-1) detects the rotation 6232 of the drawing 618. In response to detecting the rotation 6232 of the drawing 618, the device 600-2 displays the interface 604-2, as shown in Figure 6AUas depicted. In some embodiments, in response to detecting a rotation 6232 of the drawing 618, device 600-2 causes device 600-1 to display interface 601-4, as Figure 6AU depicted. In some embodiments, in response to device 600-1 detecting a rotation 6232 of the drawing 618, device 600-2 displays (or device 600-1 causes device 600-2 to display) interface 602-4, as Figure 6AU depicted.
[0364] At Figure 6AU the drawing 618 has rotated relative to the edge of the desk 621. Despite the change in position, Figure 6AU interface 602-4 in Figure 6AS includes an image of the drawing 618 in the adjustable view 6214 that appears to have changed little compared to the image of the drawing 618 in the adjustable view 6214 of interface 604-2 in Figure 6AU The adjustable view 6214 of Figure 6AS provides a perspective that makes it appear as if the drawing 618 has not rotated, similar to the position of the drawing 618 in Figure 6AU In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct (e.g., by skewing and / or rotating) the image of the drawing 618 based on its new position relative to the camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track the drawing 618 as it rotates relative to the camera 602. In some embodiments, the adjustable view 6214 of interface 604-2 is provided without any change in the position (e.g., positioning, orientation, and / or rotation) of the camera 602. The adjustable view 6216 is updated in a manner similar to the adjustable view 6214.
[0365] At Figure 6AU the digital sun 6222 maintains a position relative to the house in the image of the drawing 618 in the adjustable view 6214 that is similar to the position of the digital sun 6222 relative to the house in the image of the drawing 618 in the adjustable view 6214 of Figure 6AS In some embodiments, device 600-2 (and / or device 600-1) displays the digital sun 6222 overlaid on the image of the drawing 618 that has been corrected based on the rotation of the drawing 618.
[0366] At Figure 6AV device 600-2 displays interface 604-2, which is similar to Figure 6AUThe interface 604-2 but with different states (e.g., John's representation 622-2 and the option menu 609 have been added to the user interface 604-2). The device 600-4 is no longer used in the real-time communication session. Additionally, the device 600-2 has moved from its position in Figure 6AU to the same position that the device 600-2 has in Figure 6AQ . Thus, the device 600-2 updates Figure 6AV the adjustable view 6214 to include the same perspective as the adjustable view 6214 in Figure 6AQ . As shown, the adjustable view 6214 includes a top view perspective. Additionally, the digital sun 6222 is shown as having the same position relative to the house in the image of the drawing 618 in the adjustable view 6214 of Figure 6AQ .
[0367] At Figure 6AV , when the digital sun 6222 is displayed in the adjustable view 6214, the device 600-2 detects a movement 6218av (e.g., rotation and / or lift). In response to detecting the movement 6218av, the device 600-2 displays the interface 604-2 as depicted in Figure 6AW . In response to detecting the movement 6218aw, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) the interface 604-1 as depicted in Figure 6AW .
[0368] At Figure 6AW , the interface 604-2 includes an updated adjustable view 6214 similar to the adjustable view 6214 in Figure 6AR (which corresponds to the updated adjustable view 6216 in the interface 606-1). Notably, the device does not update the representation 622-2 in response to detecting the movement 6218aw. Thus, in some embodiments, the device 600-2 displays a dynamic representation updated based on the position of the device 600-2 and a static representation not updated based on the position of the device 600-2. The interface 604-2 also includes a boundary indicator 6226 in the adjustable view 6214, similar to the boundary indicator 6226 in Figure 6AR .
[0369] At Figure 6AW , when displaying the interface 604-2, the device 600-2 detects a movement 6218aw (e.g., rotation and / or lowering). In response to detecting the movement 6218aw, the device 600-2 displays the interface 604-2 as depicted in Figure 6AX . In response to detecting the movement 6218aw, the device 600-1 displays (and / or the device 600-2 causes the device 600-1 to display) the interface 604-1 as depicted in Figure 6AXas depicted in.
[0370] At Figure 6AX interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). Since the adjustable view 6214 is substantially the same view provided by presentation 622-2, the shaded region 6206 overlaps with the shaded region 6217. Since movement 6218aq causes the view to change to a view of user 622's face and / or non-drawing 618, device 600-2 no longer displays digital sun 6222 in adjustable view 6214. When Figure 6AX interface 604-2 is displayed at, device 600-2 (and / or device 600-1) detects a set of one or more inputs corresponding to a request for a display surface view (e.g., inputs and / or gestures similar to those described with reference to Figures 6A to 6AL ). In some such embodiments, at device 600-2, the presentation 616-1 of Figure 6C , the presentation 616-2 of Figure 6G , the preview mode 674-1 of Figure 6H , the presentation 676 of the preview mode 674-2 in Figure 6I , the presentation 676 of the preview mode interface 674-3 in Figure 6J , and the affordance presentations 648-1, 648-2, 648-3 of Figures 6N to 6Q are displayed to allow device 600-2 to control the presentation of the modified image of drawing 618 in the same manner as the inputs detected at device 600-1. In response to detecting the set of one or more inputs corresponding to a request for a display surface view, device 600-2 displays interface 604-2, as Figure 6AY depicted in. Additionally or alternatively, in response to detecting the set of one or more inputs, device 600-1 displays interface 604-2, as Figure 6AY depicted in. In some embodiments, device 600-1 detects the set of one or more inputs, as described with reference to Figures 6A to 6AL . In some embodiments, device 600-2 detects the set of one or more inputs. In such embodiments, device 600-2 detects a selection of the view affordance presentation 6236 of option menu 609, and the view affordance presentation of this option menu is similar to the view affordance presentation 607-2 of option menu 608 described with reference to Figure 6F . In response, a view menu similar to view menu 616-2 described with reference to Figure 6G includes an affordance presentation requesting a display of the surface view of a remote participant.
[0371] At Figure 6AY the adjustable view 6214 includes a surface view that is similar to, for example,Figure 6M the representation 624-1 depicted and described above. As Figure 6AY depicted in, the adjustable view 6214 includes an image of the drawing 618 displayed on the device 600-2 that is modified such that the user 623 has a downward view similar to the perspective that the user 622 has when looking down at the drawing 618 in the physical environment, as described in more detail with respect to Figures 6A to 6AL described above. It is noted that Figure 6AY the digital sun 6222 of Figure 6AQ is shown as having the same position relative to the house in the image of the drawing 618 in the adjustable view 6214 as the digital sun 6222 of
[0372] Figure 7 is a flowchart showing a method for using a computer system to manage a real-time video communication session according to some embodiments. The method 700 is performed at a computer system (e.g., 600-1, 600-2, 600-3, 600-4, 906a, 906b, 906c, 906d, 6100-1, 6100-2, 1100a, 1100b, 1100c, and / or 1100d) (e.g., a smart phone, a tablet computer, a laptop computer, and / or a desktop computer) (e.g., 100, 300, or 500) that is coupled to a display generation component (e.g., 601, 683, and / or 6101) (e.g., a display controller, a touch-sensitive display system, and / or a monitor), one or more cameras (e.g., 602, 682, and / or 6102) (e.g., an infrared camera, a depth camera, and / or a visible light camera), and one or more input devices (e.g., 601, 683, and / or 6103) (e.g., a touch-sensitive surface, a keyboard, a controller, and / or a mouse). Some of the operations in the method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0373] As described below, the method 700 provides an intuitive way to manage a real-time video communication session. The method reduces the cognitive burden on the user for managing a real-time video communication session, thereby creating a more effective human-machine interface. For battery-powered computing devices, enabling the user to manage a real-time video communication session faster and more efficiently saves power and increases the time interval between two battery charges.
[0374] In method 700, a computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays (702), via a display generation component, a real-time video communication interface (e.g., 604-1, 604-2, 6120, 6121, 6130, and / or 6131) (e.g., an interface for an incoming and / or outgoing real-time audio / video communication session) for a real-time video communication session. In some embodiments, the real-time communication session is at least between a computer system (e.g., a first computer system) and a second computer system.
[0375] The real-time video communication interface includes a representation (e.g., 622-1, 622-2, 6124, and / or 6132) (e.g., a first representation) of at least a portion of the field of view of the one or more cameras (e.g., 620, 688, 6144, 6146, and / or 6148). In some embodiments, the first representation includes an image of a physical environment (e.g., a scene and / or region of the physical environment within the field of view of the one or more cameras). In some embodiments, the representation includes a portion of the field of view of the one or more cameras (e.g., a first cropped portion). In some embodiments, the representation includes a static image. In some embodiments, the representation includes a series of images (e.g., video). In some embodiments, the representation includes a real-time (e.g., instant) video feed of the field of view of the one or more cameras (or a portion thereof). In some embodiments, the field of view is based on the physical characteristics of the one or more cameras (e.g., orientation, lens, focal length of the lens, and / or sensor size). In some embodiments, the representation is displayed in a window (e.g., a first window). In some embodiments, the representation of at least a portion of the field of view includes an image of a first user (e.g., the face of the first user). In some embodiments, the representation of at least a portion of the field of view is provided by an application (e.g., 6110) that provides the real-time video communication session. In some embodiments, the representation of at least a portion of the field of view is provided by an application (e.g., 6108) different from the application (e.g., 6110) that provides the real-time video communication session.
[0376] When displaying a real-time video communication interface, a computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) detects (704) via the one or more input devices (e.g., 601, 683, and / or 6103) one or more user inputs (e.g., 612c, 612d, 614, 612g, 612i, 612j, 6112, 6118, 6128, and / or 6138) (e.g., taps on a touch-sensitive surface, keyboard input, mouse input, touchpad input, gestures (e.g., hand gestures), and / or audio input (e.g., voice commands)), where the user input points to a surface (e.g., 619) in a scene (e.g., a physical environment) within the field of view of the one or more cameras (e.g., a physical surface; the surface of a desk and / or the surface of an object (e.g., a book, paper, tablet) resting on the desk; or the surface of a wall and / or an object on the wall (e.g., a whiteboard or blackboard); or other surfaces (e.g., a freestanding whiteboard or blackboard)). In some embodiments, the user input corresponds to a request for a view of the display surface. In some embodiments, detecting the user input via the one or more input devices includes obtaining image data of the field of view of the one or more cameras that includes gestures (e.g., hand gestures, eye poses, or other body postures). In some embodiments, the computer system determines from the image data that the gesture meets a predetermined criterion.
[0377] In response to detecting the one or more user inputs, the computer system (e.g., 600-1, 600-2, 6100-1, and / or 6100-2) displays a representation (e.g., an image and / or video), e.g., a second representation, of a surface (e.g., 624-1, 624-2, 6140, and / or 6142) via a display generation component (e.g., 601, 683, and / or 6101). In some embodiments, the representation of the surface is obtained by digitally zooming and / or panning the field of view captured by the one or more cameras. In some embodiments, the representation of the surface is obtained by moving (e.g., panning and / or rotating) the one or more cameras. In some embodiments, the second representation is displayed in a window (e.g., a second window, the same window in which the first representation is displayed, or a window different from the window in which the first representation is displayed). In some embodiments, the second window is different from the first window. In some embodiments, the second window (e.g., 6140 and / or 6142) is provided by an application (e.g., 6110) that provides a real-time video communication session (e.g., as shown in Figure 6AJ ). In some embodiments, the second window (e.g., 6114) is provided by an application different from the application that provides the real-time video communication session (e.g., as shown in Figure 6AKProvided by an application (e.g., 6108) as shown. In some embodiments, the second representation includes a cropped portion (e.g., a second cropped portion) of the field of view of the one or more cameras. In some embodiments, the second representation is different from the first representation. In some embodiments, the second representation is different from the first representation because the second representation shows a portion (e.g., a second cropped portion) of the field of view that is different from a portion (e.g., a first cropped portion) shown in the first representation (e.g., a pan view, a zoomed-out view, and / or a zoomed-in view). In some embodiments, the second representation includes an image of a portion of the scene that is not included in the first representation and / or the first representation includes an image of a portion of the scene that is not included in the second representation. In some embodiments, the surface is not shown in the first representation.
[0378] The representation of the surface (e.g., 624-1, 624-2, 6140, and / or 6142) includes an image (e.g., a photograph, a video, and / or a real-time video feed) (sometimes referred to as a representation of the modified image of the surface) of the surface (e.g., 619) captured by the one or more cameras (e.g., 602, 682, and / or 6102) that has been modified (e.g., adjusted, manipulated, corrected) based on the position (e.g., location and / or orientation) of the surface relative to the one or more cameras (e.g., to correct for distortion of the image of the surface). In some embodiments, the image of the surface displayed in the second representation is based on image data that has been modified using image processing software (e.g., skewed, rotated, flipped, and / or otherwise manipulated the image data captured by the one or more cameras). In some embodiments, the image of the surface displayed in the second representation is modified without physically adjusting the camera (e.g., without rotating the camera, without lifting the camera, without lowering the camera, without adjusting the angle of the camera, and / or without adjusting the physical components of the camera (e.g., the lens and / or the sensor)). In some embodiments, the image of the surface displayed in the second representation is modified such that the camera appears to be pointed at the surface (e.g., facing the surface, aiming at the surface, pointed along an axis perpendicular to the surface). In some embodiments, the image of the surface displayed in the second representation is corrected such that the line of sight of the camera appears to be perpendicular to the surface. In some embodiments, the image of the scene displayed in the first representation is not modified based on the position of the surface relative to the one or more cameras. In some embodiments, the representation of the surface is displayed simultaneously with the first representation (e.g., the first representation is maintained (e.g., for a user of a computer system), and the image of the surface is displayed in a separate window). In some embodiments, the image of the surface is modified instantaneously and automatically (e.g., during a real-time video communication session). In some embodiments, the image of the surface is automatically modified based on the position of the surface relative to the one or more first cameras (e.g., without user input). The representation of the surface that includes an image of the surface modified based on the position of the surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the surface regardless of the position of the surface relative to the camera without further input from the user (which provides improved visual feedback and reduces the amount of input required to perform an operation).
[0379] In some embodiments, a computer system (e.g., 600-1 and / or 600-2) receives, during a real-time video communication session, image data captured by one of the one or more cameras (e.g., 602) (e.g., a wide-angle camera). The computer system displays, via a display generation component, a representation (e.g., 622-1 and / or 622-2) (e.g., a first representation) of at least a portion of the field of view based on the image data captured by the camera. The computer system displays, via the display generation component, a representation of a surface (e.g., 624-1 and / or 624-2) (e.g., a second representation) based on the image data captured by the camera (e.g., a representation of at least a portion of the field of view of the one or more cameras and the representation of the surface is based on image data captured by a single (e.g., only one) camera of the one or more cameras). Displaying a representation of at least a portion of the field of view and a representation of the surface captured from the same camera enhances the video communication session experience by displaying content captured from the same camera at different perspectives without input from the user (which reduces the number of inputs (and / or devices) required to perform an operation).
[0380] In some embodiments, the image of the surface is modified (e.g., by the computer system) by rotating the image of the surface relative to a representation of at least a portion of the field of view of the one or more cameras (e.g., the image of the surface in 624-2 is rotated 180 degrees relative to representation 622-2). In some embodiments, the representation of the surface is rotated 180 degrees relative to a representation of at least a portion of the field of view of the one or more cameras. Rotating the image of the surface relative to a representation of at least a portion of the field of view of the one or more cameras enhances the video communication session experience because the content associated with the surface can be viewed from a perspective different from other portions of the field of view without input from the user, which provides improved visual feedback and reduces the number of inputs required to perform an operation.
[0381] In some embodiments, the image of the surface is rotated based on the position (e.g., positioning and / or orientation) of the surface (e.g., 619) relative to a user (e.g., 622) in the field of view of the one or more cameras (e.g., the position of the user). In some embodiments, a representation of the user is displayed at a first angle and the image of the surface is rotated to a second angle different from the first angle (e.g., even if the image of the user and the image of the surface are captured at the same camera angle). Rotating the image of the surface based on the position of the surface relative to a user in the field of view of the one or more cameras enhances the video communication session experience because the content associated with the surface can be viewed from a perspective based on the position of the surface without input from the user, which provides improved visual feedback and reduces the number of inputs required to perform an operation.
[0382] In some embodiments, based on determining that the surface is in a first position of the user relative to the field of view of the one or more cameras (e.g., the surface 619 is positioned in front of the user 622 on the desk 621 in Figure 6A (e.g., a predefined position) (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by at least 45 degrees relative to the representation of the user in the field of view of the one or more cameras (e.g., the image of the surface 619 in representation 624-1 is rotated 180 degrees relative to Figure 6M representation 622-1 in). In some embodiments, the image of the surface is rotated in the range of 160 degrees to 200 degrees (e.g., rotated 180 degrees). In some embodiments, based on determining that the surface is in a first position of the user relative to the field of view of the one or more cameras (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by a first amount. In some embodiments, the first amount is in the range of 160 degrees to 200 degrees (e.g., 180 degrees). In some embodiments, based on determining that the surface is in a second position of the user relative to the field of view of the one or more cameras (e.g., on one side of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated by a second amount. In some embodiments, the second amount is in the range of 45 degrees to 120 degrees (e.g., 90 degrees). When the surface is in the first position relative to the user, rotating the image of the surface by at least 45 degrees relative to the representation of the user captured in the field of view of the one or more cameras enhances the video communication session experience by adjusting the image to provide a more natural, intuitive image without further input from the user (which provides improved visual feedback and performs an operation when a set of conditions is met without further user input).
[0383] In some embodiments, the representation of at least a portion of the field of view includes the user and is displayed simultaneously with the representation of the surface (e.g., Figure 6M representation 622-1 and 624-1 in or representation 622-2 and 624-2). In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are captured by the same camera (e.g., a single camera of the one or more cameras) and displayed simultaneously. In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are displayed in separate windows that are displayed simultaneously. Including the user in the representation of at least a portion of the field of view and displaying the representation simultaneously with the representation of the surface enhances the video communication session experience by allowing the user to view the reactions of the participants while the representation of the surface is being displayed without further input from the user (which provides improved visual feedback and performs an operation when a set of conditions is met without further user input).
[0384] In some embodiments, in response to detecting the one or more user inputs and prior to a representation of the display surface, the computer system displays a preview of the image data of the field of view of the one or more cameras (e.g., as depicted in Figures 6H to 6J )(e.g., in a preview mode of a real-time video communication interface), the preview including an image of the surface that is not modified based on the position of the surface relative to the one or more cameras (sometimes referred to as a representation of the unmodified image of the surface). In some embodiments, the preview of the field of view is displayed after a representation of the image of the display surface (e.g., in response to detecting a user input corresponding to a selection of the representation of the surface). Displaying a preview of the image of the surface that is not modified based on the position of the surface relative to the one or more cameras allows the user to quickly identify the surface within the preview since distortion correction has not yet been applied, which provides improved visual feedback.
[0385] In some embodiments, displaying a preview of the image data of the field of view of the one or more cameras includes displaying a plurality of selectable options corresponding to respective portions of the field of view of the one or more cameras (e.g., the field of view within which the surface is located) (e.g., Figure 6I 636-1 and / or 636-2 of Figure 6J 636a-i of Figure 6J ). In some embodiments, the computer system detects an input selecting one of the plurality of options corresponding to respective portions of the field of view of the one or more cameras (e.g., 612i or 612j). In response to detecting an input selecting one of the plurality of options corresponding to respective portions of the field of view of the one or more cameras and based on determining that the input selecting one of the plurality of options corresponding to respective portions of the field of view of the one or more cameras points to a first option corresponding to a first portion of the field of view of the one or more cameras, the computer system displays a representation of the surface based on the first portion of the field of view of the one or more cameras (e.g., selection of 636h of Figure 6J causes the corresponding portion to be displayed) (e.g., the computer system displays a modified version of the image of the first portion of the field of view, optionally with a first distortion correction). In response to detecting an input selecting one of the plurality of options corresponding to respective portions of the field of view of the one or more cameras and based on determining that the input selecting one of the plurality of options corresponding to respective portions of the field of view of the one or more cameras points to a second option corresponding to a second portion of the field of view of the one or more cameras, the computer system displays a representation of the surface based on the second portion of the field of view of the one or more cameras (e.g., Figure 6JThe selection of 636g therein causes the corresponding portion to be displayed (e.g., the computer system displays a modified version of an image of a second portion of the field of view, optionally with a second distortion correction different from the first distortion correction), where the second option is different from the first option. Displaying multiple selectable options corresponding to the respective portions of the field of view of the one or more cameras in the preview of the image data allows the user to identify the portions of the field of view that can be displayed as representations in the video conferencing interface, which provides improved visual feedback.
[0386] In some embodiments, displaying a preview of the image data of the field of view of the one or more cameras includes displaying a preview (e.g., Figure 6I 636-1, 636-2, and / or Figure 6J 636a-i) of multiple regions (e.g., different regions, non-overlapping regions, rectangular regions, square regions, and / or quadrants) (e.g., the one or more regions may correspond to different portions of the image data of the field of view). In some embodiments, the computer system detects user input (e.g., 612i and / or 612j) corresponding to one or more of the multiple regions. In response to detecting user input corresponding to the one or more regions and based on determining that the user input corresponding to the one or more regions corresponds to a first region among the one or more regions, the computer system displays a representation of the first region in the real-time video communication interface (e.g., as described with reference to Figures 6I to 6J )(e.g., with a distortion correction based on the first region). In response to detecting user input corresponding to the one or more regions and based on determining that the user input corresponding to the one or more regions corresponds to a second region among the one or more regions, the computer system displays a representation of the second region as a representation in the real-time video communication interface (e.g., with a distortion correction based on the second region, which is different from the distortion correction based on the first region). Displaying a representation of the first region or a representation of the second region in the real-time video communication interface enhances the video communication session experience by allowing the user to effectively manage the content displayed in the real-time video communication interface (which provides improved visual feedback and reduces the amount of input required to perform an operation).
[0387] In some embodiments, the one or more user inputs include gestures (e.g., 612d) in the field of view of the one or more cameras (e.g., body postures, hand gestures, head postures, arm postures, and / or eye gestures) (e.g., a gesture performed in the field of view of the one or more cameras that points to a physical location surface). Utilizing gestures in the field of view of the one or more cameras as input enhances the video communication session experience by allowing the user to control the displayed content without physically touching the device (which provides additional control options without cluttering the user interface).
[0388] In some embodiments, a computer system displays surface view options (e.g., 610) (e.g., icons, buttons, affordances, and / or user interaction graphical user interface objects), where the one or more user inputs include inputs directed to the surface view options (e.g., 612c and / or 612g) (e.g., a tap input on a touch-sensitive surface, a click with a mouse when the cursor is over the surface view option, or an air gesture when gaze is directed at the surface view option). In some embodiments, the surface view options are displayed in a representation of at least a portion of the field of view of the one or more cameras. Displaying the surface view options enhances the video communication session experience by allowing the user to effectively manage the content displayed in the real-time video communication interface, which provides additional control options without cluttering the user interface.
[0389] In some embodiments, the computer system detects a user input corresponding to a selection of a surface view option. In response to detecting a user input corresponding to a selection of a surface view option, the computer system displays a preview of image data of the field of view of the one or more cameras (e.g., as depicted in Figures 6H to 6J (e.g., in a preview mode of the real-time video communication interface), the preview including a plurality of portions of the field of view of the one or more cameras, the plurality of portions including at least a portion of the field of view of the one or more cameras (e.g., 636-1, 636-2, and / or Figure 6I 636a-i of Figure 6J ), where the preview includes a visual indication (e.g., text, graphics, icons, and / or colors) of an active portion (e.g., the portion of the field of view being transmitted to other participants in the real-time video communication session and / or being displayed by other participants in the real-time video communication session) of the field of view (e.g., 641-1 and / or Figure 6I 640a-f of Figure 6J ). In some embodiments, the visual indication indicates that a single portion (e.g., only one portion) of the plurality of portions of the field of view is active. In some embodiments, the visual indication indicates that two or more portions of the plurality of portions of the field of view are active. Displaying a preview of a plurality of portions of the field of view of the one or more cameras, where the preview includes a visual indication of an active portion of the field of view, enhances the video communication session experience by providing the user with feedback as to which portion of the field of view is active, which provides improved visual feedback.
[0390] In some embodiments, the computer system detects a user input corresponding to a selection of a surface view option (e.g., 612c, 612d, 614, 612g, 612i, and / or 612j). In response to detecting a user input corresponding to a selection of a surface view option, the computer system displays a preview of image data of the field of view of the one or more cameras (e.g., 674-2 and / or 674-3) (e.g., asFigures 6I to 6J as described in (e.g., in the preview mode of a real-time video communication interface), the preview includes a plurality of visually distinct selectable portions (e.g., as Figures 6I to 6J described in) that are overlaid on a representation of the field of view of the one or more cameras. Displaying a preview that includes a plurality of visually distinct selectable portions overlaid on a representation of the field of view of the one or more cameras enhances the video communication session experience by providing the user with feedback as to which portions of the field of view are selectable for display as a representation during the video communication session (which provides improved visual feedback).
[0391] In some embodiments, the surface is a vertical surface in the scene (e.g., as Figure 6J depicted in) (e.g., a wall, an easel, and / or a whiteboard) (e.g., the surface is within a predetermined angle in the direction of gravity (e.g., 5 degrees, 10 degrees, or 20 degrees)). Displaying a representation of the vertical surface that includes an image of the vertical surface modified based on the position of the vertical surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the vertical surface regardless of the position of the vertical surface relative to the camera without further input from the user (which provides improved visual feedback and reduces the amount of input required to perform the operation).
[0392] In some embodiments, the surface is a horizontal surface in the scene (e.g., 619) (e.g., a table, a floor, and / or a desk) (e.g., the surface is within a predetermined angle in a plane perpendicular to the direction of gravity (e.g., 5 degrees, 10 degrees, or 20 degrees). Displaying a representation of the horizontal surface that includes an image of the horizontal surface modified based on the position of the horizontal surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the horizontal surface regardless of the position of the horizontal surface relative to the camera without further input from the user (which provides improved visual feedback and reduces the amount of input required to perform the operation).
[0393] In some embodiments, the representation of the display surface includes a first view of the display surface (e.g., Figure 6N 624-1 in) (e.g., at a first rotation angle and / or a first zoom level). In some embodiments, when displaying the first view of the display surface, the computer system displays one or more shifted view options (e.g., 648-1 and / or 648-2) (e.g., buttons, icons, affordance representations, and / or user interaction graphical user interface objects). The computer system detects user input (e.g., 650a and / or 650b) directed to a corresponding shifted view option of the one or more shifted view options. In response to detecting user input directed to the corresponding shifted view option, the computer system displays a second view of the surface that is different from the first view of the surface (e.g.,Figure 6P 624-1 in and / or Figure 6Q 624-1 in (e.g., a second rotation angle different from the first rotation angle and / or a second zoom level different from the first zoom level) (e.g., shifting the view of the surface from a first view to a second view). Providing a shifted view option to display a second view of the surface that is currently being displayed at the first view of the surface enhances the video communication session experience by allowing the user to view the content associated with the surface from different perspectives (which provides additional control options without cluttering the user interface).
[0394] In some embodiments, displaying a first view of the surface includes displaying an image of the surface modified in a first manner (e.g., as depicted in Figure 6N ), (e.g., a first distortion correction applied), and wherein displaying a second view of the surface includes displaying an image of the surface modified in a second manner different from the first mann...
Claims
1. A method, comprising: at a first computer system in communication with a display generation component, one or more cameras, and one or more input devices: detecting, via the one or more input devices, one or more first user inputs corresponding to a request for a user interface for a display application that is configured to display a visual representation of a surface within a field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component and based on determining that a first set of one or more criteria is met: a visual representation of a first portion of the field of view of the one or more cameras; and a visual indication that indicates a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras that will be presented by a second computer system as a view of the surface.
2. The method of claim 1, wherein the visual representation of the first portion of the field of view of the one or more cameras and the visual indication of the first region of the field of view are simultaneously displayed when the first computer system does not share the second portion of the field of view of the one or more cameras with the second computer system.
3. The method according to any one of claims 1 to 2, wherein the second portion of the field of view of the one or more cameras includes an image of a surface located between the one or more cameras and a user within the field of view of the one or more cameras.
4. The method according to any one of claims 1 to 2, wherein the view of the surface to be presented by the second computer system includes an image of the surface that is modified based on the position of the surface relative to the one or more cameras.
5. The method according to any one of claims 1 to 2, wherein the first portion of the field of view of the one or more cameras includes an image of a user within the field of view of the one or more cameras.
6. The method according to any one of claims 1 to 2, further comprising: after detecting a change in position of the one or more cameras, simultaneously displaying, via the display generation component: a visual representation of a third portion of the field of view of the one or more cameras; and the visual indication, wherein the visual indication indicates a second region of the field of view of the one or more cameras, the second region being a subset of the third portion of the field of view of the one or more cameras, wherein the second region indicates a fourth portion of the field of view of the one or more cameras that will be presented by the second computer system as a view of the surface.
7. The method according to any one of claims 1 to 2, further comprising: when the one or more cameras are substantially stationary and while displaying the visual representation of the first portion of the field of view of the one or more cameras and the visual indication, detecting one or more second user inputs via the one or more input devices; and in response to detecting the one or more second user inputs and while the one or more cameras remain substantially stationary, simultaneously display via the display generation component: the visual representation of the first portion of the field of view; and the visual indication, wherein the visual indication indicates a third region of the field of view of the one or more cameras, the third region being a subset of the first portion of the field of view of the one or more cameras, wherein the third region indicates a fifth portion of the field of view that is different from the second portion, and the fifth portion will be presented by the second computer system as a view of the surface.
8. The method according to any one of claims 1 to 2, further comprising: when displaying the visual representation of the first portion of the field of view of the one or more cameras and the visual indication, detecting, via the one or more input devices, a user input pointing to a control that includes a set of options for the visual indication; and in response to detecting the user input pointing to the control: display the visual indication to indicate a fourth region of the field of view of the one or more cameras, the fourth region including a sixth portion of the field of view that is different from the second portion, and the sixth portion will be presented by the second computer system as a view of the surface.
9. The method according to claim 8, further comprising: in response to detecting the user input pointing to the control: maintain the position of a first portion of a boundary of a portion of the field of view that will be presented by the second computer system as a view of the surface; and modify the position of a second portion of the boundary of the portion of the field of view that will be presented by the second computer system as a view of the surface.
10. The method according to claim 9, wherein the first portion of the visual indication corresponds to the uppermost edge of the second portion of the field of view that will be presented by the second computer system as the view of the surface.
11. The method according to any one of claims 1 to 2, wherein the first portion of the field of view of the one or more cameras and the second portion of the field of view of the one or more cameras that will be presented by the second computer system as the view of the surface are based on image data captured by a first camera.
12. The method according to any one of claims 1 to 2, further comprising: detecting, via the one or more input devices, one or more third user inputs corresponding to a request to display the user interface of the application for displaying a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more third user inputs: simultaneously display via the display generation component according to a determination that a first set of one or more criteria is met: a visual representation of a seventh portion of the field of view of the one or more cameras; and A visual indication that indicates a fifth region of the field of view of the one or more cameras, the fifth region being a subset of a seventh portion of the field of view of the one or more cameras, wherein the fifth region indicates an eighth portion of the field of view of the one or more cameras, the eighth portion to be presented by a third computer system different from the second computer system as a view of the surface.
13. The method according to claim 12, wherein visual characteristics of the visual indication are user-configurable, and wherein the first computer system displays the visual indication that indicates the fifth region as having visual characteristics based on the visual characteristics of the visual indication used during the most recent use of the one or more cameras, to be presented by a remote computer system as a view of the surface.
14. The method according to any one of claims 1 to 2, further comprising: When displaying the visual representation of the first portion of the field of view of the one or more cameras and the visual indication, detecting, via the one or more input devices, one or more fourth user inputs corresponding to a request to modify the visual characteristics of the visual indication; In response to detecting the one or more fourth user inputs: Displaying the visual indication to indicate a sixth region of the field of view of the one or more cameras, the sixth region including a ninth portion of the field of view, the ninth portion being different from the second portion and to be presented by the second computer system as a view of the surface; When displaying the visual indication to indicate that the sixth region of the field of view of the one or more cameras including the ninth portion of the field of view is to be presented by the second computer system as a view of the surface, detecting one or more user inputs corresponding to a request to share the view of the surface; and In response to detecting the one or more user inputs corresponding to the request to share the view of the surface, sharing the ninth portion of the field of view for presentation by the second computer system.
15. The method according to any one of claims 1 to 2, further comprising: In response to detecting one or more first user inputs: Based on determining that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria is different from the first set of one or more criteria: Displaying the second portion of the field of view as a view of the surface to be presented by the second computer system, wherein the second portion of the field of view includes an image of the surface modified based on the position of the surface relative to the one or more cameras.
16. The method according to any one of claims 1 to 2, further comprising: When providing the second portion of the field of view as a view of the surface for presentation by the second computer system, displaying, via the display generation component, controls for modifying a portion of the field of view of the one or more cameras to be presented by the second computer system as a view of the surface.
17. The method according to claim 16, further comprising: Determine a region corresponding to the view of the surface at which the focus is directed, and via the display generation component, display controls for modifying the field of view of the one or more cameras for a portion of the view of the surface that will be presented by the second computer system as the view of the surface; and Determine that the focus is not directed at the region corresponding to the view of the surface, and abandon displaying the controls for modifying the field of view of the one or more cameras for the portion of the view of the surface that will be presented by the second computer system as the view of the surface.
18. The method according to claim 16, wherein the second portion of the field of view includes a first boundary, and the method further comprises: Detecting one or more fifth user inputs directed at the controls for modifying the field of view of the one or more cameras for a portion of the view of the surface that will be presented by the second computer system as the view of the surface; and In response to detecting the one or more fifth user inputs: Maintaining the position of the first boundary of the second portion of the field of view; and Modifying an amount of a portion of the field of view included in the second portion of the field of view.
19. The method according to any one of claims 1 to 2, further comprises: When the one or more cameras are substantially stationary, and while displaying the visual representation and the visual indication of the first portion of the field of view of the one or more cameras, detecting one or more sixth user inputs via the one or more input devices; and In response to detecting the one or more sixth user inputs and while the one or more cameras remain substantially stationary, simultaneously displaying via the display generation component: A visual representation of an eleventh portion of the field of view of the one or more cameras, the eleventh portion being different from the first portion of the field of view of the one or more cameras; and The visual indication, wherein the visual indication indicates a seventh region of the field of view of the one or more cameras, the seventh region being a subset of the eleventh portion of the field of view of the one or more cameras, wherein the seventh region indicates a twelfth portion of the field of view different from the second portion, the twelfth portion being a view of the surface that will be presented by the second computer system.
20. The method according to any one of claims 1 to 2, wherein displaying the visual indication comprises: Determining that a set of one or more alignment criteria is satisfied, wherein the set of one or more alignment criteria includes: an alignment criterion based on alignment between a current region of the field of view of the one or more cameras indicated by the visual indication and a specified portion of the field of view of the one or more cameras, and displaying the visual indication having a first appearance; and Determining that the alignment criteria are not satisfied, and displaying the visual indication having a second appearance different from the first appearance.
21. The method according to any one of claims 1 to 2, further comprises: When the visual indication indicates an eighth region of the field of view of the one or more cameras, a target region indication indicating a first designated region of the field of view of the one or more cameras is displayed simultaneously with a visual representation of a thirteenth portion of the field of view of the one or more cameras and the visual indication, wherein the first designated region indication is based on a determined portion of the field of view of the one or more cameras of the position of the surface in the field of view of the one or more cameras.
22. The method according to claim 21, wherein the target region indication is stationary relative to the surface.
23. The method according to claim 22, wherein the target region indication is selected based on an edge of the surface.
24. The method according to claim 22, wherein the target region indication is selected based on the position of a person in the field of view of the one or more cameras.
25. The method according to claim 21, further comprising: After detecting a change in the position of the one or more cameras, the target region indication is displayed via the display generating component, wherein the target region indication indicates a second designated region of the field of view of the one or more cameras, wherein the second designated region indicates a second determined portion of the field of view of the one or more cameras, and the second determined portion is based on the position of the surface in the field of view of the one or more cameras after the change in the position of the one or more cameras.
26. The method according to any one of claims 1 to 2, further comprising: A surface view representation of the surface in a ninth region of the field of view of the one or more cameras indicated by the visual indication is displayed simultaneously with the visual representation and the visual indication of the field of view of the one or more cameras, wherein the surface view representation includes an image of the surface captured by the one or more cameras, and the image is modified based on the position of the surface relative to the one or more cameras to correct the perspective of the surface.
27. The method according to claim 26, wherein displaying the surface view representation includes displaying the surface view representation in a visual representation of a portion of the field of view of the one or more cameras that includes a person.
28. The method according to claim 27, further comprising: After displaying the surface view representation of the surface in the ninth region of the field of view of the one or more cameras indicated by the visual indication, detecting a change in the field of view of the one or more cameras indicated by the visual indication; and In response to detecting the change in the field of view of the one or more cameras indicated by the visual indication, displaying the surface view representation, wherein the surface view representation includes the surface in the ninth region of the field of view of the one or more cameras indicated by the visual indication after the change in the field of view of the one or more cameras indicated by the visual indication.
29. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 28.
30. A first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, the first computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 28.
31. A first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: a memory; and means for performing the method according to any one of claims 1 to 28.
32. A computer program product comprising one or more programs configured to be executed by one or more processors of a first computer system in communication with a display generation component, one or more cameras, and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 28.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Acceleration-based theft detection system for portable electronic devices
US20050190059A1
Methods and apparatuses for operating a portable device based on an accelerometer
US20060017692A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1