Wide angle video conferencing
By communicating with display generation components, cameras and input devices in a computer system, a method is provided to display a real-time video communication interface and respond to user input, solving the problem of complex and time-consuming user interfaces in the prior art, achieving a faster and more efficient interface, saving device energy.
Patent Information
- Application Number
- CN202510448916.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-22
- Filing Date
- 2022-09-23
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art When managing real-time video communication sessions, the user interface is complex and time-consuming, resulting in wasting user time and device energy.
By communicating with the display generation component, the camera, and the input device in a computer system, a method is provided to display a real-time video communication interface, detect user input, and respond to input to modify content displayed on the interface.
A faster and more efficient user interface is achieved, reducing user cognitive burden, saving power from battery-driven devices, and extending battery charging intervals.
Smart Images

Figure CN120017786A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on September 23, 2022, with application number 202280037932.6 and invention title "Wide-angle Video Conferencing".
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Patent Application No. 17 / 950,868, filed September 22, 2022, entitled "WIDE ANGLE VIDEO CONFERENCE"; and priority to U.S. Patent Application No. 17 / 950,900, filed September 22, 2022, entitled "WIDE ANGLE VIDEO CONFERENCE"; and priority to U.S. Patent Application No. 17 / 950,922, filed September 22, 2022, entitled "WIDE ANGLE VIDEO CONFERENCE"; and priority to U.S. Provisional Patent Application No. 63 / 392,096, filed July 25, 2022, entitled "WIDE ANGLE VIDEO CONFERENCE"; and priority to U.S. Provisional Patent Application No. 63 / 392,096, filed June 30, 2022, entitled "WIDE ANGLE VIDEO CONFERENCE". Priority is claimed to U.S. Provisional Patent Application No. 63 / 357,605 entitled “VIDEO CONFERENCE”, filed June 5, 2022; U.S. Provisional Patent Application No. 63 / 349,134 entitled “WIDE ANGLE VIDEO CONFERENCE”, filed February 8, 2022; and U.S. Provisional Patent Application No. 63 / 307,780 entitled “WIDE ANGLE VIDEO CONFERENCE”, filed September 24, 2021; and U.S. Provisional Patent Application No. 63 / 248,137 entitled “WIDE ANGLE VIDEO CONFERENCE”, filed September 24, 2021. The contents of each of these patent applications are incorporated herein by reference in their entirety. Technical Field
[0004] This disclosure relates in general to computer user interfaces, and more specifically to techniques for managing real-time video communication sessions and / or managing digital content. Background Technology
[0005] The computer system may include hardware and / or software for displaying an interface for real-time video communication sessions. Summary of the Invention
[0006] However, some technologies used to manage real-time video communication sessions using electronic devices are often cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may involve multiple keystrokes or button presses. These technologies require more time than necessary, resulting in wasted user time and device power. This latter consideration is particularly important in battery-powered devices.
[0007] Therefore, this technology provides electronic devices with faster and more efficient methods and interfaces for managing real-time video communication sessions and / or managing digital content. Such methods and interfaces optionally complement or replace other methods for managing real-time video communication sessions and / or managing digital content. These methods and interfaces reduce the cognitive burden on users and result in more efficient human-machine interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging.
[0008] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component, one or more cameras, and one or more input devices. The method includes: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the real-time video communication interface, the one or more user inputs including user input pointing to a surface in the scene within the field of view of the one or more cameras; and, in response to detecting the one or more user inputs, displaying a representation of the surface via the display generation component, wherein the representation of the surface includes images of the surface captured by the one or more cameras, modified based on the position of the surface relative to the one or more cameras.
[0009] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the real-time video communication interface, the one or more user inputs including user input pointing to a surface in the scene within the field of view of the one or more cameras; and displaying a representation of the surface via the display generation component in response to detecting the one or more user inputs, wherein the representation of the surface includes images of the surface captured by the one or more cameras, modified based on the surface's position relative to the one or more cameras.
[0010] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the real-time video communication interface, the one or more user inputs including user input pointing to a surface in the scene within the field of view of the one or more cameras; and displaying a representation of the surface via the display generation component in response to detecting the one or more user inputs, wherein the representation of the surface includes images of the surface captured by the one or more cameras, modified based on the surface's position relative to the one or more cameras.
[0011] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of at least a portion of the field of view of the one or more cameras; detecting one or more user inputs via the one or more input devices while displaying the real-time video communication interface, the one or more user inputs including user input pointing to a surface in the scene within the field of view of the one or more cameras; and displaying a representation of the surface via the display generation component in response to detecting the one or more user inputs, wherein the representation of the surface includes images of the surface captured by the one or more cameras, modified based on the surface's position relative to the one or more cameras.
[0012] According to some embodiments, a computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: means for displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of a first portion of a scene in a field of view captured by the one or more cameras; and means for acquiring image data of the field of view of the one or more cameras via the one or more cameras while displaying the real-time video communication interface, the image data including a first gesture; and means for, in response to acquiring the image data of the field of view of the one or more cameras, to: display a representation of a second portion of the scene in the field of view of the one or more cameras via the display generation component, the representation of the second portion of the scene including visual content different from the representation of the first portion of the scene, based on determining that the first gesture meets a second set of criteria different from the first set of criteria; and to continue displaying the representation of the first portion of the scene via the display generation component based on determining that the first gesture meets a second set of criteria different from the first set of criteria.
[0013] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: displaying a real-time video communication interface for a real-time video communication session via the display generation component, the real-time video communication interface including a representation of a first portion of a scene in a field of view captured by the one or more cameras; and, while displaying the real-time video communication interface, acquiring image data of the field of view of the one or more cameras via the one or more cameras, the image data including a first gesture; and, in response to acquiring the image data of the field of view of the one or more cameras: based on determining that the first gesture meets a first set of criteria, displaying via the display generation component a representation of a second portion of the scene in the field of view of the one or more cameras, the representation of the second portion of the scene including visual content different from the representation of the first portion of the scene; and based on determining that the first gesture meets a second set of criteria different from the first set of criteria, continuing to display the representation of the first portion of the scene via the display generation component.
[0014] According to some embodiments, a method is described that is performed at a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The method includes: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0015] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component. The real-time video communication interface includes: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0016] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component. The real-time video communication interface includes: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0017] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface including a real-time video communication session of multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0018] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: means for detecting one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and means for displaying a real-time video communication interface for the real-time video communication session via the display generation component in response to detecting the one or more user inputs, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0019] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting one or more sets of user input corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the one or more sets of user input, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0020] According to some embodiments, a method is described that is performed at a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The method includes: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0021] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component. The real-time video communication interface includes: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0022] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component. The real-time video communication interface includes: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0023] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a set of one or more user inputs corresponding to a request to display a user interface including a real-time video communication session of multiple participants; and, in response to detecting the set of one or more user inputs, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0024] According to some embodiments, a computer system configured to communicate with a display generation component, one or more first cameras, and one or more input devices is described. The computer system includes: means for detecting one or more user inputs corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and means for displaying a real-time video communication interface for the real-time video communication session via the display generation component in response to detecting the one or more user inputs, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0025] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more first cameras, and one or more input devices. The one or more programs include instructions for: detecting one or more sets of user input corresponding to a request to display a user interface for a real-time video communication session including multiple participants; and, in response to detecting the one or more sets of user input, displaying a real-time video communication interface for the real-time video communication session via the display generation component, the real-time video communication interface including: a first representation of the field of view of the one or more first cameras of the first computer system; a second representation of the field of view of the one or more first cameras of the first computer system, the second representation of the field of view of the one or more first cameras of the first computer system including a representation of surfaces in a first scene within the field of view of the one or more first cameras of the first computer system; a first representation of the field of view of one or more second cameras of a second computer system; and a second representation of the field of view of the one or more second cameras of a second computer system, the second representation of the field of view of the one or more second cameras of the second computer system including a representation of surfaces in a second scene within the field of view of the one or more second cameras of the second computer system.
[0026] According to some embodiments, a method is described. The method includes: at a first computer system in communication with a first display generating component and one or more sensors; while the first computer system is in a real-time video communication session with a second computer system: displaying via the first display generating component a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting via the one or more sensors a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, display via the first display generating component a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0027] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a first display generating component and one or more sensors. The one or more programs include instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system via the first display generating component; detecting a change in the position of the first computer system via the one or more sensors while displaying the representation of the first view of the physical environment; and in response to detecting the change in the position of the first computer system, displaying a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system via the first display generating component, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0028] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a first display generating component and one or more sensors. The one or more programs include instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying, via the first display generating component, a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system; while displaying the representation of the first view of the physical environment, detecting, via the one or more sensors, a change in the position of the first computer system; and in response to detecting the change in the position of the first computer system, displaying, via the first display generating component, a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0029] According to some embodiments, a computer system configured to communicate with a first display generating component and one or more sensors is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system via the first display generating component; detecting a change in the position of the first computer system via the one or more sensors while displaying the representation of the first view of the physical environment; and in response to detecting the change in the position of the first computer system, displaying a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system via the first display generating component, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0030] According to some embodiments, a computer system configured to communicate with a first display generating component and one or more sensors is described. The computer system includes means for: when the first computer system is in a real-time video communication session with a second computer system: displaying via the first display generating component a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system; detecting via the one or more sensors a change in the position of the first computer system while displaying the first view of the physical environment; and in response to detecting the change in the position of the first computer system, display via the first display generating component a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0031] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a first display generating component and one or more sensors. The one or more programs include instructions for: when the first computer system is in a real-time video communication session with a second computer system: displaying via the first display generating component a representation of a first view of the physical environment in the field of view of one or more cameras of the second computer system; detecting via the one or more sensors a change in the position of the first computer system while displaying the representation of the first view of the physical environment; and in response to detecting the change in the position of the first computer system, display via the first display generating component a representation of a second view of the physical environment in the field of view of the one or more cameras of the second computer system, the second view being different from the first view of the physical environment in the field of view of the one or more cameras of the second computer system.
[0032] According to some embodiments, a method is described. The method includes: at a computer system in communication with a display generation component: displaying, via the display generation component, a representation of physical markers in a physical environment based on a view of the physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements of a portion of the field of view of the one or more cameras; while displaying the representation of the physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras, obtaining data including new physical markers in the physical environment; and in response to obtaining the data representing the new physical markers in the physical environment, displaying the representation of the new physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras.
[0033] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include instructions for: displaying, via the display generation component, a representation of physical markers in the physical environment based on a view of the physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements in a portion of the field of view of the one or more cameras; while displaying the representation of the physical markers without displaying the physical background of the one or more elements in that portion of the field of view of the one or more cameras, obtaining data including new physical markers in the physical environment; and in response to obtaining the data representing the new physical markers in the physical environment, displaying the representation of the new physical markers without displaying the physical background of the one or more elements in that portion of the field of view of the one or more cameras.
[0034] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include instructions for: displaying, via the display generation component, a representation of physical markers in a physical environment based on a view of the physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements of a portion of the field of view of the one or more cameras; while displaying the representation of the physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras, obtaining data including new physical markers in the physical environment; and in response to obtaining data representing new physical markers in the physical environment, displaying the representation of the new physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras.
[0035] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a representation of physical markers in the physical environment based on a view of the physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements of a portion of the field of view of the one or more cameras; while displaying the representation of the physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras, obtaining data including new physical markers in the physical environment; and in response to obtaining the data representing the new physical markers in the physical environment, displaying the representation of the new physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras.
[0036] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: means for displaying a representation of physical markers in the physical environment via the display generation component based on a view of a physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements in a portion of the field of view of the one or more cameras; means for obtaining data including new physical markers in the physical environment while displaying the representation of the physical markers without displaying the physical background of the one or more elements in the portion of the field of view of the one or more cameras; and means for displaying the representation of the new physical markers without displaying the physical background of the one or more elements in the portion of the field of view of the one or more cameras in response to obtaining the data representing the new physical markers in the physical environment.
[0037] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include instructions for: displaying, via the display generation component, a representation of physical markers in the physical environment based on a view of the physical environment in the field of view of one or more cameras, wherein: the view of the physical environment includes physical markers and a physical background, and displaying the representation of the physical markers includes displaying the representation of the physical markers without displaying the physical background of one or more elements of a portion of the field of view of the one or more cameras; while displaying the representation of the physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras, obtaining data including new physical markers in the physical environment; and in response to obtaining the data representing the new physical markers in the physical environment, displaying the representation of the new physical markers without displaying the physical background of the one or more elements of that portion of the field of view of the one or more cameras.
[0038] According to some embodiments, a method is described. The method includes, at a computer system in communication with a display generating component and one or more cameras: displaying an electronic document via the display generating component; detecting handwriting of physical marks on a physical surface included in the field of view of the one or more cameras and separate from the computer system via the one or more cameras; and, in response to detecting the handwriting of physical marks on a physical surface included in the field of view of the one or more cameras and separate from the computer system, displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras.
[0039] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicating with a display generating component and one or more cameras. The one or more programs include instructions for: displaying an electronic document via the display generating component; detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras; and, in response to detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras, displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras.
[0040] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicating with a display generating component and one or more cameras. The one or more programs include instructions for: displaying an electronic document via the display generating component; detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras; and, in response to detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras, displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras.
[0041] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying an electronic document via the display generation component; detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras; and, in response to detecting handwriting of physical marks included on a physical surface separated from the computer system and within the field of view of the one or more cameras, displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras.
[0042] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes: means for displaying an electronic document via the display generation component; means for detecting handwriting of physical markings on a physical surface included in the field of view of the one or more cameras and separate from the computer system; and means for displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras in response to detecting the handwriting of physical markings on a physical surface included in the field of view of the one or more cameras and separate from the computer system.
[0043] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more cameras. The one or more programs include instructions for: displaying an electronic document via the display generating component; detecting handwriting of physical marks on a physical surface included in the field of view of the one or more cameras and separate from the computer system via the one or more cameras; and, in response to detecting handwriting of physical marks on a physical surface included in the field of view of the one or more cameras and separate from the computer system, displaying digital text in the electronic document corresponding to the handwriting in the field of view of the one or more cameras.
[0044] According to some embodiments, a method is described that is executed at a first computer system in communication with a display generation component, one or more cameras, and one or more input devices. The method includes: detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras, based on determining that a first set of one or more criteria is satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion being rendered as a view of the surface by a second computer system.
[0045] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras, based on determining that one or more criteria are satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion of which will be rendered as a view of a surface by a second computer system.
[0046] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and, in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras, based on determining that one or more criteria are satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion being rendered as a view of a surface by a second computer system.
[0047] According to some embodiments, a first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying, via the display generation component, a visual representation of a first portion of the field of view of the one or more cameras, based on determining that a first set of one or more criteria is satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion being rendered as a view of a surface by a second computer system.
[0048] According to some embodiments, a first computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices is described. The computer system includes: means for detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and means for simultaneously displaying, in response to detecting the one or more first user inputs, a visual representation of a first portion of the field of view of the one or more cameras via the display generation component, based on determining that one or more criteria are satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion being rendered as a view of the surface by a second computer system.
[0049] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a first computer system, the first computer system communicating with a display generation component, one or more cameras, and one or more input devices. The one or more programs include instructions for: detecting one or more first user inputs via the one or more input devices corresponding to a request for a user interface of a display application, the user interface being used to display a visual representation of a surface in the field of view of the one or more cameras; and in response to detecting the one or more first user inputs: simultaneously displaying via the display generation component a visual representation of a first portion of the field of view of the one or more cameras, based on determining that one or more criteria are satisfied; and a visual indication indicating a first region of the field of view of the one or more cameras, the first region being a subset of the first portion of the field of view of the one or more cameras, wherein the first region indicates a second portion of the field of view of the one or more cameras, the second portion being rendered as a view of a surface by a second computer system.
[0050] According to some embodiments, a method is described. The method includes: at a computer system communicating with a display generation component and one or more input devices: detecting a request to use a function on the computer system via the one or more input devices; and in response to detecting the request to use the function on the computer system, displaying via the display generation component a tutorial for using the function, including a virtual demonstration of the function, comprising: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0051] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicating with a display generation component and one or more input devices. The one or more programs include instructions for: detecting a request to use a function on the computer system via the one or more input devices; and, in response to detecting the request to use the function on the computer system, displaying a tutorial via the display generation component for using the function, including a virtual demonstration of the function, comprising: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0052] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicating with a display generation component and one or more input devices. The one or more programs include instructions for: detecting a request to use a function on the computer system via the one or more input devices; and, in response to detecting the request to use the function on the computer system, displaying a tutorial, including a virtual demonstration of the function, via the display generation component, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0053] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a request to use a function on the computer system via the one or more input devices; and, in response to detecting the request to use the function on the computer system, displaying a tutorial, including a virtual demonstration of the function, via the display generation component, including: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0054] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system includes: means for detecting a request to use a function on the computer system via the one or more input devices; and means for displaying a tutorial including a virtual demonstration of the function via the display generation component in response to detecting the request to use the function on the computer system, including: means for displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and means for displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0055] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices. The one or more programs include instructions for: detecting a request to use a function on the computer system via the one or more input devices; and, in response to detecting the request to use the function on the computer system, displaying a tutorial via the display generation component for using the function, including a virtual demonstration of the function, comprising: displaying a virtual demonstration having a first appearance based on determining that an attribute of the computer system has a first value; and displaying a virtual demonstration having a second appearance different from the first appearance based on determining that an attribute of the computer system has a second value.
[0056] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0057] Therefore, faster and more efficient methods and interfaces are provided for managing real-time video communication sessions, thereby improving the effectiveness, efficiency, and user satisfaction of such devices. These methods and interfaces can complement or replace other methods used for managing real-time video communication sessions. Attached Figure Description
[0058] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.
[0059] Figure 1A This is a block diagram illustrating a portable multi-functional device with a touch-sensitive display according to some embodiments.
[0060] Figure 1B This is a block diagram illustrating exemplary components for event handling according to some implementation schemes.
[0061] Figure 2 A portable multi-functional device with a touchscreen is shown according to some embodiments.
[0062] Figure 3 This is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface according to some implementation schemes.
[0063] Figure 4A An exemplary user interface for a menu on an application on a portable multifunction device, according to some implementation schemes, is shown.
[0064] Figure 4B An exemplary user interface for a multifunctional device having a touch-sensitive surface separate from the display is shown according to some embodiments.
[0065] Figure 5A A personal electronic device according to some implementation schemes is shown.
[0066] Figure 5B This is a block diagram illustrating a personal electronic device according to some implementation schemes.
[0067] Figure 5C An exemplary diagram of a communication session between electronic devices according to some implementation schemes is shown.
[0068] Figures 6A to 6AY An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0069] Figure 7 A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is shown.
[0070] Figure 8 A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is shown.
[0071] Figures 9A to 9T An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0072] Figure 10 A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is shown.
[0073] Figures 11A to 11P An exemplary user interface for managing digital content is shown according to some implementation schemes.
[0074] Figure 12 This is a flowchart illustrating a method for managing digital content according to some implementation schemes.
[0075] Figures 13A to 13K An exemplary user interface for managing digital content is shown according to some implementation schemes.
[0076] Figure 14 This is a flowchart illustrating a method for managing digital content according to some implementation schemes.
[0077] Figure 15 A flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes is shown.
[0078] Figures 16A to 16QAn exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes.
[0079] Figure 17 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes.
[0080] Figures 18A to 18N An exemplary user interface for displaying functions on a computer system is shown according to some implementation schemes.
[0081] Figure 19 This is a flowchart illustrating a method for demonstrating functions on a computer system according to some implementation schemes. Detailed Implementation
[0082] The following description illustrates exemplary methods, parameters, etc. However, it should be understood that such description is not intended to limit the scope of this disclosure, but is provided as a description of exemplary embodiments.
[0083] There is a need for electronic devices that provide efficient methods and interfaces for managing real-time video communication sessions and / or managing digital content. For example, electronic devices are needed to improve content sharing. Such technologies can reduce the cognitive burden on users sharing content and / or managing digital content in electronic documents during real-time video communication sessions, thereby increasing productivity. Furthermore, such technologies can reduce processor power and battery power that would otherwise be wasted on redundant user input.
[0084] under, Figures 1A to 1B , Figure 2 , Figure 3 , Figures 4A to 4B and Figures 5A to 5C Description of an exemplary device is provided for performing techniques for managing real-time video communication sessions and / or managing digital content. Figures 6A to 6AY An exemplary user interface for managing real-time video communication sessions is shown. Figures 7 to 8 and Figure 15 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Figures 6A to 6AY The user interface in the document is used to display including Figures 7 to 8 and Figure 15 The process described below is the process in the middle. Figures 9A to 9T An exemplary user interface for managing real-time video communications is shown. Figure 10 This is a flowchart illustrating a method for managing real-time video communication according to some implementation schemes. Figures 9A to 9T The user interface in the document is used to display including Figure 10 The process described below is the process in the middle. Figures 11A to 11P An exemplary user interface for managing digital content is shown. Figure 12 This is a flowchart illustrating a method for managing digital content according to some implementation schemes. Figures 11A to 11P The user interface in the document is used to display including Figure 12 The process described below is the process in the middle. Figures 13A to 13K An exemplary user interface for managing digital content is shown according to some implementation schemes. Figure 14 This is a flowchart illustrating a method for managing digital content according to some implementation schemes. Figures 13A to 13K The user interface in the document is used to display including Figure 14 The process described below is the process in the middle. Figures 16A to 16O An exemplary user interface for managing real-time video communication sessions is shown according to some implementation schemes. Figure 17 This is a flowchart illustrating a method for managing real-time video communication sessions according to some implementation schemes. Figures 16A to 16Q The user interface in the document is used to display including Figure 17 The process described below is the process in the middle. Figures 18A to 18N An exemplary user interface for displaying functions on a computer system is shown according to some implementation schemes. Figure 19 This is a flowchart illustrating a method for demonstrating functions on a computer system according to some implementation schemes. Figures 18A to 18N The user interface in the document is used to display including Figure 19 The process described below is the process in the middle.
[0085] The processes described below enhance device operability and make the user-device interface more efficient through various technologies (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), including providing users with improved visual feedback, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional display controls, performing operations when a set of conditions are met without further user input, improving the efficiency of managing digital content, improving collaboration between users in real-time communication sessions, improving the real-time communication session experience, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently.
[0086] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if a condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0087] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch may be named a second touch and similarly, a second touch may be named a first touch, without departing from the scope of the various described embodiments. In some embodiments, a first touch and a second touch are two separate references to the same touch. In some embodiments, both a first touch and a second touch are touches, but they are not the same touch.
[0088] The terminology used in the description of the various embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0089] Depending on the context, the term "if" may optionally be interpreted as meaning "when," "at," or "in response to determination" or "in response to detection." Similarly, depending on the context, the phrases "if determination..." or "if detection [the stated condition or event]" may optionally be interpreted as meaning "in response to determination..." or "in response to detection [the stated condition or event]."
[0090] This document describes implementations of electronic devices, user interfaces for such devices, and associated processes for using such devices. In some implementations, the device is a portable communication device, such as a mobile phone, that also includes other functionalities such as PDA and / or music player functionality. Exemplary implementations of portable multi-functional devices include, but are not limited to, those from Apple Inc. (Cupertino, California). Devices, iPod Equipment, and Device. Optionally, other portable electronic devices may be used, such as laptops or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). In some embodiments, the electronic device is a computer system that communicates with a display generating component (e.g., via wireless or wired communication). The display generating component is configured to provide visual output, such as a display via a CRT monitor, a display via an LED monitor, or a display via image projection. In some embodiments, the display generating component is integrated with the computer system. In some embodiments, the display generating component is separate from the computer system. As used herein, “display” content includes displaying content (e.g., video data rendered or decoded by display controller 156) by transmitting data (e.g., image data or video data) to an integrated or external display generating component via a wired or wireless connection to visually generate content.
[0091] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick.
[0092] The device typically supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, phone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camcorder applications, web browsing applications, digital music player applications, and / or digital video player applications.
[0093] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or adjusted and / or varied within the respective applications. In this way, the device's common physical architecture (such as the touch-sensitive surface) optionally utilizes a user interface that is intuitive and clear to the user to support various applications.
[0094] Now let’s turn our attention to implementation schemes for portable devices with touch-sensitive displays. Figure 1A This is a block diagram illustrating a portable multi-functional device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 is sometimes referred to as a “touchscreen” for convenience, and is sometimes referred to as or called a “touch-sensitive display system.” Device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more haptic output generators 167 for generating haptic output on device 100 (e.g., generating haptic output on a touch-sensitive surface such as the touch-sensitive display system 112 of device 100 or the touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0095] As used in this specification and claims, the term "intensity" of contact on a tactile surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on a tactile surface, or to a substitute (alternative) for the force or pressure of a contact on a tactile surface. The intensity of contact has a range of values that includes at least four different values and more typically hundreds of different values (e.g., at least 256). The intensity of contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the tactile surface are optionally used to measure the force at different points on the tactile surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the tactile surface. Alternatively, the size and / or variation of the contact area detected on the touch-sensitive surface, the capacitance and / or variation of the touch-sensitive surface near the contact, and / or the resistance and / or variation of the touch-sensitive surface near the contact may optionally be used as substitutes for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether an intensity threshold (e.g., the intensity threshold is described in units corresponding to the substitute measurement) has been exceeded. In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure) has been exceeded. Using the intensity of the contact as an attribute of user input allows the user to access additional device functions that would otherwise be inaccessible to the user on a smaller device with limited physical space, which is used (e.g., on a touch-sensitive display) to display an indication and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).
[0096] As used in this specification and claims, the term "haptic output" refers to a physical displacement of the device relative to a previous part of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, which is detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with the touch-sensitive surface, which has been physically pressed (e.g., displaced) by the user's movement, does not move. For example, even when the smoothness of the tactile surface remains unchanged, the movement of the tactile surface can optionally be interpreted or sensed by the user as the "roughness" of the tactile surface. While such interpretations of touch by users will be limited by the individualized sensory perceptions of the user, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of a user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception of a typical (or ordinary) user.
[0097] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both, including one or more signal processing and / or application-specific integrated circuits.
[0098] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls other components of device 100 to access memory 102.
[0099] Peripheral interface 118 can be used to couple the device's input and output peripherals to CPU 120 and memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0100] RF (Radio Frequency) circuit 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 108 converts electrical signals into electromagnetic signals and vice versa, and communicates with communication networks and other communication devices via these electromagnetic signals. RF circuit 108 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 108 optionally communicates wirelessly with networks and other devices, such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). RF circuit 108 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via near-field communication radio components. Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution, Pure Data (EV-DO), HSPA, HSPA+, Dual-Unit HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE...). 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence with Extended Utility (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol that has not been developed as of the date of this document submission.
[0101] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and device 100. Audio circuitry 110 receives audio data from peripheral interface 118, converts the audio data into electrical signals, and transmits the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuitry 110 converts the electrical signals into audio data and transmits the audio data to peripheral interface 118 for processing. Audio data is optionally retrieved by peripheral interface 118 from and / or transmitted to memory 102 and / or RF circuitry 108. In some embodiments, audio circuitry 110 also includes a headset jack (e.g., ...). Figure 2 (212 in the text). The headset jack provides an interface between the audio circuitry 110 and a removable audio input / output peripheral device, such as an output-only headphone or a headset with both output (e.g., a single-ear or dual-ear headphone) and input (e.g., a microphone).
[0102] I / O subsystem 106 couples input / output peripherals on device 100, such as touchscreen 112 and other input control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, depth camera controller 169, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / send electrical signals to the other input control device 116. The other input control device 116 optionally includes physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click dials, etc. In some embodiments, input controller 160 is optionally coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, and pointing device such as mouse. One or more buttons (e.g., Figure 2 Optionally, 208) includes an increase / decrease button for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include a push-button (e.g., Figure 2(Ref. 206 in the original text). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication or via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some implementations, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes a hand moving a predetermined amount and / or speed in a predetermined posture, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0103] A quick press of the push button optionally disengages the touchscreen 112 from its lock or optionally initiates a process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application 11 / 322,549 (i.e., U.S. Patent No. 7,657,849), filed December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. A long press of the push button (e.g., 206) optionally powers the device 100 on or off. The function of one or more buttons is optionally user-customizable. The touchscreen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0104] The touch-sensitive display 112 provides input and output interfaces between the device and the user. The display controller 156 receives electrical signals from and / or sends electrical signals to the touchscreen 112. The touchscreen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0105] Touchscreen 112 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or haptic contact. Touchscreen 112 and display controller 156 (along with any associated modules and / or instruction set in memory 102) detect contact on touchscreen 112 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touchscreen 112. In an exemplary embodiment, the contact point between touchscreen 112 and the user corresponds to the user's finger.
[0106] Touchscreen 112 optionally employs LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies are used in other embodiments. Touchscreen 112 and display controller 156 optionally employ any of a variety of touch sensing technologies now known or to be developed hereafter, along with other proximity sensor arrays or other elements for determining one or more points of contact with touchscreen 112, to detect contact and any movement or interruption thereof. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that from Apple Inc. (Cupertino, California). and iPod The technology used.
[0107] In some embodiments of the touchscreen 112, the touch-sensitive display optionally resembles a multi-touch touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is incorporated herein by reference in its entirety. However, the touchscreen 112 displays visual output from the device 100, while the touch-sensitive touchpad does not provide visual output.
[0108] The touch-sensitive display in some embodiments of the touchscreen 112 is described in the following applications: (1) U.S. Patent Application 11 / 381,313, filed May 2, 2006, “Multipoint Touch Surface Controller”; (2) U.S. Patent Application 10 / 840,862, filed May 6, 2004, “Multipoint Touchscreen”; (3) U.S. Patent Application 10 / 903,964, filed July 30, 2004, “Gestures For Touch Sensitive Input Devices”; (4) U.S. Patent Application 11 / 048,264, filed January 31, 2005, “Gestures For Touch Sensitive Input Devices”; and (5) U.S. Patent Application 11 / 038,590, filed January 18, 2005, “Mode-Based Graphical User Interfaces For Touch Sensitive Input”. (6) U.S. Patent Application No. 11 / 228,758, filed September 16, 2005, “Virtual Input Device Placement On A Touch Screen User Interface”; (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”; and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, “Multi-Functional Hand-Held Device”. The full text of all these applications is incorporated herein by reference.
[0109] Touchscreen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. Users optionally use any suitable object or accessory such as a stylus, finger, etc., to interact with touchscreen 112. In some embodiments, the user interface is designed to operate primarily through finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor locations or commands to perform the user-desired actions.
[0110] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. Optionally, the touchpad is a touch-sensitive surface separate from the touchscreen 112, or an extension of the touch-sensitive surface formed by the touchscreen.
[0111] The device 100 also includes a power system 162 for supplying power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.
[0112] The device 100 may optionally also include one or more optical sensors 164. Figure 1AAn optical sensor 164 is shown coupled to an optical sensor controller 158 in the I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with an imaging module 143 (also called a camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite to a touchscreen display 112 on the front of the device, allowing the touchscreen display to be used as a viewfinder for still image and / or video image acquisition. In some embodiments, the optical sensor is located on the front of the device, allowing images of the user to be optionally acquired for video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the location of the optical sensor 164 can be changed by the user (e.g., by rotating the lenses and sensor within the device housing), allowing a single optical sensor 164 to be used with the touchscreen display for both video conferencing and still image and / or video image acquisition.
[0113] The device 100 optionally also includes one or more depth camera sensors 175. Figure 1A A depth camera sensor is shown coupled to a depth camera controller 169 in I / O subsystem 106. Depth camera sensor 175 receives data from the environment to create a 3D model of an object (e.g., a face) within the scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with imaging module 143 (also referred to as a camera module), depth camera sensor 175 may optionally be used to determine depth maps of different portions of an image captured by imaging module 143. In some embodiments, the depth camera sensor is located at the front of device 100, such that user images with depth information are optionally acquired for video conferencing while the user views other video conferencing participants on a touchscreen display, and selfies with depth map data are captured. In some embodiments, depth camera sensor 175 is located at the rear of the device, or both the rear and front of device 100. In some embodiments, the position of depth camera sensor 175 may be changed by the user (e.g., by rotating a lens and sensor within the device housing), such that depth camera sensor 175 is used in conjunction with a touchscreen display for both video conferencing and still image and / or video image acquisition.
[0114] In some implementations, the depth map (e.g., a depth map image) contains information (e.g., values) relating to the distance of objects in the scene from the viewpoint (e.g., a camera, optical sensor, depth camera sensor). In one implementation of the depth map, each depth pixel defines the location of its corresponding two-dimensional pixel on the Z-axis of the viewpoint. In some implementations, the depth map consists of pixels, where each pixel is defined by a value (e.g., 0 to 255). For example, a "0" value represents the pixel furthest from the viewpoint (e.g., a camera, optical sensor, depth camera sensor) in the "3D" scene, and a "255" value represents the pixel closest to the viewpoint in the "3D" scene. In other implementations, the depth map represents the distance between objects in the scene and the plane of the viewpoint. In some implementations, the depth map includes information about the relative depth of various features of the object of interest within the field of view of the depth camera (e.g., the relative depth of the eyes, nose, mouth, and ears of a user's face). In some implementations, the depth map includes information that enables the device to determine the contour of the object of interest in the z-direction.
[0115] The device 100 may optionally also include one or more contact strength sensors 165. Figure 1A A contact strength sensor is shown coupled to a strength sensor controller 159 in I / O subsystem 106. The contact strength sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact strength sensor 165 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact strength sensor is located on the rear of device 100, opposite to the touchscreen display 112 located on the front of device 100.
[0116] The device 100 optionally also includes one or more proximity sensors 166. Figure 1AA proximity sensor 166 coupled to a peripheral device interface 118 is shown. Alternatively, the proximity sensor 166 may optionally be coupled to an input controller 160 in an I / O subsystem 106. The proximity sensor 166 may optionally perform as described in the following U.S. patent applications: No. 11 / 241,839, entitled "Proximity Detector In Handheld Device"; No. 11 / 240,788, entitled "Proximity Detector In Handheld Device"; No. 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; No. 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and No. 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals", the entire contents of which are incorporated herein by reference. In some implementations, the proximity sensor is turned off and the touchscreen 112 is disabled when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).
[0117] The device 100 may optionally also include one or more tactile output generators 167. Figure 1AA haptic output generator coupled to a haptic feedback controller 161 in I / O subsystem 106 is shown. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting electrical signals into haptic outputs on the device). A contact intensity sensor 165 receives haptic feedback generation instructions from a haptic feedback module 133 and generates a haptic output on device 100 that can be felt by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a haptic surface (e.g., haptic display system 112) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 100) or laterally (e.g., backward and forward in the same plane as the surface of device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.
[0118] The device 100 may optionally also include one or more accelerometers 168. Figure 1A An accelerometer 168 coupled to a peripheral device interface 118 is shown. Alternatively, the accelerometer 168 may be coupled to an input controller 160 in an I / O subsystem 106. The accelerometer 168 may optionally perform as described in the following U.S. patent publications: U.S. Patent Publication No. 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication No. 20060017692, entitled "Methods and Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from one or more accelerometers. Device 100 may optionally include, in addition to the accelerometer 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for acquiring information about the location and orientation (e.g., portrait or landscape) of device 100.
[0119] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application program (or instruction set) 136. Furthermore, in some embodiments, memory 102 ( Figure 1A ) or 370 ( Figure 3 Storage device / global internal state 157, such as Figure 1A and Figure 3 As shown in the figure. Device / global internal state 157 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, indicating what applications, views or other information occupy various areas of the touch screen display 112; sensor state, including information obtained from the device's various sensors and input control devices 116; and position information relating to the device's position and / or orientation.
[0120] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.
[0121] The communication module 128 facilitates communication with other devices via one or more external ports 124 and includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is connected to… (Trademark of Apple Inc.) The same or similar and / or compatible multi-pin (e.g., 30-pin) connectors used in Apple Inc. devices.
[0122] The contact / motion module 130 optionally detects contact with the touchscreen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., touchpads or physical click-based rotary dials). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether a contact has occurred (e.g., detecting a finger press event), determining the contact intensity (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or a contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple-finger contact). In some implementations, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.
[0123] In some implementations, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by a user (e.g., determining whether the user has “clicked” an icon). In some implementations, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a specific physical actuator and can be adjusted without changing the physical hardware of device 100). For example, the mouse “click” threshold of a touchpad or touchscreen can be set to any threshold in a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some specific implementations, the user of the device is provided with software settings for adjusting one or more intensity thresholds in a set (e.g., by adjusting the individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using system-level clicks on the “intensity” parameter).
[0124] The touch / motion module 130 optionally detects gesture input performed by the user. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event.
[0125] The graphics module 132 includes various known software components for rendering and displaying graphics on the touchscreen 112 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0126] In some implementations, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes from an application or the like to specify the graphic to be displayed, and, if necessary, also receives coordinate data and other graphic attribute data, and then generates screen image data for output to the display controller 156.
[0127] The haptic feedback module 133 includes various software components for generating instructions which are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.
[0128] Optionally, the text input module 134, a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).
[0129] GPS module 135 determines the location of the device and provides that information for use in various applications (e.g., to phone 138 for use in location-based dialing; to camera 143 as image / video metadata; and to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).
[0130] Application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:
[0131] • Contacts module 137 (sometimes called address book or contact list);
[0132] • Telephone module 138;
[0133] • Video conferencing module 139;
[0134] • Email client module 140;
[0135] • Instant Messaging (IM) module 141;
[0136] Fitness support module 142;
[0137] • Camera module 143 for still images and / or video images;
[0138] • Image management module 144;
[0139] • Video player module;
[0140] Music player module;
[0141] • Browser module 147;
[0142] • Calendar module 148;
[0143] • Widget module 149, which optionally includes one or more of the following: weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets acquired by the user, and user-created widgets 149-6;
[0144] • Widget creator module 150 for creating user-created widgets 149-6;
[0145] • Search module 151;
[0146] • Video and music player module 152, which combines a video player module and a music player module;
[0147] • Notes module 153;
[0148] • Map module 154; and / or
[0149] • Online video module 155.
[0150] Examples of other applications 136 that may be optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital rights management, speech recognition, and speech duplication.
[0151] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the contact module 137 is optionally used to manage an address book or contact list (e.g., in the application internal state 192 of the contact module 137 stored in memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0152] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, telephone module 138 is optionally used to input character sequences corresponding to telephone numbers, access one or more telephone numbers in contact module 137, modify entered telephone numbers, dial corresponding telephone numbers, initiate conversations, and disconnect or hang up when a conversation is completed. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0153] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contact module 137, and telephone module 138, video conferencing module 139 includes executable instructions to initiate, conduct, and terminate video conferences between the user and one or more other participants based on user instructions.
[0154] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands. Combined with image management module 144, email client module 140 makes it very easy to create and send emails containing still images or video images captured by camera module 143.
[0155] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for: inputting a character sequence corresponding to an instant message, modifying previously input characters, transmitting a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving an instant message, and viewing a received instant message. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photographs, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Services (EMS). As used herein, "instant message" refers to both telephone-based messages (e.g., messages sent using SMS or MMS) and internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0156] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, fitness support module 142 includes executable instructions for creating fitness activities (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (executive devices); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness activities; and displaying, storing, and transmitting fitness data.
[0157] In conjunction with the touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, the camera module 143 includes executable instructions for: capturing still images or videos (including video streams) and storing them in memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from memory 102.
[0158] Incorporating the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0159] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as links to attachments and other files on web pages.
[0160] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions to create, display, modify, and store calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.
[0161] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is optionally a microapplication downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or a user-created microapplication (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).
[0162] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 can optionally be used by the user to create widgets (e.g., to convert user-specified portions of a webpage into widgets).
[0163] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.
[0164] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions allowing users to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0165] Incorporating the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the note-taking module 153 includes executable instructions for creating and managing notes, to-do items, etc., according to user instructions.
[0166] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and map-related data (e.g., driving directions, data related to shops and other points of interest at or near a specific location, and other location-based data) according to user instructions.
[0167] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, the online video module 155 includes instructions for performing the following actions: allowing users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen or on an external display connected via external port 124), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 141 is used instead of the email client module 140 to send links to specific online videos. Further descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are incorporated herein by reference in their entirety.
[0168] Each of the above modules and applications corresponds to an executable set of instructions for performing one or more of the functions described above and the methods described in this patent application (e.g., computer-implemented methods and other information processing methods as described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module (e.g., Figure 1A (e.g., video and music player module 152). In some embodiments, memory 102 optionally stores subgroups of the aforementioned modules and data structures. Additionally, memory 102 optionally stores other modules and data structures not described above.
[0169] In some implementations, device 100 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push-buttons, dials, etc.) on device 100 can be optionally reduced.
[0170] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 from any user interface displayed on device 100 to the main menu, main desktop menu, or root menu. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push-button or other physical input control device, rather than a touchpad.
[0171] Figure 1B This is a block diagram illustrating exemplary components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 This includes an event classifier 170 (e.g., in operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 137 to 151, 155, 380 to 390).
[0172] Event classifier 170 receives event information and determines the application 136-1 and application view 191 of application 136-1 to which the event information should be delivered. Event classifier 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information should be delivered.
[0173] In some implementations, the application internal state 192 includes additional information such as one or more of the following: recovery information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 136-1, a state queue for enabling the user to return to the previous state or view of the application 136-1, and a repeat / undo queue for the user's previous actions.
[0174] Event monitor 171 receives event information from peripheral device interface 118. The event information includes information about sub-events (e.g., user touches on touch-sensitive display 112 as part of a multi-touch gesture). Peripheral device interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via audio circuitry 110). The information received by peripheral device interface 118 from I / O subsystem 106 includes information from touch-sensitive display 112 or touch-sensitive surfaces.
[0175] In some implementations, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other implementations, peripheral device interface 118 transmits event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or receiving input for a predetermined duration).
[0176] In some implementations, the event classifier 170 also includes a hit view determination module 172 and / or an activity event recognizer determination module 173.
[0177] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.
[0178] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a procedural level within the application's procedural or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events identified as correct input is optionally determined at least in part based on the hit view of the initial touch that initiates a touch-based gesture.
[0179] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.
[0180] The activity event recognizer determination module 173 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 173 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 173 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, higher views in the hierarchy will still remain actively participating views.
[0181] Event assigner module 174 assigns event information to event identifiers (e.g., event identifier 180). In embodiments that include active event identifier determination module 173, event assigner module 174 delivers event information to the event identifier determined by active event identifier determination module 173. In some embodiments, event assigner module 174 stores event information in an event queue, which is retrieved by the corresponding event receiver 182.
[0182] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or part of another module (such as contact / motion module 130) stored in memory 102.
[0183] In some implementations, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other implementations, one or more of the event recognizers 180 are part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some implementations, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handlers 190 optionally utilize or invoke the data updater 176, the object updater 177, or the GUI updater 178 to update the application's internal state 192. Alternatively, one or more application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0184] The corresponding event recognizer 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies the event based on the event information. The event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event recognizer 180 also includes at least one subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0185] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events, such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a longitudinal orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device orientation).
[0186] Event comparator 184 compares event information with predefined event or sub-event definitions and determines the event or sub-event based on the comparison, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (187-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off of a predetermined duration (touch end). In another example, event 2 (187-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) on the displayed object for a predetermined duration, movement of the touch on the touch-sensitive display 112, and lifting off the touch (end of touch). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0187] In some implementations, event definition 187 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub-event and the object that triggered the hit test.
[0188] In some implementations, the definition of the corresponding event (187) also includes a delay action that delays the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.
[0189] When the corresponding event recognizer 180 determines that the sub-event sequence does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.
[0190] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing how or how event recognizers can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0191] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some implementations, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and delaying) the sub-event to the corresponding hit view. In some implementations, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with that flag retrieves the flag and executes a predefined process.
[0192] In some implementations, event delivery instruction 188 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and executes a predetermined process.
[0193] In some implementations, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video player module. In some implementations, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates parts of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on touch-sensitive display.
[0194] In some implementations, event handler 190 includes, or has access to, a data updater 176, an object updater 177, and a GUI updater 178. In some implementations, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other implementations, they are included in two or more software modules.
[0195] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 100 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.
[0196] Figure 2A portable multifunction device 100 with a touchscreen 112 is shown according to some embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 100. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.
[0197] Device 100 optionally also includes one or more physical buttons, such as a "main desktop" or menu button 204. As previously described, menu button 204 is optionally used to navigate to any application 136 of a set of applications optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touchscreen 112.
[0198] In some embodiments, device 100 includes a touchscreen 112, a menu button 204, a push-button 206 for powering on / off and locking the device, one or more volume control buttons 208, a SIM card slot 210, a headset jack 212, and a docking / charging external port 124. The push-button 206 is optionally used to power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In another embodiment, device 100 also accepts voice input via microphone 113 for activating or deactivating certain functions. Device 100 also optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on the touchscreen 112, and / or one or more haptic output generators 167 for generating haptic outputs for a user of device 100.
[0199] Figure 3This is a block diagram of an exemplary multi-functional device with a display and a touch-sensitive surface according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 300 includes an input / output (I / O) interface 330 with a display 340, which is typically a touchscreen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, and a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to the reference above). Figure 1A The tactile output generator 167 and sensor 359 (e.g., optical sensor, accelerometer, proximity sensor, touch sensor and / or contact intensity sensor (similar to the one mentioned above)) are described. Figure 1A The contact strength sensor 165 is described above. The memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. The memory 370 optionally includes one or more storage devices located remotely from the CPU 310. In some embodiments, the memory 370 stores information related to the portable multifunction device 100. Figure 1A The memory 370 stores programs, modules, and data structures similar to those in the memory 102 of the portable multifunction device 100, or subsets thereof. Additionally, the memory 370 optionally stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunction device 100. For example, the memory 370 of the device 300 optionally stores a drawing module 380, a rendering module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while the portable multifunction device 100 (… Figure 1A The memory 102 may optionally not store these modules.
[0200] Figure 3Each of the elements described above is optionally stored in one or more memory devices of the previously mentioned memory devices. Each of the modules described above corresponds to a set of instructions for performing the functions described above. The modules or computer programs described above (e.g., instruction sets or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subgroup of the modules and data structures described above. In addition, memory 370 optionally stores additional modules and data structures not described above.
[0201] Now let’s turn our attention to the implementation of the user interface, which is optionally implemented on, for example, a portable multifunction device 100.
[0202] Figure 4A An exemplary user interface for an application menu on a portable multifunction device 100 according to some embodiments is shown. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements or a subset or superset thereof:
[0203] • Signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0204] • Time 404;
[0205] Bluetooth indicator 405;
[0206] • Battery status indicator 406;
[0207] • Tray tray 408 with icons for commonly used applications, such as:
[0208] ○ The telephone module 138 has an icon 416 labeled "telephone", which optionally includes an indicator 414 indicating the number of missed calls or voicemail messages;
[0209] ○ An icon 418 labeled "Mail" in the email client module 140, which optionally includes an indicator 410 for the number of unread emails;
[0210] ○ The icon 420 labeled "Browser" in browser module 147; and
[0211] ○ The icon 422 marked "iPod" for the video and music player module 152 (also known as the iPod module 152, a trademark of Apple Inc.); and
[0212] • Icons of other applications, such as:
[0213] ○The icon 424 of the IM module 141 marked as "Message";
[0214] ○The calendar module 148 has an icon 426 labeled "Calendar";
[0215] ○ The icon 428 of the image management module 144 is labeled "Photo".
[0216] ○ The icon 430 of camera module 143, which is labeled "camera";
[0217] ○ The icon 432 of the online video module 155, which is labeled "Online Video";
[0218] ○ The icon 434 labeled "Stock Market" in the Stock Market widget 149-2;
[0219] ○The icon 436 of the map module 154 that is labeled "map";
[0220] ○The weather widget 149-1 has icon 438 labeled "weather";
[0221] ○ The alarm clock widget 149-4 has an icon 440 labeled "clock";
[0222] ○ The icon 442 of the fitness support module 142 is labeled "fitness support";
[0223] ○ The icon 444 labeled "Notes" in Notes module 153; and
[0224] ○ The icon 446, labeled "Settings," is used to set the settings of the device 100 and its various applications 136.
[0225] It should be pointed out that, Figure 4A The icon labels shown are merely exemplary. For example, icon 422 of video and music player module 152 is labeled "Music" or "Music Player". Other labels may be optionally used for various application icons. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to that particular application icon.
[0226] Figure 4B A touch-sensitive surface 451 (e.g., separate from the display 450 (e.g., touchscreen display 112)) is shown. Figure 3 Devices such as tablets or touchpads (e.g., 355) Figure 3An exemplary user interface on the device 300. The device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of the sensors 359) for detecting the intensity of contact on the tactile surface 451 and / or one or more tactile output generators 357 for generating tactile outputs for the user of the device 300.
[0227] While some examples of input on a reference touchscreen display 112 (which combines a touch-sensitive surface and a display) are given below, in some implementations the device detects input on a touch-sensitive surface separate from the display, such as... Figure 4B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a spindle (e.g., on the display (e.g., 450)). Figure 4B The spindle corresponding to 453 in the middle (e.g., Figure 4B (452 in the example). According to these embodiments, the device detects the position corresponding to the corresponding position on the display (e.g., in the example). Figure 4B In the middle, 460 corresponds to 468 and 462 corresponds to 470) is in contact with the touch-sensitive surface 451 (e.g., Figure 4B (460 and 462 in the text). Thus, when the touch-sensitive surface (e.g., ...) Figure 4B 451) and the display of a multi-functional device (e.g., Figure 4B When 450 is separated from the touch-sensitive surface, user input detected by the device on that touch-sensitive surface (e.g., touches 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may optionally be used for other user interfaces described herein.
[0228] Additionally, while the examples below are primarily given with reference to finger input (e.g., finger touch, single-finger tap, finger swipe), it should be understood that in some implementations, one or more of these finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the swipe path (e.g., instead of movement of the touch). Similarly, a tap gesture may optionally be replaced by a mouse click while the cursor is over the location of the tap gesture (e.g., instead of detection of touch, followed by cessation of touch detection). Likewise, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.
[0229] Figure 5AAn exemplary personal electronic device 500 is illustrated. Device 500 includes a body 502. In some embodiments, device 500 may include components relative to devices 100 and 300 (e.g., Figures 1A to 4B Some or all of the features described herein. In some embodiments, device 500 has a touch-sensitive display 504, referred to below as touchscreen 504. As an alternative to or complement to touchscreen 504, device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, touchscreen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., touch). The one or more intensity sensors of touchscreen 504 (or touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of device 500 can respond to touches based on the intensity of the touch, meaning that touches of different intensities can invoke different user interface operations on device 500.
[0230] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is incorporated herein by reference in its entirety.
[0231] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) may be physical. Examples of physical input mechanisms include push-buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch straps, bangles, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow a user to wear device 500.
[0232] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include a reference retrieval system. Figure 1A , Figure 1B and Figure 3 Some or all of the components described herein. Device 500 has a bus 512 that operatively couples I / O portion 514 to one or more computer processors 516 and memory 518. I / O portion 514 may be connected to display 504, which may have touch-sensitive components 522 and optionally have an intensity sensor 524 (e.g., a contact intensity sensor). Furthermore, I / O portion 514 may be connected to communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 may optionally be a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 508 may optionally be a button.
[0233] In some examples, the input mechanism 508 is optionally a microphone. The personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.
[0234] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform, for example, the techniques described below, including processes 700, 800, 1000, 1200, 1400, 1500, 1700, and 1900. Figures 7 to 8 , Figure 10 , Figure 12 , Figure 14 , Figure 15 , Figure 17 and Figure 19 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state storage such as flash memory, solid-state drives, etc. Personal electronic devices are not limited to... Figure 5B It can be the components and configurations, or it can include other components or additional components in a variety of configurations.
[0235] As used herein, the term "power indication" refers to an indicator optionally displayed on devices 100, 300, and / or 500 ( Figure 1A , Figure 3 and Figures 5A to 5C A user-interactive graphical user interface object on a display screen. For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) optionally each constitute a functional representation.
[0236] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some specific implementations that include a cursor or other positional marker, the cursor acts as a "focus selector," such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the cursor is positioned on a touch-sensitive surface (e.g., a...). Figure 3 The touchpad 355 or Figure 4B When an input (e.g., a press input) is detected on the touch-sensitive surface 451 of the display, the specific user interface element is adjusted according to the detected input. This applies to touchscreen displays (e.g., those capable of enabling direct interaction with user interface elements on the touchscreen display) Figure 1A The touch-sensitive display system 112 or Figure 4AIn some embodiments of the touchscreen 112, a touch detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by touch) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, that particular user interface element is adjusted according to the detected input. In some embodiments, focus moves from one area of the user interface to another without corresponding movement of the cursor or movement of a touch on the touchscreen display (e.g., moving focus from one button to another using tab keys or arrow keys); in these embodiments, the focus selector moves according to the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user-controlled user interface element (or a touch on the touchscreen display) that delivers the user-expected interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of the focus selector (e.g., a cursor, touch, or selection box) above the corresponding button will indicate to the user that they expect to activate the corresponding button (rather than other user interface elements shown on the device's display).
[0237] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted away, before or after contact begins to move, before contact ends, before or after contact intensity is detected to increase and / or before or after contact intensity decreases). The characteristic intensity of the contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half maximum value of the contact intensity, the 90% maximum value of the contact intensity, etc. In some embodiments, the duration of the contact is used when determining the characteristic intensity (e.g., when the characteristic intensity is the average value of the contact intensity over time). In some implementations, the feature intensity is compared to a set of one or more intensity thresholds to determine whether the user has performed an action. For example, the set of one or more intensity thresholds may optionally include a first intensity threshold and a second intensity threshold. In this example, contact with a feature intensity not exceeding the first threshold results in a first action, contact with a feature intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a feature intensity exceeding the second threshold results in a third action. In some implementations, a comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more actions (e.g., whether to perform the corresponding action or abandon performing the corresponding action) rather than to determine whether to perform the first action or the second action.
[0238] Figure 5C An exemplary diagram depicts a communication session between electronic devices 500A, 500B, and 500C. Devices 500A, 500B, and 500C are similar to electronic device 500, and each device shares one or more data connections 510 (such as an internet connection, Wi-Fi connection, cellular connection, short-range communication connection, and / or any other such data connection or network) to facilitate real-time communication of audio and / or video data between the respective devices for a sustained period of time. In some embodiments, the exemplary communication session may include a shared data session, whereby data is transmitted from one or more electronic devices to other electronic devices to enable simultaneous output of corresponding content at the electronic devices. In some embodiments, the exemplary communication session may include a video conferencing session, whereby audio and / or video data are transmitted between devices 500A, 500B, and 500C, enabling users of the respective devices to communicate in real time using the electronic devices.
[0239] exist Figure 5C In this design, device 500A represents an electronic device associated with user A. Device 500A (via data connection 510) communicates with devices 500B and 500C, which are associated with users B and C, respectively. Device 500A includes a camera 501A for capturing video data of the communication session, and a display 504A (e.g., a touchscreen) for displaying content associated with the communication session. Device 500A also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0240] Device 500A displays a communication UI 520A via a display 504A. This communication UI is a user interface used to facilitate communication sessions (e.g., video conferencing sessions) between device 500B and device 500C. The communication UI 520A includes video feeds 525-1A and 525-2A. Video feed 525-1A is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during a communication session. Video feed 525-2A is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during a communication session.
[0241] The communication UI 520A includes a camera preview 550A, which is a representation of video data captured by camera 501A at device 500A. Camera preview 550A indicates to user A the expected video feed to be displayed at corresponding devices 500B and 500C.
[0242] The communication UI 520A includes one or more controls 555A for controlling one or more aspects of a communication session. For example, controls 555A may include controls for muting audio in the communication session, changing the camera view of the communication session (e.g., changing the camera used to capture video of the communication session, adjusting zoom values), terminating the communication session, applying visual effects to the camera view of the communication session, and activating one or more modes associated with the communication session. In some embodiments, one or more controls 555A are optionally displayed in the communication UI 520A. In some embodiments, one or more controls 555A are displayed separately from the camera preview 550A. In some embodiments, one or more controls 555A are displayed to cover at least a portion of the camera preview 550A.
[0243] exist Figure 5CIn this context, device 500B represents an electronic device associated with user B, who communicates with devices 500A and 500C (via data connection 510). Device 500B includes a camera 501B for capturing video data of the communication session, and a display 504B (e.g., a touchscreen) for displaying content associated with the communication session. Device 500B also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0244] Device 500B displays a communication UI 520B similar to the communication UI 520A of device 500A via touchscreen 504B. Communication UI 520B includes video feeds 525-1B and 525-2B. Video feed 525-1B is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during a communication session. Video feed 525-2B is a representation of video data captured at device 500C (e.g., using camera 501C) and transmitted from device 500C to devices 500A and 500B during a communication session. Communication UI 520B also includes: a camera preview 550B, which is a representation of video data captured at device 500B via camera 501B; and one or more controls 555B similar to controls 555A, which are used to control one or more aspects of the communication session. Camera preview 550B indicates to user B the expected video feed that user B will see displayed at the corresponding devices 500A and 500C.
[0245] exist Figure 5C In this context, device 500C represents an electronic device associated with user C, who communicates with devices 500A and 500B (via data connection 510). Device 500C includes a camera 501C for capturing video data of the communication session, and a display 504C (e.g., a touchscreen) for displaying content associated with the communication session. Device 500C also includes other components such as a microphone (e.g., 113) for recording audio of the communication session and a speaker (e.g., 111) for outputting the audio of the communication session.
[0246] Device 500C displays a communication UI 520C, similar to the communication UI 520A of device 500A and the communication UI 520B of device 500B, via a touchscreen 504C. The communication UI 520C includes video feeds 525-1C and 525-2C. Video feed 525-1C is a representation of video data captured at device 500B (e.g., using camera 501B) and transmitted from device 500B to devices 500A and 500C during a communication session. Video feed 525-2C is a representation of video data captured at device 500A (e.g., using camera 501A) and transmitted from device 500A to devices 500B and 500C during a communication session. The communication UI 520C also includes: a camera preview 550C, which is a representation of video data captured at device 500C via camera 501C; and one or more controls 555C similar to controls 555A and 555B, which are used to control one or more aspects of the communication session. The camera preview 550C indicates to user C the expected video feed to be displayed at the corresponding devices 500A and 500B.
[0247] Although Figure 5C The diagram depicts a communication session between three electronic devices, but such a session can be established between two or more electronic devices, and the number of devices participating in the session can change as electronic devices join or leave. For example, if one electronic device leaves the communication session, audio and video data from the device that has stopped participating are no longer represented on the participating devices. For instance, if device 500B stops participating in the communication session, there is no data connection 510 between devices 500A and 500C, and there is no data connection 510 between devices 500C and 500B. Furthermore, device 500A does not include video feed 525-1A, and device 500C does not include video feed 525-1C. Similarly, if a device joins the communication session, a connection is established between the joining device and the existing devices, and video and audio data are shared among all devices, enabling each device to output data transmitted from other devices.
[0248] Figure 5C The implementation scheme depicted in the diagram represents a communication session between multiple electronic devices, including Figures 6A to 6AY , Figures 9A to 9T , Figures 11A to 11P , Figures 13A to 13K and Figures 16A to 16Q The exemplary communication session is depicted in the image. In some implementations, Figures 6A to 6AY , Figures 9A to 9T , Figures 13A to 13K and Figures 16A to 16QThe communication session depicted in the figure includes two or more electronic devices, even if other electronic devices participating in the communication session are not depicted in the figure.
[0249] Now let’s turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as portable multifunction devices 100, 300 or 500).
[0250] Figures 6A to 6AY Exemplary user interfaces for managing real-time video communication sessions are shown according to some implementation schemes. The user interfaces in these figures are used to illustrate including... Figures 7 to 8 and Figure 15 The process described below is the process in the middle.
[0251] Figures 6A to 6AY An exemplary user interface for managing a real-time video communication session is shown from the perspective of different users (e.g., users participating in a real-time video communication session from different devices, different types of devices, devices with different applications installed, and / or devices with different operating system software).
[0252] refer to Figure 6A Device 600-1 corresponds to user 622 (e.g., "John"), who in some embodiments is a participant in a real-time video communication session. Device 600-1 includes a display (e.g., a touch-sensitive display) 601 and a camera 602 (e.g., a front-facing camera) having a field of view 620. In some embodiments, camera 602 is configured to capture image data and / or depth data of the physical environment within the field of view 620. The field of view 620 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view. In some embodiments, camera 602 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600-1 may include multiple cameras. Therefore, while device 600-1 is described herein for using camera 602 to capture image data during a real-time video communication session, it should be understood that device 600-1 may use multiple cameras to capture image data.
[0253] refer to Figure 6ADevice 600-2 corresponds to user 623 (e.g., “Jane”), who in some embodiments is a participant in a real-time video communication session. Device 600-2 includes a display (e.g., a touch-sensitive display) 683 and a camera 682 (e.g., a front-facing camera) having a field of view 688. In some embodiments, camera 682 is configured to capture image data and / or depth data of the physical environment within the field of view 688. The field of view 688 is sometimes referred to herein as the available field of view, the entire field of view, or the camera field of view. In some embodiments, camera 682 is a wide-angle camera (e.g., a camera including a wide-angle lens or a lens with a relatively short focal length and a wide field of view). In some embodiments, device 600-2 may include multiple cameras. Therefore, while device 600-2 for capturing image data using camera 682 during a real-time video communication session is described herein, it should be understood that device 600-2 may use multiple cameras to capture image data.
[0254] As shown, user 622 (“John”) is positioned (e.g., seated) in front of desk 621 (and device 600-1) in environment 615. In some examples, user 622 is positioned in front of desk 621 such that user 622 is captured within the field of view 620 of camera 602. In some embodiments, one or more objects proximate to user 622 are positioned such that these objects are captured within the field of view 620 of camera 602. In some embodiments, both user 622 and objects proximate to user 622 are captured simultaneously within the field of view 620. For example, as shown, drawing 618 is positioned on surface 619 in front of user 622 (relative to camera 602) such that both user 622 and drawing 618 are captured within the field of view 620 of camera 602 and displayed in representations 622-1 (displayed by device 600-1) and 622-2 (displayed by device 600-2).
[0255] Similarly, user 623 (“Jane”) is positioned (e.g., seated) in front of desk 686 (and device 600-2) in environment 685. In some examples, user 623 is positioned in front of desk 686 such that user 623 is captured within the field of view 688 of camera 682. As shown, user 623 is displayed in representation 623-1 (displayed by device 600-1) and representation 623-2 (displayed by device 600-2).
[0256] Generally, during operation, devices 600-1 and 600-2 capture image data, which is then exchanged between devices 600-1 and 600-2 and used by devices 600-1 and 600-2 to display various representations of content during a real-time video communication session. While each of devices 600-1 and 600-2 is shown, the examples described primarily pertain to the user interface displayed on device 600-1 and / or user input detected by that device. It should be understood that in some examples, electronic device 600-2 operates in a manner similar to that of electronic device 600-1 during a real-time video communication session. In some examples, devices 600-1 and 600-2 display similar user interfaces and / or enable operations similar to those described below.
[0257] As will be described in further detail below, in some examples, such representations include images that have been modified during a real-time video communication session to provide an improved perspective of surfaces and / or objects within the field of view (also referred to herein as the "field of view") of the cameras of devices 600-1 and 600-2. Any known image processing techniques may be used to modify the images, including but not limited to image rotation and / or distortion correction (e.g., image skew). Therefore, although image data may be captured from a camera having a specific position relative to the user, the representation can provide a perspective showing the user (and / or surfaces or objects in the user's environment) from a viewpoint different from the perspective from which the image data is captured by the camera. Figures 6A to 6AY The embodiments disclosed involve displaying elements and detecting input (including hand gestures) at device 600-1 to control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2). In some embodiments, device 600-2 displays similar elements and detects similar input (including hand gestures) to control how image data captured by camera 602 is displayed (at device 600-1 and / or device 600-2).
[0258] refer to Figure 6A Device 600-1 displays a video conferencing interface 604-1 on display 601. Video conferencing interface 604-1 includes a representation 622-1, which in turn includes images (e.g., frames of a video stream) of the physical environment (e.g., a scene) within the field of view 620 of camera 602. In some examples, the image representing 622-1 includes the entire field of view 620. In other examples, the image representing 622-1 includes a portion of the entire field of view 620 (e.g., a cropped portion or subset). As shown, in some examples, the image representing 622-1 includes a surface 619 near user 622 on which user 622 and / or drawing 618 are located.
[0259] The video conferencing interface 604-1 also includes a representation 623-1, which in turn includes an image of the physical environment within the field of view 688 of the camera 682. In some examples, the image of representation 623-1 includes the entire field of view 688. In other examples, the image of representation 623-1 includes a portion of the entire field of view 688 (e.g., a cropped portion or subset). As shown, in some examples, the image of representation 623-1 includes user 623. As shown, representation 623-1 is displayed with a larger magnitude than representation 622-1. In this way, user 622 can better observe user 623 and / or interact with that user during a real-time communication session.
[0260] Device 600-2 displays a video conferencing interface 604-2 on display 683. The video conferencing interface 604-2 includes representation 622-2, which in turn includes an image of the physical environment within the field of view 620 of camera 602. The video conferencing interface 604-2 also includes representation 623-2, which in turn includes an image of the physical environment within the field of view 688 of camera 682. As shown, representation 622-2 is displayed with a larger magnitude than representation 623-2. In this way, user 623 can better observe user 622 and / or interact with that user during a real-time communication session.
[0261] exist Figure 6A At this point, device 600-1 displays interface 604-1. While displaying interface 604-1, device 600-1 detects input 612a (e.g., swipe input) corresponding to a request to display the settings interface. In response to detecting input 612a, device 600-1 displays settings interface 606, as shown below. Figure 6B As depicted in the figure, in some implementations, the settings interface 606 overlays the interface 604-1.
[0262] In some implementations, the settings interface 606 includes one or more display options for controlling settings of device 600-1 (e.g., volume, display brightness, and / or Wi-Fi settings). For example, the settings interface 606 includes a view display option 607-1 that, when selected, causes device 600-1 to display a view menu, such as... Figure 6B As shown in the image.
[0263] like Figure 6B As shown, when displaying the settings interface 606, device 600-1 detects input 612b. In some embodiments, input 612b is a tap gesture on the view display 607-1. In response to detecting input 612b, device 600-1 displays view menu 616-1, as shown... Figure 6C As shown in the image.
[0264] Generally, the view menu 616-1 includes one or more power indicators that can be used to manage (e.g., control) the way the indicator is displayed during a real-time video communication session. For example, selecting a particular power indicator can cause device 600-1 to display or stop displaying the indicator in an interface (e.g., interface 604-1 or interface 604-2).
[0265] View menu 616-1 includes, for example, a surface view display 610, which, when selected, causes device 600-1 to display a representation including a modified image of the surface. In some embodiments, when surface view display 610 is selected, the user interface directly transitions to... Figure 6M The user interface. Additionally or alternatively, Figures 6D to 6L (As described below) shows that it is possible to Figure 6M The user interface in the middle displays other user interfaces before it, as well as those used to initiate such... Figure 6M Other inputs are shown in the process of displaying the user interface. For example, when displaying the view menu 616-1, device 600-1 detects input 612c corresponding to the selection of the surface view display representation 610. In some examples, input 612c is a touch input. In response to detecting input 612c, device 600-1 displays representation 624-1, as shown. Figure 6M As shown in the diagram. Further in response to the detection of input 612c, device 600-2 displays representation 624-2. As described, in some embodiments, an image is modified during a real-time video communication session to provide an image with a specific viewpoint. Therefore, in some examples, representation 624-1 is provided by generating an image from image data captured by camera 602, modifying the image (or a portion of the image), and displaying representation 624-1 with the modified image. In some embodiments, any known image processing techniques are used to modify the image, including but not limited to image rotation and / or distortion correction (e.g., image skew). In some embodiments, the image representing 624-2 is also provided in this manner.
[0266] In some implementations, the image representing 624-1 is modified to provide a desired viewing angle (e.g., a surface view). In some implementations, the image representing 624-1 is modified based on the position of surface 619 relative to camera 602. For example, device 600-1 may rotate the image representing 624-1 by a predetermined amount (e.g., 45 degrees, 90 degrees, or 180 degrees) so that surface 619 can be viewed more intuitively in representation 624-1. Figure 6MAs shown, for example, where camera 602 captures surface 619 from a viewpoint facing user 622, the image representing 624-1 is rotated 180 degrees to provide a viewpoint from user 622's perspective. Therefore, during a real-time video communication session, devices 600-1, 600-2 display surface 619 from user 622's perspective (and via extended drawing 618) during the real-time communication session. In some examples, the image representing 624-2 is also provided in this manner.
[0267] In some implementations, to ensure that user 623 maintains the view of user 622 when representing a modified image including surface 619 of 624-2, device 600-2 maintains the display of representation 622-2. For example... Figure 6M As shown, maintaining the display of representation 622-2 in this manner may include adjusting the size and / or position of representation 622-2 within interface 604-2. Optionally, in some embodiments, device 600-2 stops displaying representation 622-2 to provide a larger-sized representation 624-2. Optionally, in some embodiments, device 600-1 stops displaying representation 622-1 to provide a larger-sized representation 624-1.
[0268] Indicates 624-1 and 624-2 include images of drawing 618 that are modified relative to the position (e.g., positioning and / or orientation) of drawing 618 relative to camera 602. For example, as Figure 6A As depicted, prior to modification, the image was shown as having a specific orientation (e.g., inverted) in representations 622-1, 622-2. As a result of the image modification, the image of drawing 618 is rotated and / or skewed so that the perspective of representations 624-1, 624-2 appears to be from the perspective of user 622. In this way, the modified image of drawing 618 provides a perspective different from that of representations 624-1, 624-2, in order to give user 623 (and / or user 622) a more natural and direct view of drawing 618. Therefore, user 623 can view drawing 618 more easily and intuitively during a real-time video communication session.
[0269] As described, a modified image representation including the surface is provided in response to the selection of a surface image energy representation (e.g., surface view energy representation 610). In some examples, a modified view representation including the surface is provided in response to the detection of other types of input.
[0270] refer to Figure 6DIn some examples, a modified image representation of a surface is provided in response to one or more gestures. For example, device 600-1 may use camera 602 to detect a gesture and, in response to detecting the gesture, determine whether the gesture meets a set of criteria (e.g., a set of gesture criteria). In some embodiments, the criteria include the requirement that the gesture is a pointing gesture, and optionally, the pointing gesture has a specific orientation and / or points to a surface and / or object. For example, refer to... Figure 6D Device 600-1 detects gesture 612d and determines that gesture 612d is a pointing gesture towards drawing 618. In response, device 600-1 displays a representation including a modified image of surface 619, as shown in the reference. Figure 6M As mentioned above.
[0271] In some implementations, this set of standards includes a requirement that the gesture be performed for at least a threshold amount of time. For example, refer to Figure 6E In response to the detection of a gesture, device 600-1 overlays a graphical object 626 onto representation 622-1, indicating that device 600-1 has detected that the user is currently performing a gesture, such as 612d. As shown in some embodiments, device 600-1 magnifies representation 622-1 to help user 622 better view the detected gesture and / or graphical object 626.
[0272] In some embodiments, the graphic object 626 includes a timer 628 (e.g., a digital timer, a ring that fills over time, and / or a bar that fills over time) indicating the amount of time for which gesture 612d has been detected. In some embodiments, the timer 628 also (or alternatively) indicates a threshold amount of time for which gesture 612d will continue to be provided to satisfy that set of criteria. In response to gesture 612d satisfying the threshold amount of time (e.g., 0.5 seconds, 2 seconds, and / or 5 seconds), device 600-1 displays a representation 624-1 including a modified image of the surface. Figure 6M As described above.
[0273] In some examples, graphic object 626 indicates the type of gesture currently detected by device 600-1. In some examples, graphic object 626 is the outline of a hand performing the detected gesture type and / or an image of the detected gesture type. Graphic object 626 may, for example, include a hand performing a pointing gesture in response to device 600-1 detecting that user 622 is performing a pointing gesture. Additionally or alternatively, graphic object 626 may optionally indicate a zoom level (e.g., a zoom level for displaying or showing a representation of a second part of the scene).
[0274] In some examples, a representation of a modified image is provided in response to one or more voice inputs. For example, during a real-time communication session, device 600-1 receives voice input, such as... Figure 6DVoice input 614 (“View my drawing”). In response, device 600-1 displays a representation 624-1 including a modified image of the surface. Figure 6M As described above.
[0275] In some examples, voice input received by device 600-1 may include references to any surface and / or object recognizable by device 600-1, and in response, device 600-1 provides a representation of a modified image including the referenced object or surface. For example, device 600-1 may receive voice input referencing a wall (e.g., the wall behind user 622). In response, device 600-1 provides a representation of a modified image including the wall.
[0276] In some implementations, voice input may be used in combination with other types of input, such as gestures (e.g., gesture 612d). Thus, in some implementations, device 600-1 displays a modified image of a surface (or object) in response to detecting both gesture and voice input corresponding to a request to provide a modified image of the surface.
[0277] In some implementations, surface view representation is provided in other ways. (See reference) Figure 6F For example, the video conferencing interface 604-1 includes an options menu 608. The options menu 608 includes a set of power indicators that can be used to control the device 600-1 during a real-time video communication session, including a view power indicator 607-2.
[0278] When displaying option menu 608, device 600-1 detects input 612f corresponding to a selection of view display 607-2. In response to detecting input 612f, device 600-1 displays view menu 616-2, as follows: Figure 6G As shown in the image. View menu 616-2 can be used to control how the representation is displayed during a real-time video communication session, such as relative to... Figure 6C As mentioned above.
[0279] Although the options menu 608 is shown as persistently displayed in the video conferencing interface 604-1 in all the accompanying drawings, the options menu 608 may be hidden and / or redisplayed by the device 600-1 at any point during a live video communication session. For example, in response to the detection of one or more user inputs and / or periods of inactivity, the options menu 608 may be displayed and / or removed from the display.
[0280] Although the input pointing to the surface has been described as causing device 600-1 to display a representation including a modified image of the surface (e.g., in response to the detection of the surface), Figure 6C The input is 612c, and the device 600-1 displays 624-1, such as... Figure 6M (as shown in the diagram), but in some embodiments, detecting input pointing at the surface can cause device 600-1 to enter preview mode, for example, before displaying representation 624-1 (e.g., Figures 6H to 6J ).
[0281] Figure 6H An example of device 600-1 operating in preview mode is shown. Generally, preview mode can be used to selectively provide a portion or region of a represented image to one or more other users during a live video communication session.
[0282] In some implementations, before operating in preview mode, device 600-1 detects input (e.g., input 612c) pointing to the surface view indicator 610. In response, device 600-1 initiates preview mode. When operating in preview mode, device 600-1 displays preview interface 674-1. Preview interface 674-1 includes a left scroll indicator 634-2, a right scroll indicator 634-1, and preview 636.
[0283] In some implementations, selecting a left scrolling power indicator causes device 600-1 to alter (e.g., replace) preview 636. For example, selecting either left scrolling power indicator 634-2 or right scrolling power indicator 634-1 causes device 600-1 to cycle through various images (the user's image, an unmodified image of the surface, and / or a modified image of surface 619), allowing the user to select a specific viewpoint to share when exiting preview mode, for example, in response to detecting input directed at preview 636. Additionally or alternatively, these techniques can be used to cycle through and / or select specific surfaces (e.g., vertical and / or horizontal surfaces) and / or specific portions (e.g., cropped portions or subsets) within the field of view.
[0284] As shown in the figure, in some implementations, the preview user interface 674-1 is displayed on device 600-1, but not on device 600-2. For example, device 600-2 displays the video conferencing interface 604-2 (including representation 622-2), while device 600-1 displays the preview interface 674-1. In this way, the preview user interface 674-1 allows user 622 to select a view before sharing it with user 623.
[0285] Figure 6IAn example of device 600-1 operating in preview mode is shown. As depicted, when device 600-1 is operating in preview mode, device 600-1 displays a preview interface 674-2. In some embodiments, preview interface 674-2 includes a representation 676 having regions 636-1, 636-2. In some embodiments, representation 676 includes an image that is the same as or substantially similar to the image included in representation 622-1. Optionally, as shown, representation 676 is larger than... Figure 6A Representation 622-1. The position of representation 676 differs from that of representation 622-1. Adjusting the size and / or position of the representation in preview interface 674-2, compared to the size and / or position of a representation including a similar or identical image in video conferencing interface 604-1, allows user 622 to better view the image before sharing it with user 623.
[0286] In some implementations, regions 636-1 and 636-2 correspond to the respective portions representing 676. For example, as shown, region 636-1 corresponds to the upper portion representing 676 (e.g., the portion including the upper body of user 622), and region 636-2 corresponds to the lower portion representing 676 (e.g., the portion including the lower body of user 622 and / or the portion including drawing 618).
[0287] In some embodiments, region 636-1 and region 636-2 are displayed as distinct regions (e.g., non-overlapping regions). In some embodiments, region 636-1 and region 636-2 overlap. Additionally or alternatively, one or more graphic objects 638-1 (e.g., lines, boxes, and / or dashed lines) may distinguish (e.g., visually distinguish) region 636-1 from region 636-2.
[0288] In some implementations, the preview interface 674-2 includes one or more graphical objects to indicate whether a region is active or inactive. Figure 6I In the example, the preview interface 674-2 includes graphical objects 641a and 641b. In some implementations, the appearance (e.g., shape, size, and / or color) of the graphical objects 641a and 641b indicates whether the corresponding area is active or inactive.
[0289] When active, the area is shared with one or more other users in a real-time video communication session. For example, see reference. Figure 6IThe graphical user interface object 641 indicates that region 636-1 is active. Therefore, image data corresponding to region 636-1 is displayed by device 600-2 in representation 622-2. In some examples, device 600-1 shares only the image data of the active region. In some implementations, device 600-1 shares all image data and instructs device 600-2 to display the image only based on the portion of the image data corresponding to the active region 636-1.
[0290] When display interface 674-2 is active, device 600-1 detects input 612i at a location corresponding to region 636-2. In some embodiments, input 612i is touch input. In response to detecting input 612i, device 600-1 activates region 636-2. Therefore, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, region 636-1 remains active in response to input 612i (e.g., user 623 can see user 622, for example, in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates region 636-1 in response to input 612i (e.g., user 623 can no longer see user 622, for example, in representation 622-2).
[0291] Although described relative to the preview mode with a representation including two regions 636-1 and 636-2 Figure 6I For example, see the example provided, but in some implementations, other numbers of regions may be used. For example, refer to... Figure 6J Device 600-1 is operating in preview mode, where preview interface 674-3 includes representation 676, which includes areas 636a-636i.
[0292] In some implementations, multiple regions are active (and / or can be activated). For example, as shown, device 600-1 displays regions 636a-636i, where regions 636a-f are active. Therefore, device 600-2 displays representation 622-2.
[0293] In some implementations, device 600-1 modifies an image of a surface having any type of orientation, including any angle relative to gravity (e.g., between zero and ninety degrees). For example, when the surface is horizontal (e.g., a surface in a plane ranging from 70 to 110 degrees to the direction of gravity), device 600-1 can modify an image of that surface. As another example, when the surface is vertical (e.g., a surface in a plane at up to 30 degrees to the direction of gravity), device 600-1 can modify an image of that surface.
[0294] When displaying interface 674-3, device 600-1 detects input 612j at a location corresponding to region 636h. In response to the detection of input 612j, device 600-1 activates region 636-2. Therefore, device 600-2 displays a representation including a modified image of surface 619, such as representation 624-2. In some embodiments, regions 636a-f remain active in response to input 612j (e.g., user 623 can see user 622, for example, in representation 622-2). Optionally, in some embodiments, device 600-1 deactivates regions 636a-f in response to input 612j (e.g., user 623 can no longer see user 622, for example, in representation 622-2).
[0295] Figures 6K to 6L An exemplary animation is shown that can be displayed by device 600-1 and / or device 600-2. For example... Figures 6A to 6I As discussed herein, device 600-1 may display a representation including a modified image. In some embodiments, device 600-1 and / or device 600-2 display animations of transitions between views and / or modifications to the image over time. This animation may include, for example, translation, rotation, and / or other modifications to the image to provide a modified image. Additionally or alternatively, the animation occurs in response to the detection of input directed at the surface (e.g., selection of the surface view display 610, gesture, and / or voice input).
[0296] Figure 6K An exemplary animation is shown in which device 600-2 translates and rotates an image representing 642a. During the animation, the image representing 642a is panned downwards to view surface 619 from a more "top-down" perspective. The animation also includes rotating the image representing 642a so that surface 619 is viewed from the perspective of user 622. Although four frames of the animation are shown, the animation may include any number of frames. Optionally, in some embodiments, device 600-1 translates and rotates an image representing (e.g., representing 622-1).
[0297] Figure 6L An example is shown where device 600-2 zooms in and rotates an image representing 642a. During the animation, representation 642a is zoomed in until the desired zoom level is achieved. The animation also includes rotating representation 642a until the image of drawing 618 is oriented to the viewpoint of user 622, as described. Although four frames of the animation are shown, the animation may include any number of frames. Optionally, in some embodiments, device 600-1 zooms in and rotates an image representing (e.g., representation 622-1).
[0298] Figures 6N to 6R An example of a modified image of the surface being further modified during a real-time communication session is shown.
[0299] Figure 6N An example of a real-time communication session is shown, in which the user provides various inputs. For example, while displaying interface 678, device 600-1 detects input 677 corresponding to the rotation of device 600-1. Figure 6O As depicted, in response to the detection of input 677, device 600-1 modifies interface 678 to compensate for (e.g., rotation of camera 602). Figure 6O As shown, device 600-1 arranges representations 623-1 and 624-1 of interface 678 in a vertical configuration. Furthermore, representation 624-1 is rotated according to the rotation of device 600-1, such that the viewing angle of representation 624-1 remains in the same orientation relative to user 622. Additionally, the viewing angle of representation 624-2 remains in the same orientation relative to user 623.
[0300] Further reference Figure 6N In some examples, device 600-1 displays control indicator 648-1, 648-2 to modify the image of indicator 624-1. Control indicator 648-1, 648-2 can respond, for example, to an option menu 608 (e.g., ...). Figure 6B The display is generated by selecting one or more inputs to indicate the power level.
[0301] As shown in the figure, in some embodiments, device 600-1 displays a representation 624-1 including a modified image of a surface. When a rotation display 648-1 is selected, device 600-1 rotates the image of representation 624-1. For example, when displaying interface 678, device 600-1 detects an input 650a corresponding to the selection of rotation display 648-1. In response to input 650a, device 600-1 changes the orientation of the image of representation 624-1 from a first orientation (…). Figure 6N (as shown) modified to the second orientation ( Figure 6O (As shown in the diagram). In some embodiments, this indicates that the image of 624-1 is rotated by a predetermined amount (e.g., 90 degrees).
[0302] The scaling capability (648-2) modifies the scaling level of the image represented by (624-1) when selected. For example, as... Figure 6N As depicted, an image representing 624-1 is displayed at a first zoom level (e.g., "1X"). When displaying a zoom capability representation 648-2, device 600-1 detects an input 650b corresponding to a selection of the zoom capability representation 648-2. In response to input 650b, device 600-1 modifies the zoom level of the image representing 624-1 from the first zoom level (e.g., "1X") to a second zoom level (e.g., "2X"), as shown. Figure 6Q As shown in the image.
[0303] Additionally or alternatively, in some embodiments, the video conferencing interface 604-1 includes an option to display a magnified view of at least a portion of an image representing 624-1, such as Figure 6R As shown in the diagram. For example, when displaying representation 624-1, device 600-1 may detect input 654 (e.g., a gesture pointing to a surface and / or object) corresponding to a request for a magnified view of a portion of the image displayed in representation 624-1. In response to the detection of input 654, device 600-1 displays the magnified portion 652-1 at a greater zoom level than the second portion 652-2 of representation 624-1. In some embodiments, the magnified portion of the image representing 624-1 is determined based on the location of input 654. In some embodiments, in response to the detection of input 650c ( Figure 6R and Figure 6Q The device 600-1 stops displaying control power indicators 648-1 and 648-2.
[0304] Figures 6S to 6AC An example is shown where the device modifies the image of the representation in response to user input. As described in more detail below, device 600-1 may modify the image of the representation (e.g., representation 622-1) in video conferencing interface 604-1 in response to non-touch user input (including gestures and / or audio input), thereby improving the way users interact with the device to manage and / or modify the representation during a real-time video communication session.
[0305] Figures 6S to 6T An example is shown where the device, in response to a gesture, obscures at least a portion of the displayed image. Figure 6S As shown, device 600-1 detects a gesture 656a corresponding to a request to modify at least a portion of the image representing 622-1. In some examples, gesture 656a is an upward gesture (e.g., a "shh" gesture) made by user 622 near the mouth of user 622. Figure 6T As shown, in response, device 600-1 replaces representation 622-1 with representation 622-1' that includes a modified image, which includes a modified portion 658-1 (e.g., the background of user 622's physical environment). In some examples, modifying portion 658-1 in this way includes blurring, graying out, or otherwise obscuring portion 658-1. In some examples, device 600-1 does not modify portion 658-2 in response to gesture 656a.
[0306] Figures 6U to 6V An example is shown where the device zooms in on a portion of an image in response to a detected gesture. (e.g.) Figure 6UAs shown, in some embodiments, device 600-1 detects a pointing gesture 656b corresponding to a request for at least a portion of the magnified representation 622-1. As illustrated, the pointing gesture 656b points to object 660.
[0307] like Figure 6V As depicted, in response to a pointing gesture 656b, device 600-1 replaces representation 622-1 with representation 622-1' comprising a modified image by magnifying a portion of the image comprising representation 622-1 of object 660. In some embodiments, the magnification is based on the position of object 660 (e.g., relative to camera 602) and / or the size of object 660.
[0308] Figures 6W to 6X An example is shown where the device zooms in on a portion of the displayed view in response to a detected gesture. Figure 6W As shown, in some embodiments, device 600-1 detects a framing gesture 656c corresponding to a request for at least a portion of magnification representation 622-1. As shown, since framing gesture 656c at least partially frames, surrounds, and / or outlines the object 660, it is pointing at object 660.
[0309] like Figure 6X As depicted, in response to framing gesture 656c, device 600-1 modifies the image of representation 622-1 by magnifying a portion of the image including object 660. In some embodiments, the magnification is based on the position of object 660 (e.g., relative to camera 602) and / or the size of object 660. Additionally or alternatively, after magnifying a portion of the image of representation 622-1, device 600-1 may track the movement of framing gesture 656c. In response, device 600-1 may pan to different portions of the image.
[0310] Figures 6Y to 6Z An example is shown where the device pans the displayed image in response to a detected gesture. Figure 6Y As shown, device 600-1 detects a pointing gesture 656d corresponding to a request for a view of an image representing 622-1 that has been translated in a specific direction (e.g., horizontally). As shown, pointing gesture 656d points to the left of user 622.
[0311] like Figure 6Z As shown, in response to pointing gesture 656d, device 600-1 replaces representation 622-1 with representation 622-1', which includes a modified image of representation 622-1 translated in the direction of pointing gesture 656d (e.g., to the left of user 622).
[0312] Although in some implementation schemes, such as Figure 6Z As shown, due to the translation operation, a portion of user 622 (e.g., user 622's right shoulder) can be excluded from the image representing 622-1'. However, in some embodiments, device 600-1 may adjust the scaling level of the image representing 622-1' during translation to ensure that user 622 remains fully in the image.
[0313] Figures 6AA to 6AB An example is shown where the device modifies the scaling level of the representation in response to detecting a pinch and / or unfold gesture. Figure 6AA As shown, in some embodiments, device 600-1 detects an unfolding gesture 656e, wherein user 622 increases the distance between the thumb and index finger of user 622's right hand.
[0314] like Figure 6AB As depicted, in response to an unfolding gesture 656e, device 600-1 replaces representation 622-1 with representation 622-1' by magnifying a portion of the image representing 622-1. In some embodiments, the magnification is based on the position of the unfolding gesture 656e (e.g., relative to camera 602) and / or the magnitude of the unfolding gesture 656e. In some embodiments, the portion of the image is magnified according to a predetermined zoom level.
[0315] refer to Figure 6AA In some embodiments, in response to detecting an expand gesture 656e, device 600-1 displays a zoom indicator 662 indicating the zoom level of the image representing 622-1'. Once user 622 has completed the expand gesture 656e and device 600-1 has zoomed in on the portion representing 622-1', device 600-1 updates the display of the zoom indicator 662 to indicate the current zoom level of the image representing 622-1'. In some embodiments, the zoom indicator 662 is dynamically updated as user 622 performs gesture 656e.
[0316] Although the scaling level of an image is described in this document in relation to increasing the scaling level in response to an unfolding gesture 656e, in some examples, the scaling level of an image is decreased in response to a gesture (e.g., other types of gestures, such as a pinch gesture).
[0317] Figure 6ACVarious gestures that can be used to modify a represented image are illustrated. In some implementations, for example, a user can use gestures to indicate a zoom level. For instance, gesture 664 can be used to indicate that the zoom level of the represented image should be "1X", and in response to detecting gesture 664, device 600-1 can modify the represented image to have a "1X" zoom level. Similarly, gesture 666 can be used to indicate that the zoom level of the represented image should be "2X", and in response to detecting gesture 666, device 600-1 can modify the represented image to have a "2X" zoom level. While for... Figure 6AC Two scaling levels (e.g., "1X" and "2X" scaling levels) are described, but in some embodiments, device 600-1 can use the same or different gestures to modify the represented image to other scaling levels (e.g., 0.5X, 3X, 5X, or 10X). In some embodiments, device 600-1 can modify the represented image to three or more different scaling levels. In some embodiments, the scaling levels are discrete or continuous.
[0318] As another example, a gesture by which the user curls their fingers can be used to adjust the zoom level. For example, gesture 668 (e.g., a gesture by which the user's fingers curl in a direction away from the camera 668b (e.g., when the back of the hand 668a is oriented towards the camera)) can be used to indicate that the zoom level of the image should be increased (e.g., magnified). Gesture 670 (e.g., a gesture by which the user's fingers curl in a direction towards the camera 670b (e.g., when the palm of the hand 668a is oriented towards the camera)) can be used to indicate that the zoom level of the image should be decreased (e.g., reduced).
[0319] Figures 6AD to 6AE This illustrates an example of a user using two devices to participate in a real-time video communication session.
[0320] For example, such as Figure 6AD As shown, user 623 is using additional device 600-3 during a real-time video communication session. In some embodiments, devices 600-2 and 600-3 simultaneously display representations including images with different views. For example, while device 600-3 displays representation 622-2, device 600-2 displays representation 624-2.
[0321] In some implementations, device 600-2 is positioned on desk 686 in front of user 623 in a manner corresponding to the position of surface 619 relative to user 622. Therefore, user 623 can view representation 624-2 (including an image of surface 619) in a manner similar to how user 622 views surface 619 in the physical environment.
[0322] like Figure 6AEAs shown, during a real-time communication session, user 623 can modify the image displayed in representation 624-2 via mobile device 600-2. In response to user 623 changing the orientation of device 600-2, device 600-2 modifies the image in representation 624-2, for example, in a manner corresponding to the change in device 600-2's orientation. For example, in response to user 623 tilting device 600-2, device 600-2 translates upwards to display other portions of surface 619. In this way, user 623 can change the orientation of device 600-2 (in any direction) to view various portions of surface 619 that are not otherwise displayed when device 600-2 is in a different orientation.
[0323] Figures 6AF to 6AL A diagram is shown for accessing the reference. Figures 6A to 6AE Various user interface implementations are shown and described. Figures 6AF to 6AL In the embodiments depicted, these interfaces are displayed using a laptop computer (e.g., John's device 6100-1 and / or Jane's device 6100-2). It should be understood that... Figures 6AF to 6AL The embodiments shown can be implemented using different devices, such as tablet computers (e.g., John's tablet computer 600-1 and / or Jane's device 600-2). Similarly, Figures 6A to 6AE The embodiments shown can be implemented using different devices (such as John's device 6100-1 and / or Jane's device 6100-2). Therefore, for the sake of brevity, the above description will not be repeated below. Figures 6A to 6AE The various operations or features described. For example, relative to Figures 6A to 6AE The applications, interfaces (e.g., 604-1 and / or 604-2), and displayed elements (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, and / or 624-2) discussed are similar to those relative to... Figures 6AF to 6AL The applications, interfaces (e.g., 6121 and / or 6131) and displayed elements (e.g., 6124, 6132, 6122, 6134, 6116, 6140 and / or 6142) discussed are included. Therefore, for the sake of brevity, details of these applications, interfaces, and displayed elements may not be repeated below.
[0324] Figure 6AFJohn's device 6100-1 is depicted, which includes a display 6101, one or more cameras 6102, and a keyboard 6103 (in some embodiments, the keyboard includes a touchpad). John's device 6100-1 displays a main screen via the display 6101, including a camera application icon 6108 and a video conferencing application icon 6110. The camera application icon 6108 corresponds to a camera application operating on John's device 6100-1, which is used to access camera 6102. The video conferencing application icon 6110 corresponds to a video conferencing application operating on John's device 6100-1, which is used to initiate and / or participate in conferencing similar to those described in the above references. Figures 6A to 6AE The real-time video communication sessions discussed (e.g., video calls and / or video chats). John's device 6100-1 also displays a taskbar 6104, which includes various application icons, including a subset of icons displayed in the dynamic area 6106. The icons displayed in the dynamic area 6106 represent applications active (e.g., launched, opened, and / or used) on John's device 6100-1. Figure 6AF In this context, neither the camera application nor the video conferencing application is currently active. Therefore, icons representing the camera application or video conferencing application are not displayed in the dynamic area 6106, and John's device 6100-1 is not participating in the real-time video communication session.
[0325] exist Figure 6AF In this context, John's device 6100-1 detects input such as selection of the camera application icon 6108 as indicated by cursor 6112 (e.g., cursor input caused by clicking the mouse, tapping on the touchpad, and / or other such input). In response, John's device 6100-1 launches the camera application and displays the camera application window 6114, as... Figure 6AG As shown. In Figure 6AG In the embodiment depicted, a camera application is used to access camera 6102 to generate a surface view 6116, which is similar to, for example... Figure 6M The representation 624-1 is depicted and described above. In some embodiments, the camera application may have different modes (e.g., user-selectable modes), such as, for example, an extended field-of-view mode (which provides an extended field of view for camera 6102) and a surface view mode (which provides... Figure 6AG The surface view shown is thus represented by surface view 6116. Therefore, surface view 6116 represents the result of being acquired using camera 6102 and modified (e.g., zoomed in, rotated, cropped, and / or skewed) by a camera application. Figure 6AGThe image data is shown in the surface view 6116. Additionally, because John's laptop computer has launched the camera application, the camera application icon 6108-1 is displayed in the dynamic area 6106 of the taskbar 6104, indicating that the camera application is active. In some embodiments, when application icons (e.g., 6108-1) are added to the dynamic area of the taskbar, these application icons are displayed with animated effects (e.g., bouncing).
[0326] exist Figure 6AG In the process, John's device 6100-1 detects input 6118 to select the video conferencing application icon 6110. In response, John's device 6100-1 launches the video conferencing application, displays the video conferencing application icon 6110-1 in the dynamic area 6106, and displays the video conferencing application window 6120, as shown. Figure 6AH As shown in the diagram. The video conferencing application window 6120 includes a video conferencing interface 6121 similar to interface 604-1, and includes Jane's video feed 6122 (similar to representation 623-1) and John's video feed 6124 (similar to representation 622-1). In some embodiments, John's device 6100-1 displays the video conferencing application window 6120 with video conferencing interface 6121 after detecting one or more additional inputs following input 6118. For example, such inputs could be inputs initiating or accepting requests to participate in a video call with Jane's laptop.
[0327] exist Figure 6AH In this embodiment, John's device 6100-1 displays a video conferencing application window 6120 partially overlaid on the camera application window 6114. In some implementations, in response to the detection of a selection of the camera application icon 6108, a selection of icon 6108-1, and / or input on the camera application window 6114, John's device 6100-1 can bring the camera application window 6114 to the foreground (e.g., partially overlaid on the video conferencing application window 6120). Similarly, in response to the detection of a selection of the video conferencing application icon 6110, a selection of icon 6110-1, and / or input on the video conferencing application window 6120, the video conferencing application window 6120 can be brought to the foreground (e.g., partially overlaid on the camera application window 6114).
[0328] exist Figure 6AHIn the illustration, John's device 6100-1 is shown participating in a real-time video communication session with Jane's device 6100-2. Therefore, Jane's device 6100-2 is depicted displaying a video conferencing application window 6130, which is similar to the video conferencing application window 6120 on John's device 6100-1. The video conferencing application window 6130 includes a video conferencing interface 6131 similar to interface 604-2, and includes John's video feed 6132 (similar to representation 622-2) and Jane's video feed 6134 (similar to representation 623-2).
[0329] exist Figure 6AH In the embodiments depicted, a video conferencing application is used to access camera 6102 to generate video feeds 6124 and 6132. Therefore, video feeds 6124 and 6132 represent views of image data acquired using camera 6102 and modified (e.g., zoomed in and / or cropped) by the video conferencing application to produce the images (e.g., videos) shown in video feeds 6124 and 6132. In some embodiments, the camera application and the video conferencing application may use different cameras to provide the respective video feeds.
[0330] The video conferencing application window 6120 includes menu options 6126, which can be selected to display different options for sharing content in a real-time video communication session. Figure 6AH In the process, John's device 6100-1 detects input 6128 for selecting menu option 6126 and, in response, displays the shared menu 6136, such as... Figure 6AI As shown in the diagram. The sharing menu 6136 includes sharing options 6136-1, 6136-2, and 6136-3. Sharing option 6136-1 is an option to share content from the camera application. Sharing option 6136-2 is an option to share content from the desktop of John's device 6100-1. Sharing option 6136-3 is an option to share content from the presentation application. In response to detecting input 6138 on sharing option 6136-1, John's device 6100-1 begins sharing content from the camera application, as shown in the diagram. Figure 6AJ and Figure 6AK As shown in the image.
[0331] exist Figure 6AJ In this process, John's device 6100-1 updates the video conferencing interface 6121 to include a surface view 6140, which is shared with Jane's device 6100-2 during the live video communication session. Figure 6AJIn the implementation described, John's device 6100-1 shares the video feed generated using the camera application (shown as surface view 6116 in the camera application window 6114), and displays a representation of the video feed as surface view 6140 in the video conferencing application window 6120. Additionally, John's laptop computer emphasizes the display of surface view 6140 in the video conferencing interface 6121 (e.g., by displaying the surface view at a larger size than other video feeds) and reduces the display size of Jane's video feed 6122. Figure 6AJ In this configuration, John's device 6100-1 displays a surface view 6140 simultaneously in the video conferencing application window 6120 along with John's video feed 6124 and Jane's video feed 6122. In some implementations, the display of John's video feed 6124 and / or Jane's video feed 6122 in the video conferencing application window 6120 is optional.
[0332] Jane's device 6100-2 updates the video conferencing interface 6131 to show the surface video feed 6142, which is a surface view (from the camera application) shared by John's device 6100-1. Figure 6AJ As shown, Jane's device 6100-2 adds a surface video feed 6142 to the video conferencing interface 6131 so that the surface video feed is displayed simultaneously with Jane's video feed 6134 and John's video feed 6132. Optionally, the size of John's video feed has been adjusted to accommodate the addition of the surface video feed 6142. In some embodiments, Jane's device 6100-2 replaces John's video feed 6132 and / or Jane's video feed 6134 with the surface video feed 6142.
[0333] Figure 6AK An alternative embodiment depicting sharing content from a camera application in response to the detection of input 6138 on sharing option 6136-1 is shown. Figure 6AKIn this configuration, John's laptop displays a camera application window 6114 with a surface view 6116 (optionally minimized or hidden, such as a video conferencing application window 6120). John's device 6100-1 also displays John's video feed 6115 (similar to John's video feed 6124) and Jane's video feed 6117 (similar to Jane's video feed 6122), indicating that John's laptop is sharing the surface view 6116 with Jane's device 6100-2 in a real-time video communication session (e.g., video chat provided by a video conferencing application). In some embodiments, the display of John's video feed 6115 and / or Jane's video feed 6117 is optional. Figure 6AJ In the embodiment shown, Jane's device 6100-2 shows a surface video feed 6142, which is a surface view (from a camera application) shared by John's device 6100-1.
[0334] Figure 6AL The field of view of camera 6102 and its application are shown. Figures 6AF to 6AK The diagram illustrates portions of the field of view of the video conferencing application and camera application in the implementation scheme described herein. For example, in Figure 6AL The image shows a silhouette view of John's laptop computer 6100 within John's physical environment. Dashed lines 6145-1 and dotted lines 6147-2 represent the outer dimensions of the field of view of camera 6102, which in some embodiments is a wide-angle camera. The collective field of view of camera 6102 is indicated by shaded areas 6144, 6146, and 6148. The portion of the camera's field of view used by the camera application (e.g., for surface view 6116) is indicated by dotted lines 6147-1 and 6147-2 and shaded areas 6146 and 6148. In other words, surface view 6116 (and surface view 6140) is generated by the camera application using a portion of the camera's field of view represented by shaded areas 6146 and 6148 between dotted lines 6147-1 and 6147-2. The portion of the camera's field of view used by the video conferencing application (e.g., for John's video feed 6124) is indicated by dashed lines 6145-1 and 6145-2 and shaded areas 6144 and 6146. In other words, John's video feed 6124 is generated by the video conferencing application using the portion of the camera's field of view represented by the shaded areas 6144 and 6146 between dashed lines 6145-1 and 6145-2. The shaded area 6146 represents the overlap of the portion of the camera's field of view used to generate the video feed for the respective camera and video conferencing application.
[0335] Figure 6AM to Figure 6AY A control reference is shown. Figures 6A to 6ALVarious user interfaces and views, and / or implementations for interacting with them, are shown and described. Figure 6AM to Figure 6AY In the embodiments depicted, these interfaces are shown using tablet computers (e.g., John's tablet computer 600-1 and / or Jane's device 600-2) and computers (e.g., Jane's computer 600-4). Figure 6AM to Figure 6AY The embodiments shown are optionally implemented using different devices, such as laptop computers (e.g., John's device 6100-1 and / or Jane's device 6100-2)). Similarly, Figures 6A to 6AL The implementation shown can optionally be implemented using different devices (such as Jane's computer 6100-2). Therefore, for the sake of brevity, the above description will not be repeated below. Figures 6A to 6AL The various operations or features described.
[0336] Additionally, the relative position provided by one or more cameras (e.g., 602, 682, and / or 6102) Figures 6A to 6AL The applications, interfaces (e.g., 604-1, 604-2, 6121 and / or 6131), and fields of view (e.g., 620, 688, 6145-1, and 6147-2) discussed are similar to those provided by a camera (e.g., 602) relative to... Figure 6AM to Figure 6AY The applications, interfaces (e.g., 604-4), and fields of view (e.g., 620) discussed are described below. Therefore, for the sake of brevity, details of these applications, interfaces, and fields of view may not be repeated below. Additionally, the control detected by device 600-1 and its relationship to... Figures 6A to 6AL Options and requests (e.g., input and / or hand gestures) for views associated with the elements displayed in the discussion (e.g., 622-1, 622-2, 623-1, 623-2, 624-1, 624-2, 6121, and / or 6131) are optionally detected by devices 600-2 and / or 600-4 to control the view relative to... Figure 6AM to Figure 6AY The discussed elements (e.g., 622-1, 622-4, 623-1, 623-4, 6214, and / or 6216) are associated with views (e.g., user 623 optionally provides input to cause device 600-1 and / or device 600-2 to provide a representation 624-1 including a modified image of the surface). Additionally, Figure 6AM to Figure 6AY Devices 600-1 and 600-2 are described and depicted as being in a lateral orientation. In some embodiments, devices 600-1 and / or device 600-2 are in a longitudinal orientation, similar to... Figure 6O Device 600-1. Therefore, for the sake of brevity, the details of these options and requests detected by device 600-2 may not be repeated below.
[0337] Figures 6AM to 6AJExemplary user interfaces for controlling a physical environment are shown and described. The user interfaces in these figures are used to illustrate the following descriptions, including... Figure 15 The process of... In... Figure 6AM At this location, devices 600-1 and 600-4 display interfaces 604-1 and 604-4, respectively. Interface 604-1 includes representation 622-1, and interface 604-4 includes representation 622-4. Representations 622-1 and 622-4 include images of image data from a portion of the field of view 620 (specifically, the shaded area 6206). As shown, representations 622-1 and 622-4 include images of the head and upper torso of user 622, but do not include images of drawing 618 on desk 621. Interfaces 604-1 and 604-4 include representations 623-1 and 623-4, respectively, which include images of user 223 within the field of view 6204 of camera 6202. Interfaces 604-1 and 604-4 also include an options menu 609 (similar to...). Figures 6A to 6AE The discussion includes options menu 608 for controlling image data captured by 602 and / or by camera 6202, including Figures 6F to 6G This option menu allows devices 600-1 and 600-4 to manage how image data is displayed.
[0338] exist Figure 6ANDuring a real-time video communication session, user 623 brings device 600-2 to the vicinity of device 600-4. As depicted, in response to the detection of device 600-2 (e.g., via wireless communication), device 600-4 displays an add notification 6210a. Similarly, in response to the detection of device 600-4, device 600-2 displays an add notification 6210b via display 683 (e.g., a touch-sensitive display). In some embodiments, devices 600-2 and 600-4 use specific device standards to trigger the display of add notifications 6210a and 6210b. In some embodiments, the specific device standards include standards for a specific location (e.g., orientation, orientation, and / or angle) of device 600-2, which trigger the display of add notifications 6210a and / or 6210b when met. In such embodiments, a specific location (e.g., orientation, orientation, and / or angle) of device 600-2 includes criteria such as device 600-2 having a specific angle or angle range (e.g., an angle or angle range indicating that the device is horizontal and / or flat on desk 686) and / or display 683 facing upwards (e.g., opposite to facing downwards towards desk 686). In some embodiments, a specific device criterion includes the criterion that device 600-2 is near device 600-4 (e.g., within a threshold distance of device 600-4). In some embodiments, device 600-2 wirelessly communicates with device 600-4 to convey the location and / or proximity of device 600-2 (e.g., using location data and / or short-range wireless communication, such as Bluetooth and / or NFC). In some embodiments, a specific device criterion includes the criterion that device 600-2 and device 600-4 are associated with the same user (e.g., being used by the same user and / or logged into the same user). In some implementations, specific device standards include standards for device 600-2 having specific states (e.g., unlocked and / or the display powered on, the opposite of locked and / or the display powered off).
[0339] exist Figure 6AN At this point, connection notifications 6210a-6210b include an indication to include device 600-4 in a real-time video communication session. For example, addition notifications 6210a-6210b include an indication to add a representation (for display on device 600-2) that includes an image of field of view 620 captured by camera 602. In some embodiments, addition notifications 6210a-6210b include an indication to add a representation (for display on device 600-1) that includes an image of field of view 6204 captured by camera 6202.
[0340] exist Figure 6ANIn this context, notifications 6210a and 6210b include accept enable representations 6212a and 6212b, which, when selected, add (e.g., connect) device 600-2 to the live video communication session. Notifications 6210a and 6210b also include reject enable representations 6213a and 6213b, which, when selected, clear notifications 6210a and 6210b respectively, without adding device 600-2 to the live video communication session. When accept enable representation 6212b is displayed, device 600-2 detects input 6250an directed to accept enable representation 6212b (e.g., a tap, mouse click, or other selection input). In response to detecting input 6250an, device 600-2 displays interface 604-2, as shown... Figure 6AO As depicted in the text.
[0341] exist Figure 6AO At this point, interface 604-2 is similar to interface 604-2 described herein (e.g., refer to...). Figures 6A to 6AE ) and video conferencing interface 6131 as described herein (e.g., reference ) Figures 6AH to 6AK ), but with different states. For example, Figure 6AO Interface 604-2 does not include representations of 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, or options menu 609. In some implementations, Figure 6AO The interface 604-2 includes one or more of the following: representations 622-2 and 623-2, John's video feed 6132 and Jane's video feed 6134, and / or options menu 609.
[0342] exist Figure 6AO At this point, interface 604-2 includes an adjustable view 6214 of the video feed captured by camera 602 (similar to John's video feed 6132 and representation 622-2, but including a different portion of the field of view 620). The adjustable view 6214 is associated with a portion of the field of view 620 corresponding to the shadow region 6217. In some embodiments, Figure 6AO Interface 604-2 includes representations 622-4 and 623-4 and / or an options menu 609. In some embodiments, in response to input detected at device 600-2 and / or device 600-4, representations 622-4 and 623-4 and / or the options menu 609 are moved from interface 604-4 to interface 604-2 so that they are displayed simultaneously with the adjustable view 6214. In such embodiments, display 6201 acts as an auxiliary display (e.g., an extended display) and / or vice versa for display 604-1.
[0343] exist Figure 6AO In response to Figure 6ANInput 6250an was detected at the device, and device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, such as Figure 6AO As depicted in the text. Figure 6AO The interface 604-1 is similar to Figure 6AN The interface 604-1 has a different state (e.g., 623-1 and 622-1 are smaller in size and in different positions). Interface 604-1 includes an adjustable view 6216, which is similar to the adjustable view 6214 displayed at device 600-2 (e.g., adjustable view 6216 is associated with a portion of the field of view 620 corresponding to the shaded area 6217). When input described herein is detected by device 600-2 (e.g., movement of device 600-2), the adjustable view 6216 is updated to include an image similar to that of adjustable view 6214. Displaying the adjustable view 6216 allows user 622 to see what portion of the field of view 620 user 624 is currently viewing, as user 623 optionally controls what view is displayed within the field of view 620, as described in more detail below.
[0344] exist Figure 6AO At point 604-2, when displaying interface 600-2, device 600-2 detects movement 6218ao of device 600-2. In response to detecting movement 6218ao, device 600-2 displays... Figure 6AP Interface 602-4. Additionally, in response to the detection of movement 6218ao, device 600-2 causes device 600-1 to display... Figure 6AP The interface is 604-1.
[0345] exist Figure 6AP At this point, interface 602-4 includes an updated adjustable view 6214. (And...) Figure 6AO Compared to the adjustable view 6214, Figure 6AP The adjustable view 6214 represents different views within the field of view 620. For example, Figure 6AP The shaded area 6217 has been relative to Figure 6AOThe shadow area 6217 moves. It is noteworthy that camera 602 does not move. In some embodiments, the movement 6218ao of device 600-2 corresponds to the amount of change (e.g., proportional to) the adjustable view 6214. For example, in some embodiments, the amount of angle by which device 600-2 rotates (e.g., relative to gravity) corresponds to the amount of change (e.g., the amount by which image data is translated to include the new angle of the view). In some embodiments, the direction of movement (e.g., movement 6218ao) of device 600-2 (e.g., tilting downwards and / or rotating downwards) corresponds to the direction of change (e.g., translating downwards) of the adjustable view 6214. In some embodiments, the acceleration and / or velocity of movement (e.g., movement 6218) corresponds to the velocity of change of the adjustable view 6214. In some embodiments, device 600-2 (and / or device 600-1) displays from... Figure 6AO Adjustable view 6214 in Figure 6AP The gradual transition of the adjustable view 6214 (e.g., a series of views). Additionally or alternatively, such as... Figure 6AP As depicted, device 600-2 is placed flat on desk 686. In some embodiments, device 600-2 displays in response to detecting a specific location or a location within a predetermined range (e.g., horizontal and / or upward display). Figure 6AP Adjustable view 6214. As depicted, Figure 6AO Mobile 6218ao in the middle does not enable device 600-2 update Figure 6AP The representations are 622-4 and 623-4 (and / or 623-1 and 622-1 on device 600-1).
[0346] exist Figure 6AP At this location, the image of drawing 618 in view 6214 can be adjusted to be in the same position as... Figure 6AO The adjustable view 6214 shows the image of drawing 618 from different perspectives. For example, Figure 6AP The adjustable view 6214 includes a top view perspective, while Figure 6AO The adjustable view 6214 includes a perspective that combines a side view and a top view. In some embodiments, it includes... Figure 6AP The image of the drawing in the adjustable view 6214 is based on the use of a reference. Figures 6A to 6AL The described similar techniques have been used to modify (e.g., skew and / or magnify) the image data. In some implementations, including... Figure 6AO The image of the drawing in the adjustable view 6214 is based on an image that has not been modified (e.g., skewed and / or magnified) and / or has been compared with... Figure 6APThe image of drawing 618 in the adjustable view 6214 is modified in a different way (e.g., to a lesser extent) (e.g., with) Figure 6AP The image data is skewed and / or magnified less (compared to the amount of skew and / or magnification applied in the image). Providing a top-down view offers greater convenience in collaboration and content sharing because it gives user 623 a view of the drawing that would be similar to the view user 623 would have if user 623 were sitting opposite user 622 looking down at the surface 619 of desk 621.
[0347] exist Figure 6AP At the same location, the adjustable view 6216 of interface 604-1 has also been updated in a similar manner. In some embodiments, the images of adjustable view 6216 and / or adjustable view 6214 are modified based on the position of surface 619 relative to camera 602, as shown in reference [reference needed]. Figures 6A to 6AL As described above. In such embodiments, devices 600-1 and / or 600-2 rotate the image of the adjustable view 6214 by an amount (e.g., 45 degrees, 90 degrees, or 180 degrees) so that the image of drawing 618 can be viewed more intuitively in the adjustable view 6216 and / or the adjustable view 6214 (e.g., displaying the image of drawing 618 so that the house is facing up, opposite to being upside down).
[0348] exist Figure 6AP At this point, user 623 uses stylus 6220 to apply numerical markers to adjustable view 6214. For example, in the display Figure 6AP When the adjustable view 6214 is accessed, device 600-2 detects input corresponding to a request to add numeric markers to the adjustable view 6214 (e.g., using stylus 6220). In response to detecting input corresponding to a request to add numeric markers to the adjustable view 6214, device 600-2 displays interface 604-2, as shown. Figure 6AO As depicted herein. Additionally or alternatively, in response to detecting input corresponding to a request to add a numerical marker to the adjustable view 6214, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown in the image. Figure 6AQ As depicted in the text.
[0349] exist Figure 6AQAt the same time, interface 602-4 includes a digital sun 6222 in an adjustable view 6214, and interface 602-1 includes a digital sun 6223 in an adjustable view 6214. Displaying the digital sun at both devices allows users 623 and 622 to collaborate over a video communication session. Additionally, as depicted, the digital sun 6222 has a position relative to the image of drawing 618. As described in more detail below, even if device 600-1 detects further movement and / or if drawing 618 moves on surface 619, the digital sun 6222 maintains its position relative to the image of drawing 618. In some embodiments, device 600-2 stores data corresponding to the relationship between digital markers (e.g., digital sun 6223) detected in the image data and objects (e.g., a house) in order to determine where (and / or whether) the digital sun 6222 should be displayed. In some embodiments, device 600-2 stores data corresponding to the relationship between the digital marker (e.g., digital sun 6223) and the position of device 600-2 in order to determine where (and / or whether) the digital sun 6222 should be displayed. In some embodiments, device 600-2 detects digital markers applied to other views within the field of view 620. For example, digital markers may be applied to an image of the user's head, such as... Figure 6AR An image of the head of user 622 in the adjustable view 6214.
[0350] exist Figure 6AQ At this point, interface 604-2 includes control and power indicator 648-1 (similar to...). Figure 6N The control indicator 648-1 is used to modify the image in the adjustable view 6214. Rotating the indicator 648-1, when selected, causes device 600-1 (and / or device 600-2) to rotate the image in the adjustable view 6214, similar to how the control indicator 648-1 modifies... Figure 6N The image in the figure represents 624-1.
[0351] exist Figure 6AQ In some implementations, interface 604-2 includes something similar to Figure 6N The scaling display representation in 648-2 is a scaling display representation. In such an implementation, the scaling display representation modifies the image in the adjustable view 6214, similar to how the scaling display representation 648-2 is modified. Figure 6N The image in the image represents 624-1. Control indicators 648-1 and 648-2 can respond, for example, to an option menu 609 (e.g., ...). Figure 6AM The display is generated by selecting one or more inputs to indicate the power level.
[0352] exist Figure 6AQIn some implementations, the digital sun 6222 is projected onto the physical surface of drawing 618, similar to how mark 956 is projected onto surface 908b. Figures 9K to 9N As described herein. In such embodiments, electronic devices (e.g., projectors and / or light-emitting projectors) are used to project and / or render images of the digital sun 6222 within a physical environment 915. For example, the electronic devices may use relative to Figures 9K to 9N The technique described above, based on the relative position of the digital sun 6222 with respect to the drawing 618, allows the projection of the digital sun to be displayed next to the drawing 618.
[0353] exist Figure 6AQ When the digital sun 6222 is displayed in the adjustable view 6214, device 600-2 detects movement 6218aq (e.g., rotation and / or lifting). In response to the detected movement 6218aq, device 600-2 displays interface 604-2, as shown... Figure 6AR As depicted in the text. In response to the detection of movement 6218aq, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown. Figure 6AR As depicted in the text.
[0354] exist Figure 6AR At this point, interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). With Figure 6AQ Compared to the adjustable view 6214, Figure 6AR The adjustable view 6214 represents different views within the field of view 620. For example, Figure 6AP The shaded area 6217 has been relative to Figure 6AQ The shadow area 6217 moves. In some embodiments, the direction of movement 6218aq (e.g., tilting upwards) corresponds to the direction of change in the view (e.g., panning upwards). Additionally, shadow area 6217 overlaps with shadow area 6206, as depicted by the darker shadow area 6224. The darker shadow area 6224 is a schematic representation of the updated adjustable view 6214 based on a portion of the image data used to represent 622-4. Because movement 6218aq has resulted in a change of view (e.g., a change to a view of the user 622's face and / or non-drawing 618), device 600-2 no longer displays the digital sun 6222 in the adjustable view 6214.
[0355] exist Figure 6ARAt this location, the adjustable view 6214 includes a boundary indicator 6226. The boundary indicator 6226 indicates that a boundary has been reached. In some embodiments, the boundary is configured (e.g., by a user) to set restrictions on what portions of the field of view 620 (or the environment captured by camera 602) are used for display. For example, user 622 may restrict what portions are available to user 623. In some embodiments, the boundary is defined by the physical limitations (e.g., image sensor and / or lens) of the camera 602 providing the field of view 620. Figure 6AR At this point, the shadow area 6217 has not yet reached the limit of the field of view 620. Thus, the boundary indicator 6226 provides a configurable setting based on the limitation of which portion of the field of view 620 is used for display. (Temporarily switch to...) Figure 6AT In response to determining that the viewing angle provided in the adjustable view 6214 has reached the edge of the field of view 620, a boundary indicator 6226 is displayed.
[0356] exist Figure 6AR At this location, the boundary indicator 6226 is depicted with crosshairs. In some embodiments, the safety boundary indicator 6226 is a visual effect (e.g., blurring and / or fading) applied to the adjustable view 6214 (and / or adjustable view 6216). In some embodiments, the boundary indicator 6226 is displayed along the edge of the adjustable view 6214 (and / or 6216) to indicate the location of the boundary. Figure 6AR At this location, the boundary indicator 6226 is displayed along the top and side edges to indicate that the user cannot see above and / or further to the side of the boundary indicator 6226. When in Figure 6AR When displaying interface 604-2, device 600-2 detects movement 6218ar (e.g., rotation and / or lowering). In response to detecting movement 6218ar, device 600-2 displays interface 604-2, as shown below. Figure 6AS As depicted in the text. In response to the detection of movement 6218ar, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as shown in the text. Figure 6AS As depicted in the text.
[0357] exist Figure 6AS At this point, interface 604-2 includes an updated adjustable view 6214, which includes an image of drawing 618. Figure 6AS At this location, equipment 600-2 is located in relation to equipment 600-2. Figure 6AO In similar positions. Thus, Figure 6AS The adjustable view 6214 includes... Figure 6AO The adjustable view 6214 has the same viewing angle as the image of drawing 618 in the adjustable view 6214. It is worth noting that device 600-2 in... Figure 6AS The adjustable view 6214 displays the digital sun 6222. Figure 6AS The position of the sun 6222 in the diagram is similar to that of the house in drawing 618. Figure 6AQ The position of the digital sun 6222 relative to the house in drawing 618 is shown, except for slight differences based on different views. In this way, the digital sun 6222 appears fixed in physical space, as if it were drawn next to drawing 618. Fixing the position of the digital marker in physical space facilitates better collaboration between users, because users can digitally draw or write in the view, move the device to see different views, and then move the device back to re-display the digital drawing or text and the background in which it was created.
[0358] For clarity, it has been... Figures 6AS to 6AU The shaded areas 6217 and 6206, as well as the field of view 620, are omitted. In some embodiments, 622-1 and adjustable views 6214 and 6216 correspond to... Figure 6AO The shaded areas 6217 and 6206 and the view associated with the field of view 620.
[0359] exist Figure 6AS At this point, device 600-2 (and / or device 600-1) detects movement of drawing 618 and maintains the image of drawing 618 displayed in adjustable view 6214. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to modify (e.g., scale, skew, and / or rotate) image data to maintain the image of drawing 618 displayed in adjustable view 6214. When displaying interface 604-2, device 600-2 (and / or device 600-1) detects horizontal movement 6230 of drawing 618. In response to detecting horizontal movement 6230 of drawing 618, device 600-2 displays interface 604-2, as shown... Figure 6AT As depicted herein. In some embodiments, in response to detecting horizontal movement 6230 of drawing 618, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-2, as shown in the image. Figure 6AT As depicted herein. In some embodiments, in response to device 600-1 detecting horizontal movement 6230 of drawing 618, device 600-2 displays (and / or device 600-1 causes device 600-2 to display) interface 602-4, such as Figure 6AT As depicted in the text.
[0360] exist Figure 6AT At this point, drawing 618 has been moved to the edge of desk 621, which is further away from camera 602 (e.g., and closer to the side of the camera). Regardless of the change in position, Figure 6ATThe interface 602-4 includes an image of drawing 618 in an adjustable view 6214, which looks similar to Figure 6AS The image of drawing 618 in the adjustable view 6214 of interface 602-4 remains almost unchanged compared to the previous view. For example, the adjustable view 6214 provides a perspective that makes it appear as if drawing 618 is still directly in front of camera 602, similar to... Figure 6AS The position of drawing 618 in the image. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct (e.g., by skew and / or magnification) the image of drawing 618 based on the new position relative to camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it moves relative to camera 602. In some embodiments, providing Figure 6AT The adjustable view 6214 of the interface 604-2 does not change the position (e.g., positioning, orientation, and / or rotation) of the camera 602.
[0361] exist Figure 6AT At this point, device 600-2 displays boundary indicator 6226 in adjustable view 6216 (similar to adjustable view 6214 displayed by device 600-1 in adjustable view 6214). As described above relative to Figure 6AR The boundary indicator 6226 discussed indicates that the limits of the field of view or physical space have been reached. Figure 6AT At this point, device 600-2 displays a boundary indicator 6226 in adjustable view 6214 to indicate that the edge of the field of view 620 has been reached. The boundary indicator 6226 is along the right edge of adjustable view 6214 (and adjustable view 6216), indicating that the view to the right of the current view is outside the field of view of camera 602.
[0362] exist Figure 6AT At this point, the house in the image of drawing 618 in the adjustable view 6214 remains relative to the digital sun 6222. Figure 6AS The corresponding position of the house in the image of drawing 618 in the adjustable view 6214 is similar to the corresponding position of the house. In some embodiments, device 600-2 (and / or device 600-1) displays a digital sun 6222 overlaid on an image of drawing 618 that has been corrected based on the new position of drawing 618.
[0363] Return to temporarily Figure 6AS When displaying interface 602-4, device 600-2 (and / or device 600-1) detects rotation 6232 of drawing 618. In response to detecting rotation 6232 of drawing 618, device 600-2 displays interface 604-2, as shown. Figure 6AUAs depicted herein. In some embodiments, in response to detecting a rotation 6232 of the drawing 618, device 600-2 causes device 600-1 to display interface 601-4, as shown in the image. Figure 6AU As depicted herein. In some embodiments, in response to device 600-1 detecting a rotation 6232 of drawing 618, device 600-2 displays (or device 600-1 causes device 600-2 to display) interface 602-4, as shown in the image. Figure 6AU As depicted in the text.
[0364] exist Figure 6AU At this point, drawing 618 has been rotated relative to the edge of desk 621. Regardless of the change in position, Figure 6AU The interface 602-4 includes an image of drawing 618 in an adjustable view 6214, which looks similar to Figure 6AS The image of drawing 618 in the adjustable view 6214 of interface 604-2 is almost unchanged compared to that in the interface. That is, Figure 6AU The adjustable view 6214 provides a perspective that makes it appear as if drawing 618 has not been rotated, similar to... Figure 6AS The position of drawing 618 in the image. In some embodiments, device 600-2 (and / or device 600-1) uses image correction software to correct (e.g., by skew and / or rotation) the image of drawing 618 based on the new position relative to camera 602. In some embodiments, device 600-2 (and / or device 600-1) uses object detection software to track drawing 618 as it rotates relative to camera 602. In some embodiments, providing Figure 6AU The adjustable view 6214 of interface 604-2 remains unchanged while the position (e.g., positioning, orientation, and / or rotation) of camera 602 is not altered. Adjustable view 6216 is updated in a similar manner to adjustable view 6214.
[0365] exist Figure 6AU At this point, the house in the image of drawing 618 in the adjustable view 6214 remains relative to the digital sun 6222. Figure 6AS The house is positioned similarly to the house in the image of drawing 618 in the adjustable view 6214. In some embodiments, device 600-2 (and / or device 600-1) displays a digital sun 6222 overlaid on an image of drawing 618 that has been rotated and corrected based on drawing 618.
[0366] exist Figure 6AV At this location, device 600-2 displays interface 604-2, which is similar to... Figure 6AUThe interface 604-2 has different states (e.g., John's representation 622-2 and option menu 609 have been added to the user interface 604-2). Device 600-4 is no longer used in real-time communication sessions. Additionally, device 600-2 has been removed from its... Figure 6AU The position in the middle is moved to device 600-2. Figure 6AQ The same location is present in the middle. Thus, device 600-2 is updated. Figure 6AV Adjustable view 6214 to include with Figure 6AQ The adjustable view 6214 has the same viewing angle. As shown, the adjustable view 6214 includes a top view perspective. Additionally, the digital sun 6222 is displayed with the digital sun 6222 relative to... Figure 6AQ The same location of the house in the image of drawing 618 in the adjustable view 6214.
[0367] exist Figure 6AV When the digital sun 6222 is displayed in the adjustable view 6214, device 600-2 detects movement 6218av (e.g., rotation and / or lifting). In response to the detected movement 6218av, device 600-2 displays interface 604-2, as shown... Figure 6AW As depicted in the text. In response to the detection of movement 6218aw, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown. Figure 6AW As depicted in the text.
[0368] exist Figure 6AW At this point, interface 604-2 includes something similar to Figure 6AR The adjustable view 6214 is an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). It is worth noting that the device does not update representation 622-2 in response to the detection of movement 6218aw. Therefore, in some embodiments, device 600-2 displays a dynamic representation based on updates to the location of device 600-2 and a static representation not based on updates to the location of device 600-2. Interface 604-2 also includes a boundary indicator 6226 in the adjustable view 6214, similar to... Figure 6AR The boundary indicator is 6226.
[0369] exist Figure 6AW At the location of display interface 604-2, device 600-2 detects movement 6218aw (e.g., rotation and / or lowering). In response to detecting movement 6218aw, device 600-2 displays interface 604-2, as shown below. Figure 6AX As depicted in the text. In response to the detection of movement 6218aw, device 600-1 displays (and / or device 600-2 causes device 600-1 to display) interface 604-1, as shown. Figure 6AXAs depicted in the text.
[0370] exist Figure 6AX At this point, interface 604-2 includes an updated adjustable view 6214 (which corresponds to the updated adjustable view 6216 in interface 606-1). Because the adjustable view 6214 is essentially the same view provided by representation 622-2, the shaded area 6206 overlaps with the shaded area 6217. Because movement 6218aq causes the view to change to the view of user 622's face and / or the view of non-drawing 618, device 600-2 no longer displays the digital sun 6222 in the adjustable view 6214. When in Figure 6AX When displaying interface 604-2, device 600-2 (and / or device 600-1) detects one or more sets of inputs (e.g., similar to reference) corresponding to a request for a view of the display surface. Figures 6A to 6AL The input and / or gestures mentioned above. In some such embodiments, it is displayed at device 600-2. Figure 6C 616-1 Figure 6G 616-2, Figure 6H Preview mode 674-1 Figure 6I In the preview mode 674-2, the representation is 676. Figure 6J In the preview mode interface 674-3, 676 is represented. Figures 6N to 6Q The device displays 648-1, 648-2, and 648-3 to allow device 600-2 to control the representation of the modified image of drawing 618 in the same manner as the input detected at device 600-1. In response to the detection of one or more inputs corresponding to a request to display a surface view, device 600-2 displays interface 604-2, such as... Figure 6AY As depicted herein. Additionally or alternatively, in response to the detection of one or more of this set of inputs, device 600-1 displays interface 604-2, as shown in the image. Figure 6AY As depicted in [reference]. In some implementations, device 600-1 detects one or more of this set of inputs, as shown in [reference]. Figures 6A to 6AL As described above. In some embodiments, device 600-2 detects one or more of the group of inputs. In such embodiments, device 600-2 detects a selection of a view display representation 6236 of option menu 609, the view display representation of which is similar to a reference. Figure 6F The view display of the option menu 608 is shown in 607-2. In response, with reference... Figure 6G The view menu 616-2, similar to the view menu mentioned above, includes an indication that requests to display a surface view of a remote participant.
[0371] exist Figure 6AY At this location, the adjustable view 6214 includes a surface view, which is similar to, for example... Figure 6M The descriptions in the text and the above-mentioned representations are as follows: 624-1. Figure 6AY The adjustable view 6214 depicted includes an image modified to give user 623 a downward-viewing perspective of the image of drawing 618 displayed on device 600-2, similar to the perspective user 622 has when looking down at drawing 618 in a physical environment, such as relative to... Figures 6A to 6AL As described in more detail. It is worth noting that... Figure 6AY The digital sun 6222 is displayed as having the same Figure 6AQ The digital sun 6222 is the same as the position of the house in the image of drawing 618 in the adjustable view 6214.
[0372] Figure 7 This is a flowchart illustrating a method for managing real-time video communication sessions using a computer system, according to some implementation schemes. Method 700 is performed at a computer system (e.g., 600-1, 600-2, 600-3, 600-4, 906a, 906b, 906c, 906d, 6100-1, 6100-2, 1100a, 1100b, 1100c and / or 1100d) (e.g., a smartphone, tablet, laptop computer and / or desktop computer) (e.g., 100, 300 or 500), the computer system being coupled with display generating components (e.g., 601, 683 and / or 6101) (e.g., a display controller, touch-sensitive display system and / or monitor), one or more cameras (e.g., 602, 682 and / or 6102) (e.g., an infrared camera, a depth camera and / or a visible light camera) and one or more input devices (e.g., 601, 683 and / or 6103) (e.g., a touch-sensitive surface, a keyboard, a controller and / or a mouse). Some operations in method 700 may be arbitrarily combined, the order of some operations may be arbitrarily changed, and some operations may be arbitrarily omitted.
[0373] As described below, method 700 provides an intuitive way to manage real-time video communication sessions. This method reduces the cognitive burden on users managing real-time video communication sessions, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to manage real-time video communication sessions faster and more efficiently, saving power and increasing the time interval between battery charges.
[0374] In method 700, the computer system (e.g., 600-1, 600-2, 6100-1 and / or 6100-2) displays (702) a real-time video communication interface (e.g., 604-1, 604-2, 6120, 6121, 6130 and / or 6131) for a real-time video communication session via a display generation component (e.g., an interface for incoming and / or outgoing real-time audio / video communication sessions). In some embodiments, the real-time communication session is at least between the computer system (e.g., a first computer system) and a second computer system.
[0375] The real-time video communication interface includes a representation (e.g., 622-1, 622-2, 6124, and / or 6132) of at least a portion of the field of view (e.g., 620, 688, 6144, 6146, and / or 6148) of the one or more cameras (e.g., a first representation). In some embodiments, the first representation includes an image of the physical environment (e.g., a scene and / or region of the physical environment within the field of view of the one or more cameras). In some embodiments, the representation includes a portion of the field of view of the one or more cameras (e.g., a first cropped portion). In some embodiments, the representation includes still images. In some embodiments, the representation includes a series of images (e.g., video). In some embodiments, the representation includes a real-time (e.g., instantaneous) video feed of the field of view (or a portion thereof) of the one or more cameras. In some embodiments, the field of view is based on the physical characteristics of the one or more cameras (e.g., orientation, lens, lens focal length, and / or sensor size). In some embodiments, the representation is displayed in a window (e.g., a first window). In some embodiments, the representation of at least a portion of the field of view includes an image of a first user (e.g., the face of the first user). In some embodiments, the representation of at least a portion of the field of view is provided by an application providing a real-time video communication session (e.g., 6110). In some embodiments, the representation of at least a portion of the field of view is provided by an application different from the application providing the real-time video communication session (e.g., 6110) (e.g., 6108).
[0376] When displaying a real-time video communication interface, the computer system (e.g., 600-1, 600-2, 6100-1 and / or 6100-2) detects (704) user input (e.g., 612c, 612d, 614, 612g, 612i, 612j, 6112, 6118, 6128 and / or 6138) via one or more input devices (e.g., taps on a touch-sensitive surface, keyboard input, mouse input, touchpad input). The user input refers to one or more user inputs, including gestures (e.g., hand gestures) and / or audio inputs (e.g., voice commands), pointing to a surface (e.g., 619) in a scene (e.g., physical environment) within the field of view of the one or more cameras (e.g., a physical surface; the surface of a desk and / or the surface of an object placed on the desk (e.g., a book, paper, tablet); or the surface of a wall and / or the surface of an object on the wall (e.g., a whiteboard or blackboard); or other surfaces (e.g., a freestanding whiteboard or blackboard)). In some embodiments, the user input corresponds to a request for a view of the displayed surface. In some embodiments, detecting the user input via the one or more input devices includes obtaining image data including gestures (e.g., hand gestures, eye gestures, or other body postures) within the field of view of the one or more cameras. In some embodiments, the computer system determines from the image data that the gesture meets predetermined criteria.
[0377] In response to the detection of one or more user inputs, the computer system (e.g., 600-1, 600-2, 6100-1 and / or 6100-2) displays a representation (e.g., image and / or video, e.g., a second representation) of a surface (e.g., 624-1, 624-2, 6140 and / or 6142) via a display generation component (e.g., 601, 683 and / or 6101). In some embodiments, the representation of the surface is obtained by digitally scaling and / or panning the field of view captured by the one or more cameras. In some embodiments, the representation of the surface is obtained by moving (e.g., panning and / or rotating) the one or more cameras. In some embodiments, the second representation is displayed in a window (e.g., a second window, the same window in which the first representation is displayed, or a window different from the window in which the first representation is displayed). In some embodiments, the second window is different from the first window. In some embodiments, the second window (e.g., 6140 and / or 6142) is provided by a real-time video communication session (e.g., as shown in the image). Figure 6AJ The application (e.g., 6110) shown in the diagram provides the second window. In some implementations, the second window (e.g., 6114) is provided by an application different from the one providing the real-time video communication session (e.g., such as...). Figure 6AKThe application (e.g., 6108) shown in the diagram is provided. In some embodiments, the second representation includes a cropped portion of the field of view of the one or more cameras (e.g., a second cropped portion). In some embodiments, the second representation differs from the first representation. In some embodiments, the second representation differs from the first representation because the second representation displays a portion of the field of view that is different from a portion (e.g., a first cropped portion) displayed in the first representation (e.g., a pan view, a zoomed-out view, and / or a magnified view). In some embodiments, the second representation includes an image of a portion of the scene not included in the first representation and / or the first representation includes an image of a portion of the scene not included in the second representation. In some embodiments, the surface is not displayed in the first representation.
[0378] The representation of the surface (e.g., 624-1, 624-2, 6140, and / or 6142) includes an image (e.g., a photograph, video, and / or live video feed) of the surface (e.g., 619) captured by the one or more cameras (e.g., 602, 682, and / or 6102) that has been modified (e.g., to correct image distortion of the surface) based on the surface's position (e.g., positioning and / or orientation) relative to the one or more cameras (e.g., adjustment, manipulation, correction) (sometimes referred to as a modified image representation of the surface). In some embodiments, the image of the surface displayed in the second representation is based on image data modified using image processing software (e.g., skewed, rotated, flipped, and / or otherwise manipulated image data captured by the one or more cameras). In some embodiments, the image of the surface displayed in the second representation is modified without physically adjusting the camera (e.g., without rotating the camera, raising the camera, lowering the camera, adjusting the camera angle, and / or adjusting the camera's physical components (e.g., the lens and / or sensor)). In some embodiments, the image of the surface displayed in the second representation is modified such that the camera appears to be pointing at the surface (e.g., facing the surface, aiming at the surface, pointing along an axis perpendicular to the surface). In some embodiments, the image of the surface displayed in the second representation is corrected such that the camera's line of sight appears perpendicular to the surface. In some embodiments, the image of the scene displayed in the first representation is not modified based on the position of the surface relative to the one or more cameras. In some embodiments, the representation of the surface is displayed simultaneously with the first representation (e.g., maintaining the first representation (e.g., for the user of the computer system), and the image of the surface is displayed in a separate window). In some embodiments, the image of the surface is modified automatically on the fly (e.g., during a real-time video communication session). In some embodiments, the image of the surface is modified automatically based on the position of the surface relative to the one or more first cameras (e.g., without user input). Displaying a representation of the surface that includes an image of the surface modified based on the position of the surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the surface regardless of its position relative to the camera without requiring further input from the user (which provides improved visual feedback and reduces the amount of input required to perform actions).
[0379] In some implementations, a computer system (e.g., 600-1 and / or 600-2) receives image data captured by one of the one or more cameras (e.g., 602) (e.g., a wide-angle camera) during a real-time video communication session. The computer system displays a representation of at least a portion of the field of view (e.g., 622-1 and / or 622-2) (e.g., a first representation) via a display generation component based on the image data captured by the camera. The computer system displays a representation of a surface (e.g., 624-1 and / or 624-2) (e.g., a second representation) via a display generation component based on the image data captured by the camera (e.g., a representation of at least a portion of the field of view of the one or more cameras, and the representation of the surface is based on image data captured by a single (e.g., only one) camera among the one or more cameras). Displaying representations of at least a portion of the field of view captured from the same camera and representations of surfaces enhances the video communication session experience by displaying content captured by the same camera from different perspectives without user input (which reduces the amount of input (and / or device) required to perform operations).
[0380] In some embodiments, the image of the surface is modified (e.g., by a computer system) by rotating an image of the surface relative to at least a portion of the field of view of the one or more cameras (e.g., the image of the surface in 624-2 is rotated 180 degrees relative to representation 622-2). In some embodiments, the representation of the surface is rotated 180 degrees relative to at least a portion of the field of view of the one or more cameras. Rotating the image of the surface relative to at least a portion of the field of view of the one or more cameras enhances the video communication session experience because the content associated with the surface can be viewed from a different perspective than the rest of the field of view without user input, providing improved visual feedback and reducing the amount of input required to perform operations.
[0381] In some implementations, the image of the surface (e.g., 619) is rotated based on the position (e.g., orientation and / or orientation) of the user (e.g., 622) (e.g., the user's position) relative to the field of view of the one or more cameras. In some implementations, the user's representation is displayed at a first angle and the image of the surface is rotated to a second angle different from the first angle (e.g., even if the image of the user and the image of the surface are captured at the same camera angle). Rotating the image of the surface based on the user's position relative to the one or more cameras enhances the video communication session experience because content associated with the surface can be viewed from a perspective based on the surface's position without user input, providing improved visual feedback and reducing the amount of input required to perform operations.
[0382] In some implementations, the surface is determined to be in a first position of the user relative to the field of view of the one or more cameras (e.g., surface 619 is positioned at...). Figure 6A In a location (e.g., in front of the user 622 on desk 621, at a predefined position) (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane), the image of the surface is rotated at least 45 degrees relative to the representation of the user in the field of view of the one or more cameras (e.g., the image of surface 619 in 624-1 is rotated relative to the user's representation in the field of view of the one or more cameras). Figure 6M The notation in 622-1 indicates a 180-degree rotation. In some embodiments, the image of the surface is rotated within a range of 160 to 200 degrees (e.g., a 180-degree rotation). In some embodiments, the image of the surface is rotated by a first amount based on determining that the surface is in a first position of the user relative to the one or more cameras in the field of view (e.g., in front of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane). In some embodiments, the first amount is in the range of 160 to 200 degrees (e.g., 180 degrees). In some embodiments, the image of the surface is rotated by a second amount based on determining that the surface is in a second position of the user relative to the one or more cameras in the field of view (e.g., to the side of the user, between the user and the one or more cameras, and / or in a substantially horizontal plane). In some embodiments, the second amount is in the range of 45 to 120 degrees (e.g., 90 degrees). When the surface is in a first position relative to the user, the image of the surface is rotated at least 45 degrees relative to the user's representation captured in the field of view of the one or more cameras. This adjusts the image to provide a more natural and intuitive picture without further input from the user, which enhances the video communication session experience by providing improved visual feedback and performing actions when a set of conditions are met without further user input.
[0383] In some implementations, the representation of at least a portion of the field of view includes the user and is related to the representation of the surface (e.g., Figure 6M The representations 622-1 and 624-1 or 622-2 and 624-2 are displayed simultaneously. In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are captured by the same camera (e.g., a single camera among the one or more cameras) and displayed simultaneously. In some embodiments, the representation of at least a portion of the field of view and the representation of the surface are displayed in separate windows that are displayed simultaneously. Including the user in the representation of at least a portion of the field of view and displaying the representation simultaneously with the representation of the surface enhances the video communication session experience by allowing the user to see the participant's reaction while the representation of the surface is displayed without further input from the user (this provides improved visual feedback and allows actions to be performed when a set of conditions are met without further user input).
[0384] In some implementations, in response to detecting the one or more user inputs and before their representation on the display surface, the computer system displays a preview of image data of the field of view of the one or more cameras (e.g., ...). Figures 6H to 6J (As depicted herein) (e.g., in preview mode of a real-time video communication interface), the preview includes an image of the surface that is not modified based on the surface's position relative to the one or more cameras (sometimes referred to as a representation of the surface's unmodified image). In some embodiments, a preview of the field of view is displayed after the image representation of the surface is displayed (e.g., in response to detecting user input corresponding to a selection of the surface representation). Displaying a preview including an image of the surface that is not modified based on the surface's position relative to the one or more cameras allows the user to quickly identify the surface within the preview, as distortion correction has not yet been applied, providing improved visual feedback.
[0385] In some implementations, a preview of image data displaying the field of view of the one or more cameras includes multiple selectable options for displaying a corresponding portion of the field of view (e.g., a surface within it) of the one or more cameras. Figure 6I 636-1 and / or 636-2, or Figure 6J (636a-i). In some embodiments, the computer system detects input selecting one of a plurality of options corresponding to a corresponding portion of the field of view of the one or more cameras (e.g., 612i or 612j). In response to detecting input selecting one of a plurality of options corresponding to a corresponding portion of the field of view of the one or more cameras and determining that the input selecting one of a plurality of options corresponding to a corresponding portion of the field of view of the one or more cameras points to a first option corresponding to a first portion of the field of view of the one or more cameras, the computer system displays a representation of a surface based on the first portion of the field of view of the one or more cameras (e.g., for...). Figure 6J The selection of 636h in the image causes the corresponding portion to be displayed (e.g., the computer system displays a modified version of an image of a first portion of the field of view, optionally corrected for a first distortion). In response to detecting an input selecting one of the plurality of options corresponding to a corresponding portion of the field of view of the one or more cameras, and based on determining that the input selecting one of the plurality of options corresponding to a corresponding portion of the field of view of the one or more cameras points to a second option corresponding to a second portion of the field of view of the one or more cameras, the computer system displays a representation of the surface based on the second portion of the field of view of the one or more cameras (e.g., for...). Figure 6JThe selection of 636g in the image allows the display of the corresponding portion (e.g., a modified version of the image of the second portion of the field of view, optionally with a second distortion correction different from the first distortion correction), where the second option differs from the first option. Displaying multiple selectable options corresponding to the respective portions of the field of view of the one or more cameras in the preview of the image data allows users to identify the portions of the field of view that can be displayed as a representation in the video conferencing interface, providing improved visual feedback.
[0386] In some implementations, displaying a preview of image data showing the field of view of the one or more cameras includes displaying a preview (e.g., Figure 6I 636-1, 636-2 and / or Figure 6J Multiple regions (e.g., different regions, non-overlapping regions, rectangular regions, square regions, and / or quadrants) of 636a-i) (e.g., the one or more regions may correspond to different portions of the image data of the field of view). In some embodiments, the computer system detects user input corresponding to one or more of the multiple regions (e.g., 612i and / or 612j). In response to detecting user input corresponding to the one or more regions and determining that the user input corresponding to the one or more regions corresponds to a first region among the one or more regions, the computer system displays a representation of the first region in a real-time video communication interface (e.g., as shown in reference). Figures 6I to 6J (e.g., with distortion correction based on a first region). In response to detecting user input corresponding to the one or more regions and determining that the user input corresponding to the one or more regions corresponds to a second region within the one or more regions, the computer system displays a representation of the second region as a representation in the real-time video communication interface (e.g., with distortion correction based on the second region, which differs from distortion correction based on the first region). Displaying a representation of the first region or a representation of the second region in the real-time video communication interface enhances the video communication session experience by allowing users to effectively manage the content displayed in the real-time video communication interface (which provides improved visual feedback and reduces the amount of input required to perform operations).
[0387] In some implementations, the one or more user inputs include gestures (e.g., 612d) within the field of view of the one or more cameras (e.g., body poses, hand gestures, head poses, arm poses, and / or eye poses) (e.g., gestures pointing to a physical surface performed within the field of view of the one or more cameras). Utilizing gestures within the field of view of the one or more cameras as input enhances the video communication session experience by allowing users to control the displayed content without physically touching the device (which provides additional control options without cluttering the user interface).
[0388] In some embodiments, the computer system displays surface view options (e.g., 610) (e.g., icons, buttons, indicator lights, and / or user-interactive graphical interface objects), wherein the one or more user inputs include inputs pointing to the surface view options (e.g., 612c and / or 612g) (e.g., tap input on a touch-sensitive surface, a mouse click while the cursor is over the surface view option, or an air gesture while gazing at the surface view option). In some embodiments, the surface view options are displayed in a representation of at least a portion of the field of view of the one or more cameras. Displaying surface view options enhances the video communication session experience by allowing users to effectively manage the content displayed in the real-time video communication interface (providing additional control options without cluttering the user interface).
[0389] In some implementations, the computer system detects user input corresponding to a selection of a surface view option. In response to detecting user input corresponding to a selection of a surface view option, the computer system displays a preview of image data of the field of view of the one or more cameras (e.g., as shown in the image). Figures 6H to 6J (as depicted herein) (e.g., in preview mode of a real-time video communication interface), the preview includes multiple portions of the field of view of the one or more cameras, the multiple portions including at least a portion of the field of view of the one or more cameras (e.g., Figure 6I 636-1, 636-2 and / or Figure 6J (636a-i), where the preview includes the field of view (e.g., Figure 6I 641-1 and / or Figure 6J Visual indications (e.g., text, graphics, icons, and / or colors) of the active portion of the field of view (e.g., the portion being transmitted to and / or displayed by other participants in the live video communication session) of the camera (640a-f). In some embodiments, the visual indications indicate that a single portion (e.g., only one portion) of the multiple portions of the field of view is active. In some embodiments, the visual indications indicate that two or more portions of the multiple portions of the field of view are active. Displaying a preview of multiple portions of the field of view of the one or more cameras (where the preview includes visual indications of the active portion of the field of view) enhances the video communication session experience by providing the user with feedback about which portion of the field of view is active (this provides improved visual feedback).
[0390] In some implementations, the computer system detects user input corresponding to a selection of a surface view option (e.g., 612c, 612d, 614, 612g, 612i, and / or 612j). In response to detecting user input corresponding to a selection of a surface view option, the computer system displays a preview of image data of the field of view of the one or more cameras (e.g., 674-2 and / or 674-3) (e.g., as shown in the image). Figures 6I to 6J (as described in the above) (e.g., in preview mode of a real-time video communication interface), the preview includes multiple selectable, visually distinct portions (e.g., as described in the above) overlaid on a representation of the field of view of the one or more cameras. Figures 6I to 6J (As described in the document). Displaying previews of multiple selectable, visually distinct portions overlaid on the representation of the field of view of the one or more cameras enhances the video communication session experience by providing the user with feedback on which portions of the field of view are selectable for display as a representation during the video communication session (which provides improved visual feedback).
[0391] In some implementations, the surface is a vertical surface in the scene (e.g., such as...). Figure 6J The vertical surface (as depicted in the image) (e.g., a wall, easel, and / or whiteboard) (e.g., the surface is within a predetermined angle (e.g., 5 degrees, 10 degrees, or 20 degrees) in the direction of gravity). Displaying a representation of the vertical surface, including an image of the vertical surface modified based on the position of the vertical surface relative to the one or more cameras, enhances the video communication session experience by providing a clearer view of the vertical surface regardless of its position relative to the camera, without requiring further input from the user (this provides improved visual feedback and reduces the amount of input required to perform actions).
[0392] In some implementations, the surface is a horizontal surface in the scene (e.g., 619) (e.g., a table, floor, and / or desk) (e.g., the surface is within a predetermined angle (e.g., 5 degrees, 10 degrees, or 20 degrees) in a plane perpendicular to the direction of gravity). Displaying a representation of the horizontal surface that includes an image modified based on the position of the horizontal surface relative to the one or more cameras enhances the video communication session experience by providing a clearer view of the horizontal surface regardless of its position relative to the camera, without requiring further input from the user (which provides improved visual feedback and reduces the amount of input required to perform actions).
[0393] In some implementations, the representation of the display surface includes a first view of the display surface (e.g., Figure 6N (e.g., with a first rotation angle and / or a first zoom level). In some embodiments, while displaying a first view of the surface, the computer system displays one or more shift view options (e.g., 648-1 and / or 648-2) (e.g., buttons, icons, indicator lights, and / or user-interactive graphical user interface objects). The computer system detects user input pointing to a corresponding shift view option among the one or more shift view options (e.g., 650a and / or 650b). In response to detecting user input pointing to a corresponding shift view option, the computer system displays a second view of the surface that differs from the first view of the surface (e.g., ...). Figure 6P 624-1 and / or Figure 6Q (e.g., a second rotation angle different from the first rotation angle and / or a second zoom level different from the first zoom level) (e.g., shifting the view of the surface from the first view to the second view). Providing a shift view option to display a second view of the surface currently displayed at the first view of the surface enhances the video communication session experience by allowing users to view the content associated with the surface from different perspectives (which provides additional control options without cluttering the user interface).
[0394] In some implementations, displaying a first view of the surface includes displaying an image of the surface modified in a first manner (e.g., such as...). Figure 6N (as depicted in the image) (e.g., with first distortion correction applied), and wherein the second view showing the surface includes an image showing the surface modified in a second manner different from the first manner (e.g., as shown in the image). Figure 6P and / or Figure 6Q (As depicted in the text) (e.g., a second distortion correction is applied) (e.g., the computer system alters (e.g., shifts) the distortion correction applied to the image of the surface based on the view of the surface to be displayed (e.g., orientation and / or scaling). Displaying an image of the surface modified in the first manner and displaying a second view of the surface, including displaying an image of the surface modified in the second manner, enhances the video communication session experience by allowing users to automatically view the modified content without further input from the user (which provides improved visual feedback and reduces the amount of input required to perform operations).
[0395] In some implementations, a representation of the surface is displayed at a first zoom level (e.g., as shown in the image). Figure 6N (As depicted in the text). In some embodiments, while displaying a representation of the surface at a first zoom level, the computer system detects user input (e.g., 650b and / or 654) corresponding to a request to change the zoom level of the surface representation (e.g., selection of zoom options (e.g., buttons, icons, display indicators, and / or user interface elements)). In response to detecting user input corresponding to a request to change the zoom level of the surface representation, the computer system displays the surface representation at a second zoom level different from...
Claims
1. A method for displaying a user interface, comprising: At a computer system in communication with a display generating component, one or more cameras, and one or more input devices: displaying, via the display generating component, a real-time video communications interface for a real-time video communications session, the real-time video communications interface including a representation of at least a portion of a field of view of the one or more cameras; while displaying the real-time video communications interface, detecting via the one or more input devices one or more user inputs, the one or more user inputs comprising a user input corresponding to a request to display a view of a surface in a scene in a field of view of a first camera of the one or more cameras; as well as In response to detecting the one or more user inputs and when the first camera captures image data of the scene in the field of view of the first camera, simultaneously displaying via the display generation component: a representation of the surface, wherein the representation of the surface comprises an image of the surface based on the image data captured by the first camera of the one or more cameras, wherein the image of the surface is modified based on a position of the surface relative to the one or more cameras; as well as a representation of at least a portion of the scene based on the image data captured by the first camera of the one or more cameras, unmodified based on the position of the surface relative to the one or more cameras, wherein: modifying the image of the surface by rotating the image of the surface relative to the representation of at least a portion of the scene, the representation of at least a portion of the scene being based on the image data captured by the first camera of the one or more cameras and not modified based on the position of the surface relative to the first camera of the one or more cameras; as well as The image of the surface is rotated based on a position of the surface relative to a user in the field of view of the first camera of the one or more cameras.
2. The method according to claim 1, wherein: the representation of at least a portion of the field of view of the one or more cameras corresponding to a representation of a first portion of a field of view of the first camera of the one or more cameras; detecting, while displaying the real-time video communications interface, a user input corresponding to a request to display a view of the surface in a scene in a field of view of a first camera of the one or more cameras, and wherein the user input corresponding to a request to display a view of the surface in the scene in the field of view of the first camera of the one or more cameras corresponds to a first request to display a view of the surface in the scene in the field of view of the first camera; and The method further comprises: in response to detecting a user input corresponding to a first request to display a view of the surface in the scene in the field of view of the first camera and when the first camera captures the image data of the scene in the field of view of the first camera, displaying, via the one or more display generation components, a representation of a second portion of the field of view of the first camera, the representation of the second portion of the field of view of the first camera being based on the image data captured by the first camera, wherein the representation of the second portion of the field of view of the first camera includes a portion of the first portion of the field of view of the first camera that was not included in the representation of the first camera; and detecting, via the one or more input devices, a user input corresponding to a second request to display a view of the surface while the representation of the second portion of the field of view of the first camera is displayed and the first camera captures the image data of the scene in the field of view of the first camera, wherein: in response to a user input corresponding to a second request to display a view of the surface, displaying a representation of the surface, the representation of the surface comprising an image of the surface modified based on image data captured by the first camera of the one or more cameras and based on a position of the surface relative to the first camera of the one or more cameras; and The representation of the surface comprising an image of the surface modified based on image data captured by the first camera of the one or more cameras and based on a position of the surface relative to the first camera of the one or more cameras includes a third portion of the field of view of the first camera that is different from the first portion of the field of view of the first camera and the second portion of the field of view of the first camera.
3. The method according to any one of claims 1 to 2, wherein: Based on determining that the surface is in a first position relative to a user in the field of view of the first camera of the one or more cameras, the image of the surface is rotated at least 45 degrees relative to the representation of the user in the field of view of the first camera of the one or more cameras.
4. The method of any one of claims 1-2, wherein the representation of the at least a portion of the scene based on the image data captured by the first camera includes a user and is displayed simultaneously with the representation of the surface.
5. The method according to any one of claims 1 to 2, further comprising: In response to detecting the one or more user inputs: Prior to displaying the representation of the surface, a preview of image data captured by the first camera of the one or more cameras is displayed, the preview comprising an image of the surface that is not modified based on the position of the surface relative to the one or more cameras.
6. The method of claim 5, wherein displaying a preview of the image data of the field of view of the one or more cameras comprises displaying a plurality of selectable options corresponding to respective portions of the field of view of the first camera of the one or more cameras, further comprising: detecting an input selecting one of the plurality of selectable options corresponding to a respective portion of the field of view of the first of the one or more cameras; as well as In response to detecting the input selecting one of the plurality of selectable options corresponding to a respective portion of the field of view of the first of the one or more cameras: displaying the representation of the surface based on the first portion of the field of view of the first camera of the one or more cameras based on determining that the input selecting one of the plurality of selectable options corresponding to the respective portion of the field of view of the first camera of the one or more cameras is directed to a first option corresponding to a first portion of the field of view of the first camera of the one or more cameras; as well as Based on determining that the input selecting one of the multiple selectable options corresponding to the corresponding portion of the field of view of the first camera of the one or more cameras is directed to a second option corresponding to a second portion of the field of view of the first camera of the one or more cameras, the representation of the surface is displayed based on the second portion of the field of view of the first camera of the one or more cameras, wherein the second option is different from the first option.
7. The method of claim 5, wherein displaying a preview of the image data captured by the first camera of the one or more cameras comprises displaying a plurality of regions of the preview, further comprising: detecting user input corresponding to one or more of the plurality of regions; as well as In response to detecting the user input corresponding to the one or more regions: based on determining that the user input corresponding to the one or more regions corresponds to a first region among the one or more regions, displaying a representation of the first region in the real-time video communications interface; as well as Based on determining that the user input corresponding to the one or more regions corresponds to a second region of the one or more regions, a representation of the second region is displayed as a representation in the real-time video communication interface.
8. The method of any one of claims 1-2, wherein the one or more user inputs comprise gestures in the field of view of the one or more cameras.
9. The method according to any one of claims 1 to 2, further comprising: Surface view options are displayed, wherein the one or more user inputs include input directed to the surface view options.
10. The method according to claim 9, further comprising: detecting a user input corresponding to a selection of the surface view option; as well as In response to detecting the user input corresponding to a selection of the surface view option, displaying a preview of the image data captured by the first camera of the one or more cameras, the preview comprising multiple portions of the field of view of the first camera of the one or more cameras, the multiple portions comprising the at least a portion of the scene based on the image data captured by the first camera, wherein the preview includes a visual indication of an active portion of the field of view.
11. The method according to claim 9, further comprising: detecting a user input corresponding to a selection of the surface view option; as well as In response to detecting the user input corresponding to a selection of the surface view option, displaying a preview of the image data captured by the first camera of the one or more cameras, the preview comprising a plurality of selectable visually distinguishable portions overlaid on a representation of the field of view of the first camera of the one or more cameras.
12. The method of any one of claims 1-2, wherein displaying the representation of the surface comprises displaying a first view of the surface, further comprising: while displaying the first view of the surface, displaying one or more shifted view options; detecting a user input directed to a corresponding shift view option of the one or more shift view options; as well as In response to detecting the user input directed to the corresponding shifted view option, a second view of the surface different from the first view of the surface is displayed.
13. The method of claim 12, wherein displaying the first view of the surface comprises displaying an image of the surface modified in a first manner, and wherein displaying the second view of the surface comprises displaying an image of the surface modified in a second manner different from the first manner.
14. The method of any of claims 1-2, wherein displaying the representation of the surface at a first zoom level further comprises: while displaying the representation of the surface at the first zoom level, detecting user input corresponding to a request to change a zoom level of the representation of the surface; as well as In response to detecting the user input corresponding to a request to change a zoom level of the representation of the surface, the representation of the surface is displayed at a second zoom level different than the first zoom level.
15. The method according to any one of claims 1-2, further comprising: while displaying the real-time video communications interface, displaying a selectable control option that, when selected, causes display of the representation of the surface, Wherein the one or more user inputs include a user input corresponding to a selection of the selectable control option.
16. The method of claim 15, wherein: The real-time video communication session is provided by a first application running at the computer system; and The selectable control option is associated with a second application different from the first application.
17. The method according to claim 16, further comprising: In response to detecting the one or more user inputs, displaying a user interface of the second application, wherein the one or more user inputs include the user input corresponding to a selection of the selectable control option.
18. The method according to claim 16, further comprising: The user interface of the second application is displayed before the real-time video communication interface for the real-time video communication session is displayed.
19. The method according to any one of claims 1-2, wherein: providing the real-time video communication session using a third application running at the computer system; and The representation of the surface is provided by a fourth application different from the third application.
20. The method of claim 19, wherein the representation of the surface is displayed using a user interface of the fourth application displayed in the real-time video communications session.
21. The method of claim 19, further comprising: Displaying, via the display generation component, a graphical element corresponding to the fourth application includes displaying the graphical element in an area having a group of one or more graphical elements corresponding to applications other than the fourth application.
22. The method of any one of claims 1-2, wherein displaying the representation of the surface comprises displaying, via the display generation component, an animation of a transition from the display of the representation of at least a portion of a field of view of the one or more cameras to the display of the representation of the surface.
23. The method of any one of claims 1-2, wherein the computer system is in communication with a second computer system, the second computer system is in communication with a second display generating component, further comprising: displaying on the display generating component the representation of at least a portion of the scene based on the image data captured by the first camera; as well as The representation of the surface is caused to be displayed on the second display generating component.
24. The method according to claim 23, further comprising: In response to detecting a change in orientation of the second computer system, updating the display of the representation of the surface displayed at the second display generating component from displaying a first view of the surface to displaying a second view of the surface different from the first view.
25. A method according to any one of claims 1-2, wherein displaying the representation of the surface includes displaying an animation of a transition from the display of the representation of the at least a portion of the field of view of the one or more cameras to the display of the representation of the surface, wherein the animation includes translating the view of the field of view of the one or more cameras and rotating the view of the field of view of the one or more cameras.
26. A method according to any one of claims 1-2, wherein displaying the representation of the surface includes displaying an animation of a transition from the display of the representation of the at least a portion of the field of view of the one or more cameras to the display of the representation of the surface, wherein the animation includes scaling the view of the field of view of the one or more cameras and rotating the view of the field of view of the one or more cameras.
27. The method of any one of claims 1-2, wherein the image of the surface modified based on the position of the surface relative to the first camera of the one or more cameras is an image of the surface corrected so that the line of sight of the first camera appears perpendicular to the surface.
28. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs comprising instructions for executing the method according to any one of claims 1-27.
29. A computer system configured to communicate with a display generating component, one or more cameras, and one or more input devices, the computer system comprising: one or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1-27.
30. A computer system configured to communicate with a display generation component, one or more cameras, and one or more input devices, comprising: Memory; as well as Apparatus for performing the method according to any one of claims 1-27.
31. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more cameras, and one or more input devices, the one or more programs including instructions for executing the method according to any one of claims 1-27.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Acceleration-based theft detection system for portable electronic devices
US20050190059A1
Methods and apparatuses for operating a portable device based on an accelerometer
US20060017692A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1