User interface for managing visual content in media
Through communication between the computer system and the display generation component, the user interface of media content management is simplified, the problem of complex and inefficient interfaces in the prior art is solved, faster and more efficient media content management and user interaction are achieved, and equipment energy is saved.
Patent Information
- Application Number
- CN202510387883.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-10
- Filing Date
- 2022-04-15
- Publication Date
- 2025-07-15
AI Technical Summary
The user interface for managing visual content in the prior art is complex and inefficient, resulting in a waste of user time and device energy, especially in battery-driven devices.
A method for a computer system to communicate with the display generation component is provided, and media representation and media capture display are simultaneously displayed through the camera user interface, corresponding user interface objects are displayed according to the detected text standards, and media capture and management operations are performed in response to user input, simplifying user interaction.
Faster and more efficient media content management is achieved, reducing user cognitive burden, saving device power, extending battery life, and improving user experience.
Smart Images

Figure CN120318808A_ABST
Abstract
Description
[0001] This application is a divisional application of an application with international application number PCT / US2022 / 025096, international filing date of April 15, 2022, date of entry into the Chinese national phase of September 28, 2023, national application number 202280026616.9, and invention title "USER INTERFACE FOR MANAGING VISUAL CONTENT IN MEDIA".
[0002] Cross - Reference to Related Applications
[0003] This application claims the benefit of U.S. Patent Application Serial No. 63 / 176,847, filed April 19, 2021, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA"; U.S. Provisional Patent Application Serial No. 63 / 197,497, filed June 6, 2021, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA"; U.S. Patent Application Serial No. 17 / 484,844, filed September 24, 2021, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA"; U.S. Provisional Patent Application Serial No. 17 / 484,714, filed September 24, 2021, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA"; U.S. Patent Application Serial No. 17 / 484,856, filed September 24, 2021, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA"; and U.S. Provisional Patent Application Serial No. 63 / 318,677, filed March 10, 2022, entitled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA". The content of these patent applications is hereby incorporated by reference in its entirety. TECHNICAL FIELD
[0004] The present disclosure generally relates to computer user interfaces, and more particularly to techniques for managing visual content in media. BACKGROUND ART
[0005] Smartphones and other personal electronic devices allow users to capture and view content in media. Users can capture various types of media, including video and image data. Users can store the captured media on a smartphone or other personal electronic device. SUMMARY OF THE INVENTION
[0006] However, some techniques for managing visual content in media are generally cumbersome and inefficient. For example, some prior arts use complex and time-consuming user interfaces, which may include multiple button presses or keystrokes. The prior arts require more time than necessary, which results in wasted user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0007] Accordingly, the present technology provides faster and more efficient methods and interfaces for an electronic device to manage visual content in media. Such methods and interfaces optionally supplement or replace other methods for managing visual content in media. Such methods and interfaces reduce the cognitive burden imposed on the user and result in a more effective human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.
[0008] According to some embodiments, a method is described. The method is performed at a computer system in communication with a display generation component. The method includes: displaying, via the display generation component, a camera user interface that includes simultaneously displaying a media representation and a media capture enabling representation; and when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations according to a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing displaying the first user interface object according to a determination that the corresponding set of criteria is not met; detecting a first input directed to the camera user interface when the media representation is displayed; and in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system according to a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text according to a determination that the first input corresponds to a selection of the first user interface object.
[0009] According to some embodiments, a non-transitory computer-readable storage device is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, where the computer system communicates with a display generation component, and the one or more programs include instructions for the following operations: displaying, via the display generation component, a camera user interface that includes simultaneously displaying a media representation and a media capture enabling representation; and when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations according to a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing displaying the first user interface object according to a determination that the corresponding set of criteria is not met; detecting a first input directed to the camera user interface when the media representation is displayed; and in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system according to a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text according to a determination that the first input corresponds to a selection of the first user interface object.
[0010] According to some embodiments, a transitory computer-readable storage device is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, where the computer system communicates with a display generation component, and the one or more programs include instructions for the following operations: displaying, via the display generation component, a camera user interface that includes simultaneously displaying a media representation and a media capture enabling representation; and when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations according to a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing displaying the first user interface object according to a determination that the corresponding set of criteria is not met; detecting a first input directed to the camera user interface when the media representation is displayed; and in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system according to a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text according to a determination that the first input corresponds to a selection of the first user interface object.
[0011] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a camera user interface that includes simultaneously displaying a media representation and a media capture enabling representation; and when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations based on a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing displaying the first user interface object based on a determination that the corresponding set of criteria is not met; detecting a first input directed to the camera user interface when the media representation is displayed; and in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system based on a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text based on a determination that the first input corresponds to a selection of the first user interface object.
[0012] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; a memory that stores one or more programs configured to be executed by the one or more processors; means for displaying, via the display generation component, a camera user interface that includes simultaneously displaying a media representation and a media capture enabling representation; and means for, when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations based on a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing displaying the first user interface object based on a determination that the corresponding set of criteria is not met; means for detecting a first input directed to the camera user interface when the media representation is displayed; and means for, in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system based on a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text based on a determination that the first input corresponds to a selection of the first user interface object.
[0013] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component. The one or more programs include instructions for: displaying, via the display generation component, a camera user interface that includes a simultaneous display of a media representation and a media capture enabling representation; and when the media representation and the media capture enabling representation are simultaneously displayed: displaying, via the display generation component, a first user interface object corresponding to one or more text management operations based on a determination that a corresponding set of criteria is met, where the corresponding set of criteria includes criteria that are met when corresponding text is detected in the media representation; and foregoing display of the first user interface object based on a determination that the corresponding set of criteria is not met; detecting, when the media representation is displayed, a first input directed to the camera user interface; and in response to detecting the first input directed to the camera user interface: initiating capture of media to be added to a media library associated with the computer system based on a determination that the first input corresponds to a selection of the media capture enabling representation; and displaying, via the display generation component, a plurality of options for managing the corresponding text based on a determination that the first input corresponds to a selection of the first user interface object.
[0014] According to some embodiments, a method is described. The method is performed at a computer system in communication with a display generation component and one or more input devices. The method includes: displaying, via the display generation component, a first representation of a previously captured media item; detecting, when the first representation of the previously captured media item is displayed, an input corresponding to a request to display a second representation of the previously captured media item via the one or more input devices; in response to detecting the input corresponding to the request to display the second representation of the previously captured media item, displaying, via the display generation component, the second representation of the previously captured media item; and when the second representation of the previously captured media item is displayed: displaying, via the display generation component, a visual indication corresponding to a text portion included in the second representation of the previously captured media item, the visual indication not being displayed when the first representation of the previously captured media item is displayed, based on a determination that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria.
[0015] According to some embodiments, a non-transitory computer-readable storage device is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with a display generation component and one or more input devices, and the one or more programs include instructions for the following operations: displaying, via the display generation component, a first representation of a previously captured media item; when displaying the first representation of the previously captured media item, detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the previously captured media item; in response to detecting the input corresponding to the request to display the second representation of the previously captured media item, displaying, via the display generation component, the second representation of the previously captured media item; and when displaying the second representation of the previously captured media item: displaying, via the display generation component, a visual indication corresponding to a text portion included in the second representation of the previously captured media item, based on determining that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria, the visual indication not being displayed when the first representation of the previously captured media item is displayed.
[0016] According to some embodiments, a transitory computer-readable storage device is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with a display generation component and one or more input devices, and the one or more programs include instructions for the following operations: displaying, via the display generation component, a first representation of a previously captured media item; when displaying the first representation of the previously captured media item, detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the previously captured media item; in response to detecting the input corresponding to the request to display the second representation of the previously captured media item, displaying, via the display generation component, the second representation of the previously captured media item; and when displaying the second representation of the previously captured media item: displaying, via the display generation component, a visual indication corresponding to a text portion included in the second representation of the previously captured media item, based on determining that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria, the visual indication not being displayed when the first representation of the previously captured media item is displayed.
[0017] According to some embodiments, a computer system is described that is configured to communicate with a display generation component and one or more input devices. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a first representation of a previously captured media item; when displaying the first representation of the previously captured media item, detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the previously captured media item; in response to detecting the input corresponding to the request to display the second representation of the previously captured media item, displaying, via the display generation component, the second representation of the previously captured media item; and when displaying the second representation of the previously captured media item: displaying, via the display generation component, a visual indication corresponding to the text portion included in the second representation, based on determining that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria, the visual indication not being displayed when the first representation of the previously captured media item is displayed.
[0018] According to some embodiments, a computer system is described that is configured to communicate with a display generation component and one or more input devices. The computer system includes: one or more processors; a memory that stores one or more programs configured to be executed by the one or more processors; means for displaying, via the display generation component, a first representation of a previously captured media item; means for detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the previously captured media item when displaying the first representation of the previously captured media item; means for displaying, via the display generation component, the second representation of the previously captured media item in response to detecting the input corresponding to the request to display the second representation of the previously captured media item; and means for: displaying, via the display generation component, a visual indication corresponding to the text portion included in the second representation, based on determining that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria, the visual indication not being displayed when the first representation of the previously captured media item is displayed, when displaying the second representation of the previously captured media item.
[0019] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more input devices. The one or more programs include instructions for: displaying, via the display generation component, a first representation of a previously captured media item; when displaying the first representation of the previously captured media item, detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the previously captured media item; in response to detecting the input corresponding to the request to display the second representation of the previously captured media item, displaying, via the display generation component, the second representation of the previously captured media item; and when displaying the second representation of the previously captured media item: displaying, via the display generation component, a visual indication corresponding to a text portion included in the second representation of the previously captured media item, based on determining that the text portion included in the second representation of the previously captured media item meets a corresponding set of criteria, the visual indication not being displayed when the first representation of the previously captured media item is displayed.
[0020] According to some embodiments, a method is described. The method is performed at a computer system in communication with one or more cameras, one or more input devices, and a display generation component. The method includes: displaying a first user interface including a text input area; when displaying the first user interface including the text input area, detecting a request to display a camera user interface; in response to receiving the request to display the camera user interface, displaying, via the display generation component, the camera user interface, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; when the representation of the field of view and the text insertion user interface object are displayed simultaneously, detecting, via the one or more input devices, an input corresponding to a selection of the text insertion user interface object; and in response to detecting the input corresponding to the selection of the text insertion user interface object, inserting at least a portion of the detected text into the text input area.
[0021] According to some embodiments, a non-transitory computer-readable storage device is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, where the computer system communicates with one or more cameras, one or more input devices, and a display generation component, and the one or more programs include instructions for the following operations: displaying a first user interface including a text input area; when the first user interface including the text input area is displayed, detecting a request to display a camera user interface; in response to receiving the request to display the camera user interface, displaying the camera user interface via the display generation component, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; when the representation of the field of view and the text insertion user interface object are displayed simultaneously, detecting an input corresponding to a selection of the text insertion user interface object via the one or more input devices; and in response to detecting the input corresponding to the selection of the text insertion user interface object, inserting at least a portion of the detected text into the text input area.
[0022] According to some embodiments, a transitory computer-readable storage device is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, where the computer system communicates with one or more cameras, one or more input devices, and a display generation component, and the one or more programs include instructions for the following operations: displaying a first user interface including a text input area; when the first user interface including the text input area is displayed, detecting a request to display a camera user interface; in response to receiving the request to display the camera user interface, displaying the camera user interface via the display generation component, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; when the representation of the field of view and the text insertion user interface object are displayed simultaneously, detecting an input corresponding to a selection of the text insertion user interface object via the one or more input devices; and in response to detecting the input corresponding to the selection of the text insertion user interface object, inserting at least a portion of the detected text into the text input area.
[0023] According to some embodiments, a computer system configured to communicate with one or more cameras, one or more input devices, and a display generation component is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a first user interface including a text input area; when the first user interface including the text input area is displayed, detecting a request to display a camera user interface; in response to receiving the request to display the camera user interface, displaying, via the display generation component, the camera user interface, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; when the representation of the field of view and the text insertion user interface object are displayed simultaneously, detecting, via the one or more input devices, an input corresponding to a selection of the text insertion user interface object; and in response to detecting the input corresponding to the selection of the text insertion user interface object, inserting at least a portion of the detected text into the text input area.
[0024] According to some embodiments, a computer system configured to communicate with one or more cameras, one or more input devices, and a display generation component is described. The computer system includes: a memory that stores one or more programs configured to be executed by one or more processors; means for displaying a first user interface including a text input area; means for detecting a request to display a camera user interface when the first user interface including the text input area is displayed; means for displaying, in response to receiving the request to display the camera user interface, via the display generation component, the camera user interface, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; means for detecting, when the representation of the field of view and the text insertion user interface object are displayed simultaneously, via the one or more input devices, an input corresponding to a selection of the text insertion user interface object; and means for inserting, in response to detecting the input corresponding to the selection of the text insertion user interface object, at least a portion of the detected text into the text input area.
[0025] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more cameras, one or more input devices, and a display generation component. The one or more programs include instructions for: displaying a first user interface including a text input area; when the first user interface including the text input area is displayed, detecting a request to display a camera user interface; in response to receiving the request to display the camera user interface, displaying the camera user interface via the display generation component, the camera user interface including: a representation of the field of view of the one or more cameras; and displaying a text insertion user interface object based on determining that the representation of the field of view of the one or more cameras includes detected text that meets one or more criteria, the text insertion user interface object being selectable to insert at least a portion of the detected text into the text input area; when the representation of the field of view and the text insertion user interface object are displayed simultaneously, detecting an input corresponding to a selection of the text insertion user interface object via the one or more input devices; and in response to detecting the input corresponding to the selection of the text insertion user interface object, inserting at least a portion of the detected text into the text input area.
[0026] According to some embodiments, a method is described. The method is performed at a computer system in communication with a display generation component. The method includes: displaying, via the display generation component, a media user interface including a media representation; when the media user interface including the media representation is displayed, receiving a request to display additional information about a plurality of detected features in the media representation; and in response to receiving the request to display the additional information about the plurality of detected features and when the media user interface including the media representation is displayed, displaying one or more indicators of the detected features in the media, the one or more indicators including a first indicator of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: the first indicator having a first appearance based on determining that the first detected feature is of a first feature type; and the first indicator having a second appearance different from the first appearance based on determining that the first detected feature is of a second feature type different from the first feature type.
[0027] According to some embodiments, a non-transitory computer-readable storage device is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with a display generation component, and the one or more programs include instructions for: displaying, via the display generation component, a media user interface including a media representation; when displaying the media user interface including the media representation, receiving a request to display additional information about a plurality of detected features in the media representation; and in response to receiving the request to display the additional information about the plurality of detected features and when displaying the media user interface including the media representation, displaying one or more indications of the detected features in the media, the one or more indications including a first indication of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: according to determining that the first detected feature is of a first feature type, the first indication having a first appearance; and according to determining that the first detected feature is of a second feature type different from the first feature type, the first indication having a second appearance different from the first appearance.
[0028] According to some embodiments, a transitory computer-readable storage device is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with a display generation component, and the one or more programs include instructions for: displaying, via the display generation component, a media user interface including a media representation; when displaying the media user interface including the media representation, receiving a request to display additional information about a plurality of detected features in the media representation; and in response to receiving the request to display the additional information about the plurality of detected features and when displaying the media user interface including the media representation, displaying one or more indications of the detected features in the media, the one or more indications including a first indication of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: according to determining that the first detected feature is of a first feature type, the first indication having a first appearance; and according to determining that the first detected feature is of a second feature type different from the first feature type, the first indication having a second appearance different from the first appearance.
[0029] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a media user interface including a media representation; receiving, when displaying the media user interface including the media representation, a request to display additional information about a plurality of detected features in the media representation; and in response to receiving the request to display the additional information about the plurality of detected features and when displaying the media user interface including the media representation, displaying one or more indications of the detected features in the media, the one or more indications including a first indication of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: based on determining that the first detected feature is of a first feature type, the first indication having a first appearance; and based on determining that the first detected feature is of a second feature type different from the first feature type, the first indication having a second appearance different from the first appearance.
[0030] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes: one or more processors; a memory that stores one or more programs configured to be executed by the one or more processors; means for displaying, via the display generation component, a media user interface including a media representation; means for receiving, when displaying the media user interface including the media representation, a request to display additional information about a plurality of detected features in the media representation; and means for displaying, in response to receiving the request to display the additional information about the plurality of detected features and when displaying the media user interface including the media representation, one or more indications of the detected features in the media, the one or more indications including a first indication of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: based on determining that the first detected feature is of a first feature type, the first indication having a first appearance; and based on determining that the first detected feature is of a second feature type different from the first feature type, the first indication having a second appearance different from the first appearance.
[0031] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component. The one or more programs include instructions for: displaying, via the display generation component, a media user interface including a media representation; receiving, when the media user interface including the media representation is displayed, a request to display additional information about a plurality of detected features in the media representation; and in response to receiving the request to display the additional information about the plurality of detected features and when the media user interface including the media representation is displayed, displaying one or more indicators of the detected features in the media, the one or more indicators including a first indicator of a first detected feature displayed at a first position in the media representation, the first position corresponding to the position of the first detected feature in the media representation, including: according to determining that the first detected feature is of a first feature type, the first indicator having a first appearance; and according to determining that the first detected feature is of a second feature type different from the first feature type, the first indicator having a second appearance different from the first appearance.
[0032] According to some embodiments, a method is described. The method is performed at a computer system in communication with one or more cameras, a display generation component, and one or more input devices. The method includes: receiving a request to display a representation of the field of view of the one or more cameras; in response to receiving the request to display the representation of the field of view of the one or more cameras: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located in the field of view of the one or more cameras; and automatically displaying, via the display generation component, a plurality of indicators of the translated text, the plurality of indicators of the translated text including a first indicator of the translation of a first portion of the text and a second indicator of the translation of a second portion of the text; when the first indicator and the second indicator are displayed via the display generation component, receiving, via the one or more input devices, a request to select a corresponding indicator of a plurality of translated portions; and in response to receiving the request to select the corresponding indicator, according to determining that the request is a request to select the first indicator, displaying, via the display generation component, a first translated user interface object that includes the first portion of the text and the translation of the first portion of the text and does not include the translation of the second portion of the text.
[0033] According to some embodiments, a non-transitory computer-readable storage device is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with one or more cameras, a display generation component, and one or more input devices, and the one or more programs include instructions for the following operations: receiving a request to display a representation of the field of view of the one or more cameras; in response to receiving the request to display the representation of the field of view of the one or more cameras: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located in the field of view of the one or more cameras; and automatically displaying, via the display generation component, multiple indications of the converted text, the multiple indications of the converted text including a first indication of the conversion of a first portion of the text and a second indication of the conversion of a second portion of the text; when the first indication and the second indication are displayed via the display generation component, receiving, via the one or more input devices, a request to select a corresponding indication of multiple converted portions; and in response to receiving the request to select the corresponding indication, and based on determining that the request is a request to select the first indication, displaying, via the display generation component, a first converted user interface object that includes the first portion of the text and the conversion of the first portion of the text and does not include the conversion of the second portion of the text.
[0034] According to some embodiments, a transitory computer-readable storage device is described. The transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with one or more cameras, a display generation component, and one or more input devices, and the one or more programs include instructions for the following operations: receiving a request to display a representation of the field of view of the one or more cameras; in response to receiving the request to display the representation of the field of view of the one or more cameras: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located in the field of view of the one or more cameras; and automatically displaying, via the display generation component, multiple indications of the converted text, the multiple indications of the converted text including a first indication of the conversion of a first portion of the text and a second indication of the conversion of a second portion of the text; when the first indication and the second indication are displayed via the display generation component, receiving, via the one or more input devices, a request to select a corresponding indication of multiple converted portions; and in response to receiving the request to select the corresponding indication, and based on determining that the request is a request to select the first indication, displaying, via the display generation component, a first converted user interface object that includes the first portion of the text and the conversion of the first portion of the text and does not include the conversion of the second portion of the text.
[0035] According to some embodiments, a computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: receiving a request to display a representation of the field of view of the one or more cameras; in response to receiving the request to display the representation of the field of view of the one or more cameras: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located within the field of view of the one or more cameras; and automatically displaying, via the display generation component, a plurality of indications of the translated text, the plurality of indications of the translated text including a first indication of the translation of a first portion of the text and a second indication of the translation of a second portion of the text; when the first indication and the second indication are being displayed via the display generation component, receiving, via the one or more input devices, a request to select the respective indication of the plurality of translated portions; and in response to receiving the request to select the respective indication, and based on determining that the request is a request to select the first indication, displaying, via the display generation component, a first translated user interface object that includes the first portion of the text and the translation of the first portion of the text and does not include the translation of the second portion of the text.
[0036] According to some embodiments, a computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices is described. The computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: means for receiving a request to display a representation of the field of view of the one or more cameras; means for, in response to receiving the request to display the representation of the field of view of the one or more cameras, performing the following operations: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located within the field of view of the one or more cameras; and automatically displaying, via the display generation component, a plurality of indications of the translated text, the plurality of indications of the translated text including a first indication of the translation of a first portion of the text and a second indication of the translation of a second portion of the text; means for, when the first indication and the second indication are being displayed via the display generation component, receiving, via the one or more input devices, a request to select the respective indication of the plurality of translated portions; and means for, in response to receiving the request to select the respective indication, and based on determining that the request is a request to select the first indication, displaying, via the display generation component, a first translated user interface object that includes the first portion of the text and the translation of the first portion of the text and does not include the translation of the second portion of the text.
[0037] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras, a display generation component, and one or more input devices. The one or more programs include instructions for: receiving a request to display a representation of the field of view of the one or more cameras; in response to receiving the request to display the representation of the field of view of the one or more cameras: displaying, via the display generation component, the representation of the field of view of the one or more cameras, wherein the representation includes text located in the field of view of the one or more cameras; and automatically displaying, via the display generation component, a plurality of indications of transformed text, the plurality of indications of transformed text including a first indication of a transformation of a first portion of the text and a second indication of a transformation of a second portion of the text; when the first indication and the second indication are displayed via the display generation component, receiving, via the one or more input devices, a request to select a respective indication of a plurality of transformed portions; and in response to receiving the request to select the respective indication, and based on determining that the request is a request to select the first indication, displaying, via the display generation component, a first transformed user interface object that includes the first portion of the text and the transformation of the first portion of the text and does not include the transformation of the second portion of the text.
[0038] According to some embodiments, a method is described that is executed at a computer system that communicates with a display generation component. The method includes: when displaying a user interface that includes a media representation, detecting a request to display additional information corresponding to the media representation; and in response to detecting the request to display additional information corresponding to the media representation: based on determining that detected text in the media representation has a first set of attributes, displaying, via the display generation component, a first user interface object that, when selected, causes the computer system to perform a first operation based on the detected text; and based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, displaying, via the display generation component, a second user interface object that, when selected, causes the computer system to perform a second operation different from the first operation based on the detected text.
[0039] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: when displaying a user interface including a media representation, detecting a request to display additional information corresponding to the media representation; and in response to detecting the request to display additional information corresponding to the media representation: displaying, via the display generation component, a first user interface object based on determining that the detected text in the media representation has a first set of attributes, the first user interface object, when selected, causing the computer system to perform a first operation based on the detected text; and displaying, via the display generation component, a second user interface object based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, the second user interface object, when selected, causing the computer system to perform a second operation different from the first operation based on the detected text.
[0040] According to some embodiments, a transitory computer-readable storage medium is described. The transitory computer storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: when displaying a user interface including a media representation, detecting a request to display additional information corresponding to the media representation; and in response to detecting the request to display additional information corresponding to the media representation: displaying, via the display generation component, a first user interface object based on determining that the detected text in the media representation has a first set of attributes, the first user interface object, when selected, causing the computer system to perform a first operation based on the detected text; and displaying, via the display generation component, a second user interface object based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, the second user interface object, when selected, causing the computer system to perform a second operation different from the first operation based on the detected text.
[0041] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when displaying a user interface including a media representation, detecting a request to display additional information corresponding to the media representation; and in response to detecting the request to display additional information corresponding to the media representation: displaying a first user interface object via the display generation component based on determining that the detected text in the media representation has a first set of attributes, the first user interface object, when selected, causing the computer system to perform a first operation based on the detected text; and displaying a second user interface object via the display generation component based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, the second user interface object, when selected, causing the computer system to perform a second operation different from the first operation based on the detected text.
[0042] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and includes: means for detecting a request to display additional information corresponding to a media representation when displaying a user interface including the media representation; and means for performing the following operations in response to detecting the request to display additional information corresponding to the media representation: displaying a first user interface object via the display generation component based on determining that the detected text in the media representation has a first set of attributes, the first user interface object, when selected, causing the computer system to perform a first operation based on the detected text; and displaying a second user interface object via the display generation component based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, the second user interface object, when selected, causing the computer system to perform a second operation different from the first operation based on the detected text.
[0043] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a generation component, the one or more programs including instructions for: when displaying a user interface including a media representation, detecting a request to display additional information corresponding to the media representation; and in response to detecting the request to display additional information corresponding to the media representation: displaying, via the display generation component, a first user interface object based on determining that the detected text in the media representation has a first set of attributes, the first user interface object, when selected, causing the computer system to perform a first operation based on the detected text; and displaying, via the display generation component, a second user interface object based on determining that the detected text in the media representation has a second set of attributes different from the first set of attributes, the second user interface object, when selected, causing the computer system to perform a second operation different from the first operation based on the detected text.
[0044] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are optionally included in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0045] Accordingly, faster and more efficient methods and interfaces for managing visual content in media are provided for a device, thereby improving the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces may supplement or replace other methods for managing visual content in media. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To better understand the various described embodiments, reference should be made to the following detailed description taken in conjunction with the accompanying drawings, in which like reference numerals indicate corresponding parts in all the figures.
[0047] Figure 1A is a block diagram showing a portable multifunctional device having a touch-sensitive display in accordance with some embodiments.
[0048] Figure 1B is a block diagram showing exemplary components for event handling in accordance with some embodiments.
[0049] Figure 2 shows a portable multifunctional device having a touch screen in accordance with some embodiments.
[0050] Figure 3 is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface in accordance with some embodiments.
[0051] Figure 4A Shows an exemplary user interface for a menu of an application on a portable multifunctional device according to some embodiments.
[0052] Figure 4B Shows an exemplary user interface for a multifunctional device having a touch-sensitive surface separate from a display according to some embodiments.
[0053] Figure 5A Shows a personal electronic device according to some embodiments.
[0054] Figure 5B Is a block diagram showing a personal electronic device according to some embodiments.
[0055] Figures 6A to 6Z Shows an exemplary user interface for managing visual content in media according to some embodiments.
[0056] Figures 7A to 7L Shows an exemplary user interface for visual indicators for managing visual content in media according to some embodiments.
[0057] Figure 8 Is a flowchart showing a method for managing visual content in media according to some embodiments.
[0058] Figure 9 Is a flowchart showing visual indicators for managing visual content in media according to some embodiments.
[0059] Figures 10A to 10AD Shows an exemplary user interface for inserting visual content into media according to some embodiments.
[0060] Figure 11 Is a flowchart showing a user interface for inserting visual content into media according to some embodiments.
[0061] Figures 12A to 12L Shows an exemplary user interface for identifying visual content in media according to some embodiments.
[0062] Figure 13 Is a flowchart showing a method for identifying visual content in media according to some embodiments.
[0063] Figures 14A to 14N Shows an exemplary user interface for converting visual content in media according to some embodiments.
[0064] Figure 15 Is a flowchart showing a method for converting visual content in media according to some embodiments.
[0065] Figures 16A to 16O An exemplary user interface for managing user interface objects for visual content in media is shown.
[0066] Figure 17 Is a flowchart showing a method for managing user interface objects for visual content in media according to some embodiments. Detailed Description
[0067] The following description sets forth exemplary methods, parameters, etc. However, it should be recognized that such description is not intended to limit the scope of the present disclosure, but rather is provided as a description of exemplary embodiments.
[0068] There is a need for an electronic device to provide an efficient method and interface for managing visual content. For example, there is a need for an electronic device and / or computer system to allow a user to manage visual content included in objects captured by one or more cameras of the computer system, such as a marker or a restaurant menu. Such techniques can reduce the cognitive burden on the user managing the visual content, thereby increasing productivity. In addition, such techniques can reduce the processor power and battery power otherwise wasted on redundant user input.
[0069] Below Figures 1A to 1B , Figure 2 , Figure 3 , Figures 4A to 4B and Figures 5A to 5B Provides a description of an exemplary device for performing techniques for managing visual content.
[0070] Figures 6A to 6Z An exemplary user interface for managing visual content in media is shown. Figure 8 Is a flowchart showing a method for managing visual content according to some embodiments. Figures 6A to 6Z The user interface in is used to illustrate the processes described below, including Figure 8 the processes in.
[0071] Figures 7A to 7L An exemplary user interface for visual indicators for managing visual content in media is shown. Figure 9 Is a flowchart showing a method for managing visual indicators for visual content in media according to some embodiments. Figures 7A to 7L The user interface in is used to illustrate the processes described below, including Figure 9 the processes in.
[0072] Figures 10A to 10AD An exemplary user interface for inserting visual content in media is shown. Figure 11 Is a flowchart showing a method for inserting visual content in media. Figures 10A to 10ADThe user interface in is used to illustrate the processes described below, including Figure 11 the processes in.
[0073] Figures 12A to 12L illustrates an exemplary user interface for identifying visual content in media. Figure 13 is a flowchart showing a method for identifying visual content in media. Figures 12A to 12L The user interface in is used to illustrate the processes described below, including Figure 13 the processes in.
[0074] Figures 14A to 14N illustrates an exemplary user interface for converting visual content in media. Figure 15 is a flowchart showing a method for converting visual content in media according to some embodiments. Figures 14A to 14N The user interface in is used to illustrate the processes described below, including Figure 15 the processes in.
[0075] Figures 16A to 16O illustrates an exemplary user interface for managing user interface objects for visual content in media according to some embodiments. Figure 17 is a flowchart showing a method for managing user interface objects for visual content in media according to some embodiments. Figures 16A to 16O The user interface in is used to illustrate the processes described below, including Figure 17 the processes in.
[0076] The processes described below enhance the operability of the device and make the user-device interface more effective through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including by providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more effectively.
[0077] In addition, in a method where one or more of the steps described herein depend on one or more conditions being satisfied, it should be understood that the method can be repeated in multiple iterations such that, during the repetition process, all the conditions that determine the steps in the method are satisfied in different iterations of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if the condition is not satisfied), one of ordinary skill in the art will know to repeat the stated steps until both the condition being satisfied and the condition not being satisfied (in no particular order) occur. Thus, a method described as having one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of the corresponding one or more conditions and is thus capable of determining whether the possible conditions have been satisfied without explicitly repeating the steps of the method until all the conditions that determine the steps in the method are satisfied. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.
[0078] Although the following description uses the terms "first", "second", etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first touch could be named a second touch and similarly a second touch could be named a first touch without departing from the scope of the various described embodiments. Both the first touch and the second touch are touches, but they are not the same touch.
[0079] The terms used in the description of the various described embodiments herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various described embodiments and the appended claims, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms "comprises" and / or "comprising" when used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0080] Depending on the context, the term "if" is optionally interpreted to mean "when", "upon", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined that..." or "if [stated condition or event] is detected" is optionally interpreted to mean "when it is determined that..." or "in response to determining..." or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]".
[0081] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions such as PDA and / or music player functions. Exemplary embodiments of the portable multifunctional device include, but are not limited to, devices from Apple Inc. (Cupertino, California), devices, iPod devices, and devices. Other portable electronic devices, such as a laptop or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad), may optionally be used. It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). In some embodiments, the electronic device is a computer system that communicates (e.g., via wireless communication, via wired communication) with a display generation component. The display generation component is configured to provide a visual output, such as a display via a CRT monitor, a display via an LED monitor, or a display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes displaying content (e.g., video data rendered or decoded by display controller 156) by transmitting data (e.g., image data or video data) to an integrated or external display generation component via a wired or wireless connection to visually generate the content.
[0082] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick.
[0083] The device generally supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, gaming applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, and / or digital video player applications.
[0084] The various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or within the respective applications. Thus, a common physical architecture of the device (such as a touch-sensitive surface) optionally supports a variety of applications with a user interface that is intuitive and clear to the user.
[0085] Attention is now turned to an embodiment of a portable device having a touch-sensitive display. Figure 1A FIG. is a block diagram of a portable multifunctional device 100 having a touch-sensitive display system 112 in accordance with some embodiments. The touch-sensitive display 112 is sometimes called a "touch screen" for convenience and is sometimes referred to as or called a "touch-sensitive display system". The device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. The device 100 optionally includes one or more optical sensors 164. The device 100 optionally includes one or more contact intensity sensors 165 for detecting the intensity of a contact on the device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100). The device 100 optionally includes one or more tactile output generators 167 for generating tactile output on the device 100 (e.g., generating tactile output on a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100 or the touchpad 355 of the device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0086] As used in this specification and the claims, the “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or to a surrogate for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using a variety of methods and a variety of sensors or combinations of sensors. For example, one or more force sensors beneath or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in size of the contact area detected on the touch-sensitive surface, the capacitance and / or change in capacitance of the touch-sensitive surface near the contact, and / or the resistance and / or change in resistance of the touch-sensitive surface near the contact are optionally used as surrogates for the force or pressure of a contact on the touch-sensitive surface. In some embodiments, the surrogate measurements of contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measurements). In some embodiments, the surrogate measurements of contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of user input allows the user to access additional device functions that would otherwise be inaccessible on a smaller device with a limited footprint, the smaller device being used to display affordances and / or receive user input (e.g., via a touch-sensitive display, a touch-sensitive surface, or physical / mechanical controls such as knobs or buttons).
[0087] As used in this specification and the claims, the term "haptic output" refers to a physical displacement of a device relative to a previous position of the device detected by a user using the user's sense of touch, a physical displacement of a component of a device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device. For example, in the case of contact between the device or a component of the device and a surface that is sensitive to touch by the user (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation that corresponds to a perceived change in the physical characteristics of the device or the component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or a touchpad) is optionally interpreted by the user as a "press click" or "release click" of a physical actuation button. In some cases, the user will feel a tactile sensation, such as a "press click" or "release click," even when the physical actuation button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when there is no change in the smoothness of the touch-sensitive surface, movement of the touch-sensitive surface will optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface. Although such interpretations of touch by the user will be limited by the user's individual sensory perceptions, many sensory perceptions of touch are common to most users. Thus, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "press click," "release click," "roughness"), unless otherwise stated, the haptic output generated corresponds to a physical displacement of the device or a component thereof that would generate the stated sensory perception of a typical (or average) user.
[0088] It should be understood that device 100 is merely one example of a portable multifunctional device, and device 100 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0089] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.
[0090] The peripheral device interface 118 can be used to couple input and output peripheral devices of the device to the CPU 120 and the memory 102. One or more processors 120 run or execute various software programs and / or instruction sets stored in the memory 102 to perform various functions of the device 100 and process data. In some embodiments, the peripheral device interface 118, the CPU 120, and the memory controller 122 are optionally implemented on a single chip such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0091] The RF (Radio Frequency) circuit 108 receives and transmits RF signals, which are also referred to as electromagnetic signals. The RF circuit 108 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with the communication network and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chip sets, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 108 optionally communicates with the network and other devices via wireless communication, and these networks are such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular phone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). The RF circuit 108 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as via short-range communication radio components. The wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution-Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), WiMAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the date of submission of this document.
[0092] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral interface 118, converts the audio data into an electrical signal, and transmits the electrical signal to the speaker 111. The speaker 111 converts the electrical signal into sound waves audible to humans. The audio circuit 110 also receives the electrical signal converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signal into audio data and transmits the audio data to the peripheral interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral interface 118. In some embodiments, the audio circuit 110 also includes an earphone jack (e.g., Figure 2 212 in
[0093] ). The earphone jack provides an interface between the audio circuit 110 and a removable audio input / output peripheral device, which is such as an output-only headset or an earphone having both an output (e.g., a mono or stereo earphone) and an input (e.g., a microphone). Figure 2 The I / O subsystem 106 couples input / output peripheral devices such as the touch screen 112 and other input control devices 116 on the device 100 to the peripheral interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / transmit electrical signals to the other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some embodiments, the input controller 160 is optionally coupled to (or not coupled to) any one of the following: a keyboard, an infrared port, a USB port, and a pointing device such as a mouse. One or more buttons (e.g., Figure 2206) In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication, via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system.
[0094] Rapidly pressing the depress button optionally disengages the lock of the touch screen 112 or optionally initiates a process of unlocking the device using gestures on the touch screen, as described in U.S. Patent Application No. 11 / 322,549, filed on December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image" (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. Long pressing the depress button (e.g., 206) optionally powers on or powers off the device 100. The functions of the one or more buttons are optionally user-customizable. The touch screen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0095] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch screen 112 and / or sends electrical signals to the touch screen 112. The touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0096] The touch screen 112 has a touch-sensitive surface, sensor, or group of sensors that accepts input from the user based on tactile and / or haptic contact. The touch screen 112 and the display controller 156 (along with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of the contact) on the touch screen 112 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.
[0097] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, but in other embodiments, other display technologies are used. The touch screen 112 and the display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112 to detect contact and any movement or interruption thereof. The variety of touch sensing technologies includes, but is not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as the technology used in and the iPod as used in.
[0098] The touch-sensitive display in some embodiments of the touch screen 112 optionally resembles the multi-touch sensitive touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, the touch screen 112 displays the visual output from the device 100, while the touch-sensitive touchpad does not provide a visual output.
[0099] Some implementations of the touch-sensitive display in touchscreen 112 are described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, "Multipoint Touch Surface Controller", filed May 2, 2006; (2) U.S. Patent Application No. 10 / 840,862, "Multipoint Touchscreen", filed May 6, 2004; (3) U.S. Patent Application No. 10 / 903,964, "Gestures For Touch Sensitive Input Devices", filed Jul. 30, 2004; (4) U.S. Patent Application No. 11 / 048,264, "Gestures For Touch Sensitive Input Devices", filed Jan. 31, 2005; (5) U.S. Patent Application No. 11 / 038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices", filed Jan. 18, 2005; (6) U.S. Patent Application No. 11 / 228,758, "Virtual Input Device Placement On A Touch Screen User Interface", filed Sep. 16, 2005; (7) U.S. Patent Application No. 11 / 228,700, "Operation Of A Computer With A Touch Screen Interface", filed Sep. 16, 2005; (8) U.S. Patent Application No. 11 / 228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard", filed Sep. 16, 2005; and (9) U.S. Patent Application No. 11 / 367,749, "Multi-Functional Hand-Held Device", filed Mar. 3, 2006. All of these applications are hereby incorporated by reference in their entirety.
[0100] The touch screen 112 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen has a video resolution of approximately 160 dpi. The user optionally uses any suitable object or attachment such as a stylus, finger, etc. to contact the touch screen 112. In some embodiments, the user interface is designed to work primarily through finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts the rough finger-based input into an accurate pointer / cursor position or command for performing the action desired by the user.
[0101] In some embodiments, in addition to the touch screen, the device 100 optionally further includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device, which, unlike the touch screen, does not display a visual output. The touchpad is optionally a touch-sensitive surface separate from the touch screen 112, or an extension of the touch-sensitive surface formed by the touch screen.
[0102] The device 100 further includes a power system 162 for powering various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0103] The device 100 optionally further includes one or more optical sensors 164. Figure 1AAn optical sensor coupled to the optical sensor controller 158 in the I / O subsystem 106 is shown. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with the imaging module 143 (also called the camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite the touch screen display 112 on the front of the device, such that the touch screen display can be used as a viewfinder for still image and / or video image capture. In some embodiments, the optical sensor is located on the front of the device such that an image of the user can optionally be captured for video conferencing while the user views other video conferencing participants on the touch screen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that a single optical sensor 164 can be used with the touch screen display for both video conferencing and still image and / or video image capture.
[0104] The device 100 optionally further includes one or more depth camera sensors 175. Figure 1A A depth camera sensor coupled to the depth camera controller 169 in the I / O subsystem 106 is shown. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of an object (e.g., a face) within the scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with the imaging module 143 (also referred to as the camera module), the depth camera sensor 175 is optionally used to determine depth maps of different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located on the front of the device 100 such that an image of the user with depth information can optionally be captured for video conferencing while the user views other video conferencing participants on the touch screen display, and a selfie with depth map data can be captured. In some embodiments, the depth camera sensor 175 is located on the rear of the device, or on both the rear and front of the device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) such that the depth camera sensor 175 can be used with the touch screen display for both video conferencing and still image and / or video image capture.
[0105] The device 100 optionally further includes one or more contact intensity sensors 165. Figure 1AA contact intensity sensor coupled to an intensity sensor controller 159 in the I / O subsystem 106 is shown. The contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact intensity sensor 165 receives contact intensity information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., the touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the rear of the device 100, opposite the touchscreen display 112 located on the front of the device 100).
[0106] The device 100 optionally further includes one or more proximity sensors 166. Figure 1A A proximity sensor 166 coupled to the peripheral device interface 118 is shown. Alternatively, the proximity sensor 166 is optionally coupled to an input controller 160 in the I / O subsystem 106. The proximity sensor 166 optionally operates as described in the following U.S. patent application numbers: 11 / 241,839, entitled "Proximity Detector In Handheld Device"; 11 / 240,788, entitled "ProximityDetector In Handheld Device"; 11 / 620,702, entitled "Using Ambient Light Sensor ToAugment Proximity Sensor Output"; 11 / 586,862, entitled "Automated Response To AndSensing Of User Activity In Portable Devices"; and 11 / 638,251, entitled "MethodsAnd Systems For Automatic Configuration Of Peripherals", which U.S. patents are hereby incorporated by reference in their entirety. In some embodiments, when the multifunctional device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touchscreen 112.
[0107] The device 100 optionally further includes one or more haptic output generators 167. Figure 1AShows a haptic output generator coupled to the haptic feedback controller 161 in the I / O subsystem 106. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting an electrical signal into a haptic output on the device). The contact intensity sensor 165 receives haptic feedback generation instructions from the haptic feedback module 133 and generates a haptic output on the device 100 that can be felt by a user of the device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., the touch-sensitive display system 112), and optionally generates a haptic output by moving the touch-sensitive surface vertically (e.g., into / out of the surface of the device 100) or laterally (e.g., backward and forward in the same plane as the surface of the device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.
[0108] The device 100 optionally further includes one or more accelerometers 168. Figure 1A Shows an accelerometer 168 coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to the input controller 160 in the I / O subsystem 106. The accelerometer 168 optionally operates as described in the following U.S. Patent Publication Nos.: 20050190059, titled "Acceleration-based Theft Detection System for Portable Electronic Devices" and 20060017692, titled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", both of which are hereby incorporated by reference in their entireties. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on an analysis of data received from one or more accelerometers. The device 100 optionally further includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver in addition to the accelerometer 168 for obtaining information about the location and orientation (e.g., portrait or landscape) of the device 100.
[0109] In some embodiments, the software components stored in the memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and application programs (or instruction sets) 136. Additionally, in some embodiments, the memory 102 ( Figure 1A ) or 370 ( Figure 3 ) stores a device / global internal state 157, as shown in Figure 1A and Figure 3 . The device / global internal state 157 includes one or more of the following: an active application state, which indicates which applications (if any) are currently active; a display state, indicating what applications, views, or other information occupy the various regions of the touchscreen display 112; a sensor state, including information obtained from the various sensors and input control devices 116 of the device; and location information related to the location and / or orientation of the device.
[0110] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between the various hardware components and software components.
[0111] The communication module 128 facilitates communication with other devices via one or more external ports 124, and also includes various software components for processing data received by the RF circuit 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices, or indirectly coupled via a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or compatible with the 30-pin connector used on (a trademark of Apple Inc.) devices.
[0112] The contact / motion module 130 optionally detects contact with the touch screen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or a physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the contact intensity (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has ceased (e.g., detecting a finger lift event or contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of the contact point being represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple finger contact). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.
[0113] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by the user (e.g., determining whether the user has "clicked" an icon). In some embodiments, at least a subset of the intensity thresholds is determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and can be adjusted without changing the physical hardware of the device 100). For example, the mouse "click" threshold of a touchpad or a touch screen can be set to any threshold within a wide range of predefined thresholds without changing the touchpad or touch screen display hardware. Additionally, in some implementations, software settings are provided to the user of the device for adjusting one or more of the intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by using a system-level click on an "intensity" parameter to adjust multiple intensity thresholds at once).
[0114] The touch / motion module 130 optionally detects gesture inputs made by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timing, and / or intensities of the detected contacts). Thus, gestures are optionally detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift (lift-off) event.
[0115] The graphics module 132 includes various known software components for presenting and displaying graphics on the touch screen 112 or other display, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual attributes). As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0116] In some embodiments, the graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes for specifying the graphics to be displayed from an application, etc., and also receives coordinate data and other graphic attribute data as necessary, and then generates screen image data for output to the display controller 156.
[0117] The haptic feedback module 133 includes various software components for generating instructions that are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.
[0118] The text input module 134, which is optionally a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).
[0119] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., provided to the phone 138 for location-based dialing; provided to the camera 143 as picture / video metadata; and provided to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).
[0120] The application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:
[0121] · A contacts module 137 (sometimes referred to as an address book or contacts list);
[0122] · A phone module 138;
[0123] · An email client module 140;
[0124] · An instant messaging (IM) module 141;
[0125] · A fitness support module 142;
[0126] · A camera module 143 for still images and / or video images;
[0127] · An image management module 144;
[0128] · A video player module;
[0129] · A music player module;
[0130] · A browser module 147;
[0131] · A calendar module 148;
[0132] · A widget module 149, which optionally includes one or more of the following: a weather widget 149-1, a stock market widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6;
[0133] · A widget creator module 150 for forming user-created widgets 149-6;
[0134] · A search module 151;
[0135] · A video and music player module 152, which combines the video player module and the music player module;
[0136] · A notes module 153;
[0137] · A maps module 154; and / or
[0138] · An online video module 155.
[0139] Examples of other applications 136 that are optionally stored in the memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-supported applications, encryption, digital rights management, speech recognition, and speech replication.
[0140] In conjunction with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the contacts module 137 is optionally used to manage an address book or contact list (e.g., the application internal state 192 of the contacts module 137 stored in the memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating a phone number, email address, physical address, or other information with a name; associating an image with a name; categorizing and classifying names; providing a phone number or email address to initiate and / or facilitate communication via the phone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0141] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the phone module 138 is optionally used to input a character sequence corresponding to a phone number, access one or more phone numbers in the contacts module 137, modify an entered phone number, dial the corresponding phone number, conduct a session, and disconnect or hang up when the session is complete. As described above, the wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies.
[0142] In conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contacts module 137, and phone module 138, the video conferencing module 139 includes executable instructions to initiate, conduct, and terminate a video conference between the user and one or more other participants according to user instructions.
[0143] In conjunction with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the email client module 140 includes executable instructions to create, send, receive, and manage emails in response to user instructions. In conjunction with the image management module 144, the email client module 140 makes it very easy to create and send emails with static images or video images captured by the camera module 143.
[0144] In combination with the RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the instant messaging module 141 includes executable instructions for the following operations: inputting a character sequence corresponding to an instant message, modifying a previously input character, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0145] In combination with the RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, the fitness support module 142 includes executable instructions for creating a fitness (e.g., having time, distance, and / or calorie burn goals); communicating with fitness sensors (exercise devices); receiving fitness sensor data; calibrating sensors for monitoring fitness; selecting and playing music for fitness; and displaying, storing, and transmitting fitness data.
[0146] In combination with the touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, the camera module 143 includes executable instructions for the following operations: capturing still images or video (including video streams) and storing them in the memory 102, modifying the characteristics of the still images or video, or deleting the still images or video from the memory 102.
[0147] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or video images.
[0148] In combination with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for browsing the Internet in accordance with user instructions, including searching for, linking to, receiving, and displaying web pages or portions thereof, and linking to attachments and other files of web pages.
[0149] In combination with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions for creating, displaying, modifying, and storing calendars and data associated with the calendars (e.g., calendar entries, to-do items, etc.) in accordance with user instructions.
[0150] In combination with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is a miniature application (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) optionally downloaded and used by a user or a miniature application created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (HyperText Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).
[0151] In combination with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 is optionally used by a user to create widgets (e.g., transforming a user-specified portion of a web page into a widget).
[0152] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, search module 151 includes executable instructions for searching memory 102 for text, music, sound, images, video, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.
[0153] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch screen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0154] In combination with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, note module 153 includes executable instructions for creating and managing notes, to-do lists, and the like according to user instructions.
[0155] In combination with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and data associated with the maps (e.g., driving directions, data related to stores and other points of interest at or near a particular location, and other location-based data) according to user instructions.
[0156] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, e-mail client module 140, and browser module 147, the online video module 155 includes instructions for performing the following operations: allowing a user to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touch screen or on an external display connected via the external port 124), send an e-mail with a link to a particular online video, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, the instant message module 141 is used instead of the e-mail client module 140 to send a link to a particular online video. Other descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed Jun. 20, 2007, and entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," and U.S. Patent Application No. 11 / 968,067, filed Dec. 31, 2007, and entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of both of which are hereby incorporated by reference in their entireties.
[0157] Each of the above modules and applications corresponds to a set of executable instructions for performing one or more of the above functions and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., the instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, the video player module is optionally combined with the music player module into a single module (e.g., Figure 1A the video and music player module 152 in). In some embodiments, the memory 102 optionally stores a subgroup of the above modules and data structures. In addition, the memory 102 optionally stores additional modules and data structures not described above.
[0158] In some embodiments, device 100 is a device in which the operation of a predefined set of functions on the device is performed uniquely via a touch screen and / or a touchpad. By using the touch screen and / or the touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on device 100 is optionally reduced.
[0159] The predefined set of functions performed uniquely via the touch screen and / or the touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by the user, navigates device 100 from any user interface displayed on device 100 to a main menu, a home menu, or a root menu. In such embodiments, the touchpad is used to implement a "menu button". In some other embodiments, the menu button is a physical push button or other physical input control device rather than the touchpad.
[0160] Figure 1B is a block diagram showing exemplary components for event handling according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 ) includes an event classifier 170 (e.g., in operating system 126) and corresponding application 136-1 (e.g., any one of the aforementioned applications 137 to 151, 155, 380 to 390).
[0161] Event classifier 170 receives event information and determines the application 136-1 to which the event information is to be delivered and the application view 191 of application 136-1. Event classifier 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, application 136-1 includes an application internal state 192 that indicates one or more current application views displayed on the touch-sensitive display 112 when the application is active or executing. In some embodiments, the device / global internal state 157 is used by event classifier 170 to determine which application(s) is / are currently active, and the application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0162] In some embodiments, the application internal state 192 includes additional information, such as one or more of the following: recovery information to be used when application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by application 136-1, a state queue for enabling the user to return to a previous state or view of application 136-1, and a repeat / undo queue of previous actions taken by the user.
[0163] The event monitor 171 receives event information from the peripheral interface 118. The event information includes information about sub-events (e.g., a user touch on the touch-sensitive display 112, as part of a multi-touch gesture). The peripheral interface 118 transmits information that it receives from the I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via the audio circuitry 110). The information that the peripheral interface 118 receives from the I / O subsystem 106 includes information from the touch-sensitive display 112 or a touch-sensitive surface.
[0164] In some embodiments, the event monitor 171 sends requests to the peripheral interface 118 at predetermined intervals. In response, the peripheral interface 118 transmits event information. In other embodiments, the peripheral interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or receiving an input for longer than a predetermined duration).
[0165] In some embodiments, the event classifier 170 further includes a hit view determination module 172 and / or an active event recognizer determination module 173.
[0166] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where within one or more of the views a sub-event has occurred. A view consists of controls and other elements that a user can see on the display.
[0167] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, within which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events identified as correct inputs is optionally determined at least in part based on the hit view of the initial touch that begins a touch-based gesture.
[0168] The hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest view in the hierarchical structure that should handle the sub-event. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view generally receives all sub-events related to the same touch or input source for which it was identified as the hit view.
[0169] The active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, the active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, the active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views and, thus, determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with a particular view, higher views in the hierarchy will still remain as actively participating views.
[0170] The event dispatcher module 174 distributes event information to event recognizers (e.g., event recognizer 180). In embodiments that include the active event recognizer determination module 173, the event dispatcher module 174 delivers the event information to the event recognizer determined by the active event recognizer determination module 173. In some embodiments, the event dispatcher module 174 stores the event information in an event queue, which is retrieved by the corresponding event receiver 182.
[0171] In some embodiments, the operating system 126 includes the event classifier 170. Alternatively, the application 136-1 includes the event classifier 170. In yet another embodiment, the event classifier 170 is an independent module or is part of another module (such as the contact / motion module 130) stored in the memory 102.
[0172] In some embodiments, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events that occur within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other embodiments, one or more of the event recognizers 180 are part of an independent module that is a higher-level object such as a user interface toolkit or from which application 136-1 inherits methods and other properties. In some embodiments, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from event classifier 170. Event handler 190 optionally utilizes or invokes data updater 176, object updater 177, or GUI updater 178 to update the application's internal state 192. Alternatively, one or more of the application views 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included within a corresponding application view 191.
[0173] A corresponding event recognizer 180 receives event information (e.g., event data 179) from event classifier 170 and identifies an event based on the event information. Event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event recognizer 180 further includes at least a subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0174] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events such as a touch or a touch movement. Depending on the sub-event, the event information further includes additional information such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally further includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device's attitude).
[0175] Event comparator 184 compares the event information with predefined event or sub - event definitions and, based on the comparison, determines an event or sub - event or determines or updates the status of an event or sub - event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains the definition of an event (e.g., a predefined sequence of sub - events), such as event 1 (187 - 1), event 2 (187 - 2), and others. In some embodiments, sub - events in an event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi - touch. In one example, the definition of event 1 (187 - 1) is a double - tap on a displayed object. For example, a double - tap includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift - off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on the displayed object, and a second lift - off (touch end) of a predetermined duration. In another example, the definition of event 2 (187 - 2) is a drag on a displayed object. For example, a drag includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch - sensitive display 112, and lift - off of the touch (touch end). In some embodiments, an event also includes information for one or more associated event handlers 190.
[0176] In some embodiments, event definition 187 includes the definition of an event for a corresponding user interface object. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with a sub - event. For example, in an application view that displays three user interface objects on touch - sensitive display 112, when a touch is detected on touch - sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub - event). If each displayed object is associated with a corresponding event handler 190, event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub - event and the object that triggered the hit test.
[0177] In some embodiments, the definition of a corresponding event (187) also includes a delay action that delays the delivery of event information until it has been determined that the sub - event sequence does or does not correspond to the event type of the event recognizer.
[0178] When the corresponding event recognizer 180 determines that the sub - event sequence does not match any event in the event definition 186, the corresponding event recognizer 180 enters an event - impossible, event - failed, or event - ended state, after which subsequent sub - events of the touch - based gesture are ignored. In such a case, other event recognizers (if any) for which the hit view remains active continue to track and process the ongoing sub - events of the touch - based gesture.
[0179] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists indicating how the event delivery system should perform sub - event delivery to the actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists indicating how event recognizers interact with each other or can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists indicating whether sub - events are delivered to different levels in the view or the programmatic hierarchy.
[0180] In some embodiments, when one or more specific sub - events of an event are recognized, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferring the sending of) sub - events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a token associated with the discerned event, and the event handler 190 associated with that token retrieves the token and executes a predefined process.
[0181] In some embodiments, the event delivery instruction 188 includes a sub - event delivery instruction that delivers event information about the sub - event without activating the event handler. Instead, the sub - event delivery instruction delivers the event information to the event handler associated with the sub - event sequence or to the actively participating view. The event handler associated with the sub - event sequence or with the actively participating view receives the event information and executes a predetermined process.
[0182] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates the phone numbers used in contact module 137 or stores video files used in the video player module. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the positions of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on the touch-sensitive display.
[0183] In some embodiments, event handler 190 includes data updater 176, object updater 177, and GUI updater 178, or has access to the data updater, the object updater, and the GUI updater. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0184] It should be understood that the above discussion of event handling for user touches on the touch-sensitive display also applies to other forms of user input for operating multifunctional device 100 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in cooperation with single or multiple keyboard presses or holds; contact movement on a touchpad, such as tapping, dragging, scrolling, etc.; stylus input; movement of the device; voice commands; detected eye movement; biometric input; and / or any combination thereof are optionally used as inputs corresponding to sub-events that define the events to be discriminated.
[0185] Figure 2FIG. 0 shows a portable multifunctional device 100 having a touch screen 112, according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this and other embodiments described below, a user is capable of selecting one or more of these graphics by making gestures on the graphics, such as by using one or more fingers 202 (not drawn to scale in the figures) or one or more styli 203 (not drawn to scale in the figures). In some embodiments, selection of one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gestures optionally include one or more taps, one or more swipes (from left to right, right to left, up, and / or down), and / or rolling of a finger that has made contact with the device 100 (from right to left, left to right, up, and / or down). In some implementations or in some situations, inadvertently contacting a graphic does not select the graphic. For example, when the gesture corresponding to selection is a tap, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application.
[0186] The device 100 optionally further includes one or more physical buttons, such as a “home” or menu button 204. As previously mentioned, the menu button 204 is optionally used to navigate to any of a set of applications 136 that are optionally executed on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen 112.
[0187] In some embodiments, the device 100 includes a touch screen 112, a menu button 204, a push button 206 for powering the device on / off and for locking the device, one or more volume adjustment buttons 208, a subscriber identity module (SIM) card slot 210, an earphone jack 212, and a docking / charging external port 124. The push button 206 is optionally used to power the device on / off by pressing the button and holding the button in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In an alternative embodiment, the device 100 also receives voice input for activating or deactivating certain functions via a microphone 113. The device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of contacts on the touch screen 112, and / or one or more tactile output generators 167 for generating tactile output for a user of the device 100.
[0188] Figure 3is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, in accordance with some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a gaming system, or a control device (e.g., a home controller or an industrial controller). Device 300 generally includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes termed a chipset) that interconnects system components and controls the communication between them. Device 300 includes an input / output (I / O) interface 330 having a display 340, which is typically a touchscreen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to haptic output generator 167 described above with reference to Figure 1A ), sensors 359 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors (similar to contact intensity sensor 165 described above with reference to Figure 1A ). Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU 310. In some embodiments, memory 370 stores programs, modules, and data structures similar to or a subset of the programs, modules, and data structures stored in memory 102 of portable multifunctional device 100 ( Figure 1A ). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunctional device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while memory 102 of portable multifunctional device 100 ( Figure 1A ) optionally does not store these modules.
[0189] Figure 3Each of the above elements in [element name] is optionally stored in one or more of the previously mentioned memory devices of the memory device. Each of the above modules corresponds to a set of instructions for performing the above functions. The above modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, the memory 370 optionally stores a subgroup of the above modules and data structures. Additionally, the memory 370 optionally stores additional modules and data structures not described above.
[0190] Attention is now turned to embodiments of a user interface that is optionally implemented on, for example, the portable multifunctional device 100.
[0191] Figure 4A An exemplary user interface of an application menu on the portable multifunctional device 100 according to some embodiments is shown. A similar user interface is optionally implemented on the device 300. In some embodiments, the user interface 400 includes the following elements or subsets or supersets thereof:
[0192] · A signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0193] · Time 404;
[0194] · A Bluetooth indicator 405;
[0195] · A battery status indicator 406;
[0196] · A tray 408 with icons for commonly used applications, such icons as:
[0197] ○ An icon 416 marked "Phone" for the phone module 138, which icon 416 optionally includes an indicator 414 of the number of missed calls or voicemails;
[0198] ○ An icon 418 marked "Mail" for the email client module 140, which icon 418 optionally includes an indicator 410 of the number of unread emails;
[0199] ○ An icon 420 marked "Browser" for the browser module 147; and
[0200] ○ An icon 422 marked "iPod" for the video and music player module 152 (also referred to as the iPod (trademark of Apple Inc.) module 152); and
[0201] · Icons for other applications, such as:
[0202] ○ The icon 424 of the IM module 141 labeled "Message";
[0203] ○ The icon 426 of the calendar module 148 labeled "Calendar";
[0204] ○ The icon 428 of the image management module 144 labeled "Photo";
[0205] ○ The icon 430 of the camera module 143 labeled "Camera";
[0206] ○ The icon 432 of the online video module 155 labeled "Online Video";
[0207] ○ The icon 434 of the stock market widget 149-2 labeled "Stock Market";
[0208] ○ The icon 436 of the map module 154 labeled "Map";
[0209] ○ The icon 438 of the weather widget 149-1 labeled "Weather";
[0210] ○ The icon 440 of the alarm clock widget 149-4 labeled "Clock";
[0211] ○ The icon 442 of the fitness support module 142 labeled "Fitness Support";
[0212] ○ The icon 444 of the note module 153 labeled "Note"; and
[0213] ○ The icon 446 of the settings application or module labeled "Settings", which provides access to the settings of the device 100 and its various applications 136.
[0214] It should be noted that Figure 4A The icon labels shown are merely exemplary. For example, the icon 422 of the video and music player module 152 is labeled "Music" or "Music Player". Other labels are optionally used for the various application icons. In some embodiments, the label of the corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.
[0215] Figure 4B A device (e.g., Figure 3 is shown having a touch-sensitive surface 451 (e.g., Figure 3Exemplary user interfaces on device 300). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting the intensity of contacts on the touch-sensitive surface 451 and / or one or more tactile output generators 357 for generating tactile output for a user of device 300.
[0216] Although some examples below will be given with reference to input on a touch screen display 112 (where the touch-sensitive surface and the display are combined), in some embodiments, the device detects input on a touch-sensitive surface separate from the display, as Figure 4B shown. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451 in ) has a main axis corresponding to the main axis (e.g., Figure 4B 453 in ) on the display (e.g., Figure 4B 452 in ). According to these embodiments, the device detects contacts (e.g., Figure 4B 460 and 462 in ) with the touch-sensitive surface 451 at positions corresponding to corresponding positions on the display (e.g., in Figure 4B 460 corresponds to 468 and 462 corresponds to 470). Thus, when the touch-sensitive surface (e.g., Figure 4B 451 in ) is separate from the display of the multifunctional device (e.g., Figure 4B 450 in ), user input detected by the device on the touch-sensitive surface (e.g., contacts 460 and 462 and their movements) is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0217] Additionally, while the examples below are mainly given with reference to finger input (e.g., finger contact, single-finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click when the cursor is above the position of the tap gesture (e.g., instead of detecting a contact, followed by stopping detection of the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or a mouse and finger contact are optionally used simultaneously.
[0218] Figure 5AAn exemplary personal electronic device 500 is shown. The device 500 includes a body 502. In some embodiments, the device 500 may include some or all of the features described with respect to devices 100 and 300 (e.g., Figures 1A to 4B ). In some embodiments, the device 500 has a touch-sensitive display screen 504 hereinafter referred to as a touch screen 504. As an alternative or addition to the touch screen 504, the device 500 has a display and a touch-sensitive surface. As in the case of devices 100 and 300, in some embodiments, the touch screen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 504 (or the touch-sensitive surface) may provide output data representative of the intensity of the touch. The user interface of the device 500 may respond to the touch based on the intensity of the touch, meaning that touches of different intensities may invoke different user interface operations on the device 500.
[0219] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial Number PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial Number PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.
[0220] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) can be in physical form. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) may allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear device 500.
[0221] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include some or all of the components described with reference to Figure 1A , Figure 1B and Figure 3 . Device 500 has a bus 512 that operatively couples the I / O section 514 to one or more computer processors 516 and a memory 518. The I / O section 514 may be connected to a display 504 that may have a touch-sensitive component 522 and optionally an intensity sensor 524 (e.g., a contact intensity sensor). Additionally, the I / O section 514 may be connected to a communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 may include an input mechanism 506 and / or 508. For example, input mechanism 506 is optionally a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 508 is optionally a button.
[0222] In some examples, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.
[0223] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform, for example, the techniques described below, including processes 800, 900, 1100, 1300, 1500, and 1700. A computer-readable storage medium can be any medium that tangibly contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. In some examples, the storage medium is a transitory computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. The personal electronic device 500 is not limited to Figure 5B the components and configurations thereof, but may include other components or additional components in a variety of configurations.
[0224] As used herein, the term "enabling representation" refers to a user-interactive graphical user interface object optionally displayed on the display screen of devices 100, 300, and / or 500 ( Figure 1A , Figure 3 and Figures 5A to 5B ). For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) each optionally constitute an enabling representation.
[0225] As used herein, the term "focus selector" refers to an input element for indicating the current part of the user interface with which the user is interacting. In some embodiments including a cursor or other position marker, the cursor acts as the "focus selector" such that when an input (e.g., a press input) is detected on a touch-sensitive surface (e.g., Figure 3 the touchpad 355 in Figure 4B ) above a particular user interface element (e.g., a button, a window, a slider, or other user interface element), the particular user interface element is adjusted in accordance with the detected input. In embodiments including a touchscreen display capable of enabling direct interaction with user interface elements on the touchscreen display (e.g., Figure 1A the touch-sensitive display system 112 in Figure 4AIn some specific implementations of the touch screen 112), the detected contact on the touch screen serves as a "focus selector", such that when an input (e.g., a press input made by the contact) is detected at the position of a specific user interface element (e.g., a button, a window, a slider, or other user interface element) on the touch screen display, the specific user interface element is adjusted according to the detected input. In some specific implementations, the focus moves from one area of the user interface to another area of the user interface without a corresponding movement of the cursor or a movement of the contact on the touch screen display (e.g., moving the focus from one button to another button by using the tab key or arrow keys); in these specific implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is generally a user interface element (or a contact on the touch screen display) that is controlled by the user in order to deliver the interaction with the user interface that the user anticipates (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or a touch screen), the position of the focus selector (e.g., a cursor, a contact, or a selection box) above the corresponding button will indicate that the user desires to activate the corresponding button (rather than other user interface elements shown on the device display).
[0226] As used in the specification and claims, the term "feature intensity" of a contact refers to a feature of the contact based on one or more intensities of the contact. In some embodiments, the feature intensity is based on a plurality of intensity samples. The feature intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected during a predefined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after detecting the contact, before detecting the lift-off of the contact, before or after detecting the start of movement of the contact, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The feature intensity of the contact is optionally based on one or more of the following: the maximum value of the intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the value at the top 10% of the intensity of the contact, the half-maximum value of the intensity of the contact, the 90% maximum value of the intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the feature intensity (e.g., when the feature intensity is the average value of the intensity of the contact over time). In some embodiments, the feature intensity is compared with a set of one or more intensity thresholds to determine whether the user has performed an operation. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact with a feature intensity not exceeding the first threshold results in a first operation, a contact with a feature intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second operation, and a contact with a feature intensity exceeding the second threshold results in a third operation. In some embodiments, the comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to forgo performing the corresponding operation) rather than to determine whether to perform a first operation or a second operation.
[0227] In some embodiments, a portion of the gesture is identified for determining the feature intensity. For example, the touch-sensitive surface optionally receives a continuous swipe contact that transitions from a starting position and reaches an ending position where the contact intensity increases. In this example, the feature intensity of the contact at the ending position is optionally based on only a portion of the continuous swipe contact rather than the entire swipe contact (e.g., only the portion of the swipe contact at the ending position). In some embodiments, a smoothing algorithm is optionally applied to the intensity of the swipe contact before determining the feature intensity of the contact. For example, the smoothing algorithm optionally includes one or more of the following: an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swipe contact for the purpose of determining the feature intensity.
[0228] Optionally, characterize the contact intensity on the touch-sensitive surface relative to one or more intensity thresholds such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device will perform an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device will perform an operation different from an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, when a contact is detected with a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold, contacts below the nominal contact detection intensity threshold are no longer detected), the device will move the focus selector based on the movement of the contact on the touch-sensitive surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally speaking, unless otherwise stated, these intensity thresholds are consistent between different sets of user interface figures.
[0229] An increase in the contact characteristic intensity from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in the contact characteristic intensity from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in the contact characteristic intensity from an intensity below the contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in the contact characteristic intensity from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting the contact being lifted from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.
[0230] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or contacts), where the corresponding press input is detected at least in part based on the detected intensity of the contact (or contacts) increasing above a press input intensity threshold. In some embodiments, a corresponding operation is performed in response to detecting the intensity of the corresponding contact increasing above the press input intensity threshold (e.g., the "down stroke" of the corresponding press input). In some embodiments, the press input includes the intensity of the corresponding contact increasing above the press input intensity threshold and the intensity of the contact subsequently decreasing below the press input intensity threshold, and a corresponding operation is performed in response to detecting the intensity of the corresponding contact subsequently decreasing below the press input threshold (e.g., the "up stroke" of the corresponding press input).
[0231] In some embodiments, the device employs hysteresis to avoid spurious inputs sometimes referred to as "jitter", where the device defines or selects a hysteresis intensity threshold having a predefined relationship to a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Thus, in some embodiments, a press input includes the intensity of a corresponding contact increasing above the press input intensity threshold and the intensity of that contact subsequently decreasing below the hysteresis intensity threshold corresponding to the press input intensity threshold, and a corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases below the hysteresis intensity threshold (e.g., the "upstroke" of the corresponding press input). Similarly, in some embodiments, a press input is detected only when the device detects that the contact intensity increases from an intensity equal to or lower than the hysteresis intensity threshold to an intensity equal to or higher than the press input intensity threshold and optionally the contact intensity subsequently decreases to an intensity equal to or lower than the hysteresis intensity threshold, and a corresponding operation is performed in response to detecting the press input (e.g., depending on the context, the contact intensity increases or the contact intensity decreases).
[0232] For ease of explanation, optionally, a description of an operation performed in response to a press input associated with a press input intensity threshold or in response to a gesture including a press input is triggered in response to detecting any one of the following various conditions: the contact intensity increases above the press input intensity threshold, the contact intensity increases from an intensity lower than the hysteresis intensity threshold to an intensity higher than the press input intensity threshold, the contact intensity decreases below the press input intensity threshold, and / or the contact intensity decreases below the hysteresis intensity threshold corresponding to the press input intensity threshold. Additionally, in an example where an operation is described as being performed in response to detecting that the intensity of a contact decreases below the press input intensity threshold, the operation is optionally performed in response to detecting that the intensity of the contact decreases below the hysteresis intensity threshold corresponding to and less than the press input intensity threshold.
[0233] As used herein, an "installed application" refers to a software application that has been downloaded to an electronic device (e.g., devices 100, 300, and / or 500) and is ready to be launched (e.g., made open) on the device. In some embodiments, a downloaded application becomes an installed application using an installer that extracts program parts from the downloaded software package and integrates the extracted parts with the operating system of the computer system.
[0234] As used herein, the term "open application" or "executing application" refers to a software application that has maintained state information (e.g., as part of the device / global internal state 157 and / or the application internal state 192). An open or executing application is optionally any one of the following types of applications:
[0235] · The active application currently being displayed on the display screen of the device on which the application is being used;
[0236] · Background applications (or background processes), which are not currently being displayed but one or more processes of the application are being processed by one or more processors; and
[0237] · Suspended or hibernated applications that are not running but have state information stored in the memory (volatile and non-volatile respectively)
[0238] and that can be used to resume the execution of the application.
[0239] As used herein, the term "closed application" refers to a software application that does not maintain state information (e.g., the state information of a closed application is not stored in the memory of the device). Thus, closing an application includes stopping and / or removing the application process of the application and removing the state information of the application from the memory of the device. Generally speaking, when in a first application, opening a second application does not close the first application. When the second application is being displayed and the first application stops being displayed, the first application becomes a background application.
[0240] Attention is now turned to embodiments of a user interface ("UI") implemented on an electronic device (such as the portable multifunctional device 100, device 300, or device 500) and associated processes.
[0241] Figures 6A to 6Z Exemplary user interfaces for managing visual content in media are shown in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including Figure 8 the processes in.
[0242] Figure 6A A computer system 600 (e.g., an electronic device) is shown displaying a camera user interface that includes a live preview 630 that optionally extends from the top of the display of the computer system 600 to the bottom of the display of the computer system 600. In some embodiments, the computer system 600 optionally includes one or more features of the device 100, device 300, or device 500. In some embodiments, the computer system 600 is a tablet, a phone, a laptop, a desktop computer, etc.
[0243] The live preview 630 is a representation of the field of view (“FOV”) of one or more cameras of the computer system 600. In some embodiments, the live preview 630 is a representation of a partial FOV. In some embodiments, the live preview 630 is based on images detected by one or more camera sensors. In some embodiments, the computer system 600 captures images using multiple camera sensors and combines them to display the live preview 630. In some embodiments, the computer system 600 captures images using a single camera sensor to display the live preview 630.
[0244] Figure 6A The camera user interface includes an indicator region 602 and a control region 606 that are positioned relative to the live preview 630 such that the indicators and controls can be displayed simultaneously with the live preview 630. The camera display region 604 is substantially free of indicators and / or controls. As Figure 6A shown, the camera user interface includes a visual boundary 608 that indicates the boundary between the indicator region 602 and the camera display region 604 and the boundary between the camera display region 604 and the control region 606.
[0245] As Figure 6A shown, the indicator region 602 includes indicators such as a flash indicator 602a and an animated image indicator 602b. The flash indicator 602a indicates whether the flash mode is on (e.g., active), off (e.g., inactive), or in another mode (e.g., auto mode). In Figure 6A this example, the flash indicator 602a indicates to the user that the flash mode is off and the flash operation will not be used when the computer system 600 is capturing media. Additionally, the animated image indicator 602b indicates whether the camera is configured to capture a single image or multiple images (e.g., in response to detecting a request to capture media). In some embodiments, the indicator region 602 overlays the live preview 630 and optionally includes a coloring (e.g., gray; translucent) overlay.
[0246] As Figure 6A shown, the camera display region 604 includes the live preview 630 and zoom controls (e.g., enabling representations) 622. The zoom controls 622 include a 0.5x zoom control 622a, a 1x zoom control 622b, and a 2x zoom control 622c. As Figure 6A shown, the 1x zoom control 622b is bolded and enlarged compared to the other zoom controls, indicating that the 1x zoom control 622b is selected and the computer system 600 is displaying the live preview 630 at a “1x” zoom level.
[0247] As Figure 6AAs shown, the control region 606 includes representations of a camera mode control (e.g., controls) 620, a shutter control 610, a camera switcher control 614, and a media collection 612. In Figure 6A , camera mode controls 620a - 620e are shown, and the 'photo' camera mode 620c is bolded, indicating that the computer system 600 is configured to capture photo media when the shutter control 610 is activated. Thus, when activated, the shutter control 610 causes the computer system 600 to capture media (e.g., a photo when the shutter control 610 is activated in Figure 6A ) using one or more camera sensors based on the current state of the live preview 630 and the current state of the camera application (e.g., which camera mode is selected). The captured media is stored locally at the computer system 600 and / or sent to a remote server for storage. When activated, the camera switcher control 614 causes the computer system 600 to switch to displaying the field of view of a different camera in the live preview 630, such as by switching between a rear camera sensor and a front camera sensor. Figure 6A The representation of the media collection 612 shown in Figure 7B is a representation of the media (e.g., images, videos) most recently captured by the computer system 600. In some embodiments, in response to detecting a gesture directed at the media collection 612, the computer system 600 displays a user interface similar to the one shown in Figure 6A (discussed below). In some embodiments, an indicator region 602 is overlaid on the live preview 630 and optionally includes a colored (e.g., gray; translucent) overlay. In Figure 6A , the computer system detects a tap input 650a on (and / or directed at) the shutter control 610.
[0248] As Figure 6B shown, in response to detecting the tap input 650a, the computer system 600 initiates media capture to capture the Figure 6A live preview 630 and displays a new representation in the media collection 612. In Figure 6B , the new representation is a representation of the Figure 6A live preview 630 (e.g., captured in response to detecting the tap input 650a on the shutter control 610). Additionally, the new representation is shown on top of the media collection 612 in Figure 6B because the new representation corresponds to a representation of the most recently captured media.
[0249] As Figure 6BAs shown, the live preview 630 includes a representation showing a person 640 standing behind a tree, where the head of the person 640 and a part of the body of the person 640 are not blocked by the tree. Located on the tree is a marker 642, which includes a text portion 642a (e.g., "LOST DOG") and a text portion 642b (e.g., a text paragraph starting with "LOVEABLE"). In Figure 6B the text in the text portions 642a - 642b is not visually prominent, and in Figure 6B the embodiments shown, the text in the text portions 642a and 642b is small and not easily readable by a user viewing the computer system 600. In Figure 6B the computer system 600 detects an expansion input 650b on the live preview 630.
[0250] As Figure 6C shown, in response to detecting the expansion input 650b, the computer system 600 replaces the display of the 2x zoom control 622c in Figure 6B with the display of a 2.5x zoom control. Additionally, the computer system 600 updates the live preview 630 to reflect the change in the zoom level such that objects in the field of view of one or more cameras are displayed at a "2.5x" zoom level (e.g., as shown by the newly displayed and selected (e.g., enlarged and bolded) 2.5x zoom control 622d), rather than Figure 6B being displayed at a "1x" zoom level.
[0251] When compared with Figure 6B the text portions 642a - 642b in Figure 6C are more visually prominent (e.g., larger, more readable) than the text portions 642a - 642b in Figure 6B . In Figure 6C it is determined that the text portions 642a - 642b (and / or the text included in the text portions 642a - 642b) respectively meet a set of prominence criteria. Figure 6C The text portions 642a - 642b in Figures 7A to 7L meet the set of prominence criteria because each text portion occupies more than a threshold portion (e.g., 10%) of the live preview 630 and / or each text portion includes text larger than a threshold size (e.g., larger than 6pt font). In some embodiments, one or more text portions meet the set of prominence criteria based on other criteria, such as whether the corresponding text portion includes one or more text types (e.g., email, phone number, Quick Response ("QR") code, etc.), whether the corresponding text portion is displayed at a specific location (e.g., the center location) or near the live preview 630, whether the corresponding text portion is relevant, etc. (based on the context of the media shown as the live preview 630) (as regarding Figures 7A to 7L , Figure 8 andFigure 9 discussed in more detail).
[0252] As Figure 6C shown, since it is determined that the text portions 642a - 642b meet the set of prominence criteria (and / or since at least a portion of the text meets the set of prominence criteria), the computer system 600 displays brackets 636a around the text portions 642a - 642b in the camera display area 604 and displays text management control 680 to the right of the zoom control 622d in that camera display area.
[0253] Review Figure 6B , the brackets 636a and the text management control 680 are not displayed Figure 6B because it is determined that the text portions 642a - 642b do not meet the set of prominence criteria (and / or because no text portion meets the set of prominence criteria). In Figure 6B , it is determined that the text portions 642a - 642b do not meet the set of prominence criteria because the text portions 642a - 642b do not occupy more than a threshold portion of the live preview 630 and do not include text larger than a threshold size. In some embodiments (as Figures 6B to 6C shown), it is determined whether a corresponding text portion meets the set of prominence criteria based on how / when the text portion is currently displayed in the live preview 630 rather than just based on whether the live preview 630 includes text (and / or a text portion).
[0254] Returning to Figure 6C , the brackets 636a are positioned around the image of the dog on the marker 642 because the image of the dog is positioned between the text portions 642a - 642b. In some embodiments, multiple brackets are displayed such that one bracket is displayed around the text portion 642a and another bracket is displayed around the text portion 642b. In some embodiments, multiple brackets are displayed because it is determined that multiple text portions (e.g., "portions of the text") meet the set of prominence criteria and an object is located between the text portions. In some embodiments, when the object is not positioned between multiple portions of text, only one bracket is displayed around the multiple portions of text. In some embodiments, in a case where the text portion 642a meets the set of prominence criteria but the text portion 642b does not meet the set of prominence criteria, a bracket is displayed around the text portion 642a and no bracket is displayed around the text portion 642b (and vice versa). In some embodiments, in addition to and / or instead of displaying brackets around the corresponding portions of text, the computer system 600 indicates that a corresponding portion of text (e.g., a portion of the text) meets the set of prominence criteria by emphasizing the corresponding portion in other ways, such as by highlighting, bolding, resizing, displaying a box around the corresponding portion of text.
[0255] AsFigure 6C As shown, computer system 600 displays text type indicators 638a - 638b (e.g., underlined) to indicate that a specific text type (e.g., email, address, phone number, QR code, etc.) has been detected (e.g., by a data detector) in text portion 642b. In Figure 6C , text type indicator 638a is displayed under "123 Main Street" to indicate that an address has been detected, and text type indicator 638b is displayed under "123 - 4567" to indicate that a phone number has been detected. In some embodiments, when a text type indicator is displayed under a portion of text, the user can select the text portion and / or the text type indicator to perform an operation (e.g., as further discussed below with respect to Figures 6M to 6N ).
[0256] Figures 6C to 6D An exemplary embodiment showing computer system 600 moving in a physical environment is presented. Figures 6C to 6D It includes a graphical representation 660, which shows the original position 660a of computer system 600 (e.g., in Figures 6C to 6D ) relative to the changed position 660b of computer system 600 in the physical environment (e.g., in Figure 6D ). As Figure 6C shown, computer system 600 is in the original position 660a. In Figure 6C , the position of computer system 600 changes.
[0257] As Figure 6D shown, in response to the change in the position of computer system 600 (e.g., from the original position 660a to the changed position 660b), computer system 600 transitions the live preview in an upward direction. In Figure 6D , the live preview 630 is transitioned in the upward direction such that Figure 6C the top portion of the live preview 630 (e.g., the portion including text portion 642a) is stopped from being displayed, and the new bottom portion of the new live preview 630 is displayed (as Figure 6D shown). In Figure 6D , it is determined that text portion 642a does not meet the set of prominence criteria while text portion 642b does meet (or continues to meet) the set of prominence criteria. Here, it is determined that text portion 642a does not meet the set of prominence criteria because text portion 642a is no longer being displayed as Figure 6Da portion of the live preview 630 (e.g., in the camera display area). As shown, since the text portion 642a does not meet the set of prominence criteria and the text portion 642b meets the set of prominence criteria, the computer system 600 displays the bracket 636b around the text portion 642b (instead of the text portion 642a) and stops displaying the bracket 636a. In other words, the computer system 600 dynamically changes the bracket 636a to the bracket 636b based on a change in determining whether one or more text portions (e.g., text portions currently being displayed as part of the live preview 630) meet and / or do not meet the set of prominence criteria. Thus, one or more determinations of whether one or more text portions meet the set of prominence criteria are dynamic and can change when the live preview 630 changes in response to requests for zooming in (e.g., expansion input) / zooming out (e.g., pinch input), panning (e.g., right, left, up, down swipe input), and / or movement of the computer system 600 (e.g., forward, backward, up, down) and / or changes in one or more cameras of the computer system 600. In some embodiments, when one or more determinations of whether one or more text portions meet the set of prominence criteria change, the display of one or more brackets (e.g., brackets 636a - 636b) and / or the display of the text management control 680 change (as further described below with respect to Figures 7A to 7L , Figure 8 , Figure 9 ). In some embodiments, when the computer system 600 displays the bracket 636a around the text portion 642b (and / or in response to detecting text in the live preview 630), the computer system 600 dims and / or desaturates (e.g., vividness, color, and / or hue) portions of the live preview 630 that do not have text (e.g., a photo of a dog), while maintaining the saturation and / or brightness of the text portion 642b (and / or other portions of the text). In some embodiments, as part of dimming portions of the live preview 630 that do not have text and maintaining the brightness of the text portion 642b, the computer system 600 displays the text portion 642b with a greater amount of brightness than portions of the live preview 630 that do not have text.
[0258] As Figure 6D shown, since it is determined that the text portion 642b does (or continues to) meet the set of prominence criteria, the computer system 600 continues to display the text management control 680. In Figure 6D , the text management control 680 is displayed because at least one determination is made that the currently displayed text portion (e.g., the text portion of the live preview 630) meets the set of prominence criteria, regardless of whether another text portion (e.g., the text portion 642a) fails to continue (or not) to meet the set of prominence criteria. In Figure 6DIn this case, computer system 600 moves back to the original position 660a.
[0259] As Figure 6E shown, in response to computer system 600 being in the original position 660a, computer system 600 uses one or more of the techniques described above with respect to Figure 6C to redisplay the live preview 630. In Figure 6E this case, computer system 600 detects a tap input 650e on the text management control 680.
[0260] As Figure 6F shown, in response to detecting the tap input 650e, computer system 600 changes the display of the text management control 680. Specifically, computer system 600 displays the text management control 680 in an active and / or selected state (e.g., as shown by the text management control 680 that is bold as in Figure 6F ) and stops displaying the text management control 680 in an inactive and / or deselected state (e.g., as shown by the text management control 680 that is not bold as in Figure 6F ). Figure 6E In this case, in response to detecting the tap input 650e, computer system 600 emphasizes the text portions 642a - 642b and dims other portions of the live preview 630 (and / or other objects in the field of view of one or more cameras), such as the person 640, the image of the dog on the marker 642, and the tree shown in the live preview 630. Along with dimming other portions of the live preview 630, computer system 600 stops displaying one or more controls (e.g.,
[0261] As Figure 6F the zoom control 622 as in Figure 6E ) in the camera display area 604. Additionally, computer system 600 also dims (or stops displaying) portions of the camera user interface, such as indicators in the indicator area 602 and controls in the camera control area 606. In some embodiments, in Figure 6FSome of the dimmed indicators and / or controls in the camera user interface are not selectable (e.g., do not cause the computer system 600 to perform an action when selected). In some embodiments, in response to detecting a tap input 650e, some of the indicators and / or controls remain selectable and / or do not dim. In some embodiments, in response to detecting a tap input 650e, the computer system 600 maintains the display of some of the controls in the camera display area 604. In some embodiments, the computer system 600 emphasizes portions 642a - 642b by increasing the size of the text in text portions 642a - 642b, highlighting the text in text portions 642a - 642b, displaying a box around text portions 642a - 642b, etc. In some embodiments, dimming portions of the live preview 630 includes reducing the saturation of portions of the live preview 630 that do not have text (e.g., a photo of a dog), while maintaining the saturation of text portions 642a and 642b (e.g., using a similar technique as described above with respect to Figure 6D ).
[0262] Notably, in Figure 6F , the portion of the text that is emphasized in response to detecting the input 650e is the portion of the text enclosed by parentheses (e.g., parentheses 636a) when the input 650e is received in Figure 6E . In some embodiments, the parentheses around a portion of the text indicate to the user which text will be emphasized and / or managed when a selection of the text management control 680 occurs. In some embodiments, in response to selecting the text management control 680, one or more text portions that are displayed via the live preview 630 but do not have parentheses enclosing it when the input is received on the text management control 680 are not emphasized (e.g., in Figure 7F , "BRAND" is not emphasized when the text management control 680 is selected in Figure 7F ). In some embodiments, in response to selecting the text management control 680 (e.g., if it is determined that one or more portions of the text meet a set of prominence criteria), one or more text portions that are displayed via the live preview 630 but do not have parentheses around it when the input is received on the text management control 680 are emphasized.
[0263] As Figure 6F shows, in response to detecting a tap input 650e, the computer system 600 also displays text management options 682 and instructions 684 indicating one or more inputs / gestures available for selecting a subset of the text in text portions 642a - 642b (e.g., "SWIPE OR TAP TO SELECT TEXT"). In Figure 6FIn [the context], the text management option 682 is an option for managing the text portions 642a - 642b. Specifically, the text management option 682 includes a copy option 682a, a select all option 682b, a find option 682c, and a share option 682d. In some embodiments, in response to receiving an input directed to the copy option 682a, the computer system 600 copies the selected text (e.g., Figure 6F the text in the text portions 642a - 642b in [the context]) and / or saves the selected text in a copy / paste buffer, which allows the selected text to be pasted in response to receiving a request to paste the selected text. In some embodiments, in response to receiving an input directed to the select all option 682b, the computer system 600 selects all the text that is emphasized on the computer system 600. In some embodiments, when the computer system 600 selects all the text in the selected text, the computer system 600 highlights the selected text. In some embodiments, in response to receiving an input directed to the find option 682c, the computer system 600 searches for the selected text (e.g., Figure 6F the emphasized text portion in [the context]) via a search application (e.g., a web application, a dictionary application, a personal assistant application) and / or displays one or more definitions and resources for the emphasized and / or selected text. In some embodiments, in response to receiving an input directed to the share option 682d, the computer system 600 initiates a process of sharing the selected text via one or more applications (e.g., email, text messaging, word processing, social media applications) (e.g., one or more pre - defined applications). In some embodiments, as part of the process of initiating the sharing of the selected text, the computer system 600 displays a scrollable list of applications, where selecting an application from the scrollable list of applications causes the computer system 600 to share the selected text using the selected application. In some embodiments, the scrollable list of applications is displayed simultaneously with a part of the live preview 630 (e.g., including one or more of the text portions 642a - 642b). In Figure 6F [the context], the computer system 600 detects a tap input 650f on a part of the live preview 630 (e.g., a part in the darkened area of the live preview 630 and / or a part of the live preview 630 that does not include the text portions 642a - 642b and / or the text management control 680).
[0264] As Figure 6GAs shown, in response to detecting a tap input 650f, the computer system 600 displays the text management control 680 in an inactive state, de-emphasizes the text portions 642a - 642b, brightens other portions of the live preview 630 and the camera user interface, and stops displaying the text management options 682 and instructions 684. Additionally, in response to detecting a tap input 650f, the computer system 600 redisplay the brackets 636a because it is determined that the text portions 642a - 642b meet (or continue to meet) the set of prominence criteria. Effectively, in response to detecting a tap input 650f, the camera user interface returns to the state it was in before the tap input 650e was detected on the text management control 680. In Figure 6G the computer system 600 detects a tap input 650g on the text management control 680.
[0265] As Figure 6H shown, in response to detecting a tap input 650g, the computer system 600 uses one or more of the techniques described above with respect to Figure 6F to display the Figure 6H camera user interface. Notably, in Figure 6H (and in Figure 6F ), the computer system 600 emphasizes the text portions 642a - 642b because the text portions 642a - 642b meet the set of prominence criteria. In some embodiments, the computer system 600 dims one or more portions of the text that do not meet the set of prominence criteria in response to detecting a tap input 650g. In Figure 6H the computer system 600 detects a tap input 650h on the text portion 642b.
[0266] As Figure 6I shown, in response to detecting a tap input 650h, the computer system 600 selects the text portion 642b and relocates the text management option 682 such that the text management option 682 is displayed above the text portion 642b in Figure 6I instead of above the text portion 642a (e.g., as Figure 6H shown). The text management option 682 is relocated to indicate that the text management option is available for managing the text in the text portion 642b and not for managing the text in the text portion 642a. Thus, in other words, the computer system 600 changes the text selected to be managed using the text management option 682 in response to detecting an input (e.g., a swipe or a tap) that selects a specific portion of the text.
[0267] Notably, Figure 6I the live preview 630 of Figure 6H does not include the person 640 included in the live preview 630 ofFigure 6I is moved behind the tree in the live preview 630 and is thus out of the field of view of one or more cameras of the computer system 600. As Figure 6I shown, when the text management control 680 is displayed in an active state and / or when the text management option 682 is displayed, the live preview 630 continues to update to reflect changes in the field of view of one or more cameras of the computer system 600. In some embodiments, when the text management control 680 is displayed in an active state and / or when the text management option 682 is displayed, the live preview 630 does not continue to update. Thus, in embodiments where the live preview 630 is not updated, the computer system 600 will Figure 6I maintain the display of the portion of the person 640 that protrudes from behind the tree in the live preview 630 of Figure 6I . In
[0268] As Figure 6J shown, in response to detecting the expansion input 650i, the computer system 600 displays the live preview 630 at an increased zoom level and maintains the display of the text portion 642b and the text management option 682. In some embodiments, the computer system 600 continues to display at least a subset of the text portion 642b in response to a request to zoom in (e.g., expansion input) (and / or zoom out, pan, and / or move the computer system 600 and / or one or more cameras of the computer system 600) because the text portion 642b is selected. In some embodiments, the display of the selected text portion (e.g., text portion 642b) is static. Thus, in some embodiments where the selected text portion is static, the computer system 600 continues to display the selected text portion regardless of whether the selected text portion remains in the field of view of one or more cameras (e.g., as further described below with respect to Figures 6L to 6M ) (e.g., when the computer system 600 is moved, panned, and / or zoomed, etc.) (e.g., while the camera user interface remains displayed). In Figure 6J , the computer system 600 detects a tap input 650j on the word "Fluffy", which is a word included in the text portion 642b.
[0269] As Figure 6K shown, in response to detecting the tap input 650j, the computer system 600 selects and highlights the word "Fluffy". In Figure 6K , the text management option 682 shown in Figure 6K can be used to manage only the selected word "Fluffy". In Figure 6K , the computer system 600 detects a left swipe input 650k starting from the word "Fluffy".
[0270] AsFigure 6L As shown, in response to detecting a left-swipe input 650k, computer system 600 selects and highlights a plurality of words included in text portion 642b based on the direction of the swipe input 650k. As Figure 6L shown, the word "THE NAME FLUFFY" is highlighted to indicate that "THE NAME FLUFFY" has been selected based on the swipe input 650k. In Figure 6L , the text management options 682 shown in Figure 6L can be used to manage only the selected word "THE NAME FLUFFY".
[0271] Figures 6L to 6M An exemplary implementation is shown in which computer system 600 moves within a physical environment and computer system 600 continues to display the selected text portion (or a text portion that is a subset of the selected text portion), regardless of whether the selected text portion remains within the field of view of one or more cameras (e.g., as further described below with respect to Figures 6L to 6M ). Figures 6L to 6M includes a graphical representation 660 that shows the original position 660a of computer system 600 (e.g., in Figure 6M ) relative to a changed position 660c of computer system 600 (e.g., in Figures 6L to 6M ).
[0272] As Figure 6L shown, the tree marker 646 represents a static portion of the tree shown in the live preview 630 of Figures 6L to 6M . In Figure 6L , the tree marker 646 is shown below text portion 642b. In Figure 6L , the position of computer system 600 changes.
[0273] As Figure 6M shown, in response to a change in the positioning of computer system 600 (e.g., as shown by changed position 660c relative to original position 660a), computer system 600 updates the live preview such that the tree marker 646 is shown above text portion 642b. Notably, in Figure 6M , text portion 642b of computer system 600 is no longer within the field of view of one or more cameras, such that text portion 642b will be located at the position of live preview 630 where text portion 642b was shown as Figure 6M (e.g., as evidenced by the movement of tree marker 646 to a higher position in live preview 630). However, computer system 600 continues to display in Figure 6MThe text portion 642b is shown in the live preview 630 because a subset of the text portion 642b is selected (e.g., "named FLUFFY"). In some embodiments, the computer system 600 shows only the selected subset of the text portion 642b and does not show other portions of the unselected text portion 642b. In some embodiments, when text is selected and the camera moves (and / or zooms / pans) in the physical environment, the computer system 600 does not update the live preview 630. In some embodiments, the computer system 600 selects a different portion of the text in response to the computer system 600 and / or the camera of the computer system 600 being moved (e.g., and / or zoomed / panned) (e.g., as further described below with respect to Figures 7A to 7L and Figure 8 and Figure 9 ). In Figure 6M , the computer system 600 detects an input 650m on "123 - 4567", under which a text type indicator 638b is shown.
[0274] As Figure 6N shown, in response to detecting the input 650m and because it is determined that the input 650m is a tap input and "123 - 4567" corresponds to a phone number, the computer system 600 shows a phone dialer user interface and automatically (e.g., without user input on a keypad and / or contact information card) initiates a phone call to "123 - 4567". In some embodiments, a confirmation screen is shown before the computer system 600 initiates the phone call.
[0275] As Figure 6O shown, in response to detecting the input 650m and because it is determined that the input 650m is a press - hold input and "123 - 4567" corresponds to a phone number, the computer system 600 shows phone number management options 692, which include a call option 692a, a send message option 692b, an add to contacts option 692c, and a copy option 692d. As Figure 6O shown, the computer system 600 shows different options for managing certain specific text types (e.g., email, phone number, QR code) rather than other text types (as opposed to the text management options 682 shown when "named FLUFFY" is selected in Figure 6L and the phone number management options 692 shown when "123 - 4567" is selected in Figure 6O ). In some embodiments, in response to detecting an input pointing to the call option 692a, the computer system 600 initiates a phone call to "123 - 4567" (e.g., using as described above with respect to Figure 6Nthe similar technology described above). In some embodiments, in response to detecting an input pointing to the Send Message option 692b, the computer system 600 initiates a process for sending a message to "123-4567" (e.g., displaying a text management application). In some embodiments, in response to detecting an input pointing to the Add to Contacts option 692c, the computer system 600 initiates a process for adding a contact to a contact list that has "123-4567" as a phone number in the contact information. In some embodiments, in response to detecting an input pointing to the Copy option 692d, the computer system 600 uses one or more techniques as described above for the copy option 682a in Figure 6F to copy "123-4567".
[0276] Figures 6P to 6T An exemplary embodiment of displaying a QR code in the live preview 630 is shown. In some embodiments, the QR code can be replaced with other types of matrices and / or barcodes.
[0277] As Figure 6P shown, the computer system 600 simultaneously displays the QR code 668 and the QR code identifier 670 (e.g., "CAFE32.COM") in the live preview 630. In some embodiments, the QR code identifier identifies one or more of a website, a contact, a cellular plan, an email address, a calendar invitation / event, a location (e.g., GPS location), text, a video, a phone number, a WiFi network, an application, and / or an instance of an application, etc. The QR code identifier 670 includes an indication of the information identified by the QR code. In Figure 6P , the QR code 668 is in the field of view of one or more cameras of the computer system 600, and the QR code identifier 670 is not in the field of view. The computer system 600 displays the QR code identifier 670 because it is determined that the QR code 668 corresponds to (e.g., or identifies) a website destination belonging to "CAFE32.COM". In Figure 6P , the computer system 600 detects the input 650p1 and / or the input 650p2 in the camera display area 604.
[0278] As Figure 6QAs shown, in response to detecting input 650p1 and / or input 650p2 (and based on determining that at least one of the inputs is a tap input and / or a press-and-hold input), computer system 600 displays notification 674, which includes a preview of a website (e.g., the address "CAFE32.COM"). In some embodiments, the preview of the web address includes the full web address (e.g., "http:\\cafe32.com\menu") and / or an image from the web address. In some embodiments, computer system 600 displays notification 674 in response to detecting one or more inputs instead of navigating to the web address corresponding to QR code 668 to minimize the chance that the user inadvertently navigates to the website site corresponding to QR code 668. In some embodiments, in response to detecting input 650p1 on QR code 668, computer system 600 displays notification 674 (e.g., without automatically navigating to the website). In some embodiments, in response to detecting input 650p2 on the QR code identifier, computer system 600 automatically navigates to the website corresponding to QR code 668 (e.g., without displaying notification 674) (e.g., using one or more similar techniques as described below regarding computer system 600's response to tap input 650q). In Figure 6Q In, computer system 600 detects tap input 650q on notification 674.
[0279] As Figure 6R shown, in response to detecting tap input 650q, computer system 600 automatically navigates to the web address corresponding to the QR code (and / or opens) via web application 678.
[0280] As Figure 6S shown, computer system 600 uses one or more of the techniques described above regarding Figure 6P to simultaneously display QR code 668 and QR code identifier 670. In Figure 6S In, computer system 600 detects tap input 650s on text management control 680.
[0281] As Figure 6T shown, in response to detecting tap input 650s, computer system 600 displays QR code management options 672, which include a share option 672a, a copy link option 672b, an add to reading list option 672c, and an open link option 672d. As described above regarding Figure 6O Computer system 600 displays different options for managing certain specific text types rather than other text types. In some embodiments, in response to detecting an input directed to share option 672a, computer system 600 initiates a process for sharing the web address and / or link corresponding to the QR code (e.g., using as described regarding an input directed to Figure 6Fone or more similar techniques as described for the input of the sharing option 682d in). In some embodiments, in response to detecting an input pointing to the copy link option 672b, the computer system 600 copies the network address and / or link corresponding to the QR code (e.g., using one or more techniques as described above for Figure 6F the copy option 682a in). In some embodiments, in response to detecting an input pointing to the add to reading list option 672c, the computer system 600 initiates a process for adding the network address and / or link corresponding to the QR code to an item list (e.g., one or more articles, books, websites, etc.). In some embodiments, in response to detecting an input pointing to the open link option 672d, the computer system 600 navigates to the network address corresponding to the QR code (and / or opens) via the web application 678 (e.g., using similar techniques as described above for Figure 6R described).
[0282] In some embodiments, the QR code management option 672 includes one or more options dynamically selected based on the type of resource represented by the QR code (e.g., the QR code displayed when the text management control 680 is selected). For example, the type of resource represented by the QR code may include one or more of a link to a website, a contact, a cellular plan, an email address, a calendar invitation / event, a location (e.g., a GPS location), text, a video, a phone number, a WiFi network, an application, and / or an instance of an application, etc. In some embodiments, the QR code management option 672 includes a first set of controls when the QR code represents a first type of resource and a second set of controls when the QR code represents a second type of resource different from the first type. In some embodiments, the first set of controls has a different number of controls from the second set of controls. In some embodiments, a preview of the resource represented by the QR code is included in the QR code management option 672 (e.g., when the QR code represents a text string).
[0283] In some embodiments, the QR code management option 672 includes different sets of controls based on whether the computer system 600 is in a locked or unlocked state. In some embodiments, when the computer system 600 is in a locked state and the QR code represents a link to an application, control options for installing and / or opening the application are displayed. In some embodiments, when the computer system 600 is in an unlocked state, even if the application is installed, the link to open the application is not displayed (e.g., suppressed) to avoid communicating information about which applications are installed on the device to unauthorized users of the device. Optionally, instead of displaying the link to open the application, the device displays an option to use a portion of the available application without downloading the full application. In some embodiments, the computer system 600 displays a different set of controls (e.g., based on whether the computer system 600 is in a locked or unlocked state) to limit the information given to unauthorized users (e.g., information that can be used to determine whether the application represented by the QR code is installed and / or not installed on the computer system 600).
[0284] Figures 6U to 6W An exemplary scenario is shown in which the computer system 600 displays selection indicators around selected text that is divided into columns. In Figures 6U to 6W this, the computer system 600 is oriented such that text in the environment is aligned with the field of view of one or more cameras of the computer system 600. Figure 6U An example is shown in which the computer system 600 displays a live preview 630 that includes a representation of a text portion 648 (e.g., a roster of football players). In some embodiments, the computer system 600 displays a representation of previously captured media that includes a representation of the text portion 648, and uses one or more of the techniques described above with respect to Figures 6U to 6W to select words in the text portion 648.
[0285] As Figure 6U shown, the text portion 648 includes a name column 648a, a position column 648b, a state column 648c, and a grade column 648d. Each corresponding column includes text that has been detected by the computer system 600 (e.g., using one or more of the techniques described above with respect to Figures 6A to 6F ). As Figure 6U shown, the computer system 600 emphasizes the text portion 648 while reducing the visual prominence of portions of the live preview 630 that do not include text (e.g., football) (e.g., using one or more of the techniques described above with respect to Figures 6A to 6F ). Additionally, because the computer system 600 has detected the text portion 648, the computer system 600 places a box around the text 648 to emphasize the text 648. As Figure 6UAs shown, the computer system 600 displays the text management control 680 as active (e.g., as shown by the bolded text management control 680) and the text management option 682 (e.g., as described above with respect to Figure 6F ). In Figure 6U , the computer system 600 detects a first portion of the swipe input 650u on the name column 648a, which travels from the "name" header of the name column 648a to the "position" header of the position column 648b.
[0286] As Figure 6V shown, in response to detecting the first portion of the swipe input 650u, the computer system 600 displays selection indicators 696 (e.g., "gray highlighting") around all the words in the name column 648a ("Name", "Maria", "Kate", "Sarah", and "Ashley") and the "position" header of the position column 648b. The selection indicators 696 are positioned based on the position of the swipe input 650u. Since the first portion of the swipe input 650u of the computer system 600 terminates at the position of the "position" header of the position column 648b, the computer system 600 displays the selection indicators 696 around all the words up to and including (e.g., the words in the name column 648a) and including the "position" header. In some embodiments, the computer system 600 does not include the "position" header of the position column 648b because the first portion of the swipe input 650u of the computer system 600 terminates at the position of the "position" header of the position column 648b. In some embodiments, in the case where the end of the input terminates at the position of the word "DEFENDER" in the position column 648b (e.g., in the 3rd row of the position column 648b), the computer system 600 highlights all the words up to the word "DEFENDER", including all the words in the name column 648a, the "position" header of the position column 648b (e.g., on the 1st row of the position column 648b), and the word "Forward" in the 2nd row of the position column 648b.
[0287] The shape of the selection indicator 696 depends on whether the selected text (e.g., the text enclosed by the selection indicator 696) is aligned with the computer system 600. In Figure 6V , the computer system 600 displays the selection indicator 696 as a polygon with an angle of right angles (e.g., a shape with all right angles is referred to herein as a rectangle-based selection indicator). The selection indicator 696 is a rectangle-based selection indicator because it is determined that the selected text (e.g., the text of the selection indicator 696) is aligned with the computer system 600 (e.g., and / or aligned with the field of view of one or more cameras of the computer system 600) (e.g., which will be described below in connection with Figures 6X to 6Z(explained with additional details). In Figure 6V the computer system 600 detects a second portion of the swipe input 650u that is a rightward swipe input that travels from the "Location" header in the location column 648b to the "State" header in the state column 648c.
[0288] As Figure 6W shown, in response to detecting the second portion of the swipe input 650u, the computer system 600 extends the selection indicator 696 to the right such that the selection indicator 696 is shown around the words (e.g., all words) in the name column 648a and the location column 648b and also around the "State" header in the state column 648c (e.g., using one or more of the techniques described above with respect to Figures ), because the computer system recognizes the words in the name column 648a as being in the same column. As shown, the selection indicator 696 continues to be a rectangle-based selection indicator because the text portion continues to be aligned with the field of view of one or more cameras. In the computer system 600 no longer detects the swipe input 650u. However, the computer system 600 continues to show the selection indicator 696 around a portion of the text.
[0289] Figures 6X to 6Z shows an exemplary scenario where the computer system 600 shows the selection indicator around the selected text when the computer system 600 is oriented (e.g., oriented differently with respect to the corresponding text portion than the computer system 600 is oriented with respect to the corresponding text portion in Figures 6U to 6V ), such that the text in the environment is not aligned with the field of view of one or more cameras of the computer system 600. Figure 6X shows the computer system 600 displaying a live preview 630 that includes a representation of a text portion 652 (e.g., a text paragraph about football). The text portion 652 is on a piece of paper in the environment captured by the field of view of one or more cameras of the computer system 600. In some embodiments, the computer system 600 displays a representation of previously captured media that includes a representation of a text portion 648 and uses one or more of the techniques described below with respect to Figures 6U to 6W to select words in the text portion 648.
[0290] In Figure 6X the text portion 652 is not aligned with the field of view of one or more cameras. In Figure 6X the computer system 600 is oriented in a position such that the computer system 600 is not parallel to the text portion 652 and / or rotated / tilted along an axis (z-axis) in the environment (e.g., the user holds the phone at an angle and / or tilts it such that the field of view of one or more cameras is not aligned with the text portion 652). In Figure 6XIn 600, the computer system detects a swipe input 650x in a diagonal direction from the word "while" in the text portion 652 to the last period (".") of the text portion 652.
[0291] As Figure 6Y shown, in response to detecting the swipe input 650x, the computer system 600 displays a selection indicator 696 around a subset of the text portion 652 from the word "while" in the text portion 652 to the last period in the text portion 652. As Figure 6Y shown, the selection indicator 696 is a polygon having some angles that are not right angles (e.g., a shape having some acute angles and some obtuse angles is referred to herein as a non-rectangle-based selection indicator). The non-rectangle-based selection indicator is drawn by the computer system to match or appear to match (or substantially match or appear to substantially match) the orientation of the text portion 652 in the live preview 630 (e.g., as if the selection indicator 696 were a rectangle-based selection indicator on the surface containing the text portion 642, but viewed from the same perspective as Figures 6X to 6Z shown in). As described above, Figure 6Y the selection indicator 696 is not rectangle-based because it is determined that the text portion 652 is not aligned with the computer system 600 (e.g., opposite to Figures 6V to 6U the rectangle-based selection indicator of the selection indicator 696 (e.g., with respect to the orientation of the display of the computer system 600)). In Figure 6Y 600, the computer system detects a swipe input 650y from the word "while" in the text portion 652 to the word "synthetic". It is noted that the swipe input 650y moves in a diagonal direction relative to the computer system 600, but travels along a line of words in the text portion 652. In some embodiments, even if the edges of the selection indicator 696 are displayed diagonally relative to the edges of the computer system display area, the computer system places some or all of the edges of the selection indicator 696 in positions determined to be parallel or perpendicular to the text lines in the text portion 652. In some embodiments, as the angle of the camera relative to the surface containing the text portion 652 changes, the angle of the edges of the selection indicator 696 changes in the display area so as to keep the edges in positions determined to be parallel or perpendicular to the text lines in the text portion 642.
[0292] As Figure 6ZAs shown, in response to detecting a swipe input 650y, the computer system 600 expands the selection indicator 696 in the direction of the swipe input 650y such that the selection indicator 696 encompasses a subset of the text portion 652 from the word "synthesis" to the last period in the text portion 652 (e.g., when it is included in a portion of the text). Even though the selection indicator 696 has been expanded, the selection indicator 696 remains displayed as a non-rectangular-based selection indicator. Additionally, after the computer system 600 no longer detects the swipe input 650y, the selection indicator 696 continues to be displayed around the text portion.
[0293] Figures 7A to 7L An exemplary user interface for visual indicators for using a computer system to manage visual content in media is shown. The user interfaces in these figures are used to illustrate processes including those Figure 9 described below.
[0294] Figure 7A A media gallery user interface 710 is shown in which the computer system 600 simultaneously displays a thumbnail media representation 712 and a gallery region 702. The thumbnail media representation 712 includes thumbnail media representations 712a - 712c, where each of the thumbnail media representations 712a - 712c represents a different media item (e.g., media items captured at different time instances). The gallery region 702 includes a library control 702a (e.g., when selected, causes the computer system 600 to display the thumbnail media representation 712), a "for you" control 702b (e.g., when selected, causes the computer system 600 to display a thumbnail representation of dynamically generated media items based on user preferences), an album control 702c (e.g., when selected, causes the computer system 600 to display a thumbnail album representation each representing a collection of media items), and a search control 702d (e.g., when selected, causes the computer system 600 to display a search user interface including one or more controls for searching for media items). In Figure 7A , the library control 702a has been selected (e.g., as indicated by the bolded library control 702a). In Figure 7A , the computer system 600 detects a tap input 750a on the thumbnail media representation 712a.
[0295] As Figure 7BAs shown, in response to detecting a tap input 750a, computer system 600 displays a media viewer user interface 720 and stops displaying a media gallery user interface 710. The media viewer user interface 720 includes a media viewer region 724 located between an application control region 722 and an application control region 726. The media viewer region 724 includes a magnified representation 724a that represents the same media item as the thumbnail media representation 712a. The media viewer user interface 720 is substantially free of overlays, while the application control regions 722 and 726 are substantially covered with overlays.
[0296] The magnified representation 724a includes a marker 642 that includes a text portion 642a (e.g., "Lost Dog Notice") and a text portion 642b (e.g., a text paragraph starting with "Cute"), as described above with respect to Figure 6B The text of the text portions 642a - 642b is not visually prominent, and the text of the text portions 642a - 642b is small and not easily readable by a user viewing the computer system 600. Additionally, the magnified representation 724a includes a person 740 standing in front of a tree. The person 740 is wearing a hat that includes the word "Brand" (e.g., text portion 742).
[0297] The application control region 722 optionally includes an indicator of the time at which the magnified representation of the currently displayed media was taken (e.g., Figure 7B "7:54" in ) (e.g., magnified representation 724a), a cellular signal status indicator 720a indicating the status of the cellular signal, and a battery level status indicator 720b indicating the remaining battery life status of the computer system 600. The application control region 722 also includes a return control 722a (e.g., when selected, causes the computer system 600 to redisplay the media gallery user interface 710) and an edit control 722b (e.g., when selected, causes the computer system 600 to display a media edit user interface that includes one or more controls for editing the representation of the media item represented by the currently displayed magnified representation 724a).
[0298] The application control region 726 includes some of the thumbnail media representations 712 (e.g., 712a - 712c) displayed in a single row. Since the magnified representation 724a is displayed in the media viewer region 724, the thumbnail media representation 712a is displayed as being selected. Specifically, the thumbnail media representation 712a is displayed as being selected by being shown as having space from the other thumbnails (e.g., 712b and 712c). Figure 7Bis selected. Additionally, the application control region 726 includes a send control 726b (e.g., which, when selected, causes the computer system 600 to initiate a process for transmitting the media item represented by the magnified media representation), a favorites control 726c (e.g., which, when selected, causes the computer system 600 to mark / unmark the media item represented by the magnified representation 724a as a favorite media), and a recycle bin control 726d (e.g., which, when selected, causes the computer system 600 to delete (or initiate a deletion process) the media item represented by the magnified representation 724a). In Figure 7B the computer system 600 detects an expansion input 750b on the media viewer region 724 (e.g., at and / or pointed to a location on the display of the computer system 600 corresponding to the media viewer region).
[0299] As Figure 7C shown, in response to detecting the expansion input 750b, the computer system 600 updates the magnified representation 724a to reflect the change in the zoom level such that Figure 7C the magnified representation 724a of Figure 7B is displayed at a greater zoom level than the magnified representation 724a of Figure 7C At the increased zoom level, the text portions 642a - 642b of Figure 7B are larger and more visually prominent (e.g., larger and more readable) than the text portions 642a - 642b of Figure 7B In addition to updating the magnified representation 724a, the computer system 600 also expands the media viewer region 724 of Figure 7B such that the magnified representation 724a of Figure 7A occupies the portion of the display previously occupied by the application control regions 722 and 726 in
[0300] In Figure 7C it is determined that the text of the text portion 642a and the text of the text portion 642b in Figure 7C do not individually meet the set of prominence criteria (e.g., using one or more similar techniques as described above with respect to Figures 6A to 6C ). Accordingly, the computer system 600 does not display parentheses corresponding to (e.g., surrounding) the text portions 642a - 642b in Figure 7B Furthermore, because the text of the text portions 642a - 642b does not meet the set of prominence criteria, the computer system 600 does not display the text management controls 680 (e.g., as described above with respect to Figure 6B ).
[0301] In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of text portions 642a - 642b include text that occupies a predetermined amount of space (e.g., 10% to 100%) of the magnified representation 724a. In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of portions 642a - 642b include text located at or near a predetermined position (e.g., a central position) of the magnified representation 724a. In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of text portions 642a - 642b include text of a particular text type (e.g., an email, a phone number, an address, a QR code, etc.) (e.g., as described above with respect to Figures 6M to 6T ). In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of text portions 642a - 642b include text that is relevant to the context of the magnified representation 724a (e.g., the text meets a relevance threshold (e.g., the computer system 600 determines that the text is 90%, 95%, 99% relevant)).
[0302] In Figure 7C , it is determined that the main theme of the magnified representation 724a is the marker 642. That is, the context of the magnified representation 724a is the content displayed within the marker 642. In Figure 7C , it is further determined that the text portion 742 (e.g., "brand") is not relevant because it appears on the hat of the person 740 and is thus not relevant to the context of the content displayed in the magnified representation 724a. In some embodiments, the computer system 600 determines that the text portion 742 is not relevant to the context of the content displayed in the magnified representation 724a because the text portion 742 is displayed on a person or something on a person in the magnified representation 724a.
[0303] Because it is determined that the text portion 742 is not relevant, it is determined that the text portion 742 does not meet the set of prominence criteria. It is noted that even though the text portion 742 has more text than the text portions 642a - 642b, it is determined that the text portion 742 does not meet the set of prominence criteria. As Figure 7C shows, the computer system 600 does not display one or more brackets around the text portion 742 ("brand") because the text portion 742 does not meet the set of prominence criteria (e.g., due to the determination that the text portion 742 is not relevant to the context of the magnified representation 724a). In Figure 7C , the computer system 600 detects a tap input 750c on the text portion 742.
[0304] As Figure 7D shows, in response to detecting the tap input 750c, the computer system 600 maintains the display of the magnified representation 724a, asFigure 7C as shown. In Figure 7D , the computer system 600 does not update the display of the magnified representation 724a to indicate that the text portion 742 is selected because it is determined that the text portion 742 does not meet the set of prominence criteria (e.g., as described above with respect to Figure 7C ). Additionally, the computer system 600 does not update the display of the magnified representation 724a to indicate that the text portion 742 is selected because the text management control is not displayed and selected (e.g., as opposed to the computer system 600 updating the Figures 6J to 6L media representation as described above). Additionally, because the computer system 600 does not update the Figure 7D display of the magnified representation 724a, the text portions 642a - 642b continue to not meet the set of prominence criteria. Thus, as Figure 7D shown, the computer system 600 does not display brackets corresponding to the text portion 642a or 642b. In Figure 7D , the computer system 600 detects an expansion input 750d in the media viewer region 724. In some embodiments, instead of the expansion input 750d, the computer system 600 detects a directional swipe corresponding to a translation (e.g., transformation) Figure 7D of the magnified representation 724a as shown in
[0305] As Figure 7E shown, in response to detecting the expansion input 750d, the computer system 600 updates the magnified representation 724a to reflect a change in the zoom level such that Figure 7E the display of the magnified representation 724a of Figure 7D is displayed at a greater zoom level than the display of the magnified representation 724a of Figure 7E . In
[0306] it is determined that the text of the text portion 642a meets the set of prominence criteria, but the text of the text portion 642b does not meet the set of prominence criteria. As a result, the computer system 600 displays a bracket 736a at a location corresponding to the location of the text portion 642a (e.g., around the text portion 642a). However, the computer system 600 does not display a bracket 736a or any other bracket at a location corresponding to the location of the text portion 642b (e.g., because the text of the text portion 642b does not meet the set of prominence criteria). Notably, even though the text portion 742 has greater text than the text portions 642a - 642b, it is determined that the text portion 742 (e.g., "brand") continues to not meet the set of prominence criteria (e.g., due to the text portion 742 being irrelevant). In some embodiments, when the computer system 600 detects a directional swipe instead of the expansion input 750d, the computer system 600 translates the magnified representation 724a such that a different portion of the magnified representation 724a is displayed in response to receiving the expansion input 750d.As Figure 7E shown, because it is determined that the text of text portion 642a meets the set of prominence criteria, computer system 600 displays text management control 680. The text management control 680 is displayed in an inactive state (e.g., as indicated by the text management control 680 not being bolded) because the text management control 680 has not been selected (e.g., no input directed to the text management control has been detected). In Figure 7E , computer system 600 detects expansion input 750e in media viewer region 724.
[0307] As Figure 7F shown, in response to detecting expansion input 750e, computer system 600 updates the display of magnification representation 724a to reflect the change in zoom level such that Figure 7F the display of magnification representation 724a of Figure 7E is displayed at a greater zoom level than the display of magnification representation 724a of Figure 7F . In Figure 6C , it is determined that the text of text portion 642a meets the set of prominence criteria and the text of text portion 642b meets the set of prominence criteria. Thus, as described above with respect to Figure 7F , brackets 636a are displayed around the entirety of the two text portions 642a - 642b. Notably, in Figure 7F , even though text portion 742 has greater text than text portions 642a - 642b, it is determined that text portion 742 (e.g., "brand") continues not to meet the set of prominence criteria (e.g., due to text portion 742 being irrelevant). As Figure 6C shown, computer system 600 displays text type indicator 638a below "123 Main Street" to indicate that an address has been detected and displays text type indicator 638b below "123 - 4567" to indicate that a phone number has been detected (e.g., using one or more of the techniques described above with respect to Figures 6A to 6M ). In some embodiments, computer system 600 displays multiple brackets, one bracket around text portion 642a and another bracket around text portion 642b and / or other combinations of brackets (e.g., using one or more of the techniques described above with respect to Figure 7E ). In some embodiments (e.g., referring back to Figure 7F ), computer system 600 displays text type indicators below text portions regardless of whether the text portion to which the text type indicator pertains meets the set of prominence criteria. In
[0308] As Figure 7GAs shown, in response to detecting an expansion input 750f, computer system 600 updates the magnified representation 724a to reflect the change in the zoom level such that Figure 7G the display of the magnified representation 724a of Figure 7F is displayed at a greater zoom level than the display of the magnified representation 724a of Figure 7G In some embodiments, as shown in
[0309] As shown in Figure 7G the magnified representation 724a includes a subset of text portion 642a and a subset of text portion 642b. As a result of determining that the entire text portions 642a - 642b no longer meet the set of prominence criteria (e.g., and / or the magnified representation 724a that includes only the subsets of text portion 642a and text portion 642b), computer system 600 stops displaying brackets 636a around the entire text portions 642a and 642b. In Figure 7G , it is determined that a subset of the text of text portion 642b (e.g., the phone number "123 - 4567") meets the set of prominence criteria (e.g., while another subset of the text of text portion 642b does not meet the criteria). In some embodiments, because it is determined based on the input previously detected by computer system 600 (e.g., when viewing Figures 7A to 7G , computer system 600 has continued to zoom in near the phone number) that the user intends to interact with or view the phone number, it is determined that the subset of text portion 642b meets the set of prominence criteria.
[0310] In some embodiments, it is determined that Figure 7G includes a subset of text portion 642a (e.g., "dog") that meets the set of prominence criteria. In response to this determination, computer system 600 displays a set of brackets around the subset of text portion 642a simultaneously with brackets 736c.
[0311] As shown in Figure 7G computer system 600 displays only a portion of the address "123 Main Street". Accordingly, computer system 600 stops the display of text type indicator 638a. In some embodiments, computer system 600 maintains the display of text type indicator 638a below the portion of the address "123 Main Street" shown in Figure 7G . In some embodiments, computer system 600 determines that this portion of the address does not meet the set of prominence criteria because another portion of the address is not being displayed. In Figure 7G , computer system 600 detects a right - swipe 750g in the media viewer area 724.
[0312] As shown in Figure 7HAs shown, in response to detecting a 750g swipe to the right, computer system 600 pans and zooms in on representation 724a in the rightward direction. The zoomed-in representation 724a is panned such that Figure 7G the rightmost portion of the text portions 642a - 642b shown stops being displayed by computer system 600, and the leftmost portion of the text portions 642a - 642b is redisplayed by Figure 7H the computer system 600 in. As Figure 7H shown, computer system 600 does not display the entire phone number (e.g., 123 - 4567), and stops displaying the parentheses 736c and the text type indicator 638b. In some embodiments, a portion of the text type indicator 638b remains displayed below Figure 7H the portion of the phone number that continues to be displayed (e.g., "12"). In Figure 7H computer system 600 displays Figure 7H more of the address in (e.g., 123 Main Street) and redisplay the text type indicator 638a below "123 Main Street" to indicate to the user that an address has been detected.
[0313] In Figure 7H f, it is determined that another subset of text portion 642b (e.g., "$1000 REWARD") meets the set of salience criteria (e.g., no other subset of text portion 642b meets the set of salience criteria). Because it is determined that another subset of text portion 642b meets the set of salience criteria, computer system 600 displays parentheses 736d around the other subset "$1000" of text portion 642b. In some embodiments, the determination of another subset of text portion 642b (e.g., "$1000") is the most relevant text displayed based on the context of the display content of the zoomed-in representation 724a. In some embodiments, the determination Figure 7H includes a subset of text portion 642a (e.g., "LOST") that meets the set of salience criteria. In some embodiments, in response to this determination, computer system 600 displays a corresponding set of parentheses around the subset of text portion 642a simultaneously with parentheses 736e. In Figure 7H computer system 600 detects a tap input 750h on the text management control 680.
[0314] As Figure 7IAs shown, in response to detecting a tap input 750h, computer system 600 displays text management options 682, which include a copy option 682a (e.g., when selected, computer system 600 copies the text enclosed by brackets 736d), a select all option 682b (e.g., when selected, computer system 600 selects all the text enclosed by brackets 736d), a find option 682c (e.g., when selected, computer system locates the text enclosed by brackets 736d by searching (e.g., web search, dictionary search)), and a share option 682d (e.g., when selected, computer system 600 initiates a process of sharing the text enclosed by brackets 736d). In some embodiments, the various components of text management options 682 function as described above with respect to Figures 6A to 6M as described. In some embodiments, computer system 600 displays multiple text management options, where each corresponding text management option corresponds to a corresponding text portion enclosed by a corresponding pair of brackets. In some embodiments, selecting a corresponding text management option allows a user to manage the text portion corresponding to that corresponding text management option.
[0315] As Figure 7I shown, computer system 600 displays text management control 680 as being activated (e.g., as indicated by the bolded text management control 680). In Figure 7I , computer system 600 detects a tap input 750i on text management control 680.
[0316] As Figure 7J shown, in response to detecting a tap input 750i, computer system 600 uses one or more techniques as described above with respect to Figure 7H to redisplay magnified representation 724a. In Figure 7J , computer system 600 detects a downward swipe input 750j in media viewer region 724.
[0317] As Figure 7K shown, in response to detecting a downward swipe input 750j, computer system 600 pans media viewer region 724 downward (e.g., based on the swipe input) such that text portion 642b stops being displayed and computer system only displays a subset of text portion 642a. In Figure 7K , it is determined that the subset of text portion 642a (e.g., "lost") meets the set of prominence criteria. Since the subset of text portion 642a does meet the set of prominence criteria, computer system 600 displays brackets 736e around the subset of text portion 642a.
[0318] In Figure 7K , computer system 600 does not display brackets 736d because text portion 642b is not Figure 7Kis not shown as part of the magnified representation 724a. In Figure 7K the computer system 600 detects a pinch input 750k in the media viewer region 724.
[0319] As Figure 7L shown, in response to detecting the pinch input 750k, the computer system 600 updates the magnified representation 724a to reflect a change in the zoom level (e.g., a decrease in the zoom level) such that compared to the zoom level of the display of the magnified representation 724a of Figure 7K , the display of the magnified representation 724a of Figure 7L is displayed at a decreased zoom level. In Figure 7L it is determined that the text portions 642a and 642b do not meet the set of salience criteria. Accordingly (e.g., because it is determined that the text portions 642a and 642b do not meet the set of salience criteria), the computer system 600 does not display (and / or stops displaying) the text management control 680 and / or any brackets around the text portions 642a - 642b.
[0320] Although the techniques discussed above with respect to Figures 7A to 7L are discussed in the context of the computer system 600 displaying a representation of previously captured media and a media viewer user interface, the above - discussed one or more techniques are also applicable when the computer system 600 displays a representation of previously captured media and a media viewer user interface. Additionally, the techniques discussed above with respect to Figures 6A to 6Z are also applicable to the context in which the computer system 600 displays a live preview (e.g., a representation of the field of view of one or more cameras before the media has been captured) (such as the live preview 630 of Figures 7A to 7L ) and a camera user interface. Figures 6A to 6M
[0321] Although the techniques discussed above with respect to Figures 6A to 6Z are discussed in the context of the computer system 600 displaying a live preview and a camera user interface, the above - discussed one or more techniques are also applicable when the computer system 600 displays a live preview and a camera user interface. Additionally, the techniques discussed with respect to Figures 7A to 7L are also applicable to the context in which the computer system 600 displays previously captured media (such as the magnified representation 724a of the media) and a media viewer user interface. Figures 6A to 6Z
[0322] Figure 8 It is a flowchart showing a method for using a computer system to manage visual content in media according to some embodiments. Method 800 is executed at a computer system (e.g., 100, 300, 500) that communicates with a display generation component. Some operations in method 800 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0323] As described below, method 800 provides an intuitive way to manage visual content in media. This method reduces the cognitive burden on the user to manage visual content in media, thus creating a more efficient human-machine interface. For battery-powered computing devices, enabling the user to manage visual content in media faster and more efficiently saves power and increases the time interval between battery charges.
[0324] Method 800 is executed at a computer system (e.g., 600) (e.g., smartphone, desktop computer, laptop, tablet) that communicates with a display generation component (e.g., display controller, touch-sensitive display system). In some embodiments, the computer system communicates with one or more input devices (e.g., touch-sensitive surface) and / or a first camera among one or more cameras (e.g., one or more cameras on the same side or different sides of the computer system (e.g., dual camera, triple camera, quadruple camera, etc.) (e.g., front camera, rear camera)).
[0325] The computer system displays (802) a camera user interface (e.g., media capture user interface, media viewing user interface, media editing user interface) via the display generation component, and the camera user interface includes a representation (e.g., 630) of simultaneously displayed media (e.g., photo media, video media) (e.g., live media, live preview (e.g., media corresponding to the field of view of one or more cameras that have not been captured (e.g., current field of view) (e.g., in response to detecting a request to capture media (e.g., detecting a selection of a shutter enabling representation)), previously captured media (e.g., media corresponding to the field of view of one or more cameras that have been captured (e.g., previous field of view), media that has been saved and can be accessed by the user at a later time, media representation displayed in response to receiving a gesture for a thumbnail representation of the media (e.g., in a media gallery)) and a media capture enabling representation (e.g., 610) (e.g., user interface object).
[0326] When (804) simultaneously displays a media representation (e.g., 630) and a media capture enablement representation (e.g., 610) (e.g., a user interface object), and it is determined that a corresponding set of criteria is satisfied, where the corresponding set of criteria includes criteria satisfied when corresponding text (e.g., 642a, 642b) (e.g., one or more characters represented in the media) is detected in the media representation (e.g., 630), the computer system displays (806) (e.g., simultaneously with the media representation) (e.g., in the user interface) a first user interface object (e.g., 680) corresponding to one or more text management operations (e.g., simultaneously with the media representation and / or the first user interface object). In some embodiments, the plurality of options (e.g., 672, 682, 692) includes one or more options for copying the corresponding text (e.g., 682a), selecting the corresponding text (e.g., 682b), finding the corresponding text (e.g., 682c), sharing the corresponding text (e.g., 682d), and converting the corresponding text.
[0327] When (804) simultaneously displays a media representation (e.g., 630) and a media capture enablement representation (e.g., 610) (e.g., a user interface object), and it is determined that the corresponding set of criteria is not satisfied, the computer system abandons displaying (808) the first user interface object.
[0328] When displaying a media representation (e.g., 630) (e.g., when simultaneously displaying the media representation, the media capture enablement representation, and the first user interface object), the computer system detects (810) a first input (e.g., 650a, 650e, 650g, 650u) (e.g., a mouse / touchpad click / activation, keyboard input, scroll wheel input, hover gesture, tap gesture, swipe gesture) directed to a camera user interface (e.g., 602, 604, 606). In some embodiments, the first input is a non - tap gesture (e.g., a rotation gesture and / or a press - and - hold gesture).
[0329] In response to (812) detecting the first input (650a, 650e, 650g, 650u) directed to the camera user interface (e.g., the first gesture) and it being determined that the first input (e.g., 650a) corresponds to a selection of the media capture enablement representation (e.g., 610) (e.g., a gesture directed to the media capture enablement representation, a gesture at a location corresponding to the media capture enablement representation), the computer system initiates (814) the capture of media to be added to a media library (e.g., 612) associated with the computer system (e.g., 600) (e.g., without displaying options for managing the corresponding text).
[0330] In response to detecting a first input (650a, 650e, 650g, 650u) (e.g., a first gesture) directed to the camera user interface (812) and based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680), the computer system displays (816), via a display generation component, multiple options (e.g., 672, 682, 692) for managing the corresponding text (e.g., without initiating capture of media to be added to a media library (e.g., as shown at 624) associated with the computer system (e.g., 600). In some embodiments, the multiple options are displayed adjacent to the corresponding text (e.g., included in a media representation). In some embodiments, the multiple options (e.g., 672, 682, 692) include one or more options to copy the corresponding text (e.g., 682a), select the corresponding text (e.g., 682b), find the corresponding text (e.g., 682c), share the corresponding text (e.g., 682d), and transform the corresponding text (e.g., as described above with respect to Figure 6F ). In some embodiments, based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680), the first user interface object is in an active state (e.g., transitioning from being displayed in an inactive state to an active state (e.g., as described above with respect to Figure 6F ), where the first user interface object (e.g., Figure 6F 680 in) in the active state display (e.g., bold, pressed state / appearance) has a different appearance from when the first user interface object is displayed in an inactive state (e.g., Figure 6G 680 in) (e.g., not bold, released state / appearance). In some embodiments, based on determining that the first input (e.g., 650a) corresponds to a selection of a media capture enablement representation (e.g., 610), the first user interface object (e.g., 680) is in an inactive state (e.g., transitioning from being displayed in an inactive state to an active state). In some embodiments, based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680) and the first user interface object is in an inactive state (e.g., Figure 6E 680 in), the computer system displays multiple options (e.g., 672, 682, 692) for managing the corresponding text. In some embodiments, based on determining that the first input corresponds to a selection of a first user interface object (e.g., 680) and the first user interface object is in an active state (e.g., Figure 6F 680 in), the computer system abandons displaying multiple options (e.g., 672, 682, 692) for managing the corresponding text (e.g., as described above with reference to Figure 6FAs described). In some embodiments, in accordance with determining that the first input (650a) corresponds to a selection of a media capture affordance, the first user interface object (e.g., 610) continues to be displayed. In some embodiments, in accordance with determining that the first input (e.g., 650a) corresponds to a selection of a media capture affordance (e.g., 610) or a selection of the first user interface object (e.g., 680), one or more user interface objects (e.g., media capture affordance (e.g., 610), camera settings affordance, camera mode affordance (e.g., 620)) cease to be displayed or are displayed as inactive (e.g., dimmed) in the camera user interface (e.g., not responsive to user input on the respective object). Displaying multiple options for managing the corresponding text in accordance with determining that the first input corresponds to a selection of the first user interface object provides the user with the ability to manage the corresponding text quickly and efficiently without cluttering the user interface with additional user interface objects. Providing additional control of the system without cluttering the UI with additional displayed controls enhances the operability of the system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating the system / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively. Displaying multiple options for managing the corresponding text automatically provides the user with various options for different ways of managing the corresponding text when certain specified conditions are met (e.g., based on whether the first input corresponds to a selection of the first user interface object). Performing an operation when a set of conditions has been met without further user input enhances the operability of the system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating the system / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0331] In some embodiments, the first input (e.g., 650e, 650g, 650u) is a tap gesture (e.g., tap input) (e.g., a gesture at a location corresponding to the first user interface object) directed at the first user interface object (e.g., 672, 682, 692).
[0332] In some embodiments, a media representation (e.g., 630) includes corresponding text (e.g., in the case where the corresponding text is displayed when the media representation is displayed). In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u) (and when an indication that the text is selected is not displayed and / or after detecting an input / gesture corresponding to a selection of a first user interface object and / or when the first user interface object is displayed as active and / or when multiple options for managing the corresponding text are displayed), the computer system detects a second input (e.g., 650j) pointing to a camera user interface (e.g., a tap gesture and / or a swipe gesture). In some embodiments, the second input is a non-tap gesture (e.g., a rotation gesture and / or a press-and-hold gesture). In some embodiments, the first input is a non-swipe gesture (e.g., a rotation gesture, a press-and-hold gesture, a mouse / touchpad click / activation, a keyboard input, a scroll wheel input, a hover gesture, and / or a tap gesture). In some embodiments, in response to detecting the second input (e.g., 650j) pointing to the camera user interface and based on determining that the second input corresponds to a selection of one or more first portions of the corresponding text, the computer system displays an indication that one or more first portions (e.g., 642b) of the corresponding text (e.g., 642a, 642b) are selected (e.g., Figure 6K 642b in). In some embodiments, the indication is displayed around the corresponding text. In some embodiments, as part of displaying an indication that one or more portions of the corresponding text are selected, the computer system emphasizes (e.g., highlights, underlines, bolds, increases the size) one or more portions of the corresponding text. In some embodiments, when displaying an indication that a first portion of the corresponding text is selected, the computer system does not display an indication that a second portion (e.g., different from the first portion) of the corresponding text is selected. In some embodiments, based on determining that the second input (650j) corresponds to a selection of one or more portions of the corresponding text (e.g., 642a), and when a first user interface (e.g., 680) object is displayed as active (e.g., 680 as described above with respect to Figure 6F ), the computer system displays an indication that one or more first portions (e.g., 642a) of the corresponding text are selected (e.g., as described above with respect to Figure 6K and Figure 6L ). In some embodiments, based on determining that the second input (e.g., 650j) corresponds to a selection of one or more portions of the corresponding text (e.g., 642), and when a first user interface object (e.g., 680) is displayed as inactive (e.g., as described above with respect to Figure 6G ) and / or is not displayed (e.g., as described above with respect to Figure 7C , Figure 7GWhen the one or more first portions of the corresponding text are selected (e.g., as described above with respect to), the computer system does not display (e.g., discard the display) the one or more first portions of the corresponding text (e.g., 642a). Figure 7C and Figure 7G The indication that the one or more first portions of the corresponding text are selected provides the user with visual feedback as to whether the text has been selected and which text is currently selected. Providing improved visual feedback to the user enhances the operability of the computer system and makes the computer system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively. Displaying the indication that the one or more first portions of the corresponding text are selected in response to detecting a second input and based on determining that the second input corresponds to a selection of the one or more first portions of the corresponding text provides the user with additional control over selecting text without cluttering the user interface with additional user interface objects. Providing additional control over the system without cluttering the UI with additional displayed controls enhances the operability of the system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0333] In some embodiments, the second input (e.g., 650j) (e.g., a second gesture) is a tap gesture (e.g., directed at one or more portions of the corresponding text) or a swipe gesture (e.g., directed at one or more portions of the corresponding text). In some embodiments, the first input is of a first input type and the second input is of a second input type different from the first input type.
[0334] In some embodiments, in response to detecting a first input (e.g., 650e, 650g, 650u) directed to a camera user interface and based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680), the computer system displays an indication (e.g., 684) regarding (e.g., how) to select text included in a media representation (e.g., 630) (e.g., not previously displayed prior to detecting the first input) (e.g., an instruction) (e.g., an instruction that will cause the computer system to display one or more inputs for the text being selected). In some embodiments, in response to detecting a first input directed to a camera user interface and based on determining that the first input corresponds to a selection of a media capture enabling representation, the computer system does not display an indication regarding (e.g., how) to select text included in a media representation. In some embodiments, an indication (e.g., 684) regarding (e.g., how) to select text included in a media representation is displayed concurrently with multiple options (e.g., 682a, 682b, 682c, 682d) for managing the corresponding text (e.g., 642b). In some embodiments, an indication (e.g., 684) regarding text selection is displayed when the first user interface object (e.g., 680) is displayed in an active state (e.g., 680 as described above regarding Figure 6F ), and an indication (e.g., 684) regarding text selection is not displayed when the first user interface object is displayed in an inactive state (e.g., 680 as described above regarding Figure 6G ). Displaying an indication regarding how to select text included in a media representation provides the user with visual feedback regarding the steps required to select the text the user wishes to select. Providing the user with improved visual feedback enhances the operability of the computer system and makes the computer system interface more effective (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0335] In some embodiments, prior to detecting the first input (e.g., 650e, 650g, 650u), the media representation (e.g., 630) is displayed in a first appearance (e.g., Figure 6E of 630) (e.g., with a first blur value, a first dim value). In some embodiments, in response to detecting a first input (e.g., 650e, 650g, 650u) directed to a camera user interface and based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680), the computer system displays the media representation (e.g., 630) in a second appearance that is different from the first appearance (e.g., Figure 6E of 630)Figure 6F display the media representation (e.g., 630) (e.g., at a second blur value, a second dim value) (e.g., when the text is selected (e.g., in response to detecting a second input)). In some embodiments, as part of displaying the media representation in a second appearance that is different from the first appearance, the computer system blurs and / or dims at least a portion of the media representation. In some embodiments, in response to detecting a first input directed to the camera user interface and based on determining that the first input corresponds to a selection of a media capture enabling representation, the computer system displays the media representation in a third appearance that is different from the second appearance. In some embodiments, the third appearance is the first appearance. In some embodiments, the third appearance (e.g., black, solid color) is different from the first appearance (e.g., a blurred version of the field of view of one or more cameras). In some embodiments, the media representation having the third appearance is displayed for a predetermined period of time (e.g., less than one second), which is not based on whether the first user interface object is displayed in an active state. In some embodiments, when the first user interface object is displayed in an active state, the media representation having the second appearance is displayed, and when the first user interface object is displayed in an inactive state, the media representation having the second appearance is not displayed. In some embodiments, when the first user interface object is displayed in an active state, the media representation having the first appearance is not displayed, and when the first user interface object is displayed in an inactive state, the media representation having the first appearance is displayed. Displaying the media representation in a second appearance that is different from the first appearance of the representation in response to detecting the first input provides the user with visual feedback as to whether the user has selected text by de-emphasizing the less relevant portions of the media representation. Providing the user with improved visual feedback enhances the operability of the computer system and makes the computer system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0336] In some embodiments, a media presentation (e.g., 630) includes corresponding text (e.g., 642a, 642b) (e.g., in the case where the corresponding text is displayed when the media presentation is displayed). In some embodiments, based on determining that a corresponding set of criteria is met, the computer system emphasizes (e.g., highlights, displays an object (e.g., a shape, surrounding parentheses (e.g., yellow parentheses)), underlines, magnifies) one or more second portions of the corresponding text (e.g., 642a, 642b). In some embodiments, based on determining that a corresponding set of criteria is met, the computer system emphasizes one or more second portions of the corresponding text without emphasizing another portion of the corresponding text and / or another portion of the media presentation that does not include the one or more second portions of the corresponding text. Emphasizing the one or more second portions of the corresponding text provides the user with improved visual feedback as to whether a particular portion of the corresponding text included in the media meets the corresponding set of criteria. Providing the user with improved visual feedback enhances the operability of the computer system and makes the computer system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0337] In some embodiments, as part of emphasizing one or more second portions of the corresponding text, the computer system displays an indication that the corresponding text has been detected (e.g., 636a, 636b, 736c, 736d). In some embodiments, when one or more second portions of the corresponding text (e.g., 642a, 642b) are emphasized, the computer system receives a request to display a second presentation of the media (e.g., media that is the same as or different from the media represented by the media presentation). Figure 6F In some embodiments, a request to display a second presentation of the media is detected when one or more changes are detected in the field of view of one or more cameras communicating with the computer system. In some embodiments, a request to display a second presentation of the media is detected when a request to zoom in / out on the media presentation and / or pan the media presentation is detected. In some embodiments, a request to display a second presentation of the media is detected when the computer system is moved.
[0338] In some embodiments, in response to receiving a request to display a second presentation of the media (e.g., including a portion of the corresponding text and / or a second corresponding text different from the corresponding text). Figure 6FIn response to a request for the 630), the computer system converts (e.g., moves) an indication (e.g., 636a, 636b, 736c, 736d) that the corresponding text has been detected from a first location in the camera user interface to a second location in the camera user interface. In some embodiments, in response to receiving a request for a second representation of the media to be displayed, an indication that the corresponding text has been selected is modified to enclose a portion of the text that is different from the portion enclosed prior to receiving the request for the second representation of the media to be displayed. Converting the indication that the corresponding text has been detected from a first location in the camera user interface to a second location in the camera user interface in response to receiving a request for a second representation of the media to be displayed allows the user to maintain their view of the indication as the system moves between the first and second locations. Performing an operation when a set of conditions has been met without further user input enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0339] In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u) and based on determining that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680) (e.g., a first determination) (and / or when the first user interface object is displayed as active and / or when multiple options for managing the corresponding text are displayed), the media representation (e.g., 630) includes the corresponding text (e.g., 642a, 642b) and an indication that a third one or more portions of the corresponding text (e.g., 642a, 642b) are selected. In some embodiments, the computer system receives a request to display a third representation (e.g., 630) of the media (e.g., media that is the same as or different from the media represented by the media representation). In some embodiments, the request to display the third representation of the media is detected when one or more changes in the field of view of one or more cameras communicating with the computer system are detected. In some embodiments, the request to display the third representation of the media is detected when a request to zoom in / out on the media representation and / or pan the media representation is detected. In some embodiments, the request to display the third representation of the media is detected when the computer system is moved. In some embodiments, in response to receiving a request (e.g., 650c, 650d, 750e, 750f, 750g) to display a third representation (e.g., 630) of the media, the computer system displays an indication that at least a portion of the text included in the third representation of the media (e.g., 630) is selected, where the indication that at least a portion of the text included in the third representation of the media (e.g., 642a, 642b) is selected (e.g., 636a, 636b, 736c, 736d) is different from the indication that a third one or more portions of the corresponding text (e.g., 642a, 642b) are selected. In some embodiments, the text portion included in the third representation of the media includes at least a portion of the text in the third one or more portions of the text. Displaying an indication that at least a portion of the text included in the third representation of the media is selected in response to receiving a request to display the third representation provides an additional and efficient way for the user to control which portions of the text are selected without cluttering the user interface. Reducing the number of inputs required to perform an operation enhances the operability of the system and makes the user-system interface more effective (e.g., by helping the user provide appropriate inputs and reducing user errors when operating the system / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0340] In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u) and based on determining that the first input corresponds to a selection of a first user interface object (e.g., a first determination) (and / or when the first user interface object is displayed as active and / or when multiple options for managing the corresponding text are displayed), the media representation (e.g., 630) includes the corresponding text (e.g., 642b), an indication that a fourth one or more portions of the corresponding text (e.g., 642b) are selected, and the fourth one or more portions of the corresponding text (e.g., 642b) are displayed at a third location in the camera user interface (and / or on the display). In some embodiments, the computer system detects a change in the physical environment within the field of view of one or more cameras communicatively coupled to the computer system. In some embodiments, in response to detecting a change in the physical environment within the field of view of one or more cameras (e.g., 660a, 660b), the computer system continues to display the fourth one or more portions of the corresponding text (e.g., 642b) at the third location in the camera user interface (and / or on the display). In some embodiments, the selected text is frozen. In some embodiments, at least a portion of a fourth representation of the media (e.g., newly displayed in response to detecting a change in the physical environment) is displayed while maintaining the display of the fourth one or more portions of the corresponding text. In some embodiments, the computer system freezes the selected text (e.g., and / or displays the selected text in the same location and / or at the same size) while updating the media representation (e.g., live preview) to reflect the change in the physical environment. Continuing to display the fourth one or more portions of the corresponding text at the third location in the camera user interface allows the user to maintain a view of the text that has been selected by the user as the system moves between a first point and a second point. Performing an operation when a set of conditions has been met without further user input enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0341] In some embodiments, before a first input (e.g., 650e, 650g, 650u) pointing to the camera user interface is detected: the computer system (e.g., 600) communicates with one or more cameras; and the media representation (e.g., 630) is a representation (e.g., 630) (e.g., live camera preview) of one or more objects in a physical environment (e.g., physical space) in the field of view of one or more cameras. In some embodiments, receiving a request to display a fourth representation of the media (e.g., a representation of an updated camera field of view) includes detecting a change in the camera field of view. In some embodiments, the fourth representation of the media includes a change in the camera field of view. In some embodiments, when one or more objects (e.g., non-text objects) within the field of view are moving, the media representation is updated to indicate that one or more objects are moving. In some embodiments, the media representation is a live representation of the camera's field of view. Displaying a media representation of one or more objects in a physical environment in the field of view of one or more cameras (e.g., live camera preview) provides the user with greater control over the computer system (e.g., changing the field of view of the system's camera) to determine whether one or more objects in the physical space can be captured without cluttering the user interface. Providing additional control over the system without cluttering the UI with additional displayed controls enhances the system's operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0342] In some embodiments, the media representation (e.g., 630) is a first representation of the media. In some embodiments, when displaying a first user interface object, the computer system detects a request (e.g., 750k) to display a fifth representation (e.g., 630) of the media (e.g., the same or different media as the media represented by the first representation of the media). In some embodiments, when one or more changes in the field of view of one or more cameras communicating with the computer system are detected, a request to display the fifth representation of the media is detected. In some embodiments, when a request to zoom in / out on the media representation and / or pan the media representation is detected, a request to display the fifth representation of the media is detected. In some embodiments, when the computer system is moved, a request to display the fifth representation of the media is detected. In some embodiments, in response to detecting a request to display the fifth representation of the media (e.g., 750k) and based on a determination that a corresponding set of criteria is not met (e.g., corresponding text is not detected in the fifth representation of the media or corresponding text is detected but is not prominent enough), the computer system stops displaying the first user interface object (e.g., 680). In some embodiments, in response to detecting a request to display the fifth representation of the media and based on a determination that corresponding text is detected in the fifth representation of the media, the computer system continues to display the first user interface object. Stopping the display of the first user interface object when certain specified conditions are met (e.g., in response to detecting a request to display the fifth representation of the media and based on a determination that a corresponding set of criteria is not met) automatically provides the user with an indication of whether the media representation does not contain text that has been detected by the computer system. Performing an operation when a set of conditions has been met without further user input enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0343] In some embodiments, the corresponding criteria include criteria that are met when determining that the corresponding text meets a predetermined prominence criterion (e.g., the text is in a size or position in the media representation that indicates that the text is important and / or relevant) (e.g., based on the context of the media representation (e.g., importance / relevance based on the context of an image), based on the corresponding text occupying a certain amount of space on the displayed media representation, based on the corresponding text being in a specific position (e.g., the middle) on the displayed media representation, based on the corresponding text being a specific text type (e.g., email, phone number, QR code, unified access code location, etc.)) (e.g., based on as described below with respect to Figure 7C , Figures 7E to 7J and Figure 9One or more of the described techniques are determined to be relevant (e.g., when the corresponding text is on a label, a first user interface object is shown, and when detected on clothing, the first user object is not shown) (e.g., regarding how prominently / saliently the corresponding text is displayed).
[0344] In some embodiments, when a media representation (e.g., 630) is being displayed (and, in some embodiments, after detecting an input corresponding to a selection of a first user interface object and / or when the first user interface object is shown as being active and / or when multiple options for managing the corresponding text are being displayed) and based on determining that the corresponding text (e.g., 642a - 642b) includes a text portion (e.g., based on one or more regular expression patterns corresponding to different text types) that is determined to be of a corresponding text type (e.g., phone number, email), the computer system displays an indication (e.g., 638a - 638b) (e.g., an indication of a data detector) that the corresponding text type has been detected. In some embodiments, as part of displaying the indication that the corresponding text type has been detected, the computer system emphasizes (e.g., highlights, underlines, brackets) the portion of the text. In some embodiments, the indication that the corresponding text type has been detected is shown near, around, etc., the text portion of the corresponding text type. In some embodiments, based on determining that the corresponding text does not include a text portion having the corresponding text type (e.g., phone number, email), the computer system does not display (e.g., abandons displaying) the indication that the corresponding text type has been detected. Displaying the indication that a corresponding text type has been detected in a media representation provides the user with visual feedback regarding whether the media representation includes a particular text type. Providing the user with improved visual feedback enhances the operability of the computer system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0345] In some embodiments, when displaying multiple options for managing corresponding text (e.g., 680), the computer system receives a third input (e.g., 650h) (e.g., a tap input) that points to a portion of the camera user interface that does not include the corresponding text (e.g., a darkened or otherwise blurred portion of a media representation (e.g., a portion of the media representation that does not include text) (and / or a darkened portion of the camera user interface)). In some embodiments, in response to receiving the third input (e.g., 650h), the computer system stops displaying the multiple options for managing the corresponding text (e.g., 680). In some embodiments, in response to receiving the third input, one or more interface objects (e.g., a media capture enablement representation, a camera settings enablement representation, a camera mode enablement representation) are displayed (e.g., redisplayed) in the camera user interface and / or are displayed as being active (e.g., not darkened) (e.g., in response to user input on the corresponding object). Stopping the display of the multiple options for managing the corresponding text in response to receiving an input that points to a portion of the camera user interface provides the user with more control over the system without cluttering the user interface with additional user interface objects. Providing additional control over the system without cluttering the UI with additional displayed controls enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0346] In some embodiments, when a media representation (e.g., 630) and a media capture enablement representation (e.g., 610) are simultaneously displayed (e.g., before displaying a first user interface object) and based on determining that the media representation (e.g., 630) includes a first machine-readable code (e.g., a linear barcode, a matrix barcode, or a QR code), a computer system: displays a first user interface object (e.g., 680); and displays a representation of a uniform resource locator corresponding to the first machine-readable code (e.g., 668). Displaying the first user interface object and the representation of the uniform resource location improves security by notifying the location of the resource corresponding to the QR code before the user provides an input to navigate to the resource. Providing improved security reduces unauthorized execution of security operations, which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more securely and efficiently. Displaying the first user interface object and the representation of the uniform resource location when certain specified conditions are met (e.g., based on determining that the media representation includes a machine-readable code) notifies the user of the resource associated with the machine-readable code before the user selects the machine-readable code and provides the user with the uniform resource locator corresponding to the first machine-readable code. Performing an operation when a set of conditions has been met without further user input enhances the system's operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and efficiently.
[0347] In some embodiments, based on determining that a first input (e.g., 650u) corresponds to a selection of a first user interface object while a media representation (e.g., 630) includes a second machine-readable code (and when the machine-readable code is selected), the plurality of options for managing the corresponding text (e.g., 672) includes one or more options for managing information (e.g., a uniform resource locator) corresponding to the second machine-readable code. In some embodiments, based on determining that the first input corresponds to a selection of the first user interface object and the media representation does not include a machine-readable code (and / or when the machine-readable code is not selected), the plurality of options for managing the corresponding text does not include one or more options for managing information. In some embodiments, one or more of the plurality of options for managing the corresponding text displayed when the machine-readable code is selected are different from one or more of the plurality of options for managing the corresponding text displayed when text that does not include a machine-readable code is selected. Including one or more options for managing information corresponding to the machine-readable code in the plurality of options based on determining that the first input corresponds to a selection of the first user interface provides the user with additional control options (e.g., additional text management options) without cluttering the user interface. Providing additional control of the system without cluttering the UI with additional displayed controls enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the system), which in turn reduces power usage and extends the battery life of the system by enabling the user to use the system more quickly and effectively.
[0348] In some embodiments, the camera user interface includes a plurality of camera setting enabling representations (e.g., 620a - 620e) that can be selected to change the settings of one or more cameras (e.g., flash enabling representation, timer enabling representation, filter effect enabling representation, aperture number enabling representation, aspect ratio enabling representation, live photo enabling representation, etc.) (e.g., a plurality of user interface objects for accessing the corresponding camera settings). In some embodiments, the camera user interface includes a plurality of camera mode enabling representations (e.g., 620) (e.g., a plurality of user interface objects for setting the corresponding camera modes). In some embodiments, the plurality of camera setting enabling representations (e.g., 602a, 602b) are displayed simultaneously with the media capture enabling representation (e.g., 610) and / or the plurality of camera mode enabling representations (e.g., 620). In some embodiments, each camera mode (e.g., video (e.g., 620b), photo (e.g., 620c), portrait (e.g., 620b), slow motion (e.g., 620a), panorama (e.g., 620e) mode (e.g., 620)) has a plurality of settings (e.g., for the portrait camera mode: studio lighting setting, contour lighting setting, stage lighting setting), and the plurality of settings have a plurality of values (e.g., light levels for each setting) of the mode (e.g., portrait mode) in which the camera (e.g., camera sensor) is operating to capture media (including post - processing automatically performed after capture). Thus, for example, a camera mode is different from a mode that does not affect how the camera operates when capturing media or does not include a plurality of settings (e.g., flash mode with one setting that has a plurality of values (e.g., inactive, active, automatic)). In some embodiments, a camera mode allows the user to capture different types of media (e.g., photos or videos), and the settings for each mode can be optimized to capture a specific type of media corresponding to a particular mode (e.g., via post - processing), and the particular mode has specific attributes (e.g., shape (e.g., square, rectangle), speed (e.g., slow motion, time lapse), audio, video).For example, when the computer system is configured to operate in a still photo mode, one or more cameras of the computer system, when activated, utilize specific settings (e.g., flash settings, one or more filter settings) to capture a first type of media (e.g., rectangular photo); when the computer system is configured to operate in a square mode, one or more cameras of the computer system, when activated, utilize specific settings (e.g., flash settings and one or more filters) to capture a second type of media (e.g., square photo); when the computer system is configured to operate in a slow motion mode, one or more cameras of the computer system, when activated, utilize specific settings (e.g., flash settings, frames per second capture speed) to capture a third type of media (e.g., slow motion video); when the computer system is configured to operate in a portrait mode, one or more cameras of the computer system utilize specific settings (e.g., amount of a specific type of light (e.g., stage light, studio light, rim light), aperture number, blur) to capture a fifth type of media (e.g., portrait photo (e.g., photo with a blurred background)); when the computer system is configured to operate in a panoramic mode, one or more cameras of the computer system utilize specific settings (e.g., zoom, amount of field of view to capture while moving) to capture a fourth type of media (e.g., panoramic photo (e.g., wide photo)). In some embodiments, when switching between modes, the display of the representation of the field of view changes to correspond to the type of media that will be captured by that mode (e.g., when the computer system operates in a still photo mode, the representation is rectangular, and when the computer system operates in a square mode, the representation is square). A camera user interface that includes a plurality of camera setting affordance representations that can be selected to change the settings of one or more cameras provides the user with the ability to adjust multiple camera settings without having to navigate to various different user interfaces. Providing improved visual feedback to the user enhances the operability of the computer system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0349] In some embodiments, the camera user interface includes an affordance representation (e.g., 612) that, when selected, causes a representation of one or more previously captured media (e.g., 712) to be displayed (e.g., as described above with respect to Figure 6A and Figure 7A ). In some embodiments, the affordance representation includes a representation of the previously captured media. In some embodiments, in response to detecting a selection of the affordance representation (e.g., 612), a media representation (e.g., 712) in a media library associated with the computer system is displayed (e.g., as described above with respect to Figure 6A andFigure 7A The) The camera user interface that displays the affordance representation on the camera user interface provides the user with quick access to previously captured media items. Providing the user with improved visual feedback enhances the operability of the computer system and makes the user-system interface more effective (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system), which in turn reduces power usage and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0350] In some embodiments, the corresponding text includes a telephone number (e.g., detected in the corresponding text), and in response to detecting an input pointing to the telephone number, the computer system initiates a telephone call to the telephone number. In some embodiments, the corresponding text includes an email address. In some embodiments, in response to detecting an input pointing to the email address, the computer system launches (e.g., or opens) an email application that includes the email address (e.g., includes the email address in the "To" field) and / or automatically sends an email to the email address.
[0351] It should be noted that the processes described above with respect to method 800 (e.g., Figure 8 ) also apply in a similar manner to other methods described herein. For example, method 800 optionally includes one or more features of the various methods described herein with reference to methods 900, 1100, 1300, 1500, and 1700. For example, one or more indications of checkout features as described in method 1100 (e.g., Figure 11 ) may be displayed in the previously captured media item to identify the features present in the previously captured media item. For the sake of brevity, these details are not repeated hereinafter.
[0352] In some embodiments, one or more steps of the above-described method 800 may also be applied to video media representations, such as one or more live frames and / or paused frames of video media. In some embodiments, one or more steps of the above-described method 800 may be applied to media representations in user interfaces of applications different from those Figures 6A to 6Z and Figures 7A to 7L described, and these user interfaces include, but are not limited to, user interfaces corresponding to productivity applications (e.g., note-taking applications, spreadsheet applications, and / or task management applications), web applications, file viewer applications, and / or document processing applications and / or presentation applications.
[0353] Figure 9is a flowchart showing a method for managing visual indicators of visual content in media. Method 900 is executed at a computer system (e.g., 100, 300, 500) that communicates with a display generation component and one or more input devices. Some operations in method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0354] As described below, method 900 provides an intuitive way to manage visual indicators of visual content in media. The method reduces the cognitive burden on the user for managing visual indicators of visual content in media, thereby creating a more efficient human-machine interface. For battery-powered computing devices, enabling the user to manage visual indicators of visual content in media faster and more efficiently saves power and increases the time interval between battery charges.
[0355] Method 900 is executed at a computer system (e.g., a smartphone, a desktop computer, a laptop computer, a tablet computer) that communicates with a display generation component (e.g., a display controller, a touch-sensitive display system) and one or more input devices (e.g., a touch-sensitive surface).
[0356] The computer system displays (902) a first representation of a previously captured media item (e.g., photo media, video media) (e.g., a photo media or video media previously captured by receiving an input pointing to an optional user interface object for capturing media) (e.g., a photo media or video media available for the user to use, edit, and / or view later) (e.g., a representation of the previously captured media item at a first zoom level (e.g., a first portion of the previously captured media item)) (e.g., 724a (e.g., Figure 7B in 724a) (e.g., an image or video). In some embodiments, in response to receiving an input to a thumbnail representation of a previously captured media (and / or by receiving an input pointing to a representation of a different previously captured media (e.g., a swipe gesture)), a first representation of the previously captured media is displayed.
[0357] When displaying the first representation of a previously captured media item (e.g., 724a (e.g., Figure 7B in 724a)), the computer system detects (904) via one or more input devices a second representation corresponding to the displayed previously captured medi...
Claims
1. A method, comprising: at a computer system in communication with a display generation component and one or more input devices: displaying, via the display generation component, a first representation of a media item that includes a portion of text; when displaying the first representation of the media item that includes the portion of text, detecting, via the one or more input devices, an input corresponding to a request to display a second representation of the media item; in response to detecting the input corresponding to the request to display the second representation of the media item, displaying, via the display generation component, the second representation of the media item, wherein the second representation of the media item includes the portion of text at a location; and when displaying the second representation of the media item: displaying, via the display generation component, a visual indication that emphasizes a portion of the text corresponding to the portion of the text included in the second representation of the media item, based on determining that the portion of the text included in the second representation of the media item meets a corresponding set of criteria, wherein the visual indication is not displayed when the first representation of the media item is displayed, and wherein the visual indication is displayed at the location, wherein: the corresponding set of criteria includes a criterion that is met when determining that an amount of prominence of the portion of the text included in the second representation of the media item is higher than a prominence threshold, and the portion of the text included in the first representation of the media item is below the prominence threshold.
2. The method according to claim 1, further comprising: when displaying the second representation of the media item: abandoning displaying the visual indication based on determining that the portion of the text included in the second representation of the media item does not meet the corresponding set of criteria.
3. The method according to any one of claims 1 to 2, further comprising: when displaying the second representation of the media item, detecting, via the one or more input devices, an input corresponding to a request to display a third representation of the media item; and in response to detecting the input corresponding to the request to display the third representation of the media item, displaying, via the display generation component, the third representation of the media item; and when displaying the third representation of the media item: displaying, via the display generation component, a visual indication corresponding to the portion of the text included in the third representation of the media item, based on determining that a portion of the text included in the third representation of the media item meets the corresponding set of criteria.
4. The method according to any one of claims 1 to 2, wherein the first representation of the media item is a representation of the media item displayed at a first zoom level, and the second representation of the media item is a representation of the media item displayed at a second zoom level different from the first zoom level.
5. The method according to any one of claims 1 to 2, wherein the first representation of the media item is a representation of the media item displayed with a first conversion amount, and the second representation of the media item is a representation of the media item displayed with a second conversion amount different from the first conversion amount.
6. The method according to any one of claims 1 to 2, wherein the input corresponding to the request to display the second representation of the media item is an input detected on the display generation component.
7. The method according to any one of claims 1 to 2, wherein the second representation of the media item is displayed at a third zoom level, and the method further comprises: when the second representation of the media item is displayed at the third zoom level, detecting, via the one or more input devices, an input corresponding to a request to change the zoom level of the second representation of the media item; in response to detecting the input corresponding to the request to change the zoom level of the second representation of the media item, displaying, via the display generation component, a fourth representation of the media item at a fourth zoom level different from the third zoom level; and when the fourth representation of the media item is displayed at the fourth zoom level: abandoning the display of the visual indication based on determining that a portion of the text included in the fourth representation of the media item does not meet the corresponding set of criteria.
8. The method according to any one of claims 1 to 2, further comprising: when the second representation of the media item is displayed, detecting, via the one or more input devices, an input corresponding to a request to convert the second representation of the media item; in response to detecting the input corresponding to the request to convert the second representation of the media item, displaying a fifth representation of the media item, the fifth representation including a portion of the media item that is not included in the second representation of the media item; and when the fifth representation of the media item is displayed: abandoning the display of the visual indication based on determining that a second portion of the text included in the fifth representation of the media item does not meet the corresponding set of criteria.
9. The method according to any one of claims 1 to 2, wherein the amount of prominence being higher than the prominence threshold is based on the corresponding portion of the text occupying more than a threshold amount of the corresponding representation.
10. The method according to any one of claims 1 to 2, wherein the amount of prominence being higher than the prominence threshold is based on the corresponding portion of the text being displayed at a specific location in the corresponding representation.
11. The method according to any one of claims 1 to 2, further comprising: when the second representation of the media item is displayed and the visual indication is displayed, detecting, via the one or more input devices, an input corresponding to a request to change the second representation of the media item; and in response to detecting the input corresponding to the request to change the second representation of the media item, displaying a twelfth representation of the media item; and when the twelfth representation of the media item is displayed: Discard the display of the visual indication based on determining that the corresponding part of the text included in the twelfth representation of the media item does not meet the corresponding set of criteria.
12. The method according to any one of claims 1 to 2, wherein the second representation of the media item is displayed at a fifth zoom level, and the method further comprises: When the second representation of the media item and the visual indication are displayed at the fifth zoom level, detecting, via the one or more input devices, an input corresponding to a request to zoom in on the second representation of the media item; In response to detecting the input corresponding to the request to zoom in on the second representation of the media item, displaying, via the display generation component, a seventh representation of the media item at a sixth zoom level greater than the fifth zoom level, wherein the seventh representation of the media item includes a third part of the text included in the second representation of the media item, and the third part of the text is different from the part of the text included in the second representation; and When the seventh representation of the media item is displayed: Based on determining that the third part of the text meets the corresponding set of criteria, displaying, via the display generation component, a visual indication corresponding to the third part of the text, the visual indication being different from the visual indication corresponding to the part of the text.
13. The method according to any one of claims 1 to 2, wherein the second representation of the media item is displayed at a seventh zoom level, and the method further comprises: When the second representation of the media item is displayed at the seventh zoom level, detecting, via the one or more input devices, an input corresponding to a request to zoom out on the second representation of the media item; In response to detecting the input corresponding to the request to zoom out on the second representation of the media item, displaying, via the display generation component, an eighth representation of the media item at an eighth zoom level less than the seventh zoom level; and When the eighth representation of the media item is displayed at the eighth zoom level: Based on determining that the first corresponding part of the text included in the eighth representation of the media item does not meet the corresponding set of criteria, stop displaying the visual indication.
14. The method according to any one of claims 1 to 2, wherein the visual indication surrounds the part of the text included in the second representation.
15. The method according to any one of claims 1 to 2, wherein the second representation of the media item is displayed at a ninth zoom level, and the method further comprises: When the second representation of the media item and the visual indication are displayed at the ninth zoom level, detecting, via one or more input devices, an input corresponding to a request to zoom in on the second representation of the media item; And In response to detecting the input corresponding to a request to magnify the second representation of the media item, display, via the display generating component, a ninth representation of the media item at a tenth zoom level greater than the ninth zoom level, wherein the ninth representation of the media item includes a corresponding portion of text included in the second representation of the media item, and the corresponding portion of the text is different from the portion of the text included in the second representation; and When displaying the ninth representation of the media item: Stop displaying the visual indication corresponding to the portion of the text; and According to determining that a second corresponding portion of the text meets the corresponding set of criteria, display, via the display generating component, a visual indication corresponding to the second corresponding portion of the text, the visual indication being different from the visual indication corresponding to the portion of the text.
16. The method according to any one of claims 1 to 2, wherein the second representation of the media item includes a third portion of text that cannot be selected, and the method further includes: When displaying the second representation of the media item, detect, via one or more input devices, an input corresponding to a request to display a tenth representation of the media item; And In response to detecting the input corresponding to a request to display the tenth representation of the media item, display the tenth representation of the media item including the third portion of the text, wherein the third portion of the text included in the tenth representation of the media item is selectable.
17. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system, wherein the computer system communicates with a display generating component and one or more input devices, and the one or more programs include instructions for performing the method according to any one of claims 1 to 16.
18. A computer system configured to communicate with a display generating component and one or more input devices, comprising: One or more processors; And A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 16.
19. A computer system configured to communicate with a display generating component and one or more input devices, comprising: Means for performing the method according to any one of claims 1 to 16.
20. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more input devices, the one or more programs including instructions for the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1
Virtual input device placement on a touch screen user interface
US20060033724A1
Multipoint touchscreen
US20060097991A1