A user interface for managing visual content within media.

The method and interface streamline media management by conditional display of user interface objects based on detected text, enhancing efficiency and reducing power consumption in battery-operated devices.

JP7893938B2Active Publication Date: 2026-07-22APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
APPLE INC
Filing Date
2025-05-02
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing technologies for managing visual content in media are cumbersome and inefficient, often requiring multiple key presses or keystrokes, wasting user time and device energy, particularly in battery-operated devices.

Method used

A method and interface that simultaneously displays media and media capture affordances, with conditional user interface objects, allowing for efficient media capture and management based on detected text, and providing options for managing individual text or additional information based on specific criteria.

Benefits of technology

Enhances user efficiency and reduces power consumption by streamlining media management processes, improving user satisfaction and extending battery life in portable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893938000001
    Figure 0007893938000001
  • Figure 0007893938000002
    Figure 0007893938000002
  • Figure 0007893938000003
    Figure 0007893938000003
Patent Text Reader

Abstract

To provide electronic devices with faster, more efficient methods and interfaces for managing a visual content in media.SOLUTION: A method includes: displaying a camera user interface that includes concurrently displaying a representation of media and a media capture affordance; in accordance with a determination that an individual set of criteria is satisfied including a criterion that is satisfied when an individual text is detected in the representation of the media, displaying a first user interface object corresponding to one or more text management operations; detecting a first input directed to the camera user interface; in accordance with a determination that the first input corresponds to selection of the media capture affordance, initiating capture of media to be added to a media library; and displaying a plurality of options to manage the individual text.SELECTED DRAWING: Figure 6A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Patent Application No. 63 / 176,847, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on April 19, 2021; U.S. Patent Application No. 63 / 197,497, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on June 6, 2021; U.S. Patent Application No. 17 / 484,844, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on September 24, 2021; U.S. Patent Application No. 17 / 484,714, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on September 24, 2021; U.S. Patent Application No. 17 / 484,856, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on September 24, 2021; and U.S. Patent Application No. 63 / 318,677, titled "USER INTERFACES FOR MANAGING VISUAL CONTENT IN MEDIA", filed on March 10, 2022. The contents of these applications are hereby incorporated by reference in their entirety.

[0002] The present disclosure generally relates to computer user interfaces, and more specifically to techniques for managing visual content in media.

Background Art

[0003] Smartphones and other personal electronic devices allow users to capture and view content within media. Users can capture various types of media, including video and image data. Users can store the captured media on their smartphone or other personal electronic devices. [Overview of the project]

[0004] However, some technologies that use computer systems to manage content within media are generally cumbersome and inefficient. For example, some existing technologies employ complex and time-consuming user interfaces that may involve multiple key presses or keystrokes. These existing technologies take more time than necessary, wasting both the user's time and the device's energy. This latter consideration is particularly important in battery-powered devices.

[0005] Therefore, this technology provides electronic devices with a faster and more efficient method and interface for managing visual content within media. Such methods and interfaces optionally complement or replace other methods for managing visual content within media. Such methods and interfaces reduce the cognitive burden on the user and create a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the interval between battery charges.

[0006] A method is described according to several embodiments. The method is performed on a computer system that communicates with a display generation component. The method includes: displaying a camera user interface via a display generation component, the camera user interface including simultaneously displaying a representation of media and media capture affordances; displaying a first user interface object corresponding to one or more text management operations via the display generation component, in accordance with a determination that a separate set of criteria is met, including criteria that are met when individual text is detected within a representation of media, while the representation of media and media capture affordances are simultaneously displayed; refraining from displaying the first user interface object, in accordance with a determination that a separate set of criteria is not met; detecting a first input directed to the camera user interface while the representation of media is displayed; in response to the detection of the first input directed to the camera user interface, in accordance with a determination that the first input corresponds to the selection of media capture affordances, initiating the capture of media to be added to a media library associated with the computer system; and displaying a plurality of options for managing individual text via the display generation component, in accordance with a determination that the first input corresponds to the selection of the first user interface object.

[0007] According to some embodiments, non-temporary computer-readable storage is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and one or more programs, via the display generation component, display a camera user interface, the camera user interface includes simultaneously displaying a representation of media and media capture affordances, and while simultaneously displaying the representation of media and media capture affordances, the display generation component, in accordance with a determination that a separate set of criteria is met, including criteria that are met when individual text is detected within the representation of media, enables one or more text management operations. The system includes instructions to display a corresponding first user interface object, to refrain from displaying the first user interface object if a separate set of criteria is not met, to detect a first input directed to the camera user interface while displaying a media representation, to start capturing media to be added to the media library associated with the computer system in response to the detection of the first input directed to the camera user interface and in accordance with the determination that the first input corresponds to the selection of a media capture affordance, and to display multiple options for managing individual texts via a display generation component, in accordance with the determination that the first input corresponds to the selection of a first user interface object.

[0008] According to some embodiments, temporary computer-readable storage is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and one or more programs, via the display generation component, display a camera user interface, the camera user interface includes simultaneously displaying a representation of media and media capture affordances, and while simultaneously displaying the representation of media and media capture affordances, the display generation component, upon determination that a separate set of criteria is met, including criteria that are met when individual text is detected within the representation of media, one or more text management operations are performed. The command includes: displaying a first user interface object that corresponds to; refraining from displaying the first user interface object according to a determination that a separate set of criteria is not met; detecting a first input directed to the camera user interface while displaying a media representation; in response to the detection of the first input directed to the camera user interface, in accordance with a determination that the first input corresponds to the selection of a media capture affordance, initiating the capture of media to be added to the media library associated with the computer system; and, according to a determination that the first input corresponds to the selection of the first user interface object, displaying multiple options for managing individual texts via a display generation component.

[0009] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system comprises one or more processors and a memory storing one or more programs configured to run by the one or more processors, wherein one or more programs, via the display generation component, display a camera user interface, the camera user interface includes simultaneously displaying a representation of media and media capture affordances, and while simultaneously displaying a representation of media and media capture affordances, the display generation component communicates with one or more first user interfaces of which correspond to one or more text management operations, according to a determination that a separate set of criteria, including criteria that are met when individual text is detected within a representation of media, is met. The command includes displaying an object, and if a separate set of the above criteria is determined not to be met, the display of the first user interface object is withheld; detecting a first input directed to the camera user interface while displaying a media representation; and, in response to the detection of the first input directed to the camera user interface, in accordance with the determination that the first input corresponds to the selection of a media capture affordance, initiating the capture of media to be added to the media library associated with the computer system; and, in accordance with the determination that the first input corresponds to the selection of the first user interface object, displaying multiple options for managing individual texts via a display generation component.

[0010] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system described above comprises one or more processors, a memory for storing one or more programs configured to be executed by one or more processors, and means for displaying a camera user interface via a display generation component, the camera user interface including simultaneously displaying a representation of media and media capture affordances; means for displaying a first user interface object corresponding to one or more text management operations via the display generation component, in accordance with a determination that a separate set of criteria, including criteria that are satisfied when individual text is detected in the representation of media, is satisfied while the representation of media and media capture affordances are simultaneously displayed, and refraining from displaying the first user interface object in accordance with a determination that the separate set of criteria is not satisfied; means for detecting a first input directed to the camera user interface while the representation of media is displayed; and means for initiating the capture of media to be added to a media library associated with the computer system in response to the detection of a first input directed to the camera user interface, in accordance with a determination that the first input corresponds to the selection of a media capture affordance, and displaying a plurality of options for managing individual text via the display generation component, in accordance with a determination that the first input corresponds to the selection of a first user interface object.

[0011] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that communicate with a display generation component. The program includes instructions that, via a display generation component, display a camera user interface, the camera user interface includes simultaneously displaying a representation of media and media capture affordances; that, while simultaneously displaying the representation of media and media capture affordances, display a first user interface object corresponding to one or more text management operations via the display generation component, in accordance with a determination that a distinct set of criteria, including criteria that are satisfied when individual text is detected within the representation of media, is satisfied; that, in accordance with a determination that a distinct set of criteria is not satisfied, display the first user interface object; that, while displaying the representation of media, detect a first input directed to the camera user interface; that, in response to the detection of the first input directed to the camera user interface, initiate the capture of media to be added to a media library associated with the computer system, in accordance with a determination that the first input corresponds to the selection of a media capture affordance; and that, in accordance with a determination that the first input corresponds to the selection of a first user interface object, display a plurality of options for managing individual text via the display generation component.

[0012] The method is described according to several embodiments. The method is performed in a computer system that communicates with a display generation component and one or more input devices. The method includes: displaying a first representation of a previously captured media item via a display generation component; detecting inputs via one or more input devices that correspond to a request to display a second representation of a previously captured media item while the first representation of the previously captured media item is being displayed; displaying a second representation of the previously captured media item via a display generation component in response to the detection of inputs that correspond to a request to display a second representation of a previously captured media item; and, while the second representation of the previously captured media item is being displayed, displaying a visual indication via a display generation component that corresponds to a portion of the text contained in the second representation of the previously captured media item that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of the text contained in the second representation of the previously captured media item satisfies a separate set of criteria.

[0013] According to some embodiments, non-temporary computer-readable storage is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component and one or more input devices, and one or more programs include instructions to display a first representation of a previously captured media item via the display generation component, to detect inputs via one or more input devices that correspond to a request to display a second representation of the previously captured media item while the first representation of the previously captured media item is being displayed, to display a second representation of the previously captured media item via the display generation component in response to the detection of inputs that correspond to a request to display a second representation of the previously captured media item, and to display a visual indication via the display generation component that corresponds to a portion of the text contained in the second representation of the previously captured media item that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of the text contained in the second representation of the previously captured media item satisfies a separate set of criteria while the second representation of the previously captured media item is being displayed.

[0014] According to some embodiments, temporary computer-readable storage is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component and one or more input devices, and the one or more programs include instructions to display a first representation of a previously captured media item via the display generation component, to detect inputs via one or more input devices that correspond to a request to display a second representation of the previously captured media item while the first representation of the previously captured media item is being displayed, to display a second representation of the previously captured media item via the display generation component in response to the detection of inputs that correspond to a request to display a second representation of the previously captured media item, and to display a visual indication via the display generation component that corresponds to a portion of the text contained in the second representation of the previously captured media item that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of the text contained in the second representation of the previously captured media item satisfies a separate set of criteria while the second representation of the previously captured media item is being displayed.

[0015] According to some embodiments, a computer system is described that is configured to communicate with a display generation component and one or more input devices. The computer system includes one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, wherein one or more programs include instructions to display a first representation of a previously captured media item via the display generation component; while displaying the first representation of the previously captured media item, to detect inputs via one or more input devices that correspond to a request to display a second representation of the previously captured media item; to display a second representation of the previously captured media item via the display generation component in response to the detection of inputs that correspond to a request to display a second representation of the previously captured media item; and while displaying the second representation of the previously captured media item, to display a visual indication via the display generation component corresponding to a portion of the text contained in the second representation that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of the text contained in the second representation of the previously captured media item satisfies a separate set of criteria.

[0016] According to some embodiments, a computer system is described that is configured to communicate with a display generation component and one or more input devices. The computer system comprises one or more processors, a memory storing one or more programs configured to be executed by one or more processors, means for displaying a first representation of a previously captured media item via a display generation component, means for detecting inputs corresponding to a request to display a second representation of a previously captured media item via one or more input devices while the first representation of the previously captured media item is being displayed, means for displaying a second representation of the previously captured media item via a display generation component in response to the detection of inputs corresponding to a request to display a second representation of the previously captured media item, and means for displaying a visual indication via a display generation component that corresponds to a portion of text contained in the second representation of the previously captured media item that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of text contained in the second representation of the previously captured media item satisfies a separate set of criteria while the second representation of the previously captured media item is being displayed.

[0017] According to several embodiments, a computer program product is described. The computer program product comprises one or more programs configured to run by one or more processors of a computer system communicating with a display generation component and one or more input devices. The one or more programs include instructions to display a first representation of a previously captured media item via the display generation component, to detect inputs via one or more input devices that correspond to a request to display a second representation of the previously captured media item while the first representation of the previously captured media item is being displayed, to display a second representation of the previously captured media item via the display generation component in response to the detection of inputs that correspond to a request to display a second representation of the previously captured media item, and to display a visual indication via the display generation component that corresponds to a portion of the text contained in the second representation of the previously captured media item that was not displayed when the first representation of the previously captured media item was displayed, in accordance with a determination that a portion of the text contained in the second representation of the previously captured media item satisfies a separate set of criteria.

[0018] A method is described according to several embodiments. The above method is performed in a computer system communicating with one or more cameras, one or more input devices, and a display generation component. The above method includes: displaying a first user interface including a text input area; detecting a request to display a camera user interface while the first user interface including the text input area is being displayed; displaying a camera user interface via a display generation component in response to the detection of the request to display a camera user interface, wherein the camera user interface includes a representation of the field of view of one or more cameras; displaying a selectable text insertion user interface object that inserts at least a portion of detected text into the text input area, according to a determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria; detecting input via one or more input devices corresponding to the selection of the text insertion user interface object while the representation of the field of view and the text insertion user interface object are being displayed simultaneously; and inserting at least a portion of the detected text into the text input area in response to the detection of input corresponding to the selection of the text insertion user interface object.

[0019] According to some embodiments, non-temporary computer-readable storage is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with one or more cameras, one or more input devices, and a display generation component, and the one or more programs include instructions to display a first user interface including a text input area, to detect a request to display a camera user interface while the first user interface including the text input area is being displayed, to display a camera user interface via the display generation component in response to the detection of the request to display a camera user interface, the camera user interface including a representation of the field of view of one or more cameras, and to display a selectable text insertion user interface object which inserts at least a portion of the detected text into the text input area according to the determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria, and while the representation of the field of view and the text insertion user interface object are being displayed simultaneously, to detect input corresponding to the selection of the text insertion user interface object via one or more input devices, and to insert at least a portion of the detected text into the text input area in response to the detection of input corresponding to the selection of the text insertion user interface object.

[0020] According to some embodiments, temporary computer-readable storage is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with one or more cameras, one or more input devices, and a display generation component, and the one or more programs include instructions to display a first user interface including a text input area, to detect a request to display a camera user interface while the first user interface including the text input area is being displayed, to display a camera user interface via the display generation component, the camera user interface including a representation of the field of view of one or more cameras, and to display a selectable text insertion user interface object which inserts at least a portion of the detected text into the text input area according to a determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria, and while the representation of the field of view and the text insertion user interface object are being displayed simultaneously, to detect input corresponding to the selection of the text insertion user interface object via one or more input devices, and to insert at least a portion of the detected text into the text input area according to the detection of input corresponding to the selection of the text insertion user interface object.

[0021] According to some embodiments, a computer system is described that is configured to communicate with one or more cameras, one or more input devices, and an output generation component. The computer system includes one or more processors and a memory that stores one or more programs configured to be executed by one or more processors, wherein one or more programs include instructions to display a first user interface including a text input area, to detect a request to display a camera user interface while the first user interface including the text input area is being displayed, to display a camera user interface via a display generation component, to display a camera user interface which includes a representation of the field of view of one or more cameras, and to display a selectable text insertion user interface object which inserts at least a portion of the detected text into the text input area according to a determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria, and while the representation of the field of view and the text insertion user interface object are being displayed simultaneously, to detect an input corresponding to the selection of the text insertion user interface object via one or more input devices, and to insert at least a portion of the detected text into the text input area according to the detection of the input corresponding to the selection of the text insertion user interface object.

[0022] According to some embodiments, a computer system is described that is configured to communicate with one or more cameras, one or more input devices, and an output generation component. The computer system includes a memory for storing one or more programs configured to be executed by one or more processors; means for displaying a first user interface including a text input area; means for detecting a request to display a camera user interface while the first user interface including the text input area is being displayed; means for displaying a selectable text insertion user interface object via an output generation component in response to the detection of a request to display a camera user interface, which includes a representation of the field of view of one or more cameras, and inserts at least a portion of the detected text into the text input area according to a determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria; means for detecting an input corresponding to the selection of the text insertion user interface object via one or more input devices while the representation of the field of view and the text insertion user interface object are being displayed simultaneously; and means for inserting at least a portion of the detected text into the text input area in response to the detection of an input corresponding to the selection of the text insertion user interface object.

[0023] According to several embodiments, a computer program product is described. The computer program product comprises one or more programs configured to run on one or more processors of a computer system communicating with one or more cameras, one or more input devices, and a display generation component. The one or more programs include instructions to display a first user interface including a text input area, to detect a request to display a camera user interface while the first user interface including the text input area is being displayed, to display a camera user interface via the display generation component, the camera user interface including a representation of the field of view of one or more cameras, and to display a selectable text insertion user interface object that inserts at least a portion of the detected text into the text input area according to a determination that the representation of the field of view of one or more cameras includes detected text that satisfies one or more criteria, and to detect input corresponding to the selection of the text insertion user interface object via one or more input devices while the representation of the field of view and the text insertion user interface object are being displayed simultaneously, and to insert at least a portion of the detected text into the text input area according to the detection of input corresponding to the selection of the text insertion user interface object.

[0024] A method is described according to several embodiments. The method is performed on a computer system that communicates with a display generation component. The method includes: displaying a media user interface including a representation of media via a display generation component; receiving a request to display additional information about a plurality of detected features in the representation of media while the media user interface including the representation of media is being displayed; and displaying one or more indications of a plurality of detected characteristics in the media while the media user interface including the representation of media is being displayed in response to receiving a request to display additional information about a plurality of detected characteristics, wherein one or more indications of a plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location in the representation of media, the first location corresponding to the location of the first detected characteristic in the representation of media, the first indication having a first appearance according to a determination that the first detected characteristic is a characteristic of a first type, and the first indication having a second appearance according to a determination that the first detected characteristic is a characteristic of a second type different from the first type.

[0025] According to some embodiments, non-temporary computer-readable storage is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and one or more programs, via the display generation component, display a media user interface including a representation of media, and while displaying the media user interface including a representation of media, receive a request to display additional information about a plurality of detected characteristics in the representation of media, and in response to receiving a request to display additional information about a plurality of detected characteristics, while displaying the media user interface including a representation of media A non-temporary computer-readable storage medium including instructions for displaying one or more indications of a plurality of detected characteristics within the medium, wherein the one or more indications of the plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location in the representation of the medium, the first location corresponding to the location of the first detected characteristic in the representation of the medium, the first indication having a first appearance according to a determination that the first detected characteristic is a first type of characteristic, and the first indication having a second appearance according to a determination that the first detected characteristic is a second type of characteristic different from the first type of characteristic.

[0026] According to some embodiments, temporary computer-readable storage is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and one or more programs, via the display generation component, display a media user interface including a representation of media, and while displaying the media user interface including a representation of media, receive a request to display additional information about a plurality of detected characteristics in the representation of media, and in response to receiving a request to display additional information about a plurality of detected characteristics, while displaying the media user interface including a representation of media A temporary computer-readable memory containing instructions for displaying one or more indications of a plurality of detected characteristics in a medium, wherein the one or more indications of the plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location in the representation of the medium, the first location corresponding to the location of the first detected characteristic in the representation of the medium, the first indication having a first appearance according to a determination that the first detected characteristic is a characteristic of a first type, and the first indication having a second appearance according to a determination that the first detected characteristic is a characteristic of a second type different from the first type.

[0027] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs display a media user interface including a media representation via the display generation component, and while displaying the media user interface including the media representation, receive a request to display additional information regarding a plurality of detected characteristics within the media representation, and in response to receiving the request to display additional information regarding the plurality of detected characteristics, display one or more indications of one or more of the plurality of detected characteristics within the media while displaying the media user interface including the media representation. The one or more indications of the plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location within the media representation, the first location corresponding to the location of the first detected characteristic within the media representation, and according to a determination that the first detected characteristic is a first type of characteristic, the first indication has a first appearance, and according to a determination that the first detected characteristic is a second type of characteristic different from the first type of characteristic, the first indication has a second appearance different from the first appearance.

[0028] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system described above comprises one or more processors, a memory for storing one or more programs configured to be executed by the one or more processors, means for displaying a media user interface including a representation of media via a display generation component, means for receiving a request to display additional information regarding a plurality of detected characteristics in a representation of media while the media user interface including a representation of media is being displayed, and means for displaying one or more indications of a plurality of detected characteristics in media while the media user interface including a representation of media is being displayed in response to receiving a request to display additional information regarding a plurality of detected characteristics, wherein one or more indications of a plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location in the representation of media, the first location corresponding to the location of the first detected characteristic in the representation of media, the first indication having a first appearance according to a determination that the first detected characteristic is a first type of characteristic, and the first indication having a second appearance according to a determination that the first detected characteristic is a second type of characteristic different from the first type of characteristic.

[0029] According to some embodiments, a computer program product is described. The above-described computer program product comprises one or more programs configured to run on one or more processors of a computer system communicating with a display generation component, wherein one or more programs, via the display generation component, display a media user interface including a representation of media, receive a request to display additional information about a plurality of detected characteristics in the representation of media while displaying the media user interface including the representation of media, and, in response to receiving a request to display additional information about a plurality of detected characteristics, display one or more indications of a plurality of detected characteristics while displaying the media user interface including the representation of media, wherein one or more indications of a plurality of detected characteristics include a first indication of a first detected characteristic displayed at a first location in the representation of media, the first location corresponding to the location of the first detected characteristic in the representation of media, the first indication having a first appearance according to a determination that the first detected characteristic is a characteristic of a first type, and the first indication having a second appearance according to a determination that the first detected characteristic is a characteristic of a second type different from the first type.

[0030] According to some embodiments, a method is described. The method is performed in a computer system communicating with one or more cameras, a display generation component, and one or more input devices. The method includes receiving a request to display a representation of the field of view of one or more cameras, and in response to receiving the request to display a representation of the field of view of one or more cameras, displaying, via the display generation component, a representation of the field of view of one or more cameras, the representation including text within the field of view of one or more cameras, automatically displaying, via the display generation component, a plurality of indications of translated text including a first indication of a translation of a first portion of the text and a second indication of a translation of a second portion of the text, receiving, via the one or more input devices, a request to select an individual indication of a plurality of translated portions while the first indication and the second indication are being displayed via the display generation component, and in response to receiving the request to select an individual indication, displaying, via the display generation component, a first translation user interface object including the first portion of the text and the translation of the first portion of the text without including the translation of the second portion of the text according to a determination that the request is a request to select the first indication.

[0031] According to several embodiments, non-temporary computer-readable storage is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with one or more cameras, a display generation component, and one or more input devices, and one or more programs receive a request to display a representation of the field of view of one or more cameras, and in response to receiving a request to display a representation of the field of view of one or more cameras, the display generation component displays a representation of the field of view of one or more cameras, the representation includes text within the field of view of one or more cameras, and the display generation component displays a first indication of a translation of a first part of the text and text The command includes instructions to automatically display multiple indicators of the translated text, including a second indicator of the translation of the second part of the text, and to receive a request via one or more input devices to select individual indicators of the multiple translated parts while the first and second indicators are being displayed via a display generation component, and in response to the receipt of a request to select individual indicators, the command to display a first translation user interface object via the display generation component, which includes the first part of the text and the translation of the first part of the text, but does not include the translation of the second part of the text.

[0032] According to several embodiments, temporary computer-readable storage is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with one or more cameras, a display generation component, and one or more input devices, and one or more programs receive a request to display a representation of the field of view of one or more cameras, and in response to receiving a request to display a representation of the field of view of one or more cameras, the display generation component displays a representation of the field of view of one or more cameras, the representation includes text within the field of view of one or more cameras, and the display generation component displays a first indication of a translation of a first part of the text and text The command includes instructions to automatically display multiple indicators of the translated text, including a second indicator of the translation of the second part of the text, and to receive a request via one or more input devices to select individual indicators of the multiple translated parts while the first and second indicators are being displayed via a display generation component, and in response to receiving a request to select individual indicators, the command to display a first translation user interface object via the display generation component, which includes the first part of the text and the translation of the first part of the text, but does not include the translation of the second part of the text.

[0033] According to some embodiments, a computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices is described. The computer system described above includes one or more processors and a memory for storing one or more programs configured to be executed by one or more processors, and includes instructions for one or more programs to receive a request to display a representation of the field of view of one or more cameras, and in response to receiving a request to display a representation of the field of view of one or more cameras, display a representation of the field of view of one or more cameras, wherein the representation includes text within the field of view of one or more cameras, and to automatically display multiple indications of translated text, including a first indication of the translation of a first part of the text and a second indication of the translation of a second part of the text, via the display generation component, and while the first indication and the second indication are being displayed, receive a request via one or more input devices to select individual indications of the multiple translated parts, and in response to receiving a request to select individual indications, display a first translation user interface object via the display generation component, including the first part of the text and the translation of the first part of the text, but without the translation of the second part of the text, according to the determination that the request is a request to select the first indication.

[0034] According to some embodiments, a computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices is described. The computer system described above comprises one or more processors, a memory for storing one or more programs configured to be executed by one or more processors, means for receiving a request to display a representation of the field of view of one or more cameras, and in response to receiving a request to display a representation of the field of view of one or more cameras, the system displays, via a display generation component, a representation of the field of view of one or more cameras, wherein the representation includes text within the field of view of one or more cameras, and automatically displays, via the display generation component, a plurality of indications of translated text, including a first indication of the translation of a first part of the text and a second indication of the translation of a second part of the text, and, while the first indication and the second indication are being displayed, means for receiving a request via a display generation component to select individual indications of the plurality of translated parts, and in response to receiving a request to select individual indications, the system displays, via the display generation component, a first translation user interface object including the first part of the text and the translation of the first part of the text, but without the translation of the second part of the text, according to a determination that the request is a request to select the first indication.

[0035] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras, a display generation component, and one or more input devices. The program includes instructions that one or more programs receive a request to display a representation of the field of view of one or more cameras, and in response to the receipt of the request to display a representation of the field of view of one or more cameras, the program displays a representation of the field of view of one or more cameras, the representation including text within the field of view of one or more cameras, the program automatically displays multiple indicators of translated text, including a first indicator of the translation of a first part of the text and a second indicator of the translation of a second part of the text, the program receives a request via one or more input devices to select individual indicators of the multiple translated parts, and in response to the receipt of the request to select individual indicators, the program displays a first translation user interface object via the display generation component, including the first part of the text and the translation of the first part of the text, but not the translation of the second part of the text.

[0036] According to some embodiments, a method is described in which a computer system communicates with a display generation component. The method includes detecting a request to display additional information corresponding to a representation of media while displaying a user interface that includes a representation of media; and, in response to the detection of a request to display additional information corresponding to a representation of media, displaying a first user interface object via the display generation component, which, when selected, causes the computer system to perform a first action based on the detected text, according to a determination that the detected text in the representation of media has a first set of properties; and displaying a second user interface object via the display generation component, which, when selected, causes the computer system to perform a second action different from the first action, according to a determination that the detected text in the representation of media has a second set of properties different from the first set of properties.

[0037] According to some embodiments, a non-temporary computer-readable storage medium is described. The non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, and includes instructions that, while one or more programs are displaying a user interface including a representation of media, detect a request to display additional information corresponding to a representation of media, and in response to the detection of a request to display additional information corresponding to a representation of media, the display generation component, via a first user interface object, which, when selected, causes the computer system to perform a first action based on the detected text, and, upon determination that the detected text in the representation of media has a first set of properties, displays a second user interface object, which, when selected, causes the computer system to perform a second action different from the first action, via a second user interface object, via a second user interface object, which, when selected, causes the computer system to perform a second action different from the first action based on the detected text.

[0038] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, and includes instructions that, while one or more programs are displaying a user interface including a representation of media, detect a request to display additional information corresponding to a representation of media, and, in response to the detection of a request to display additional information corresponding to a representation of media, the display generation component, via a first user interface object, which, when selected, causes the computer system to perform a first action based on the detected text, and, upon determination that the detected text in the representation of media has a first set of properties, displays a second user interface object, which, when selected, causes the computer system to perform a second action different from the first action, via a second user interface object, via a second user interface object, which, when selected, causes the computer system to perform a second action different from the first action based on the detected text.

[0039] According to several embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, and the computer system comprises one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, and includes instructions for one or more programs to detect a request to display additional information corresponding to a representation of media while the one or more programs are displaying a user interface including a representation of media, and in response to the detection of a request to display additional information corresponding to a representation of media, the display generation component, via a first user interface object, which, when selected, causes the computer system to perform a first action based on the detected text, and the display generation component, via a second user interface object, which, when selected, causes the computer system to perform a second action different from the first action based on the detected text.

[0040] According to several embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and includes means for detecting a request to display additional information corresponding to a representation of media while the computer system is displaying a user interface including a representation of media; and means for displaying a first user interface object via the display generation component, in response to the detection of a request to display additional information corresponding to a representation of media, according to a determination that the detected text in the representation of media has a first set of properties, the first user interface object, when selected, causes the computer system to perform a first action based on the detected text; and means for displaying a second user interface object via the display generation component, according to a determination that the detected text in the representation of media has a second set of properties different from the first set of properties, the second user interface object, when selected, causes the computer system to perform a second action different from the first action based on the detected text.

[0041] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with a generation component, the one or more programs, while displaying a user interface including a representation of media, detect a request to display additional information corresponding to a representation of media, and, in response to the detection of a request to display additional information corresponding to a representation of media, the program includes instructions to display a first user interface object via the display generation component, which, when selected, causes the computer system to perform a first action based on the detected text, according to a determination that the detected text in the representation of media has a first set of properties, and to display a second user interface object via the display generation component, which, when selected, causes the computer system to perform a second action different from the first action based on the detected text.

[0042] The executable instructions that perform these functions are optionally contained within a non-temporary computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0043] Therefore, devices are provided with faster and more efficient ways and interfaces for managing visual content within media, thereby increasing effectiveness, efficiency, and user satisfaction with such devices. Such methods and interfaces may complement or replace other ways of managing visual content within media. [Brief explanation of the drawing]

[0044] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.

[0045] [Figure 1A] This is a block diagram showing a portable multifunctional device with a touch-sensitive display, according to several embodiments.

[0046] [Figure 1B] This is a block diagram showing exemplary components for event handling according to several embodiments.

[0047] [Figure 2] This figure shows a portable multifunctional device having a touchscreen, according to several embodiments.

[0048] [Figure 3] This is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, according to several embodiments.

[0049] [Figure 4A] The following are exemplary user interfaces for application menus on a portable multifunction device, according to several embodiments.

[0050] [Figure 4B]This document illustrates an exemplary user interface for a multifunctional device having a touch-sensitive surface separate from the display, according to several embodiments.

[0051] [Figure 5A] A personal electronic device according to several embodiments is shown.

[0052] [Figure 5B] This is a block diagram showing a personal electronic device according to several embodiments.

[0053] [Figure 6A] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6B] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6C] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6D] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6E] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6F] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6G] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6H] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6I] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6J]This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6K] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6L] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6M] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6N] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6O] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6P] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6Q] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6R] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6S] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6T] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6U] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6V] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6W]This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6X] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6Y] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments. [Figure 6Z] This section describes an exemplary user interface for managing visual content within a media, according to several embodiments.

[0054] [Figure 7A] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7B] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7C] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7D] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7E] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7F] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7G] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7H]This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7I] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7J] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7K] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments. [Figure 7L] This document illustrates an exemplary user interface for managing visual indicators for visual content within media, according to several embodiments.

[0055] [Figure 8] This flowchart illustrates a method for managing visual content within media, according to several embodiments.

[0056] [Figure 9] This flowchart illustrates how to manage visual indicators for visual content within media, according to several embodiments.

[0057] [Figure 10A] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10B] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10C] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10D] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10E] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10F] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10G] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10H] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10I] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10J] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10K] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10L] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10M] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10N] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10O] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10P] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10Q] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10R]Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10S] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10T] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10U] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10V] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10W] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10X] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10Y] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10Z] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10AA] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10AB] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10AC] Several embodiments illustrate exemplary user interfaces for inserting visual content within media. [Figure 10AD] Several embodiments illustrate exemplary user interfaces for inserting visual content within media.

[0058] [Figure 11] This flowchart illustrates a user interface for inserting visual content from media, according to several embodiments.

[0059] [Figure 12A] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12B] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12C] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12D] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12E] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12F] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12G] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12H] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12I] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12J] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12K] The following are exemplary user interfaces for identifying visual content within media, according to several embodiments. [Figure 12L]The following are exemplary user interfaces for identifying visual content within media, according to several embodiments.

[0060] [Figure 13] This flowchart illustrates a method for identifying visual content within media, according to several embodiments.

[0061] [Figure 14A] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14B] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14C] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14D] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14E] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14F] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14G] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14H] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14I] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14J] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14K]This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14L] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14M] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments. [Figure 14N] This section describes exemplary user interfaces for translating visual content within media, according to several embodiments.

[0062] [Figure 15] This flowchart illustrates a method for translating visual content within media, according to several embodiments.

[0063] [Figure 16A] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16B] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16C] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16D] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16E] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16F] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16G] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16H] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16I] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16J] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16K] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16L] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16M] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16N] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments. [Figure 16O] This document illustrates an exemplary user interface for managing user interface objects for visual content within a media, according to several embodiments.

[0064] [Figure 17] This flowchart illustrates a method for managing user interface objects for visual content within media, according to several embodiments. [Modes for carrying out the invention]

[0065] The following description includes exemplary methods, parameters, etc. However, it should be noted that such descriptions are not intended to limit the scope of this disclosure, but rather are provided as descriptions of exemplary embodiments.

[0066] There is a need for electronic devices that provide efficient methods and interfaces for managing visual content. For example, there is a need for electronic devices and / or computer systems that allow users to manage visual content contained in objects captured by one or more cameras of a computer system, such as signs or restaurant menus. Such technology can reduce the cognitive burden on users managing visual content, thereby increasing productivity. Furthermore, such technology can reduce the processor and battery power that would normally be wasted on redundant user input.

[0067] The following Figures 1A-1B, 2, 3, 4A-4B, and 5A-5B provide a description of exemplary devices that perform technologies for managing visual content.

[0068] Figures 6A to 6Z show exemplary user interfaces for managing visual content within a media. Figure 8 is a flowchart illustrating a method for managing visual content within a media according to several embodiments. The user interfaces in Figures 6A to 6Z are used to illustrate the processes described later, including the process in Figure 8.

[0069] Figures 7A to 7L show exemplary user interfaces for managing visual indicators for visual content within a media. Figure 9 is a flowchart illustrating a method for managing visual indicators for visual content within a media according to several embodiments. The user interfaces in Figures 7A to 7L are used to illustrate the processes described later, including the process in Figure 9.

[0070] Figures 10A to 10AD show exemplary user interfaces for inserting visual content into media. Figure 11 is a flowchart illustrating how to insert visual content into media. The user interfaces in Figures 10A to 10AD are used to illustrate the processes described later, including the process in Figure 11.

[0071] Figures 12A to 12L show exemplary user interfaces for identifying visual content within media. Figure 13 is a flowchart illustrating how to identify visual content within media. The user interfaces in Figures 12A to 12L are used to illustrate the processes described later, including the process in Figure 13.

[0072] Figures 14A to 14N show exemplary user interfaces for translating visual content within media. Figure 15 is a flowchart illustrating a method for translating visual content within media according to several embodiments. The user interfaces in Figures 14A to 14N are used to illustrate the processes described later, including the process in Figure 15.

[0073] Figures 16A to 16O show exemplary user interfaces for managing user interface objects for visual content in media, according to several embodiments. Figure 17 is a flowchart illustrating a method for managing user interface objects for visual content in media, according to several embodiments. The user interfaces in Figures 16A to 16O are used to illustrate processes described later, including the process in Figure 17.

[0074] The processes described below enhance the usability of the device and streamline the user-device interface by various technologies, including providing users with improved visual feedback, reducing the number of inputs required to perform operations, offering additional control options without cluttering the user interface with additional controls displayed, performing operations without requiring further user input when a set of conditions is met, and / or other techniques. These technologies also reduce power consumption and improve the device's battery life by enabling users to use the device more quickly and efficiently.

[0075] Furthermore, in any method described herein that is conditional on one or more conditions being met in one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in a specific order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for a claim of a system or computer-readable medium that includes instructions for performing a conditional operation based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency has been met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will also understand that, as with a method having conditional steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.

[0076] In the following description, terms such as “first,” “second,” etc., are used to describe various elements, but these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, without departing from the scope of the various embodiments described, the first touch may be called the second touch, and similarly, the second touch may be called the first touch. Both the first touch and the second touch are touches, but they are not the same touch.

[0077] The terminology used in the descriptions of the various embodiments described herein is intended solely to describe specific embodiments and is not intended to be limiting. In the descriptions of the various embodiments and the accompanying claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless otherwise explicitly stated in the context. Furthermore, it should be understood that, as used herein, the term “and / or” refers to and includes any and all possible combinations of one or more of the enumerated items relating to the description. It will be further understood that, as used herein, the terms “includes,” “including,” “comprises,” and / or “comprising,” specify the presence of the described features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0078] The phrase "if" can be interpreted, at will, depending on the context, as "when" or "upon," or "in response to determining" or "in response to detecting." Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" can be interpreted, at will, depending on the context, as "upon determining" or "in response to determining," or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."

[0079] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communication device, such as a mobile phone, which also includes other functions such as PDA functionality and / or music player functionality. Exemplary embodiments of portable multifunction devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Optionally, other portable electronic devices such as laptop computers or tablet computers having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad) are also used. It should also be understood that in some embodiments, the device is not a portable communication device but a desktop computer having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad). In some embodiments, the electronic device is a computer system communicating (e.g., via wired communication, via wireless communication) with a display-generating component. The display-generating component is configured to provide a visual output, such as a display via a CRT display, a display via an LED display, or a display via image projection. In some embodiments, the display-generating component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "display" content includes displaying content (e.g., video data rendered or decoded by the display controller 156) by transmitting data (e.g., image data or video data) via a wired or wireless connection to an integrated or external display generation component in order to visually generate the content.

[0080] The following discussion describes electronic devices including displays and touch-sensitive surfaces. However, it should be understood that electronic devices optionally include one or more other physical user interface devices such as a physical keyboard, mouse, and / or joystick.

[0081] The device typically supports a variety of applications, including drawing applications, presentation applications, word processing applications, website creation applications, disk authoring applications, spreadsheet applications, game applications, telephone applications, video conferencing applications, email applications, instant messaging applications, training support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, and / or digital video player applications.

[0082] Various applications running on this device optionally utilize at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface, as well as the corresponding information displayed on the device, are optionally adjusted and / or modified on an application-by-application basis and / or within individual applications. In this way, the device's common physical architecture (such as the touch-sensitive surface) optionally supports a variety of applications with intuitive and transparent user interfaces for the user.

[0083] Here, we turn our attention to embodiments of portable devices having a touch-sensitive display. Figure 1A is a block diagram of a portable multifunction device 100 having a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 may be conveniently referred to as a “touchscreen” and may be known or referred to as a “touch-sensitive display system”. Device 100 includes memory 102 (optionally including one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 (e.g., touch-sensitive surfaces such as the touch-sensitive display system 112 of Device 100) that detect the intensity of contact on Device 100. Device 100 optionally includes one or more tactile output generators 167 that generate tactile outputs on Device 100 (for example, on touch-sensitive surfaces such as the touch-sensitive display system 112 of Device 100 or the touchpad 355 of Device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0084] As used herein and in the claims, the term “strength” of contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of contact on the touch-sensitive surface (e.g., finger contact), or a proxy for the force or pressure of contact on the touch-sensitive surface. The strength of contact has a range of values, including at least four distinct values, and more typically, including several hundred (e.g., at least 256) distinct values. The strength of contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the touch-sensitive surface are optionally used to measure forces at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated force of contact. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or modification of the contact area detected on the touch-sensing surface, the capacitance and / or modification of the touch-sensing surface adjacent to the contact, and / or the resistance and / or modification of the touch-sensing surface adjacent to the contact may optionally be used as a substitute for the force or pressure of the contact on the touch-sensing surface. In some implementations, the substitute measurement of the contact force or pressure is used directly to determine whether it exceeds an intensity threshold (e.g., the intensity threshold is described in units corresponding to the substitute measurement). In some implementations, the substitute measurement of the contact force or pressure is converted into an estimate of the force or pressure, which is then used to determine whether it exceeds an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure).By using the intensity of contact as an attribute of user input, it becomes possible for users to access additional device functions that might otherwise be inaccessible (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical control such as a knob or button) on reduced-size devices where the implementation area for displaying affordances is limited.

[0085] As used herein and in the claims, the term “tactile output” refers to the physical displacement of a device relative to its previous position, the physical displacement of a component of a device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or the displacement of a component relative to the center of mass of a device, which will be detected by the user through the user’s sense of touch. For example, in a situation where a device or component of a device is in contact with a touch-sensitive user’s surface (e.g., the user’s fingers, palm, or other part of their hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical properties of the device or component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) may be optionally interpreted by the user as a “down-click” or “up-click” of a physical actuator button. In some cases, the user may perceive a tactile sensation such as a “down-click” or “up-click” even when there is no movement of a physical actuator button associated with a touch-sensitive surface that is physically pressed (e.g., displaced) by the user’s action. In another embodiment, movement of a touch-sensitive surface may be optionally interpreted or perceived by the user as "roughness" of that touch-sensitive surface, even if there is no change in the smoothness of the touch-sensitive surface. Such user interpretations of touch depend on the user's personal sensory perception, but there are many touch sensory perceptions common to the majority of users. Therefore, when a tactile output is described as corresponding to a user's specific sensory perception (e.g., "up-click," "down-click," "roughness"), unless otherwise stated, the generated tactile output corresponds to the physical displacement of the device or its components that produce the described sensory perception of a typical (or average) user.

[0086] Device 100 is merely an example of a portable multifunction device, and it should be understood that Device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of those components. The various components shown in Figure 1A are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application-specific integrated circuits.

[0087] Memory 102 optionally includes high-speed random-access memory and optionally includes one or more non-volatile memory such as magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0088] The peripheral interface 118 can be used to connect the device's input and output peripherals to the CPU 120 and memory 102. One or more processors 120 operate or execute various software programs and / or instruction sets stored in memory 102 to perform various functions for device 100 and process data. In some embodiments, the peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0089] The RF (radio frequency) circuit 108 transmits and receives RF signals, also known as electromagnetic signals. The RF circuit 108 converts electrical signals to electromagnetic signals or electromagnetic signals to electrical signals and communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, which include, but are not limited to, antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, CODEC chipsets, subscriber identity module (SIM) cards, and memory. The RF circuit 108 optionally communicates wirelessly with networks such as the Internet, also known as the World Wide Web (WWW), intranets, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), as well as with other devices. The RF circuit 108 optionally includes a well-known circuit for detecting a near-field communication (NFC) field using a short-range communication radio. Wireless communication is not limited to this, but optionally includes Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), and long-term evolution.Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, Email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), Instant messaging (e.g., Extensible Messaging and Presence Protocol) Using any of several communication standards, protocols, and technologies, including the XMPP protocol, the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), the Instant Messaging and Presence Service (IMPS), and / or the Short Message Service (SMS), or any other suitable communication protocol, including a communication protocol not yet developed as of the filing date of this specification.

[0090] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral interface 118, converts this audio data into an electrical signal, and transmits this electrical signal to the speaker 111. The speaker 111 converts the electrical signal into human audible sound waves. The audio circuit 110 also receives the electrical signal converted from the sound waves by the microphone 113. The audio circuit 110 converts the electrical signal into audio data and transmits this audio data to the peripheral interface 118 for processing. The audio data is optionally retrieved by the peripheral interface 118 from memory 102 and / or RF circuit 108 and / or transmitted to memory 102 and / or RF circuit 108. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., 212 in Figure 2). The headset jack provides an interface between the audio circuit 110 and detachable audio input / output peripherals such as output-only headphones or headsets that have both output (e.g., headphones for one or both ears) and input (e.g., a microphone).

[0091] The I / O subsystem 106 connects input / output peripherals on device 100, such as the touchscreen 112 and other input control devices 116, to the peripheral interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. One or more input controllers 160 receive electrical signals from / transmit electrical signals to other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons), dials, slider switches, joysticks, click wheels, etc. In some embodiments, one or more input controllers 160 are optionally connected to (or not connected to) one of the following: a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. One or more buttons (e.g., 208 in Figure 2) optionally include up / down buttons for volume control of speaker 111 and / or microphone 113. One or more buttons optionally include push buttons (e.g., 206 in Figure 2). In some embodiments, the electronic device is a computer system communicating with one or more input devices (e.g., via wireless communication over wired communication). In some embodiments, one or more input devices include a touch-sensitive surface (e.g., a trackpad as part of a touch-sensitive display). In some embodiments, one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking user gestures (e.g., hand gestures) as input. In some embodiments, one or more input devices are integrated with the computer system. In some embodiments, one or more input devices are separate from the computer system.

[0092] As described in U.S. Patent Application No. 11 / 322,549, “Unlocking a Device by Performing Gestures on an Unlock Image,” filed December 23, 2005, U.S. Patent No. 7,657,849, which is incorporated herein by reference in its entirety, a quick press of a push button optionally releases the lock on the touchscreen 112, or optionally initiates a process to unlock the device using gestures on the touchscreen. A longer press of a push button (e.g., 206) optionally turns power on or off the device 100. The functionality of one or more of the buttons is optionally customizable by the user. The touchscreen 112 is used to implement virtual or soft buttons and one or more soft keyboards.

[0093] The touch-sensitive display 112 provides input and output interfaces between the device and the user. The display controller 156 receives electrical signals from and / or transmits electrical signals to the touchscreen 112. The touchscreen 112 displays visual output to the user. This visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively, “graphics”). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0094] The touchscreen 112 has a touch-sensing surface, sensor, or set of sensors that accept user input based on touch and / or tactile contact. The touchscreen 112 and the display controller 156 (along with any associated modules and / or instruction sets in memory 102) detect contact (and any movement or interruption of contact) on the touchscreen 112 and translate the detected contact into interaction with user interface objects displayed on the touchscreen 112 (e.g., one or more soft keys, icons, web pages, or images). In an exemplary embodiment, the point of contact between the touchscreen 112 and the user corresponds to the user's finger.

[0095] The touchscreen 112 optionally uses LCD (liquid crystal display) technology, LPD (polymer light-emitting display) technology, or LED (light-emitting diode) technology, but other display technologies may also be used in other embodiments. The touchscreen 112 and the display controller 156 optionally, but not limited to, use any of several currently known or future-developed touch sensing technologies, including capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more points of contact with the touchscreen 112, to detect contact and any movement or interruption of it. In exemplary embodiments, projected mutual capacitive sensing technology is used, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California.

[0096] The touch-sensitive displays in some embodiments of the touchscreen 112 are optionally similar to the multi-touch-sensitive touchpads described in U.S. Patent No. 6,323,846 (Westerman et al.), No. 6,570,557 (Westerman et al.), and / or No. 6,677,932 (Westerman), and / or U.S. Patent Application Publication 2002 / 0015024(A1), which are each incorporated herein in whole by reference. However, the touchscreen 112 displays visual output from device 100, whereas the touch-sensitive touchpad does not provide visual output.

[0097] The touch-sensitive displays in some embodiments of the touchscreen 112 are described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, "Multipoint Touch Surface Controller"; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, "Multipoint Touchscreen"; (3) U.S. Patent Application No. 10 / 903,964, filed July 30, 2004, "Gestures For Touch Sensitive Input Devices"; (4) U.S. Patent Application No. 11 / 048,264, filed January 31, 2005, "Gestures For Touch Sensitive Input Devices"; (5) U.S. Patent Application No. 11 / 038,590, filed January 18, 2005, "Mode-Based Graphical User Interfaces For Touch Sensitive Input These are described in (6) U.S. Patent Application No. 11 / 228,758, filed September 16, 2005, "Virtual Input Device Placement On A Touch Screen User Interface", (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, "Operation Of A Computer With A Touch Screen Interface", (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard", and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, "Multi-Functional Hand-Held Device". All of these applications are incorporated herein by reference in their entirety.

[0098] The touchscreen 112 optionally has a video resolution greater than 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. The user optionally touches the touchscreen 112 using any suitable object or attachment such as a stylus or finger. In some embodiments, the user interface is designed to operate primarily using finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of ​​the finger on the touchscreen. In some embodiments, the device translates coarse finger input into a precise pointer / cursor position or command to perform an action desired by the user.

[0099] In some embodiments, in addition to the touchscreen, the device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display a visual output. The touchpad is optionally a touch-sensitive surface separate from the touchscreen 112 or an extension of the touch-sensitive surface formed by the touchscreen.

[0100] Device 100 also includes a power system 162 that supplies power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)), and any other components associated with generating, managing, and distributing power within the portable device.

[0101] The device 100 also optionally includes one or more optical sensors 164. Figure 1A shows an optical sensor coupled to an optical sensor controller 158 in the I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light from the environment projected through one or more lenses and converts that light into data representing an image. The optical sensor 164 works in conjunction with an imaging module 143 (also called a camera module) to optionally capture still images or video. In some embodiments, the optical sensor is located on the back of the device 100 opposite the touchscreen display 112 on the front of the device, so that the touchscreen display can be used as a viewfinder for acquiring still images and / or video. In some embodiments, the optical sensor is located on the front of the device so that the user's image is optionally acquired for video conferencing while the user is viewing other video conference participants on the touchscreen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (for example, by rotating the lens and sensor within the device housing), so that a single optical sensor 164 can be used for both video conferencing and acquiring still images and / or videos, together with the touchscreen display.

[0102] Device 100 also optionally includes one or more depth camera sensors 175. Figure 1A shows a depth camera sensor coupled to a depth camera controller 169 in the I / O subsystem 106. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of an object in the scene (e.g., a face) from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with an imaging module 143 (also called a camera module), the depth camera sensor 175 is optionally used to determine the depth map of different parts of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is positioned on the front of Device 100 to optionally acquire an image of the user with depth information for video conferencing while the user is viewing other video conference participants on a touchscreen display, and also to capture a selfie image with depth map data. In some embodiments, the depth camera sensor 175 is positioned on the back of the device, or on both the back and front of Device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (for example, by rotating the lens and sensor within the device housing), so that the depth camera sensor 175, together with the touchscreen display, can be used for both video conferencing and the acquisition of still images and / or videos.

[0103] Device 100 also optionally includes one or more contact intensity sensors 165. Figure 1A shows a contact intensity sensor coupled to an intensity sensor controller 159 in the I / O subsystem 106. The contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, pressure-power sensors, optical force sensors, capacitive touch-sensing surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of contact on a touch-sensing surface). The contact intensity sensor 165 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is positioned juxtaposed with or adjacent to a touch-sensing surface (e.g., a touch-sensing display system 112). In some embodiments, at least one contact intensity sensor is positioned on the back of Device 100, opposite the touchscreen display 112 located on the front of Device 100.

[0104] The device 100 also optionally includes one or more proximity sensors 166. Figure 1A shows a proximity sensor 166 coupled to a peripheral interface 118. Alternatively, the proximity sensor 166 is optionally coupled to an input controller 160 in the I / O subsystem 106. The proximity sensor 166 optionally functions as described in U.S. Patent Applications 11 / 241,839, “Proximity Detector In Handheld Device,” 11 / 240,788, “Proximity Detector In Handheld Device,” 11 / 620,702, “Using Ambient Light Sensor To Augment Proximity Sensor Output,” 11 / 586,862, “Automated Response To And Sensing Of User Activity In Portable Devices,” and 11 / 638,251, “Methods And Systems For Automatic Configuration Of Peripherals,” which are all incorporated herein by reference. In some embodiments, if the multifunction device is placed near the user's ear (for example, when the user is making a phone call), the proximity sensor turns off and disables the touchscreen 112.

[0105] Device 100 also optionally includes one or more tactile output generators 167. Figure 1A shows a tactile output generator coupled to a tactile feedback controller 161 in the I / O subsystem 106. The tactile output generator 167 optionally includes one or more electroacoustic devices such as a speaker or other audio component, and / or electromechanical devices that convert energy into linear motion, such as a motor, solenoid, electroactive polymer, piezoelectric actuator, electrostatic actuator, or other tactile output generating component (e.g., a component that converts an electrical signal into a tactile output on the device). The contact intensity sensor 165 receives a tactile feedback generation command from the tactile feedback module 133 and generates a tactile output on device 100 that can be sensed by the user of device 100. In some embodiments, at least one tactile output generator is positioned alongside or adjacent to a touch-sensing surface (e.g., a touch-sensing display system 112) and optionally generates a tactile output by moving the touch-sensing surface vertically (e.g., inward / outward from the surface of device 100) or horizontally (e.g., forward / backward in the same plane as the surface of device 100). In some embodiments, at least one tactile output generator sensor is positioned on the back of device 100, opposite the touchscreen display 112 which is positioned on the front of device 100.

[0106] The device 100 also optionally includes one or more accelerometers 168. Figure 1A shows an accelerometer 168 coupled to a peripheral interface 118. Alternatively, the accelerometer 168 is optionally coupled to an input controller 160 in the I / O subsystem 106. The accelerometer 168 optionally functions as described in U.S. Patent Application Publication 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and U.S. Patent Application Publication 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed on a touchscreen display in portrait or landscape orientation based on an analysis of data received from one or more accelerometers. Device 100 optionally includes, in addition to one or more accelerometers 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for acquiring information about the location and orientation of Device 100 (e.g., longitudinal or transverse).

[0107] In some embodiments, the software components stored in memory 102 include an operating system 126, a communications module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application (or instruction set) 136. Furthermore, in some embodiments, memory 102 (Figure 1A) or 370 (Figure 3) stores a device / global internal state 157, as shown in Figures 1A and 3. The device / global internal state 157 includes one or more of the following: an active application state indicating which application is active, if there is an application currently active; a display state indicating which applications, views, or other information occupy different areas of the touchscreen display 112; a sensor state including information obtained from various sensors and input control devices 116 of the device; and location information relating to the device's location and / or orientation.

[0108] An operating system 126 (for example, an embedded operating system such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware components and software components.

[0109] The communication module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by the RF circuit 108 and / or external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to connect to other devices directly or indirectly via a network (e.g., the Internet, Wi-Fi, etc.). In some embodiments, the external ports are multi-pin (e.g., 30-pin) connectors that are the same as and / or compatible with the 30-pin connector used on iPod® (a trademark of Apple Inc.) devices.

[0110] The contact / motion module 130 optionally detects contact with the touchscreen 112 and other touch-sensitive devices (e.g., a touchpad or physical click wheel) (in cooperation with the display controller 156). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a substitute for the force or pressure of the contact), determining whether there is movement of contact and tracking movement across the touch-sensitive surface (e.g., detecting one or more events of a finger dragging), and determining whether contact has been terminated (e.g., detecting a finger up event or interruption of contact). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point, represented by a series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These actions can be optionally applied to a single contact (e.g., a single finger contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.

[0111] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an action has been performed by a user (for example, to determine whether a user has "clicked" on an icon). In some embodiments, at least a subset of the intensity thresholds is determined according to software parameters (for example, the intensity thresholds can be adjusted without modifying the physical hardware of device 100, rather than being determined by the activation threshold of a particular physical actuator). For example, the mouse "click" threshold for a trackpad or touchscreen display can be set to one of a range of predefined thresholds without modifying the trackpad or touchscreen display hardware. In addition, in some implementations, the user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (for example, by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using a system-level click "intensity" parameter).

[0112] The contact / motion module 130 optionally detects gesture input from the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motion, timing, and / or intensity of detected contacts). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture involves detecting a finger down event, followed by a finger up (lift-off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface involves detecting a finger down event, followed by one or more finger drag events, and then a finger up (lift-off) event.

[0113] The graphics module 132 includes various known software components for rendering and displaying graphics on the touchscreen 112 or other display, including components that modify the visual effects of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). In this specification, the term “graphics” includes, but is not limited to, any object that can be displayed to the user, including characters, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0114] In some embodiments, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes from an application or the like, as needed, specifying the graphics to be displayed, along with coordinate data and other graphic property data, and then generates screen image data to output to the display controller 156.

[0115] The haptic feedback module 133 includes various software components for generating commands used by a tactile output generator(s) 167, and generates tactile outputs at one or more locations on the device 100 in response to the user's interaction with the device 100.

[0116] The text input module 134 is optionally a component of the graphics module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other applications that require text input).

[0117] The GPS module 135 determines the device's location and provides this information for use in various applications (for example, to the phone 138 for use in location-based dialing, to the camera 143 as picture / video metadata, and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0118] Application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof: ● Contact module 137 (sometimes called the address book or contact list), ●Telephone module 138, ●Video conferencing module 139, ● Email client module 140, ● Instant messaging (IM) module 141, ●Training support module 142, ● Camera module 143 for still images and / or video, ●Image management module 144, ●Video player module, ● Music player module, ● Browser module 147, ●Calendar module 148, ●Optionally, a widget module 149 may include one or more of the following: weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6. ●Widget creator module 150 for creating user-created widget 149-6, ● Search module 151, ● A video and music player module 152 that integrates a video player module and a music player module. ●Memo Module 153, ●Map module 154, and / or, ● Online video module 155.

[0119] Examples of other applications 136 that may be optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, Java-enabled applications, encryption, digital rights management, speech recognition, and speech duplication.

[0120] Together with the touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the contact module 137 is optionally used to manage an address book or contact list (stored, for example, in the application internal state 192 of the contact module 137 in memory 102 or memory 370), which includes adding names(s) to the address book, removing names(s) from the address book, associating telephone numbers(s) to names, email addresses(s) to names, addresses(s) to names, or other information, associating images to names, categorizing and sorting names, and providing telephone numbers or email addresses to initiate and / or facilitate communication via telephone 138, video conferencing module 139, email 140, or IM 141.

[0121] The telephone module 138 works in conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134 to optionally input a series of characters corresponding to a telephone number, access one or more telephone numbers in the contact module 137, modify an entered telephone number, dial individual telephone numbers, make a call, and disconnect and terminate a call at the end of the call. As previously mentioned, wireless communication may optionally use any of several communication standards, protocols, and technologies.

[0122] The video conferencing module 139 works in conjunction with the RF circuit 108, audio circuit 110, speaker 111, microphone 113, touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contact module 137, and telephone module 138 to include executable commands for starting, running, and ending video conferences between the user and one or more other participants in accordance with the user's commands.

[0123] The email client module 140, in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, includes executable commands for creating, sending, receiving, and managing emails in response to user commands. In conjunction with the image management module 144, the email client module 140 makes it extremely easy to create and send emails containing still or video images captured by the camera module 143.

[0124] The instant messaging module 141, in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, includes executable commands for inputting a series of characters corresponding to an instant message, modifying previously entered characters, sending individual instant messages (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for telephone-based instant messaging, or XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the instant messages sent and / or received optionally include graphics, photographs, audio files, video files, and / or other attachments, such as those supported by MMS and / or Enhanced Messaging Service (EMS). In this specification, “instant messaging” refers to both telephone-based messaging (e.g., messages sent using SMS or MMS) and internet-based messaging (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0125] In conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, the training support module 142 includes executable commands, which create training (e.g., having time, distance, and / or calorie burn goals), communicate with training sensors (sports devices), receive training sensor data, calibrate sensors used to monitor training, select and play music for training, and display, store, and transmit training data.

[0126] The camera module 143 works in conjunction with the touchscreen 112, display controller 156, optical sensor(s) 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144 to include executable commands for capturing still images or videos (including video streams) and storing them in memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from memory 102.

[0127] The image management module 144 works in conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143 to include executable commands for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or videos.

[0128] The browser module 147 works in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134 to include executable commands for browsing the internet according to user commands, including searching for, linking to, receiving, and displaying web pages or parts thereof, as well as attachments and other files linked to web pages.

[0129] The calendar module 148 works in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147 to include executable commands for creating, displaying, modifying, and storing a calendar and data associated with the calendar (e.g., calendar items, to-do lists, etc.) according to user commands.

[0130] The widget module 149 works in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147 to optionally download and use mini-applications (e.g., weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or mini-applications created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).

[0131] The widget creator module 150 works in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147 to be used by the user to optionally create widgets (for example, to turn a user-specified portion of a web page into a widget).

[0132] The search module 151 works in conjunction with the touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134 to include executable commands for searching for characters, music, sounds, images, videos, and / or other files in memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) according to user commands.

[0133] The video and music player module 152 works in conjunction with the touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, and browser module 147 to include executable commands that allow the user to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable commands for displaying, presenting, or otherwise playing videos (for example, on the touchscreen 112 or on an external display connected via the external port 124). In some embodiments, the device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0134] The memo module 153 works in conjunction with the touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134 to include executable commands for creating and managing memos, to-do lists, etc., according to user commands.

[0135] The map module 154 works in conjunction with the RF circuit 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147 to optionally receive, display, modify, and store maps and map-related data (e.g., driving directions, data on shops and other points of interest in or near a specific location, and other location-based data) in accordance with user commands.

[0136] The online video module 155, in conjunction with the touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, email client module 140, and browser module 147, includes instructions that enable the user to access, browse, receive (e.g., by streaming and / or downloading), play (e.g., on the touchscreen or on an external display connected via external port 124), send emails with links to specific online videos, and perform other management of online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 141 is used instead of the email client module 140 to send links to specific online videos. For further information regarding online video applications, please refer to U.S. Provisional Patent Application No. 60 / 936,562, “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” filed June 20, 2007, and U.S. Patent Application No. 11 / 968,067, “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” filed December 31, 2007, the entire contents of which are incorporated herein by reference.

[0137] Each of the modules and applications identified above corresponds to a set of executable instructions that perform one or more of the functions described above and the methods described in this application (e.g., the computer-based methods and other information processing methods described herein). These modules (e.g., instruction sets) do not need to be implemented as separate software programs, procedures, or modules; therefore, in various embodiments, various subsets of these modules can be optionally combined or otherwise reconfigured. For example, a video player module can optionally be combined with a music player module to form a single module (e.g., the video and music player module 152 in Figure 1A). In some embodiments, memory 102 optionally stores a subset of the modules and data structures identified above. Furthermore, memory 102 optionally stores additional modules and data structures not described above.

[0138] In some embodiments, device 100 is a device in which the operation of a default set of functions in that device is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for device 100 to operate, the number of physical input control devices (push buttons, dials, etc.) on device 100 is optionally reduced.

[0139] A default set of functions, performed only through the touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by the user, navigates the device 100 from any user interface displayed on the device 100 to a main menu, home menu, or root menu. In such embodiments, a “menu button” is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touchpad.

[0140] Figure 1B is a block diagram showing exemplary components for event processing according to several embodiments. In some embodiments, memory 102 (Figure 1A) or 370 (Figure 3) includes an event sorting unit 170 (e.g., within the operating system 126) and individual applications 136-1 (e.g., any of the aforementioned applications 137-151, 155, 380-390).

[0141] The event sorting unit 170 receives event information and determines the application 136-1 that distributes the event information, and the application view 191 of application 136-1. The event sorting unit 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, application 136-1 includes an application internal state 192 that indicates the current application view(s) displayed on the touch-sensitive display 112 when the application is active or running. In some embodiments, a device / global internal state 157 is used by the event sorting unit 170 to determine which application(s) are currently active, and the application internal state 192 is used by the event sorting unit 170 to determine the application view(s) to which the event information is distributed.

[0142] In some embodiments, the application internal state 192 includes additional information such as resume information to be used when the application 136-1 resumes execution, user interface state information that indicates or is ready to display information displayed by the application 136-1, a state queue that allows the user to return to a previous state or view of the application 136-1, and one or more redo / undo queues of previous actions performed by the user.

[0143] The event monitor 171 receives event information from the peripheral interface 118. The event information includes information about sub-events (for example, user touch as part of a multi-touch gesture on the touch-sensitive display 112). The peripheral interface 118 transmits information received from the I / O subsystem 106, or from sensors such as the proximity sensor 166, one or more accelerometers 168, and / or the microphone 113 (via the audio circuit 110). The information received by the peripheral interface 118 from the I / O subsystem 106 includes information from the touch-sensitive display 112 or the touch-sensitive surface.

[0144] In some embodiments, the event monitor 171 sends requests to the peripheral interface 118 at predetermined intervals. In response, the peripheral interface 118 transmits event information. In other embodiments, the peripheral interface 118 transmits event information only when there is a significant event (e.g., reception of input exceeding a predetermined noise threshold and / or exceeding a predetermined duration).

[0145] In some embodiments, the event sorting unit 170 also includes a hit view determination module 172 and / or an active event recognition determination module 173.

[0146] The hit view determination module 172 provides a software procedure for determining where in one or more views a sub-event occurred when the touch-sensitive display 112 is displaying two or more views. A view consists of control devices and other elements that the user can see on the display.

[0147] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of an individual application) in which a touch is detected optionally corresponds to a program level within the application's program hierarchy or view hierarchy. For example, the lowest-level view in which a touch is detected optionally refers to a hit view, and the set of events recognized as appropriate input is optionally determined at least in part based on the hit view of the initial touch that initiates a touch gesture.

[0148] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views arranged in a hierarchy, the hit view determination module 172 identifies the hit view as the lowest-level view in the hierarchy from which sub-events should be processed. In most situations, the hit view is the lowest-level view from which the initiating sub-event (e.g., the first sub-event in a series of sub-events that form an event or potential event) occurs. Once a hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source that identified it as the hit view.

[0149] The active event recognition determination module 173 determines which view(s) in the view hierarchy should receive a particular set of sub-events. In some embodiments, the active event recognition determination module 173 determines that only the hit view should receive a particular set of sub-events. In other embodiments, the active event recognition determination module 173 determines that all views, including the physical location of the sub-events, are actively involved views, and therefore all actively involved views should receive a particular set of sub-events. In other embodiments, even if the touch sub-event is entirely confined to an area associated with one particular view, higher-level views in the hierarchy still remain actively involved views.

[0150] The event dispatcher module 174 dispatches event information to an event recognition unit (e.g., an event recognition unit 180). In embodiments including an active event recognition unit determination module 173, the event dispatcher module 174 distributes the event information to the event recognition unit determined by the active event recognition unit determination module 173. In some embodiments, the event dispatcher module 174 stores event information retrieved by individual event receiving units 182 in an event queue.

[0151] In some embodiments, the operating system 126 includes an event sorting unit 170. Alternatively, application 136-1 includes an event sorting unit 170. In yet another embodiment, the event sorting unit 170 is a standalone module or part of another module stored in memory 102, such as a contact / motion module 130.

[0152] In some embodiments, application 136-1 includes a plurality of event processing units 190 and one or more application views 191, each containing instructions for handling touch events occurring within a separate view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognition units 180. Typically, a separate application view 191 includes a plurality of event recognition units 180. In other embodiments, one or more of the event recognition units 180 are part of a separate module, such as a user interface kit or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, a separate event processing unit 190 includes one or more event data 179 received from a data update unit 176, an object update unit 177, a GUI update unit 178, and / or an event sorting unit 170. The event processing unit 190 optionally uses or calls the data update unit 176, the object update unit 177, or the GUI update unit 178 to update the application's internal state 192. Alternatively, one or more of the application views 191 include one or more event processing units 190. In some embodiments, one or more of the data update unit 176, object update unit 177, and GUI update unit 178 are included in individual application views 191.

[0153] Each individual event recognition unit 180 receives event information (e.g., event data 179) from the event sorting unit 170 and identifies events from the event information. The event recognition unit 180 includes an event receiving unit 182 and an event comparison unit 184. In some embodiments, the event recognition unit 180 also includes at least a subset of metadata 183 and event distribution commands 188 (optionally including sub-event distribution commands).

[0154] The event receiving unit 182 receives event information from the event sorting unit 170. The event information includes information about sub-events, such as touches or the movement of touches. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, an event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the device's current orientation (also called the device's orientation).

[0155] The event comparison unit 184 compares event information with a predefined definition of an event or sub-event, and based on the comparison, determines an event or sub-event, or determines or updates the state of an event or sub-event. In some embodiments, the event comparison unit 184 includes an event definition 186. The event definition 186 includes definitions of events (e.g., a default set of sub-events), such as event 1 (187-1) and event 2 (187-2). In some embodiments, sub-events within event (187) include, for example, touch start, touch end, touch movement, touch cancellation, and multiple touches. In one embodiment, the definition for event 1 (187-1) is a double tap on a displayed object. A double tap includes, for example, a first touch on the displayed object for a predetermined stage (touch start), a first lift-off for the predetermined stage (touch end), a second touch on the displayed object for the predetermined stage (touch start), and a second lift-off for the predetermined stage (touch end). In another embodiment, event 2(187-2) is defined as a drag on a displayed object. A drag includes, for example, a touch (or contact) on the displayed object to a predetermined stage, movement of the touch across the touch-sensitive display 112, and lift-off of the touch (end of touch). In some embodiments, the event also includes information about one or more associated event processing units 190.

[0156] In some embodiments, the event definition 187 includes event definitions for individual user interface objects. In some embodiments, the event comparison unit 184 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view where three user interface objects are displayed on the touch-sensitive display 112, when a touch is detected on the touch-sensitive display 112, the event comparison unit 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with an individual event processing unit 190, the event comparison unit uses the results of the hit test to determine which event processing unit 190 should be activated. For example, the event comparison unit 184 selects the sub-event and the event processing unit associated with the object that triggers the hit test.

[0157] In some embodiments, the definition of an individual event 187 also includes a delay action that delays the delivery of event information until it is determined whether a series of sub-events correspond to an event type in the event recognition unit.

[0158] If an individual event recognition unit 180 determines that a series of sub-events does not match any of the events in the event definition 186, the individual event recognition unit 180 enters an event impossible, event failed, or event terminated state and thereafter ignores subsequent sub-events of the touch-based gesture. In this situation, if there are other event recognition units that remain active for the hit view, those event recognition units continue to track and process the sub-events of the ongoing touch-based gesture.

[0159] In some embodiments, an individual event recognition unit 180 includes metadata 183 having configurable properties, flags, and / or lists that indicate to the actively involved event recognition unit how the event distribution system should perform sub-event distribution. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists that indicate how the event recognition units interact with each other, or how they can interact with each other. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists that indicate how sub-events are distributed to various levels in the view hierarchy or program hierarchy.

[0160] In some embodiments, an individual event recognition unit 180 activates an event processing unit 190 associated with an event when one or more specific sub-events of an event are recognized. In some embodiments, the individual event recognition unit 180 delivers event information associated with the event to the event processing unit 190. Activating the event processing unit 190 is separate from sending (and delaying the sending of) sub-events to individual hit views. In some embodiments, the event recognition unit 180 sets a flag associated with the recognized event, and the event processing unit 190 associated with that flag captures the flag and executes a default process.

[0161] In some embodiments, the event distribution command 188 includes a sub-event distribution command that distributes event information about a sub-event without activating an event processing unit. Instead, the sub-event distribution command distributes event information to an event processing unit associated with a set of sub-events, or to an actively involved view. The event processing unit associated with the set of sub-events or the actively involved view receives the event information and executes a predetermined process.

[0162] In some embodiments, the data update unit 176 creates and updates data used in application 136-1. For example, the data update unit 176 updates telephone numbers used in contact module 137 or stores video files used in video player module. In some embodiments, the object update unit 177 creates and updates objects used in application 136-1. For example, the object update unit 177 creates new user interface objects or updates the positions of user interface objects. The GUI update unit 178 updates the GUI. For example, the GUI update unit 178 prepares display information and sends it to graphics module 132 for display on touch-sensitive display.

[0163] In some embodiments, the event processing unit(s) 190 includes or has access to a data update unit 176, an object update unit 177, and a GUI update unit 178. In some embodiments, the data update unit 176, the object update unit 177, and the GUI update unit 178 are contained in a single module of an individual application 136-1 or application view 191. In other embodiments, they are contained in two or more software modules.

[0164] The foregoing description regarding the handling of user touch events on a touch-sensitive display also applies to other forms of user input for operating the multifunction device 100 using input devices, but it should be understood that not all of these begin on the touchscreen. For example, mouse movement and mouse button presses, touch movements such as taps, drags, and scrolls on a touchpad, pen stylus input, device movement, verbal commands, detected eye movements, biometric input, and / or any combination thereof may be optionally used as inputs corresponding to sub-events that define the events to be recognized.

[0165] Figure 2 shows a portable multifunction device 100 having a touchscreen 112 according to several embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment, and in other embodiments described below, the user can select one or more of the graphics by performing gestures on the graphics using, for example, one or more fingers 202 (not shown in the figure to an exact scale) or one or more styluses 203 (not shown in the figure to an exact scale). In some embodiments, the selection of one or more graphics is performed when the user interrupts contact with that one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, right to left, upward and / or downward) and / or rolling (from right to left, left to right, upward and / or downward) with a finger in contact with the device 100. In some implementations or situations, accidental contact with a graphic does not constitute a selection of that graphic. For example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon does not arbitrarily select the corresponding application.

[0166] Device 100 also optionally includes one or more physical buttons, such as a "Home" button or a menu button 204. As previously mentioned, the menu button 204 is optionally used to navigate to any application 136 within the set of applications running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touchscreen 112.

[0167] In some embodiments, device 100 includes a touchscreen 112, a menu button 204, a push button 206 for turning the device on / off and locking the device, one or more volume buttons 208, a subscriber identification module (SIM) card slot 210, a headset jack 212, and an external port 124 for docking / charging. The push button 206 is optionally used to turn the device on / off by pressing down and holding the button down for a predetermined period of time, to lock the device by pressing down and releasing the button before a predetermined period of time has elapsed, and / or to unlock the device or initiate an unlocking process. In alternative embodiments, device 100 also accepts verbal input via a microphone 113 to activate or deactivate certain functions. Device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of contact on the touchscreen 112, and / or one or more tactile output generators 167 for generating tactile output to the user of device 100.

[0168] Figure 3 is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface according to several embodiments. The device 300 does not have to be portable. In some embodiments, the device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a game system, or a control device (e.g., a home or commercial controller). The device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 that interconnect these components. The communication buses 320 optionally include circuitry (sometimes called a chipset) that interconnects and controls communication between system components. The device 300 includes an input / output (I / O) interface 330 including a display 340, the display 340 is typically a touchscreen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 that generates tactile output on device 300 (similar to, for example, the tactile output generator(s) 167 described above with reference to Figure 1A), and a sensor 359 (e.g., light, acceleration, proximity, touch sensing, and / or a contact intensity sensor similar to the contact intensity sensor(s) 165 described above with reference to Figure 1A). The memory 370 includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 370 optionally includes one or more storage devices located remotely from the CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or subsets thereof, that are stored in memory 102 of the portable multifunction device 100 (Figure 1A). Furthermore, memory 370 optionally stores additional programs, modules, and data structures that are not present in memory 102 of the portable multifunction device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, ​​a word processing module 384, a website creation module 386, a disk authoring module 388, and / or a spreadsheet module 390, whereas memory 102 of the portable multifunction device 100 (Figure 1A) optionally does not store these modules.

[0169] Each of the elements identified above in Figure 3 is optionally stored in one or more of the memory devices described above. Each of the modules identified above corresponds to an instruction set that performs the function described above. The modules or programs (e.g., instruction sets) identified above do not need to be implemented as separate software programs, procedures, or modules, and therefore in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. In some embodiments, memory 370 optionally stores a subset of the modules and data structures identified above. Furthermore, memory 370 optionally stores additional modules and data structures not described above.

[0170] Next, we optionally turn our attention to an embodiment of a user interface implemented in, for example, a portable multi-functional device 100.

[0171] Figure 4A shows an exemplary user interface for an application menu on a portable multifunction device 100 according to several embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, the user interface 400 includes the following elements, or subsets or supersets thereof. ● Signal strength indicators (single or multiple) for wireless communication (single or multiple) such as cellular signals and Wi-Fi signals 402, ●Time 404, ●Bluetooth indicator 405, ●Battery status indicator 406, ●Tray 408 containing icons for frequently used applications, as shown below. ○Optionally including an indicator 414 for the number of missed calls or voicemail messages, an icon 416 of the phone module 138 labeled "Phone", ○Optionally including an indicator 410 for the number of unread emails, an icon 418 of the email client module 140 labeled "Mail", ○ Icon 420 of browser module 147, labeled "Browser", and ○ Icon 422 for the video and music player module 152, also known as the iPod (trademark of Apple Inc.) module 152, which is labeled "iPod", and ● Icons of other applications, such as the following: ○ Icon 424 of IM module 141, labeled "Message", ○ Icon 426 of calendar module 148, labeled "Calendar" ○ Icon 428 of image management module 144, labeled "Photo" ○ Icon 430 of camera module 143, labeled "Camera" ○ Icon 432 of online video module 155, labeled "online video", ○ Icon 434 of stock price widget 149-2, labeled "Stock Price" ○ Icon 436 of map module 154, labeled "Map" ○ Icon 438 of weather widget 149-1, labeled "Weather" ○ Icon 440 of the alarm clock widget 149-4, labeled "Clock" ○ Icon 442 of training support module 142, labeled "Training Support" ○ Icon 444 of memo module 153, labeled as "Memo", and ○ An icon 446 labeled "Settings," which provides access to the settings of the device 100 and its various applications 136, for a settings application or module.

[0172] Please note that the icon labels shown in Figure 4A are for illustrative purposes only. For example, the icon 422 for the video and music player module 152 is labeled "Music" or "Music Player". Other labels are used optionally for various application icons. In some embodiments, the label for an individual application icon includes the name of the application to which that individual application icon corresponds. In some embodiments, the label for a particular application icon is different from the name of the application to which that particular application icon corresponds.

[0173] Figure 4B shows an exemplary user interface on a device (e.g., device 300 in Figure 3) having a touch-sensitive surface 451 (e.g., tablet or touchpad 355 in Figure 3) separate from the display 450 (e.g., touchscreen display 112). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting the intensity of contact on the touch-sensitive surface 451, and / or one or more tactile output generators 357 for generating tactile output to the user of device 300.

[0174] Some of the following embodiments are given by reference to input on a touchscreen display 112 (a combination of a touch-sensing surface and a display), but in some embodiments, the device detects input on a touch-sensing surface separate from the display, as shown in Figure 4B. In some embodiments, the touch-sensing surface (e.g., 451 in Figure 4B) has a primary axis (e.g., 452 in Figure 4B) corresponding to a primary axis (e.g., 453 in Figure 4B) on the display (e.g., 450). According to these embodiments, the device detects contact with the touch-sensing surface 451 (e.g., 460 and 462 in Figure 4B) at locations corresponding to each location on the display (e.g., 460 corresponds to 468 and 462 corresponds to 470 in Figure 4B). In this way, user input (e.g., touches 460 and 462, and their movement) detected by the device on a touch-sensitive surface (e.g., 451 in Figure 4B) is used by the device to operate the user interface on the display of the multifunction device (e.g., 450 in Figure 4B) when the touch-sensitive surface is separate from the display. It should be understood that a similar method may be optionally used for other user interfaces described herein.

[0175] In addition, while the following examples are given primarily with reference to finger input (e.g., finger touch, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of touch), followed by a mouse click with cursor movement along the swipe path (e.g., instead of touch movement). As another example, a tap gesture may optionally be replaced by a mouse click (e.g., instead of touch detection and subsequent cessation of touch detection) while the cursor is located over the tap gesture location. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.

[0176] Figure 5A shows an exemplary personal electronic device 500. Device 500 includes a body 502. In some embodiments, device 500 may include some or all of the features described with respect to devices 100 and 300 (e.g., Figures 1A to 4B). In some embodiments, device 500 has a touch-sensitive display screen 504, hereafter referred to as touchscreen 504. Alternatively, in addition to touchscreen 504, device 500 may have a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, the touchscreen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors that detect the intensity of the applied contact (e.g., touch). One or more intensity sensors on the touchscreen 504 (or touch-sensitive surface) may provide output data representing the intensity of the touch. The user interface of device 500 may respond to touches based on their intensity, meaning that touches of different intensity may invoke different user interface behaviors on device 500.

[0177] For example, see, for instance, International Patent Application PCT / US2013 / 040061, “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” filed 8 May 2013, published as International Publication WO / 2013 / 169849, and International Patent Application PCT / US2013 / 069483, “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” filed 11 November 2013, published as International Publication WO / 2014 / 105276, which are all incorporated herein by reference.

[0178] In some embodiments, the device 500 has one or more input mechanisms 506 and 508. The input mechanisms 506 and 508 may be physical, if included. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, the device 500 has one or more attachment mechanisms. Such attachment mechanisms, if included, may allow the device 500 to be attached to, for example, hats, eyeglasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear the device 500.

[0179] Figure 5B shows an exemplary personal electronic device 500. In some embodiments, the device 500 may include some or all of the components described with respect to Figures 1A, 1B, and 3. The device 500 has a bus 512 that operably connects an I / O section 514 to one or more computer processors 516 and memory 518. The I / O section 514 may be connected to a display 504, which may have a touch-sensing component 522 and optionally an intensity sensor 524 (e.g., a contact intensity sensor). In addition, the I / O section 514 may be connected to a communication unit 530 that receives application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication technologies. The device 500 may include input mechanisms 506 and / or 508. The input mechanism 506 may optionally be, for example, a rotatable input device or a pressable and rotatable input device. In some embodiments, the input mechanism 508 is optionally a button.

[0180] In some embodiments, the input mechanism 508 is optionally a microphone. The personal electronic device 500 optionally includes a variety of sensors such as a GPS sensor 532, an accelerometer 534, a direction sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or a combination thereof, all of which can be operably connected to the I / O section 514.

[0181] The memory 518 of the personal electronic device 500 may include one or more non-temporary computer-readable storage media for storing computer-executable instructions, which, when executed by one or more computer processors 516, can cause the computer processors to perform, for example, the techniques described below, including processes 800, 900, 1100, 1300, 1500, and 1700. The computer-readable storage media may be any medium that can tangibly contain or store computer-executable instructions used by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a temporary computer-readable storage medium. In some embodiments, the storage medium is a non-temporary computer-readable storage medium. The non-temporary computer-readable storage medium may include, but is not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, CDs, DVDs, or optical disks based on Blu-ray technology, as well as persistent solid-state memory such as flash and solid-state drives. The personal electronic device 500 is not limited to the components and configurations shown in Figure 5B, and may include other or additional components in multiple configurations.

[0182] As used herein, the term “affordance” refers to user-interactive graphical user interface objects that are optionally displayed on the display screens of devices 100, 300, and / or 500 (Figures 1A, 3, and 5A-5B). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each optionally constitute an affordance.

[0183] As used herein, the term “focus selector” refers to an input element that indicates the current portion of the user interface with which the user is interacting. In some implementations, including a cursor or other location marker, the cursor acts as a “focus selector,” and therefore, when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 in Figure 3 or touch-sensitive surface 451 in Figure 4B) while the cursor is positioned over a particular user interface element, the particular user interface element is adjusted according to the detected input. In some implementations, including a touchscreen display that allows direct interaction with user interface elements on the touchscreen display (e.g., touch-sensitive display system 112 in Figure 1A or touchscreen 112 in Figure 4A), detected contact on the touchscreen acts as a “focus selector,” and therefore, when input (e.g., a press input by touch) is detected at a location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one area of ​​the user interface to another without corresponding cursor movement or touch movement on the touchscreen display (for example, by using the tab key or arrow keys to move focus from one button to another), and in these implementations, the focus selector moves in accordance with the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or touch on the touchscreen display) controlled by the user to communicate the user's intended interaction with the user interface (for example, by pointing to the device an element of the user interface through which the user intends to interact).For example, the position of a focus selector (e.g., cursor, touch, or selection box) over an individual button while pressure input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen) indicates that the user intends to activate that individual button (rather than other user interface elements displayed on the device's display).

[0184] As used herein and in the claims, the term “characteristic intensity” of a contact refers to the characteristics of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a set of intensity samples collected over a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detection of contact, before detection of lift-off of contact, before or after detection of the start of movement of contact, before detection of the end of contact, before or after detection of an increase in contact intensity, and / or before or after detection of a decrease in contact intensity). The characteristic intensity of a contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the top 10 percentile value of the contact intensity, the maximum half value of the contact intensity, the maximum 90 percent value of the contact intensity, and so on. In some embodiments, the duration of contact is used when determining characteristic intensity (for example, when characteristic intensity is the average intensity of contact over time). In some embodiments, characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action has been performed by the user. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this embodiment, contact with a characteristic intensity not exceeding the first threshold results in a first action, contact with a characteristic intensity above the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a characteristic intensity above the second threshold results in a third action. In some embodiments, the comparison between characteristic intensity and one or more thresholds is not used to determine whether a first action should be performed or a second action should be performed, but rather to determine whether one or more actions should be performed at all (for example, whether individual actions should be performed or whether individual actions should be withheld).

[0185] In some embodiments, a portion of the gesture is identified for the purpose of determining characteristic intensity. For example, a touch-sensitive surface optionally receives a series of swipe contacts that transition from a starting position to an ending position, where the intensity of contact increases. In this example, the characteristic intensity of the contact at the ending position is optionally based only on a portion of the series of swipe contacts (e.g., only the portion of the swipe contact at the ending position) rather than the entire swipe contact. In some embodiments, optionally, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of the following: an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or drops in the swipe contact intensity for the purpose of determining characteristic intensity.

[0186] The intensity of contact on a touch-sensitive surface is optionally characterized to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to the intensity at which the device performs an action typically associated with clicking a physical mouse button or trackpad. In some embodiments, the deep press intensity threshold corresponds to the intensity at which the device performs an action different from the action typically associated with clicking a physical mouse button or trackpad. In some embodiments, when a contact with a characteristic intensity below the light press intensity threshold (for example, above a nominal contact detection intensity threshold below which contact is no longer detected) is detected, the device moves the focus selector in accordance with the movement of the contact on the touch-sensitive surface without performing an action associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise specified, these intensity thresholds are consistent across various sets of user interface values.

[0187] An increase in the characteristic intensity of contact from an intensity below a light pressure intensity threshold to an intensity between the light and deep pressure intensity thresholds is sometimes referred to as a "light pressure" input. An increase in the characteristic intensity of contact from an intensity below a deep pressure intensity threshold to an intensity above a deep pressure intensity threshold is sometimes referred to as a "deep pressure" input. An increase in the characteristic intensity of contact from an intensity below a contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light pressure intensity threshold is sometimes referred to as the detection of contact on the touch surface. A decrease in the characteristic intensity of contact from an intensity above a contact detection intensity threshold to an intensity below a contact detection intensity threshold is sometimes referred to the detection of contact lift-off from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0188] In some embodiments described herein, one or more actions are performed in response to the detection of a gesture including an individual press input, or in response to the detection of an individual press input performed by an individual contact (or multiple contacts), wherein the individual press input is detected at least in part on the detection of an increase in the intensity of the contact (or multiple contacts) above a press input intensity threshold. In some embodiments, the individual action is performed in response to the detection of an increase in the intensity of the individual contact above a press input intensity threshold (e.g., a "downstroke" of the individual press input). In some embodiments, the press input includes an increase in the intensity of the individual contact above a press input intensity threshold, followed by a decrease in the intensity of the contact below the press input intensity threshold, and the individual action is performed in response to the detection of a subsequent decrease in the intensity of the individual contact below the press input threshold (e.g., an "upstroke" of the individual press input).

[0189] In some embodiments, the device employs intensity hysteresis to avoid accidental inputs, which may be referred to as “jitter,” and the device defines or selects a hysteresis intensity threshold that has a predetermined relationship with a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Thus, in some embodiments, the press input includes an increase in the intensity of an individual contact above the press input intensity threshold, followed by a decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and the individual action is performed in response to the detection of a subsequent decrease in the intensity of an individual contact below the hysteresis intensity threshold (e.g., an “upstroke” of the individual press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of contact from an intensity below a hysteresis intensity threshold to an intensity above a press input intensity threshold, and optionally a decrease in the intensity of contact to an intensity below the hysteresis intensity, and individual actions are performed in response to the detection of a press input (e.g., an increase in the intensity of contact or a decrease in the intensity of contact, depending on the situation).

[0190] For the sake of clarity, the description of an action performed in response to a press input associated with a press input intensity threshold, or a gesture involving a press input, is optionally triggered in response to the detection of any of the following: an increase in contact intensity above the press input intensity threshold, an increase in contact intensity from below the hysteresis intensity threshold to above the press input intensity threshold, a decrease in contact intensity below the press input intensity threshold, and / or a decrease in contact intensity below the hysteresis intensity threshold corresponding to the press input intensity threshold. Furthermore, in an example where an action is described to be performed in response to the detection of a decrease in contact intensity below the press input intensity threshold, the action is optionally performed in response to the detection of a decrease in contact intensity below a hysteresis intensity threshold corresponding to the press input intensity threshold and lower than that threshold.

[0191] In this specification, “installed application” refers to a software application that has been downloaded onto an electronic device (e.g., device 100, 300, and / or 500) and is ready to be launched on the device (e.g., opened). In some embodiments, a downloaded application becomes an installed application by an installation program that extracts the program portion from the downloaded package and integrates the extracted portion with the operating system of the computer system.

[0192] In this specification, the terms “open application” or “running application” refer to a software application that has retained state information (e.g., as part of the device / global internal state 157 and / or application internal state 192). An open or running application is optionally one of the following types of applications: ● The active application currently displayed on the display screen of the device on which the application is being used. ● Background applications (or background processes) that are not currently displayed but whose applications have one or more processes being handled by one or more processors, as well as ● An application that is not running but has state information stored in memory (volatile and non-volatile, respectively) that can be used to resume the execution of the application, either suspended or suspended.

[0193] In this specification, the term “closed application” refers to a software application that does not retain state information (for example, state information for a closed application is not stored in the device’s memory). Therefore, closing an application involves stopping and / or removing the application process for the application and removing the state information for the application from the device’s memory. Generally, opening a second application while a first application is running does not close the first application. When the second application is displayed and the first application is stopped from being displayed, the first application becomes a background application.

[0194] Next, we will focus on embodiments of user interfaces ("UI") and related processes implemented on electronic devices such as portable multifunction device 100, device 300, or device 500.

[0195] Figures 6A to 6Z show exemplary user interfaces for managing visual content within media according to several embodiments. These user interfaces are used to illustrate processes described later, including the process shown in Figure 8.

[0196] Figure 6A shows a computer system 600 (e.g., an electronic device) displaying a camera user interface, which optionally includes a live preview 630 extending from the top to the bottom of the computer system 600's display. In some embodiments, the computer system 600 optionally includes one or more characteristics of device 100, device 300, or device 500. In some embodiments, the computer system 600 is a tablet, phone, laptop, desktop, etc.

[0197] The live preview 630 is a representation of the field of view ("FOV") of one or more cameras of the computer system 600. In some embodiments, the live preview 630 is a representation of a partial FOV. In some embodiments, the live preview 630 is based on images detected by one or more camera sensors. In some embodiments, the computer system 600 uses multiple camera sensors to capture images and combines them to display the live preview 630. In some embodiments, the computer system 600 uses a single camera sensor to capture an image and display the live preview 630.

[0198] The camera user interface in Figure 6A includes an indicator area 602 and a control area 606, which are positioned relative to the live preview 630 so that the indicators and controls can be displayed simultaneously with the live preview 630. The camera display area 604 is substantially free of overlapping indicators and / or controls. As shown in Figure 6A, the camera user interface includes a visual boundary 608, which indicates the boundary between the indicator area 602 and the camera display area 604, and the boundary between the camera display area 604 and the control area 606.

[0199] As shown in FIG. 6A, the indicator region 602 includes indicators such as a flash indicator 602a and an animation image indicator 602b. The flash indicator 602a indicates whether the flash is on (e.g., active), off (e.g., inactive), or in another mode (e.g., auto mode). In FIG. 6A, the flash indicator 602a indicates that the flash mode is off and that the flash operation is not used when the computer system 600 is capturing media, indicating this to the user. The animation image indicator 602b indicates whether the camera is configured to capture a single image or multiple images (e.g., in response to detecting a request to capture media). In some embodiments, the indicator region 602 is superimposed on the live preview 630 and optionally includes a colored (e.g., gray, translucent) overlay.

[0200] As shown in FIG. 6A, the camera display region 604 includes a live preview 630 and a zoom control (e.g., an affordance) 622. The zoom control 622 includes a 0.5× zoom control 622a, a 1× zoom control 622b, and a 2× zoom control 622c. As shown in FIG. 6A, the fact that the 1× zoom control 622b is bold and enlarged compared to the other zoom controls indicates that the 1× zoom control 622b is selected and that the computer system 600 is displaying the live preview 630 at a "1×" zoom level.

[0201] As shown in Figure 6A, the control area 606 includes representations of the camera mode control (e.g., control) 620, the shutter control 610, the camera switcher control 614, and the media collection 612. In Figure 6A, the display of camera mode controls 620a to 620e, and the bolding of the "Photo" camera mode 620c, indicates that the computer system 600 is configured to capture photographic media when the shutter control 610 is active. Thus, when the shutter control 610 is enabled, it causes the computer system 600 to capture media (e.g., a photograph when the shutter control 610 is enabled in Figure 6A) using one or more camera sensors, based on the current state of the live preview 630 and the current state of the camera application (e.g., which camera mode is selected). The captured media is stored locally in the computer system 600 and / or sent to a remote server for storage. When activated, the camera switcher control 614 causes the computer system 600 to switch to displaying the fields of view of different cameras in the live preview 630, for example, by switching between the rear camera sensor and the front camera sensor. The representation of the media collection 612 shown in Figure 6A is a representation of the media (e.g., images, videos) most recently captured by the computer system 600. In some embodiments, in response to the detection of a gesture directed towards the media collection 612, the computer system 600 displays a user interface similar to the user interface shown in Figure 7B (described later). In some embodiments, the indicator area 602 is superimposed on the live preview 630 and optionally includes a colored (e.g., gray, semi-transparent) overlay. In Figure 6A, the computer system detects a tap input 650a on (and / or directed towards) the shutter control 610.

[0202] As shown in FIG. 6B, in response to detecting a tap input 650a, the computer system 600 starts capturing media, captures the live preview 630 of FIG. 6A, and displays a new representation in the media collection 612. In FIG. 6B, the new representation is the representation of the live preview 630 of FIG. 6A (e.g., captured in response to detecting the tap input 650a on the shutter control 610). Also, since the new representation corresponds to the representation of the most recently captured media, it is displayed at the top of the media collection 612 in FIG. 6B.

[0203] As shown in FIG. 6B, the live preview 630 includes a representation showing a person 640 standing behind a tree, and the head of the person 640 and a part of the body of the person 640 are not covered by the tree. On the tree, there is a sign 642 including a text portion 642a (e.g., "LOST DOG") and a text portion 642b (e.g., a paragraph of text starting with "LOVEABLE"). In FIG. 6B, the text in the text portions 642a - 642b is not visually prominent, and in the embodiment shown in FIG. 6B, the text in the text portions 642a and 642b is small and not easily readable by a user looking at the computer system 600. In FIG. 6B, the computer system 600 detects a pinch - out input 650b on the live preview 630.

[0204] As shown in FIG. 6C, in response to detecting the pinch - out input 650b, the computer system 600 replaces the display of the 2× zoom control 622c of FIG. 6B with the display of a 2.5× zoom control. Also, the computer system 600 updates the live preview 630 so that objects within the field of view of one or more cameras are displayed at a "2.5×" zoom level (e.g., as shown by the newly displayed and selected (e.g., enlarged and bolded) 2.5× zoom control 622d) instead of the "1×" zoom level of FIG. 6B to reflect the change in the zoom level.

[0205] Compared to Figure 6B, the text sections 642a-642b in Figure 6C are visually more prominent (e.g., larger and easier to read) than the text sections 642a-642b in Figure 6B. In Figure 6C, each of the text sections 642a-642b (and / or the text contained within them) was determined to satisfy the set of prominence criteria. The text sections 642a-642b in Figure 6C satisfy the set of prominence criteria because each text section occupies a portion (e.g., 10%) that exceeds the threshold of the live preview 630, and / or each text section contains text larger than the threshold size (e.g., larger than a 6pt font). In some embodiments, one or more text portions satisfy a set of prominence criteria based on other criteria, such as whether an individual text portion contains one or more types of text (e.g., email, phone number, quick response ("QR") code, whether an individual text portion is displayed in or near a specific location in the live preview 630 (e.g., a central location), or whether an individual text portion is relevant based on the context of the media displayed as the live preview 630 (as will be described in more detail in relation to Figures 7A–7L, 8, and 9).

[0206] As shown in Figure 6C, in order to determine that the text portions 642a to 642b satisfy the set of prominence criteria (and / or that at least a portion of the text satisfies the set of prominence criteria), the computer system 600 displays parentheses 636a around the text portions 642a to 642b and displays a text management control 680 to the right of the zoom control 622d in the camera display area 604.

[0207] Returning to Figure 6B, the brackets 636a and the text management control 680 were not displayed in Figure 6B because it was determined that text portions 642a-642b did not meet the set of prominence criteria (and / or no portion of the text met the set of prominence criteria). In Figure 6B, it was determined that text portions 642a-642b did not meet the set of prominence criteria because they did not occupy any portion of the live preview 630 that exceeded the threshold and did not contain any text larger than the threshold size. In some embodiments (as shown in Figures 6B-6C), the determination of whether an individual text portion meets the set of prominence criteria is based not only on whether the live preview 630 contains text (and / or text portions), but also on how / when the text portion is currently displayed within the live preview 630.

[0208] Returning to Figure 6C, since the image of the dog is positioned between text portions 642a and 642b, the bracket 636a is positioned around the image of the dog on sign 642. In some embodiments, multiple brackets are displayed, such that one bracket appears around text portion 642a and another around text portion 642b. In some embodiments, multiple brackets are displayed because multiple text portions (e.g., "Multiple portions of text") satisfy a set of prominence criteria and it has been determined that the object is positioned between the text portions. In some embodiments, when the object is not positioned between multiple portions of text, only one bracket is displayed around the multiple portions of text. In some embodiments, if text portion 642a satisfies the set of prominence criteria but text portion 642b does not, a bracket is displayed around text portion 642a and no bracket is displayed around text portion 642b (and vice versa). In some embodiments, the computer system 600 indicates that an individual text portion (e.g., a text portion) satisfies a set of prominence criteria by, in addition to displaying parentheses around the individual text portion, and / or instead, by highlighting the individual portion in other ways, such as highlighting, bolding, resizing, and boxing around the individual text portion.

[0209] As shown in Figure 6C, the computer system 600 displays text type indicators 638a-638b (e.g., underlines) to indicate that a specific type of text (e.g., email, address, telephone number, QR code, etc.) has been detected within the text portion 642b (e.g., a data detector). In Figure 6C, text type indicator 638a is displayed below "123 Main Street" indicating that an address has been detected, and text type indicator 638b is displayed below "123-4567" indicating that a telephone number has been detected. In some embodiments, when a text type indicator is displayed below a portion of text, the user can perform an action by selecting the portion of text and / or the text type indicator (e.g., as further described below in relation to Figures 6M-6N).

[0210] Figures 6C to 6D illustrate exemplary embodiments of the computer system 600 being moved in a physical environment. Figures 6C to 6D include a graphical representation 660 showing the original position 660a of the computer system 600 (e.g., in Figures 6C to 6D) relative to the changed position 660b of the computer system 600 (e.g., in Figure 6D). As shown in Figure 6C, the computer system 600 is in its original position 660a. In Figure 6C, the position of the computer system 600 has been changed.

[0211] As shown in Figure 6D, in response to the change in the position of the computer system 600 (for example, from the original position 660a to the changed position 660b), the computer system 600 shifts the live preview upward. In Figure 6D, the live preview 630 is shifted upward so that the top of the live preview 630 in Figure 6C (for example, the portion including the text portion 642a) is no longer displayed and the new bottom portion of the live preview 630 (as shown in Figure 6D) is newly displayed. In Figure 6D, it is determined that the text portion 642a does not satisfy the set of prominence criteria, while the text portion 642b (continues to) satisfy the set of prominence criteria. Here, since the text portion 642a is no longer displayed as part of the live preview 630 (for example, in the camera display area) in Figure 6D, it is determined that the text portion 642a does not satisfy the set of prominence criteria. As illustrated, since text portion 642a does not satisfy the set of prominence criteria and text portion 642b does, the computer system 600 displays parentheses 636b around text portion 642b (instead of text portion 642a) and discontinues displaying parentheses 636a. In other words, the computer system 600 dynamically changes parentheses 636a to parentheses 636b according to changes in the determination of whether one or more text portions (e.g., text portions currently displayed as part of the live preview 630) satisfy and / or do not satisfy the set of prominence criteria. Therefore, one or more determinations of whether one or more text segments satisfy the set of prominence criteria are dynamic and may change if the live preview 630 changes in response to requests for zoom in (e.g., pinch out input) / zoom out (e.g., pinch input) or pan (e.g., swipe right, left, up, down input), and / or if the live preview 630 changes in response to movement (e.g., forward, backward, up, down) of one or more cameras of the computer system 600.In some embodiments, if one or more determinations change as to whether one or more text portions satisfy a set of prominence criteria, the display of one or more brackets (e.g., brackets 636a-636b) and / or the display of the text management control 680 changes (as further described below in relation to Figures 7A-7L, 8, and 9). In some embodiments, while the computer system 600 displays brackets 636a around text portion 642b (and / or in response to the detection of text in the live preview 630), the computer system 600 darkens and / or reduces the saturation (e.g., saturation, tint, and / or hue) of portions of the live preview 630 that do not have text (e.g., a picture of a dog), while maintaining the saturation and / or brightness of text portion 642b (and / or other portions of text). In some embodiments, the computer system 600 displays the text portion 642b with greater brightness than the portion of the live preview 630 without text, as a portion of which the text portion 642b is kept luminous.

[0212] As shown in Figure 6D, the computer system 600 continues to display the text management control 680 because it has been determined that text portion 642b (continues to) satisfy the set of prominence criteria. In Figure 6D, the text management control 680 is displayed because at least one determination has been made that the currently displayed text portion (e.g., in the live preview 630) satisfies the set of prominence criteria, regardless of whether another text portion (e.g., text portion 642a) continues to satisfy (or does not satisfy) the set of prominence criteria. In Figure 6D, the computer system 600 returns to its original position 660a.

[0213] As shown in Figure 6E, depending on the computer system 600 being in its original position 660a, the computer system 600 redisplays the live preview 630 using one or more techniques as described above in relation to Figure 6C. In Figure 6E, the computer system 600 detects a tap input 650e on the text management control 680.

[0214] As shown in Figure 6F, in response to the detection of a tap input 650e, the computer system 600 changes the display of the text management control 680. Specifically, the computer system 600 displays the text management control 680 in an active and / or selected state (for example, as shown in Figure 6F where the text management control 680 is bolded), and stops displaying the text management control 680 in an inactive and / or deselected state (for example, as shown in Figure 6E where the text management control 680 is not bolded).

[0215] As shown in Figure 6F, in response to the detection of a tap input 650e, the computer system 600 highlights the text portions 642a-642b and darkens other portions of the live preview 630 (and / or other objects in the field of view of one or more cameras), such as the person 640, the image of the dog on the sign 642, and the tree displayed in the live preview 630. Along with darkening other portions of the live preview 630, the computer system 600 discontinues the display of one or more controls in the camera display area 604 (e.g., the zoom control 622 in Figure 6E). The computer system 600 also darkens (or discontinues the display of) multiple portions of the camera user interface, such as multiple indicators in the indicator area 602 and multiple controls in the camera control area 606. In some embodiments, some of the indicators and / or controls darkened in the camera user interface in Figure 6F are not selectable (e.g., the computer system 600 does not perform an action when selected). In some embodiments, some of the indicators and / or controls remain selectable and / or are not dimmed in response to the detection of the tap input 650e. In some embodiments, the computer system 600 maintains the display of some controls in the camera display area 604 in response to the detection of the tap input 650e. In some embodiments, the computer system 600 highlights portions 642a-642b by increasing the size of the text within those portions, highlighting the text within those portions, displaying a box around those portions, etc. In some embodiments, dimming portions of the live preview 630 includes reducing the saturation of portions of the live preview 630 that do not have text (e.g., a picture of a dog) while maintaining the saturation of the text portions 642a and 642b (e.g., using similar techniques as described above in relation to Figure 6D).

[0216] In particular, in Figure 6F, the portion of text highlighted in response to the detection of input 650e is the portion of text enclosed in parentheses (e.g., parentheses 636a) when input 650e was received in Figure 6E. In some embodiments, parentheses around a portion of text indicate to the user which text is highlighted and / or managed by the user when a selection is made in the text management control 680. In some embodiments, one or more portions of text that do not have parentheses surrounding them, but are displayed via the live preview 630 when input is received on the text management control 680, are not highlighted in response to a selection by the text management control 680 (for example, in Figure 7F below, "BRAND" is not highlighted when the text management control 680 is selected in Figure 7F). In some embodiments, one or more portions of text that do not have parentheses surrounding them, but are displayed via the live preview 630 when input is received on the text management control 680, are highlighted in response to a selection by the text management control 680 (for example, if it is determined that one or more portions of text satisfy a set of prominence criteria).

[0217] As shown in Figure 6F, upon detection of a tap input 650e, the computer system 600 also displays a text management option 682 and a command 684 (e.g., "Swipe or tap to select text") indicating one or more inputs / gestures that can be used to select a subset of text from text portions 642a-642b. In Figure 6F, the text management option 682 is an option for managing text portions 642a-642b. In particular, the text management option 682 includes a copy option 682a, a select all option 682b, a search option 682c, and a share option 682d. In some embodiments, upon receiving input directed to the copy option 682a, the computer system 600 copies the selected text (e.g., the text in text portions 642a-642b in Figure 6F) and / or stores the selected text in a copy / paste buffer, which allows the selected text to be pasted upon receiving a request to paste the selected text. In some embodiments, upon receiving input directed to the Select All option 682b, the computer system 600 selects all of the highlighted text on the computer system 600. In some embodiments, when the computer system 600 has selected all of the text within the selected text, the computer system 600 highlights the selected text. In some embodiments, upon receiving input directed to the Search option 682c, the computer system 600 searches for the selected text (e.g., the highlighted text portion in Figure 6F) through a search application (e.g., a web application, a dictionary application, a personal assistant application) and / or displays one or more definitions and resources for the highlighted and / or selected text.In some embodiments, upon receiving input directed to the sharing option 682d, the computer system 600 initiates a process of sharing selected text through one or more applications (e.g., email, text messaging, word processing, social media applications) (e.g., one or more predetermined applications). In some embodiments, as part of initiating the process of sharing selected text, the computer system 600 displays a scrollable list of applications, and the selection of an application from the scrollable list of applications causes the computer system 600 to share the selected text using the selected application. In some embodiments, the scrollable list of applications is displayed simultaneously with a portion of the live preview 630 (e.g., including one or more of the text portions 642a-642b). In Figure 6F, the computer system 600 detects a tap input 650f on a portion of the live preview 630 (e.g., a portion within the darkened area of ​​the live preview 630, and / or a portion of the live preview 630 that does not include the text portions 642a-642b and / or the text management control 680).

[0218] As shown in Figure 6G, upon detection of tap input 650f, the computer system 600 displays the text management control 680 in an inactive state, does not highlight text portions 642a-642b, brightens the rest of the live preview 630 and the camera user interface, and stops displaying text management options 682 and instructions 684. Also, because it has been determined that text portions 642a-642b (continue to) satisfy the set of prominence criteria, upon detection of tap input 650f, the computer system 600 redisplays parentheses 636a. In effect, upon detection of tap input 650f, the camera user interface is returned to the state it was in before tap input 650e was detected on the text management control 680. In Figure 6G, the computer system 600 detects tap input 650g on the text management control 680.

[0219] As shown in Figure 6H, in response to the detection of a tap input 650g, the computer system 600 displays the camera interface in Figure 6H using one or more techniques as described above in relation to Figure 6F. In particular, in Figure 6H (and Figure 6F), since the text portions 642a to 642b satisfy the set of prominence criteria, the computer system 600 highlights the text portions 642a to 642b. In some embodiments, in response to the detection of a tap input 650g, the computer system 600 darkens one or more portions of the text that do not satisfy the set of prominence criteria. In Figure 6H, the computer system 600 detects a tap input 650h on the text portion 642b.

[0220] As shown in Figure 6I, upon detection of a tap input 650h, the computer system 600 selects text portion 642b and repositions the text management option 682 so that it appears over text portion 642b in Figure 6I, instead of over text portion 642a (as shown in Figure 6H, for example). The repositioning of the text management option 682 indicates that the text management option can be used to manage the text in text portion 642b, but not over the text in text portion 642a. In other words, upon detection of an input (e.g., a swipe or tap) that selects a particular portion of text, the computer system 600 uses the text management option 682 to change the text selected to be managed.

[0221] In particular, the live preview 630 in Figure 6I does not include the person 640 that was included in the live preview 630 in Figure 6H. This is because the person 640 has moved behind the tree in the live preview 630 in Figure 6I and is therefore not within the field of view of one or more cameras of the computer system 600. As shown in Figure 6I, while the text management control 680 is displayed in an active state and / or while the text management option 682 is displayed, the live preview 630 continues to update to reflect changes in the field of view of one or more cameras of the computer system 600. In some embodiments, the live preview 630 continues not to update while the text management control 680 is displayed in an active state and / or while the text management option 682 is displayed. Therefore, in embodiments where the live preview 630 is not updated, the computer system 600 maintains the display of the portion of the person 640 protruding from behind the tree in the live preview 630 in Figure 6I. In Figure 6I, the computer system 600 detects a pinch-out input 650i.

[0222] As shown in FIG. 6J, in response to detecting a pinch release input 650i, the computer system 600 displays a live preview 630 at an increased zoom level and maintains the display of the text portion 642b and the text management option 682. In some embodiments, since the text portion 642b is selected, the computer system 600 continues to display at least a subset of the text portion 642b in response to a request for zoom in (e.g., pinch release input) (and / or zoom out, pan, and / or movement of one or more cameras of the computer system 600 and / or the computer system 600). In some embodiments, the display of the selected text portion (e.g., text portion 642b) is static. Thus, in some embodiments where the selected text portion is static, the computer system 600 continues to display the selected text portion regardless of whether the selected text portion remains within the field of view of one or more cameras (e.g., when the computer system 600 is moved, panned, and / or zoomed) (e.g., while the camera user interface is still being displayed). In FIG. 6J, the computer system 600 detects a tap input 650j on the word "Fluffy" included in the text portion 642b.

[0223] As shown in FIG. 6K, in response to detecting the tap input 650j, the computer system 600 selects and highlights the word "Fluffy". In FIG. 6K, only the selected word "Fluffy" can be managed by the text management option 682 displayed in FIG. 6K. In FIG. 6K, the computer system 600 detects a leftward swipe input 650k starting from the word "Fluffy".

[0224] As shown in Figure 6L, upon detection of a leftward swipe input 650k, the computer system 600 selects and highlights multiple words contained in the text portion 642b based on the direction of the swipe input 650k. As shown in Figure 6L, the word "THE NAME FLUFFY" is highlighted, indicating that "THE NAME FLUFFY" has been selected based on the swipe input 650k. In Figure 6L, only the selected word "THE NAME FLUFFY" can be managed by the text management option 682 shown in Figure 6L.

[0225] Figures 6L–6M illustrate exemplary embodiments in which the computer system 600 moves in the physical environment while the computer system 600 continues to display the selected text portion (or a subset of the text portion) regardless of whether the selected text portion remains within the field of view of one or more cameras (as will be further described below, for example, in relation to Figures 6L–6M). Figures 6L–6M include a graphical representation 660 showing the original position 660a of the computer system 600 (for example, in Figures 6L–6M) relative to the changed position 660c of the computer system 600 (for example, in Figure 6M).

[0226] As shown in Figure 6L, the tree mark 646 represents the static portion of the tree displayed in the live preview 630 in Figures 6L to 6M. In Figure 6L, the tree mark 646 is displayed below the text portion 642b. In Figure 6L, the position of the computer system 600 has been changed.

[0227] As shown in Figure 6M, in response to a change in the position of the computer system 600 (for example, as indicated by the changed position 660c relative to the original position 660a), the computer system 600 updates the live preview so that the tree mark 646 is displayed above the text portion 642b. In particular, in Figure 6M, the text portion 642b of the computer system 600 is no longer within the field of view of one or more cameras, and the text portion 642b is positioned where it would be displayed in the live preview 630 of Figure 6M (for example, as evidenced by the tree mark 646 moving to a higher position in the live preview 630). However, because a subset of the text portion 642b (for example, "THE NAME FLUFFY") is selected, the computer system 600 continues to display the text portion 642b in the live preview 630 of Figure 6M. In some embodiments, the computer system 600 displays only the selected subset of the text portion 642b without displaying the rest of the unselected text portion 642b. In some embodiments, when text is selected and the camera is moved (and / or zoomed / panned) in the physical environment, the computer system 600 does not update the live preview 630. In some embodiments (for example, as further described below in relation to Figures 7A-7L, 8, and 9), the computer system 600 selects a different portion of the text in response to the computer system 600 and / or the camera of the computer system 600 being moved (for example, and / or zoomed / panned) (for example, when the text is displayed towards the bottom of the tree in the live preview 630). In Figure 6M, the computer system 600 detects input 650m on "123-4567" with text type indication 638b displayed below.

[0228] As shown in Figure 6N, upon detection of input 650m, and after determining that input 650m is a tap input and that "123-4567" corresponds to a telephone number, the computer system 600 displays the telephone dialer user interface and automatically initiates a telephone call to "123-4567" (for example, without requiring user input on the keypad and / or contact information card). In some embodiments, a confirmation screen is displayed before the computer system 600 initiates the telephone call.

[0229] As shown in Figure 6O, upon detection of input 650m, and because it has been determined that input 650m is a long press input and that "123-4567" corresponds to a telephone number, the computer system 600 displays telephone number management options 692, including call option 692a, message sending option 692b, add contact option 692c, and copy option 692d. As shown in Figure 6O, the computer system 600 displays options for managing several specific types of text (e.g., email, telephone numbers, QR codes) that differ from the management of other types of text (as shown by the text management option 682 displayed in Figure 6L when "THE NAME FLUFFY" is selected, in contrast to the management of other types of text). In some embodiments, upon detection of input directed to call option 692a, the computer system 600 initiates a telephone call to "123-4567" (for example, using a similar technique as described above in relation to Figure 6N). In some embodiments, upon detection of input directed to the message sending option 692b, the computer system 600 initiates the process of sending a message to "123-4567" (for example, displaying a text management application). In some embodiments, upon detection of input directed to the add contact option 692c, the computer system 600 initiates the process of adding a contact to a contact list that has "123-4567" as a phone number in the contact information. In some embodiments, upon detection of input directed to the copy option 692d, the computer system 600 copies "123-4567" using one or more techniques as described above in relation to the copy option 682a in Figure 6F.

[0230] Figures 6P to 6T show exemplary embodiments in which a QR code is displayed in the live preview 630. In some embodiments, the QR code can be replaced with other types of matrix and / or barcodes.

[0231] As shown in Figure 6P, the computer system 600 displays the QR code 668 simultaneously with the QR code identifier 670 (e.g., "CAFE32.COM") in the live preview 630. In some embodiments, the QR code identifier identifies one or more of the following: a website, contacts, a cellular plan, an email address, a calendar invitation / event, a location (e.g., a GPS location), text, video, a phone number, a WiFi network, an application, and / or an instance of an application. The QR code identifier 670 includes an indication of the information identified by the QR code. In Figure 6P, the QR code 668 is within the field of view of one or more cameras of the computer system 600, while the QR code identifier 670 is not within the field of view of one or more cameras of the computer system 600. Because it has been determined that the QR code 668 corresponds to a target website belonging to "CAFE32.COM" (e.g., or identifies the target website described above), the computer system 600 displays the QR code identifier 670. In Figure 6P, the computer system 600 detects input 650p1 and / or input 650p2 in the camera display area 604.

[0232] As shown in Figure 6Q, in response to the detection of input 650p1 and / or input 650p2 (and based on the determination that at least one of the inputs is a tap input and / or a long press input), the computer system 600 displays a notification 674 that includes a preview of a website (e.g., the address "CAFE32.COM"). In some embodiments, the web address preview includes the full web address (e.g., "http:\\cafe32.com\menu") and / or an image from the web address. In some embodiments, in response to the detection of one or more inputs, the computer system 600 displays a notification 674 instead of navigating to the web address corresponding to the QR code 668 to minimize the possibility that the user unintentionally navigates to a website corresponding to the QR code 668. In some embodiments, in response to the detection of input 650p1 on the QR code 668, the computer system 600 displays a notification 674 (e.g., without automatically navigating to a website). In some embodiments, upon detection of input 650p2 for a QR code identifier, the computer system 600 automatically navigates to the website corresponding to the QR code 668 (for example, without displaying notification 674) (for example, using one or more similar techniques as described below in relation to the computer system 600's response to tap input 650q). In Figure 6Q, the computer system 600 detects tap input 650q on notification 674.

[0233] As shown in Figure 6R, upon detection of the tap input 650q, the computer system 600 automatically navigates (and / or opens) to a web address corresponding to the QR code via the web application 678.

[0234] As shown in Figure 6S, the computer system 600 displays the QR code 668 simultaneously with the QR code identifier 670 using one or more techniques as described above in relation to Figure 6P. In Figure 6S, the computer system 600 detects a tap input 650s on the text management control 680.

[0235] As shown in Figure 6T, in response to the detection of tap input 650s, the computer system 600 displays QR code management options 672, including a share option 672a, a copy link option 672b, a reading list add option 672c, and an open link option 672d. As described in relation to Figure 6O above, the computer system 600 displays different options for managing certain types of text than for managing other types of text. In some embodiments, in response to the detection of input directed to the share option 672a, the computer system 600 initiates the process of sharing the web address and / or link corresponding to the QR code (for example, using one or more similar techniques described above in relation to the input directed to the share option 682d in Figure 6F). In some embodiments, in response to the detection of input directed to the copy link option 672b, the computer system 600 copies the web address and / or link corresponding to the QR code (for example, using one or more techniques as described above in relation to the copy option 682a in Figure 6F). In some embodiments, upon detection of input directed to the reading list addition option 672c, the computer system 600 initiates the process of adding a web address and / or link corresponding to a QR code to the list of items (e.g., one or more articles, books, websites, etc.). In some embodiments, upon detection of input directed to the open link option 672d, the computer system 600 navigates to (and / or opens) the web address corresponding to a QR code via the web application 678 (e.g., using a similar technique as described above in relation to Figure 6R).

[0236] In some embodiments, the QR code management option 672 includes one or more options that are dynamically selected based on the type of resource represented by the QR code (e.g., the QR code displayed when the text management control 680 is selected). For example, the type of resource represented by the QR code may include one or more of the following: a link to a website, contacts, a cellular plan, an email address, a calendar invitation / event, a location (e.g., a GPS location), text, video, a phone number, a WiFi network, an application, and / or an instance of an application. In some embodiments, the QR code management option 672 includes a first set of controls when the QR code represents a first type of resource, and a second set of controls when the QR code represents a second type of resource different from the first type. In some embodiments, the first set of controls has a different amount of controls than the second set of controls. In some embodiments, a preview of the resource represented by the QR code is included in the QR code management option 672 (e.g., when the QR code represents a string of text).

[0237] In some embodiments, the QR code management option 672 includes different sets of controls based on whether the computer system 600 is locked or unlocked. In some embodiments, if the computer system 600 is locked and the QR code represents a link to an application, control options to install and / or open the application are displayed. In some embodiments, if the computer system 600 is unlocked, even if the application is installed, the link to open the application is not displayed (e.g., suppressed) to avoid providing information to an unauthorized user of the device about which applications are installed on the device. Optionally, instead of displaying the link to open the application, the device displays an option to use a portion of the application that is available without downloading the entire application. In some embodiments, the computer system 600 displays different sets of controls (e.g., based on whether the computer system 600 is locked or unlocked) to restrict the information given to an unauthorized user (e.g., information that can be used to determine whether and / or whether the application represented by the QR code is installed on the computer system 600).

[0238] Figures 6U–6W illustrate exemplary scenarios in which the computer system 600 displays selection indicators around selected text separated into columns. In Figures 6U–6W, the computer system 600 is oriented so that the text in the environment is aligned with the field of view of one or more cameras of the computer system 600. Figure 6U shows the computer system 600 displaying a live preview 630 containing a representation of a text portion 648 (e.g., a soccer player roster). In some embodiments, the computer system 600 displays a representation of previously captured media containing a representation of a text portion 648, and one or more techniques described below in relation to Figures 6U–6W are used to select words within the text portion 648.

[0239] As shown in Figure 6U, the text section 648 includes the Name column 648a, the Position column 648b, the State column 648c, and the Grade column 648d. The individual columns contain text detected by the computer system 600 (for example, using one or more techniques as described above in relation to Figures 6A-6F). As shown in Figure 6U, the computer system 600 highlights the text section 648 while reducing the visual prominence of the portion of the live preview 630 that does not contain text (for example, the soccer ball) (for example, using one or more techniques as described above in relation to Figures 6A-6F). Also, because the computer system 600 detected the text section 648, the computer system 600 places a box around the text 648 to highlight it. As shown in Figure 6U, the computer system 600 displays the text management control 680 in an active state (for example, as indicated by the text management control 680 being bolded) and displays the text management option 682 (for example, as described above in relation to Figure 6F). In Figure 6U, the computer system 600 detects the first portion of the swipe input 650u on the name column 648a, which has moved from the header "name" of the name column 648a to the header "position" of the position column 648b.

[0240] As shown in Figure 6V, in response to the detection of a first portion of the swipe input 650u, the computer system 600 displays selection indicators 696 (e.g., "highlighted in gray") around all the words in the name column 648a ("Name", "Maria", "Kate", "Sarah", and "Ashley") and around the header "position" in the position column 648b. The selection indicators 696 are positioned based on the location of the swipe input 650u. Because the first portion of the swipe input 650u of the computer system 600 ends at the location of the header "position" in the position column 648b, the computer system 600 displays selection indicators 696 around all words up to the header "position" (e.g., including the words in the name column 648a) and including the header "position". In some embodiments, because the first portion of the swipe input 650u of the computer system 600 ends at the location of the header "position" in the position column 648b, the computer system 600 does not include the header "position" in the position column 648b. In some embodiments, where the end of the input is at the location of the word "DEFENDER" in position column 648b (for example, the third row of position column 648b), the computer system 600 highlights all words up to the word "DEFENDER," including all words in name column 648a, the header "position" in position column 648b (for example, the first row of position column 648b), and the word "Forward" in the second row of position column 648b.

[0241] The shape of the selection indicator 696 depends on whether the selected text (e.g., the text enclosed by the selection indicator 696) is aligned with the computer system 600. In Figure 6V, the computer system 600 displays the selection indicator 696 as a polygon with right angles (e.g., a shape with all right angles, referred to herein as a rectangle-based selection indicator). Since it has been determined that the selected text (e.g., the text of the selection indicator 696) is aligned with the computer system 600 (e.g., aligned with the field of view of one or more cameras of the computer system 600), the selection indicator 696 is a rectangle-based selection indicator (e.g., described below with additional details in relation to Figures 6X to 6Z). In Figure 6V, the computer system 600 detects a second portion of the swipe input 650u, which is a rightward swipe input, moving from the header "position" in the position column 648b to the header "state" in the state column 648c.

[0242] As shown in Figure 6W, since the computer system recognizes the words in the name column 648a as being in the same column, in response to the detection of the second portion of the swipe input 650u, the computer system 600 expands the selection indicator 696 to the right (for example, using one or more techniques as described above in relation to Figures 6U to 6W) so that the selection indicator 696 is displayed around the words in the name column 648a and the position column 648b (for example, all of the words) and also around the header "state" in the state column 648c. As shown in Figure 6W, the selection indicator 696 remains a rectangle-based selection indicator because the portion of the text is still aligned with the field of view of one or more cameras. In Figure 6W, the computer system 600 no longer detects the input swipe input 650u. However, the computer system 600 continues to display the selection indicator 696 around the portion of the text.

[0243] Figures 6X–6Z illustrate exemplary scenarios in which, when the computer system 600 is oriented (for example, oriented towards individual text portions, as differs from how the computer system 600 in Figures 6U–6V is oriented towards individual text portions), the computer system 600 displays selection indicators around selected text so that the text in the environment is not aligned with the field of view of one or more cameras of the computer system 600. Figure 6X shows the computer system 600 displaying a live preview 630 containing a representation of text portion 652 (for example, a paragraph of text relating to soccer). Text portion 652 is on paper in the environment, captured by the field of view of one or more cameras of the computer system 600. In some embodiments, the computer system 600 displays a representation of previously captured media, including a representation of text portion 648, and one or more techniques, described below in relation to Figures 6U–6W, are used to select words within text portion 648.

[0244] In Figure 6X, the text portion 652 is not aligned with the field of view of one or more cameras. In Figure 6X, the computer system 600 is oriented to a location where the computer system 600 is not parallel to the text portion 652 and / or rotated / tilted along an axis (z-axis) in the environment (for example, the user is holding the phone at an angle and / or tilt such that the field of view of one or more cameras is not aligned with the text portion 652). In Figure 6X, the computer system 600 detects a diagonal swipe input 650x that moves from the word "while" in the text portion 652 to the last period (".") in the text portion 652.

[0245] As shown in Figure 6Y, in response to the detection of a swipe input 650x, the computer system 600 displays a selection indicator 696 around a subset of text portion 652, from the word "while" in text portion 652 to the last period in text portion 652. As shown in Figure 6Y, the selection indicator 696 is a polygon with several angles that are not right angles (for example, a shape with several acute and several obtuse angles, referred to herein as a non-rectangular-based selection indicator). The non-rectangular-based selection indicator is drawn by the computer system so that it matches or appears to match (or substantially matches or substantially appears to match) the orientation of text portion 652 in the live preview 630 (for example, if the selection indicator 696 was a rectangle-based selection indicator on the surface containing text portion 642, but so that the surface containing text portion 642 is viewed from the same viewpoint as seen in Figures 6X-6Z). As described above, since it was determined that the text portion 652 was not aligned with the computer system 600, the selection indicator 696 in Figure 6Y is non-rectangular based (in contrast to the selection indicators 696 in Figures 6V to 6U, which are rectangle-based selection indicators (for example, relative to the orientation of the computer system 600's display)). In Figure 6Y, the computer system 600 detects a swipe input 650y that moves from the word "while" to the word "synthetic" in the text portion 652. In particular, the swipe input 650y moves diagonally with respect to the computer system 600, but along the line of words in the text portion 652. In some embodiments, even if the edge of the selection indicator 696 is displayed diagonally with respect to the edge of the computer system's display area, some or all of the edge of the selection indicator 696 is positioned where the computer system determines it to be parallel or perpendicular to the line of text in the text portion 652.In some embodiments, as the angle of the camera relative to the surface containing the text portion 642 changes, the angle of the edge of the selection indicator 696 shifts within the display area, maintaining the edge where it is determined to be parallel or perpendicular to the lines of text in the text portion 652.

[0246] As shown in Figure 6Z, upon detection of a swipe input 650y, the computer system 600 expands the selection indicator 696 in the direction of the swipe input 650y so that the selection indicator 696 surrounds a subset of the text portion 652 from the word "synthetic" to the last period (for example, if "while" is included in the text portion). Even after expansion, the selection indicator 696 still appears as a non-rectangular base 696 selection indicator. Furthermore, after the computer system 600 no longer detects a swipe input 650y, the selection indicator 696 continues to appear around the portion of the text.

[0247] Figures 7A to 7L illustrate exemplary user interfaces for managing visual indicators for visual content in media using a computer system, according to several embodiments. These user interfaces are used to illustrate processes described later, including the process shown in Figure 9.

[0248] Figure 7A shows a computer system 600 that simultaneously displays a media gallery user interface 710, which includes thumbnail media representations 712 and a gallery area 702. The thumbnail media representations 712 include thumbnail media representations 712a to 712c, each of which represents a different media item (e.g., a media item captured in different instances within a time period). The gallery area 702 includes a library control 702a (for example, causing the computer system 600 to display the thumbnail media representations 712 when selected), a "for you" control 702b (for example, causing the computer system 600 to display dynamically generated thumbnail representations of media items based on user preferences when selected), an album control 702c (for example, causing the computer system 600 to display thumbnail album representations, each representing a collection of media items when selected), and a search control 702d (for example, causing the computer system 600 to display a search user interface, which includes one or more controls for searching for media items when selected). In Figure 7A, library control 702a is selected (as indicated by the bolding of library control 702a, for example). In Figure 7A, the computer system 600 detects a tap input 750a on the thumbnail media representation 712a.

[0249] As shown in Figure 7B, upon detection of a tap input 750a, the computer system 600 displays the media viewer user interface 720 and discontinues displaying the media gallery user interface 710. The media viewer user interface 720 includes a media viewer area 724 located between the application control area 722 and the application control area 726. The media viewer area 724 includes an enlarged representation 724a that represents the same media item as the thumbnail media representation 712a. The media viewer user interface 720 has virtually no controls superimposed on it, while the application control area 722 and the application control area 726 have virtually no controls superimposed on them.

[0250] The enlarged representation 724a includes a sign 642 containing text portion 642a (e.g., "LOST DOG") and text portion 642b (e.g., a paragraph of text beginning with "LOVEABLE"), as described above in relation to Figure 6B. The text in text portions 642a and 642b is not visually prominent, and the text in text portions 642a and 642b is small and cannot be easily read by a user viewing the computer system 600. Furthermore, the enlarged representation 724a includes a person 740 standing in front of a tree. Person 740 is wearing a hat containing the word "BRAND" (e.g., text portion 742).

[0251] The application control area 722 optionally includes an indicator of the time (e.g., "7:54" in Figure 7B) when the currently displayed enlarged media representation (e.g., enlarged representation 724a) was taken, a cellular signal status indicator 720a showing the status of the cellular signal, and a battery level status indicator 720b showing the remaining battery life of the computer system 600. The application control area 722 also includes a back control 722a (for example, when selected, causing the computer system 600 to redisplay the media gallery user interface 710) and an edit control 722b (for example, when selected, causing the computer system 600 to display a media editing user interface that includes one or more controls for editing the representation of the media item represented by the currently displayed enlarged representation 724a).

[0252] The application control area 726 includes a portion of the thumbnail media representations 712 (e.g., 712a-712c) displayed on a single line. Because the enlarged representation 724a is displayed in the media viewer area 724, the thumbnail media representation 712a is displayed as selected. In particular, the thumbnail media representation 712a is displayed as selected in Figure 7B by having space between it and the other thumbnails (e.g., 712b and 712c). The application control area 726 also includes a send control 726b (for example, to initiate the process of sending the media item represented by the enlarged media representation to the computer system 600 when selected), a favorites control 726c (for example, to mark / unmark the media item represented as a favorite media by the enlarged representation 724a to the computer system 600 when selected), and a trash can control 726d (for example, to delete (or initiate the process of deleting) the media item represented by the enlarged representation 724a to the computer system 600 when selected). In Figure 7B, the computer system 600 detects a pinch-out input 750b on the media viewer area 724 (for example, at a location on the computer system 600's display corresponding to the media viewer area 724, and / or directed to a location on the computer system 600's display corresponding to the media viewer area 710).

[0253] As shown in Figure 7C, in response to the detection of the pinch-release input 750b, the computer system 600 updates the enlarged representation 724a to reflect the change in zoom level, so that the display of the enlarged representation 724a in Figure 7C is displayed at a higher zoom level than the display of the enlarged representation 724a in Figure 7B. At the increased zoom level, the text portions 642a-642b in Figure 7C are larger and visually more prominent (e.g., larger, more readable) than the text portions 642a-642b in Figure 7B. In addition to updating the enlarged representation 724a, the computer system 600 also enlarges the media viewer area 724 in Figure 7B so that the enlarged representation 724a in Figure 7B occupies the portion of the display previously occupied by the application control areas 722 and 726 in Figure 7A.

[0254] In Figure 7C, it was determined (for example, using one or more similar techniques as described above in relation to Figures 6A to 6C) that the text in text section 642a and the text in text section 642b in Figure 7C do not satisfy the set of prominence criteria, respectively. Therefore, the computer system 600 does not display the parentheses corresponding to text sections 642a to 642b in Figure 7B (for example, the parentheses enclosing text sections 642a to 642b). Furthermore, because the text in text sections 642a to 642b does not satisfy the set of prominence criteria (for example, as described above in relation to Figure 6B), the computer system 600 does not display the text management control 680.

[0255] In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of the text portions 642a to 642b contain text that occupies a predetermined amount of space (e.g., 10% to 100%) of the enlarged representation 724a. In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of the portions 642a to 642b contain text that is located within or near a predetermined location (e.g., the central location) of the enlarged representation 724a. In some embodiments (e.g., as described above in relation to Figures 6M to 6T), the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of the text portions 642a to 642b contain text of a particular type (e.g., email, telephone number, address, QR code, etc.). In some embodiments, the set of prominence criteria includes criteria that are satisfied when it is determined that one or more of the text portions 642a to 642b contain text relevant to the context of the expanded expression 724a (for example, the text meets a relevance threshold (e.g., the computer system 600 determined that the text is 90%, 95%, and 99% relevant)).

[0256] In Figure 7C, it was determined that the main subject of the enlarged representation 724a is the sign 642. That is, the context of the enlarged representation 724a is the content displayed within the sign 642. In Figure 7C, the text portion 742 (e.g., "BRAND") appears on the hat of person 740 and is therefore not related to the context of what is displayed in the enlarged representation 724a, and thus is further determined to be irrelevant. In some embodiments, the text portion 742 appears on a person or something on a person in the enlarged representation 724a, so the computer system 600 determines that the text portion 742 is not related to the context of what is displayed in the enlarged representation 724a.

[0257] Because text portion 742 was determined to be irrelevant, it was determined that text portion 742 does not satisfy the set of prominence criteria. In particular, even though text portion 742 has more text than text portions 642a to 642b, it was determined that text portion 742 does not satisfy the set of prominence criteria. As shown in Figure 7C, because text portion 742 does not satisfy the set of prominence criteria (for example, because it was determined that text portion 742 is not relevant to the context of the enlarged expression 724a), the computer system 600 does not display one or more parentheses around text portion 742 ("BRAND"). In Figure 7C, the computer system 600 detects a tap input 750c on text portion 742.

[0258] As shown in Figure 7D, in response to the detection of tap input 750c, the computer system 600 maintains the display of the enlarged representation 724a, as shown in Figure 7C. In Figure 7D, because it was determined that text portion 742 does not satisfy the set of prominence criteria (for example, as described above in relation to Figure 7C), the computer system 600 does not update the display of the enlarged representation 724a to indicate that text portion 742 is selected. Also, because the text management control is not displayed and selected, the computer system 600 does not update the display of the enlarged representation 724a to indicate that text portion 742 is selected (for example, in contrast to the computer system 600 updating the media representation in Figures 6J to 6L as described above). Also, because the computer system 600 does not update the display of the enlarged representation 724a in Figure 7D, text portions 642a to 642b continue to not satisfy the set of prominence criteria. Therefore, as shown in Figure 7D, the computer system 600 does not display parentheses corresponding to the text portion 642a or 642b. In Figure 7D, the computer system 600 detects a pinch-out input 750d in the media viewer area 724. In some embodiments, instead of a pinch-out input 750d, the computer system 600 detects a directional swipe corresponding to a request to pan (e.g., translate) the enlarged representation 724a shown in Figure 7D.

[0259] As shown in Figure 7E, in response to the detection of the pinch-release input 750d, the computer system 600 updates the enlarged representation 724a to reflect the change in zoom level, such that the enlarged representation 724a in Figure 7E is displayed at a greater zoom level than the enlarged representation 724a in Figure 7D. In Figure 7E, it is determined that the text in text portion 642a satisfies the set of prominence criteria, but the text in text portion 642b does not. As a result, the computer system 600 displays parentheses 736a at the location corresponding to the location of text portion 642a (for example, the location surrounding text portion 642a). However, the computer system 600 does not display parentheses 736a or any other parentheses at the location corresponding to the location of text portion 642b (for example, because the text in text portion 642b does not satisfy the set of prominence criteria). In particular, even if text portion 742 (e.g., "BRAND") has larger text than text portions 642a-642b (for example, due to the absence of text portion 742), it was determined that text portion 742 still does not satisfy the set of prominence criteria. In some embodiments, when computer system 600 detects a directional swipe instead of a pinch-out input 750d, computer system 600 pans the enlarged representation 724a so that different portions of the enlarged representation 724a are displayed in response to the reception of the pinch-out input 750d.

[0260] As shown in Figure 7E, the computer system 600 displays the text management control 680 because it was determined that the text in text section 642a satisfies the set of prominence criteria. The text management control 680 is displayed in an inactive state (as indicated by the fact that the text management control 680 was not made bold) because it is not selected (for example, no input directed to the text management control was detected). In Figure 7E, the computer system 600 detects a pinch-out input 750e in the media viewer area 724.

[0261] As shown in Figure 7F, in response to the detection of the pinch-release input 750e, the computer system 600 updates the display of the enlarged representation 724a to reflect the change in zoom level, so that the display of the enlarged representation 724a in Figure 7F is displayed at a larger zoom level than the display of the enlarged representation 724a in Figure 7E. In Figure 7F, it was determined that the text of text portion 642a satisfies the set of prominence criteria, and the text of text portion 642b satisfies the set of prominence criteria. Therefore, the parentheses 636a are displayed around the entirety of both text portions 642a-642b, as described above in relation to Figure 6C. In particular, in Figure 7F, it was determined that text portion 742 (e.g., "BRAND") still does not satisfy the set of prominence criteria, even though it has larger text than text portions 642a-642b (for example, due to the absence of text portion 742). As shown in Figure 7F, the computer system 600 displays a text type indicator 638a below "123 MAIN STREET" (for example, using one or more techniques as described above in relation to Figure 6C) to indicate that an address has been detected, and a text type indicator 638b below "123-4567" to indicate that a telephone number has been detected. In some embodiments, the computer system 600 displays multiple brackets (for example, using one or more techniques as described above in relation to Figures 6A-6M), including one bracket around text portion 642a and another bracket around text portion 642b, and / or other combinations of brackets. In some embodiments (for example, referring again to Figure 7E), the computer system 600 displays a text type indicator below a text portion, regardless of whether the text portion to which the text type indicator belongs satisfies a set of prominence criteria. In Figure 7F, the computer system 600 detects a pinch-out input 750f in the media viewer area 724.

[0262] As shown in Figure 7G, in response to the detection of a pinch-release input 750f, the computer system 600 updates the magnified representation 724a to reflect the change in zoom level, such that the magnified representation 724a in Figure 7G is displayed at a greater zoom level than the magnified representation 724a in Figure 7F. In some embodiments, as shown in Figure 7G, the input displaying the magnified representation 724a corresponds to a directional swipe input.

[0263] As shown in Figure 7G, the enlarged representation 724a includes a subset of text portion 642a and a subset of text portion 642b. As a result of determining that the entire text portions 642a and 642b no longer satisfy the set of prominence criteria (for example, and / or that the enlarged representation 724a includes only subsets of text portions 642a and 642b), the computer system 600 discontinues displaying parentheses 636a around the entire text portions 642a and 642b. In Figure 7G, it was determined that a subset of the text in text portion 642b (for example, the telephone number "123-4567") satisfies the set of prominence criteria (for example, that another subset of the text in text portion 642b does not satisfy the criteria). In some embodiments, based on input previously detected by the computer system 600 (for example, looking at Figures 7A to 7G, the computer system 600 continues to zoom in near the telephone number), it was determined that the user intended to interact with or view the telephone number, and therefore it was determined that a subset of the text portion 642b satisfies the set of prominence criteria.

[0264] In some embodiments, it is determined that Figure 7G contains a subset of text portion 642a (e.g., "DOG") that satisfies a set of prominence criteria. In response to this determination, the computer system 600 displays a set of parentheses around the subset of text portion 642a, along with the parentheses 736c.

[0265] As shown in Figure 7G, the computer system 600 displays only a portion of the address "123 MAIN STREET". As a result, the computer system 600 stops displaying the text type indication 638a. In some embodiments, the computer system 600 maintains the display of the text type indication 638a below the portion of the address "123 MAIN STREET" displayed in Figure 7G. In some embodiments, since the rest of the address is not displayed, the computer system 600 determines that the portion of the address does not meet the set of prominence criteria. In Figure 7G, the computer system 600 detects a rightward swipe 750g in the media viewer area 724.

[0266] As shown in Figure 7H, in response to the detection of a rightward swipe 750g, the computer system 600 pans the enlarged representation 724a to the right. The enlarged representation 724a is panned so that the computer system 600 discontinues displaying the rightmost portion of the text portions 642a-642b shown in Figure 7G, and the computer system 600 redisplays the leftmost portion of the text portions 642a-642b in Figure 7H. As shown in Figure 7H, the computer system 600 does not display the entire telephone number (e.g., 123-4567) and discontinues displaying the parentheses 736c and the text type indication 638b. In some embodiments, a portion of the text type indication 638b remains displayed below the portion of the telephone number (e.g., "12") that continues to be displayed in Figure 7H. In Figure 7H, the computer system 600 displays more of the address (e.g., 123 MAIN STREET) in Figure 7H and redisplays the text type indication 638a below "123 MAIN STREET" to indicate to the user that the address has been detected.

[0267] In Figure 7H, it was determined that another subset of text portion 642b (e.g., "$1000 REWARD") satisfies the set of prominence criteria (for example, without any other subset of text portion 642b satisfying the set of prominence criteria). Because it was determined that another subset of text portion 642b satisfies the set of prominence criteria, the computer system 600 displays parentheses 736d around the other subset of text portion 642b, "$1000 REWARD". In some embodiments, based on the context of the displayed content of the enlarged representation 724a, it was determined that the other subset of text portion 642b (e.g., "$1000 REWARD") is the most relevant text displayed. In some embodiments, it was determined that Figure 7H includes a subset of text portion 642a that satisfies the set of prominence criteria (e.g., "LOST"). In some embodiments, in response to this determination, the computer system 600 displays separate sets of parentheses around a subset of the text portion 642a, simultaneously with the parentheses 736e. In Figure 7H, the computer system 600 detects a tap input 750h on the text management control 680.

[0268] As shown in Figure 7I, upon detection of tap input 750h, the computer system 600 displays a text management option 682, which includes a copy option 682a (for example, causing the computer system 600 to copy the text enclosed in parentheses 736d when selected), a select all option 682b (for example, causing the computer system 600 to select all the text enclosed in parentheses 736d when selected), a search option 682c (for example, causing the computer system to search for the text enclosed in parentheses 736d via a search (e.g., web search, dictionary search) when selected), and a share option 682d (for example, causing the computer system 600 to initiate a process to share the text enclosed in parentheses 736d when selected). In some embodiments, the various components of the text management option 682 function as described above in relation to Figures 6A-6M. In some embodiments, the computer system 600 displays multiple text management options, where each text management option corresponds to a separate portion of a separate pair of parenthesized texts. In some embodiments, the selection of individual text management options allows users to manage portions of text that correspond to those individual text management options.

[0269] As shown in Figure 7I, the computer system 600 displays the text management control 680 as enabled (for example, as indicated by the bolding of the text management control 680). In Figure 7I, the computer system 600 detects a tap input 750i on the text management control 680.

[0270] As shown in Figure 7J, in response to the detection of a tap input 750i, the computer system 600 redisplays the enlarged representation 724a using one or more techniques as described above in relation to Figure 7H. In Figure 7J, the computer system 600 detects a downward swipe input 750j within the media viewer area 724.

[0271] As shown in Figure 7K, upon detection of a downward swipe input 750j, the computer system 600 pans the media viewer area 724 downward (for example, based on the swipe input) so that the display of text portion 642b is stopped and the computer system displays only a subset of text portion 642a. In Figure 7K, it is determined that the subset of text portion 642a (for example, "LOST") satisfies the set of prominence criteria. Because the subset of text portion 642a satisfies the set of prominence criteria, the computer system 600 displays parentheses 736e surrounding the subset of text portion 642a.

[0272] In Figure 7K, the computer system 600 does not display the parentheses 736d because the text portion 642b is not displayed as part of the enlarged representation 724a in Figure 7K. In Figure 7K, the computer system 600 detects a pinch input 750k within the media viewer area 724.

[0273] As shown in Figure 7L, in response to the detection of a pinch input 750k, the computer system 600 updates the enlarged representation 724a to reflect the change in zoom level (e.g., a decrease in zoom level) so that the display of the enlarged representation 724a in Figure 7L is displayed at a reduced zoom level compared to the display of the enlarged representation 724a in Figure 7K. In Figure 7L, it is determined that text portions 642a and 642b do not satisfy the set of prominence criteria. Therefore (e.g., because it is determined that text portions 642a and 642b do not satisfy the set of prominence criteria), the computer system 600 does not display (and / or stops displaying) the text management control 680 and / or any parentheses surrounding text portions 642a-642b.

[0274] The techniques described above in relation to Figures 7A to 7L were described in the context of a computer system 600 displaying a representation of previously captured media and a media viewer user interface. However, one or more of the techniques described above in relation to Figures 6A to 6Z can also be applied while the computer system 600 is displaying previously captured media and a media viewer user interface. Furthermore, the techniques described above in relation to Figures 7A to 7L can also be applied in the context of a computer system 600 displaying a live preview (e.g., a representation of the field of view of one or more cameras before the media is captured) and a camera user interface, such as the live preview 630 in Figures 6A to 6M.

[0275] The techniques described above in relation to Figures 6A to 6Z were described in the context of a computer system 600 displaying a live preview and a camera user interface, but one or more of the techniques described above in relation to Figures 7A to 7L can also be applied while the computer system 600 is displaying a live preview and a camera user interface. Furthermore, the techniques described in relation to Figures 6A to 6Z can also be applied in the context of a computer system 600 displaying previously captured media, such as an enlarged media representation 724a, and a media viewer user interface.

[0276] Figure 8 is a flowchart illustrating a method for managing visual content in media using a computer system according to several embodiments. Method 800 is performed in a computer system (e.g., 100, 300, 500) communicating with a display generation component. Some operations in Method 800 are arbitrarily combined, the order of some operations is arbitrarily changed, and some operations are arbitrarily omitted.

[0277] As described later, Method 800 provides an intuitive way to manage visual content within media. This method reduces the cognitive burden on the user managing visual content within media, thereby creating a more efficient human-machine interface. In the case of battery-powered computing devices, it saves power and extends the interval between battery charges by enabling users to manage visual content within media more quickly and efficiently.

[0278] Method 800 is performed in a computer system (e.g., 600) (e.g., a smartphone, desktop computer, laptop, tablet) that is communicating with a display generation component (e.g., a display control, a touch-sensitive display system). In some embodiments, the computer system is communicating with a first camera of one or more input devices (e.g., a touch-sensitive surface) and / or one or more cameras (e.g., one or more cameras on the same or different sides of the computer system (e.g., a dual camera, a triple camera, a quad camera, etc.) (e.g., a front camera, a back camera)).

[0279] The computer system displays camera user interfaces (e.g., media capture user interface, media browsing user interface, media editing user interface) via a display generation component (802), the camera user interfaces include simultaneously displaying representations (e.g., 630) of media (e.g., photographic media, video media) (e.g., live media; live preview (e.g., media corresponding to representations of one or more camera fields of view (e.g., current field of view) that have not been captured in response to detection of a request to capture media (e.g., detection of a selection of shutter affordances)); previously captured media (e.g., media corresponding to representations of one or more camera fields of view (e.g., previous field of view) that have been captured); saved media items that can be accessed by the user later; representations of media displayed in response to the reception of a gesture on a thumbnail representation of media (e.g., in a media gallery)) and media capture affordances (e.g., 610) (e.g., user interface objects).

[0280] While simultaneously displaying a media representation (e.g., 630) and a media capture affordance (e.g., 610) (e.g., a user interface object) (804), according to the determination that a separate set of criteria is met, the computer system displays, via a display generation component, a first user interface object (e.g., 680) that corresponds to one or more text management operations (e.g., simultaneously with the media representation) (e.g., in the user interface) (e.g., simultaneously with the representation of the media and / or the first user interface object) (804), the separate set of criteria described above is met when a separate text (e.g., 642a, 642b) (e.g., one or more characters represented in the media) is detected in the media representation (e.g., 630). In some embodiments, multiple options (e.g., 672, 682, 692) include one or more options for copying individual text (e.g., 682a), selecting individual text (e.g., 682b), searching for individual text (e.g., 682c), sharing individual text (e.g., 682d), and translating individual text.

[0281] While displaying a media representation (e.g., 630) and a media capture affordance (e.g., 610) (e.g., a user interface object) simultaneously (804), the computer system refrains from displaying the first user interface object (808) in accordance with the determination that a separate set of criteria is not met.

[0282] While displaying a media representation (e.g., 630) (e.g., while simultaneously displaying a media representation, media capture affordances, and a first user interface object), the computer system detects a first input (e.g., 650a, 650e, 650g, 650u) (e.g., mouse / trackpad click / activation, keyboard input, scroll wheel input, hover gesture, tap gesture, swipe gesture) directed to the camera user interface (e.g., 602, 604, 606) (810). In some embodiments, the first input is a non-tap gesture (e.g., a rotation gesture and / or a long-press gesture).

[0283] In response to the detection of a first input (650a, 650e, 650g, 650u) directed to the camera user interface (e.g., a first gesture) (812), and in accordance with the determination that the first input (e.g., 650a) corresponds to the selection of a media capture affordance (e.g., 610) (e.g., a gesture directed to a media capture affordance, a gesture in a location corresponding to a media capture affordance), the computer system begins capturing media to be added to a media library (e.g., 612) associated with the computer system (e.g., 600) (814), without displaying options for managing individual texts.

[0284] In response to the detection of a first input (650a, 650e, 650g, 650u) (e.g., a first gesture) directed to the camera user interface (812), and in accordance with the determination that the first input (e.g., 650e, 650g, 650u) corresponds to the selection of a first user interface object (e.g., 680), the computer system displays a number of options for managing individual texts (e.g., 672, 682, 692) via a display generation component (816) (e.g., without initiating the capture of media to be added to a media library (e.g., 624) associated with the computer system (e.g., 600). In some embodiments, the number of options are displayed adjacent to the individual texts (e.g., contained in the media representation). In some embodiments, the multiple options (e.g., 672, 682, 692) (as described above in relation to Figure 6F, for example) include one or more options for copying individual texts (e.g., 682a), selecting individual texts (e.g., 682b), searching for individual texts (e.g., 682c), sharing individual texts (e.g., 682d), and translating individual texts. In some embodiments, following the determination that a first input (e.g., 650e, 650g, 650u) corresponds to the selection of a first user interface object (e.g., 680), the first user interface object is in an active state (e.g., transitioned from an inactive state to an active state), and the first user interface object displayed in the active state (e.g., 680 in Figure 6F) (e.g., bolded and pressed / appeared) has a different appearance than when the first user interface object is displayed in an inactive state (e.g., 680 in Figure 6G) (e.g., not bolded and released / appeared).In some embodiments, following the determination that a first input (e.g., 650a) corresponds to the selection of a media capture affordance (e.g., 610), the first user interface object (e.g., 680) is in an inactive state (e.g., transitioning from an inactive state to an active state). In some embodiments, following the determination that a first input (e.g., 650e, 650g, 650u) corresponds to the selection of the first user interface object (e.g., 680) and the first user interface object is in an inactive state (e.g., 680 in Figure 6E), the computer system displays several options (e.g., 672, 682, 692) for managing individual texts. In some embodiments, (for example, as described above with reference to Figure 6F) the first input corresponds to the selection of a first user interface object (e.g., 680), and upon determination that the first user interface object is in an active state (e.g., 680 in Figure 6F), the computer system refrains from displaying several options for managing individual texts (e.g., 672, 682, 692). In some embodiments, upon determination that the first input (650a) corresponds to the selection of a media capture affordance, the first user interface object (e.g., 610) continues to be displayed. In some embodiments, upon determination that a first input (e.g., 650a) corresponds to the selection of a media capture affordance (e.g., 610) or a first user interface object (e.g., 680), one or more interface objects (e.g., media capture affordance (e.g., 610), camera setting affordance(s), camera mode affordance(s)(e.g., 620)) disappear from the camera user interface or are displayed as inactive (e.g., dimmed) (e.g., not responding to user input for individual objects).Displaying multiple options for managing individual texts based on whether a first input corresponds to a selection of a first user interface object allows the user to quickly and efficiently manage individual texts without cluttering the user interface with additional user interface objects. Providing additional system controls without cluttering the UI with additional displayed controls improves system usability (for example, by helping the user provide appropriate input when operating / interacting with the system and reducing user errors), makes the user-system interface more efficient, and reduces system power consumption and improves battery life by allowing the user to use the system more quickly and efficiently. Automatically providing the user with various options for different ways of managing individual texts by displaying multiple options for managing individual texts when certain conditions are met (for example, based on whether a first input corresponds to a selection of a first user interface object). By performing actions when a set of conditions is met without requiring further user input, the system's usability is enhanced, the user-system interface becomes more efficient (for example, by assisting the user in providing appropriate input when operating / interacting with the system and reducing user errors), and in addition, power consumption is reduced and battery life is improved by enabling users to use the system more quickly and efficiently.

[0285] In some embodiments, the first input (e.g., 650e, 650g, 650u) is a tap gesture (e.g., a tap gesture) directed towards a first user interface object (e.g., 672, 682, 692) (e.g., a gesture at a location corresponding to the first user interface object).

[0286] In some embodiments, the media representation (e.g., 630) includes individual text (for example, when individual text is displayed when the media representation is displayed). In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u) (and while not displaying an indication that text is selected, and / or after detecting an input / gesture corresponding to the selection of the first user interface object, and / or while the first user interface object is displayed as being in an active state, and / or while displaying multiple options for managing the individual text), the computer system detects a second input (e.g., 650j) directed to the camera user interface (e.g., a tap gesture and / or a swipe gesture). In some embodiments, the second input is a non-tap gesture (e.g., a rotate gesture and / or a long-press gesture). In some embodiments, the first input is a non-swipe gesture (e.g., a rotate gesture, a long press gesture, a mouse / trackpad click / enable, keyboard input, scroll wheel input, a hover gesture, and / or a tap gesture). In some embodiments, in response to the detection of a second input (e.g., 650j) directed towards the camera user interface, and according to the determination that the second input corresponds to the selection of one or more first parts of a particular text, the computer system displays an indication (e.g., 642b in Figure 6K) that one or more first parts (e.g., 642b) of a particular text (e.g., 642a, 642b) have been selected. In some embodiments, the indication is displayed around the particular text. In some embodiments, as part of displaying the indication that one or more parts of a particular text have been selected, the computer system highlights one or more parts of the particular text (e.g., highlight, underline, bold, enlarge).In some embodiments, while an indication is displayed that a first portion of a particular text is selected, the computer system does not display an indication that a second portion of the particular text (e.g., different from the first portion) is selected. In some embodiments, upon determination that a second input (650j) corresponds to the selection of one or more portions of a particular text (e.g., 642a), and while the first user interface (e.g., 680) object is displayed as being in an active state (e.g., 680 as described above in relation to Figure 6F), the computer system displays an indication that one or more first portions of a particular text (e.g., 642a) is selected (e.g., as described above in relation to Figures 6K and 6L). In some embodiments, upon determination that a second input (e.g., 650j) corresponds to the selection of one or more parts (e.g., 642) of individual text, and while the first user interface object (e.g., 680) is displayed as inactive (e.g., as described above in relation to Figure 6G) and / or not displayed (e.g., as described above in relation to Figures 7C, 7G), the computer system does not display an indication that one or more first parts (e.g., 642a) of individual text are selected (e.g., it refrains from displaying it) (e.g., as discussed above in relation to Figures 7C-7G). Displaying an indication that one or more first parts of individual text are selected provides the user with visual feedback on whether text is selected and which text is currently selected.Providing users with improved visual feedback enhances the usability of the computer system, makes the computer system interface more efficient, and reduces power consumption and battery life by enabling users to use the computer system more quickly and efficiently. Providing users with additional control to select text without cluttering the user interface with additional user interface objects is achieved by displaying an indication that one or more parts of a particular text are selected, in response to the detection of a second input and in accordance with the determination that the second input corresponds to the selection of one or more parts of a particular text. Providing additional control of the system without cluttering the UI with additional displayed controls enhances the usability of the system, makes the user-system interface more efficient, and reduces power consumption and battery life by enabling users to use the system more quickly and efficiently.

[0287] In some embodiments, the second input (e.g., 650j) (e.g., second gesture) is a tap gesture (e.g., directed towards one or more parts of individual text) or a swipe gesture (e.g., directed towards one or more parts of individual text). In some embodiments, the first input is a first type of input, and the second input is a second type of input that is different from the first type of input.

[0288] In some embodiments, upon detection of a first input (e.g., 650e, 650g, 650u) directed to the camera user interface, and in accordance with the determination that the first input (e.g., 650e, 650g, 650u) corresponds to the selection of a first user interface object (e.g., 680), the computer system displays an indication (e.g., 684) (e.g., an instruction) (e.g., an instruction indicating one or more inputs that cause the computer system to display the text as selected) regarding the selection (e.g., how to select) of text contained in a media representation (e.g., 630). In some embodiments, upon detection of a first input directed to the camera user interface, and in accordance with the determination that the first input corresponds to the selection of a media capture affordance, the computer system does not display an indication (e.g., how to select) of text contained in a media representation. In some embodiments, an indication (e.g., 684) regarding the selection of text contained in a media representation (e.g., how to select it) is displayed simultaneously with several options (e.g., 682a, 682b, 682c, 682d) for managing individual texts (e.g., 642b). In some embodiments, the text selection indication (e.g., 684) is displayed when the first user interface object (e.g., 680) is active (e.g., 680 as described above in relation to Figure 6F), and the text selection indication (e.g., 684) is not displayed when the first user interface object is inactive (e.g., 680 as described above in relation to Figure 6G). By displaying an indication of how to select the text contained in a media representation, the user is provided with visual feedback regarding the steps required to select the text they wish to select.Providing users with improved visual feedback enhances the usability of the computer system, makes the computer system interface more efficient (for example, by helping users provide appropriate input when operating / interacting with the computer system and reducing user errors), reduces the power consumption of the computer system, and improves battery life by enabling users to use the computer system more quickly and efficiently.

[0289] In some embodiments, before detecting a first input (e.g., 650e, 650g, 650u), the media representation (e.g., 630) is displayed with a first appearance (e.g., 630 in Figure 6E) (e.g., a first blur value, a first dim value). In some embodiments, upon detection of a first input (e.g., 650e, 650g, 650u) directed towards the camera user interface, and in accordance with the determination that the first input (e.g., 650e, 650g, 650u) corresponds to the selection of a first user interface object (e.g., 680), the computer system displays the media representation (e.g., 630) with a second appearance (e.g., 630 in Figure 6F) (e.g., a second blur value, a second dim value) that differs from the first appearance (e.g., 630 in Figure 6E). In some embodiments, as part of displaying a media representation in a second appearance different from a first appearance, the computer system blurs and / or darkens at least a portion of the media representation. In some embodiments, in response to detection of a first input directed to the camera user interface and according to a determination that the first input corresponds to a selection of media capture affordances, the computer system displays a media representation in a third appearance different from a second appearance. In some embodiments, the third appearance is the first appearance. In some embodiments, the third appearance (e.g., black, solid) is different from the first appearance (e.g., a blurred version of the field of view of one or more cameras). In some embodiments, the media representation having the third appearance is displayed for a predetermined period (e.g., less than one second) that is not based on whether the first user interface object is displayed in an active state. In some embodiments, the media representation having the second appearance is displayed when the first user interface object is displayed in an active state and not displayed when the first user interface object is displayed in an inactive state.In some embodiments, a representation of media having a first appearance is not displayed when the first user interface object is displayed in an active state, but is displayed when the first user interface object is displayed in an inactive state. Displaying a representation of media having a second appearance different from the first appearance of the representation in response to the detection of a first input provides the user with visual feedback on whether the text has been selected by the user by de-emphasizing less relevant parts of the representation of media. Providing the user with improved visual feedback improves the usability of the computer system, makes the computer system interface more efficient (for example, by helping the user provide appropriate input when operating / interacting with the computer system and reducing user errors), reduces the power consumption of the computer system and improves battery life by enabling the user to use the computer system more quickly and efficiently.

[0290] In some embodiments, the media representation (e.g., 630) includes individual text (e.g., 642a, 642b) (for example, when individual text is displayed when the media representation is displayed). In some embodiments, according to the determination that a distinct set of criteria is met, the computer system highlights one or more second portions of the individual text (e.g., highlight, display an object (e.g., shape, surrounding brackets (e.g., yellow brackets)), underline, enlarge). In some embodiments, according to the determination that a distinct set of criteria is met, the computer system highlights one or more second portions of the individual text without highlighting other portions of the individual text and / or other portions of the media representation that do not include one or more second portions of the individual text. By highlighting one or more second portions of the individual text, the computer system provides the user with improved visual feedback regarding whether a particular portion of the individual text contained in the media meets a distinct set of criteria. Providing users with improved visual feedback enhances the usability of the computer system, makes the computer system interface more efficient (for example, by helping users provide appropriate input when operating / interacting with the computer system and reducing user errors), reduces the power consumption of the computer system, and improves battery life by enabling users to use the computer system more quickly and efficiently.

[0291] In some embodiments, as part of highlighting one or more second portions of individual text, the computer system displays an indication (e.g., 636a, 636b, 736c, 736d) that the individual text has been detected. In some embodiments, while one or more second portions of individual text (e.g., 642a, 642b) are highlighted, the computer system receives a request to display a second representation (e.g., 630 in Figure 6F) of media (e.g., the same or different media as the media represented by the representation of media). In some embodiments, a request to display a second representation of media is detected when one or more changes in the field of view of one or more cameras communicating with the computer system are detected. In some embodiments, a request to display a second representation of media is detected when a request to zoom out / zoom in and / or pan the representation of media is detected. In some embodiments, a request to display a second representation of media is detected when the computer system is moved.

[0292] In some embodiments, upon receiving a request to display a second representation of the media (e.g., 630 in Figure 6F) (including, for example, a portion of a particular text and / or a second particular text different from the particular text), the computer system translates (e.g., moves) the indication that the particular text has been detected (e.g., 636a, 636b, 736c, 736d) from a first position in the camera user interface to a second position in the camera user interface. In some embodiments, upon receiving a request to display the second representation of the media, the indication that the particular text has been selected is modified to enclose a portion of the text different from the portion of the text that was enclosed before the request to display the second representation of the media was received. By translating the indication that the particular text has been detected from a first position in the camera user interface to a second position in the camera user interface upon receiving a request to display the second representation of the media, the user can maintain viewing of the indication while the system moves between the first and second positions. By performing actions when a set of conditions is met without requiring further user input, the system's usability is enhanced, the user-system interface becomes more efficient (for example, by assisting the user in providing appropriate input when operating / interacting with the system and reducing user errors), and in addition, power consumption is reduced and battery life is improved by enabling users to use the system more quickly and efficiently.

[0293] In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u), and in accordance with a determination (first determination) that the first input (e.g., 650e, 650g, 650u) corresponds to a selection of a first user interface object (e.g., 680), (and / or while the first user interface object is displayed as being in an active state and / or while displaying multiple options that manage individual texts), a representation of the media (e.g., 630) includes individual texts (e.g., 642a, 642b) and an indication that one or more third portions of the individual texts (e.g., 642a, 642b) have been selected. In some embodiments, the computer system receives a request to display a third representation (e.g., 630) of the media (e.g., the same or different media as the media represented by the representation of the media). In some embodiments, a request to display a third representation of the media is detected when one or more changes in the field of view of one or more cameras communicating with the computer system are detected. In some embodiments, a request to display a third representation of the media is detected when a request to zoom out / zoom in and / or pan the representation of the media is detected. In some embodiments, a request to display a third representation of the media is detected when the computer system is moved. In some embodiments, in response to the receipt of a request (e.g., 650c, 650d, 750e, 750f, 750g) to display a third representation of the media (e.g., 630), the computer system displays an indication (e.g., 636a, 636b, 736c, 736d) that at least a portion of the text contained in the third representation of the media (e.g., 630) has been selected, which is different from an indication (e.g., 636a, 636b, 736c, 736d) that at least a portion of the text contained in the third representation of the media (e.g., 642a, 642b) has been selected, where one or more third portions of the individual texts (e.g., 642a, 642b) have been selected.In some embodiments, a portion of text contained in a third representation of the media includes at least a portion of text within one or more third portions of the text. Upon receiving a request to display the third representation, an indication is displayed that at least a portion of the text contained in the third representation of the media has been selected, providing the user with an additional efficient way to control which portion of the text is selected without disrupting the user interface. Reducing the number of inputs required to perform an action improves the usability of the computer system (for example, by helping the user provide appropriate input when operating / interacting with the computer system, thereby reducing user errors), makes the user-system interface more efficient, and, in addition, reduces the power consumption of the computer system and improves battery life by enabling the user to use the computer system more quickly and efficiently.

[0294] In some embodiments, after detecting a first input (e.g., 650e, 650g, 650u) and in accordance with a determination (e.g., a first determination) that the first input corresponds to a selection of a first user interface object (and / or while the first user interface object is displayed as being in an active state and / or while displaying multiple options for managing individual texts), a media representation (e.g., 630) includes individual text (e.g., 642b) and an indication that one or more fourth portions of individual text (e.g., 642b) have been selected and that one or more fourth portions of individual text (e.g., 642b) are displayed in a third position within the camera user interface (and / or on the display). In some embodiments, a computer system detects changes in the physical environment within the field of view of one or more cameras communicating with the computer system. In some embodiments, in response to the detection of a change in the physical environment within the field of view of one or more cameras (e.g., 660a, 660b), the computer system continues to display one or more fourth portions of the individual text (e.g., 642b) at a third position within the camera user interface (and / or on the display). In some embodiments, the selected text is frozen. In some embodiments, at least a portion of the fourth representation of the media is displayed (e.g., newly displayed in response to the detection of a change in the physical environment) while maintaining the display of one or more fourth portions of the individual text. In some embodiments, the computer system freezes the selected text (e.g., and / or displays the selected text in the same location and / or at the same size) while updating the representation of the media (e.g., live preview) to reflect the change in the physical environment. By continuing to display one or more fourth portions of the individual text at a third position within the camera user interface, the user can maintain viewing of the text selected by the user while the system moves between the first and second points.Performing an action when a set of conditions is met without requiring further user input improves system usability (for example, by helping the user provide appropriate input when operating / interacting with the system, thereby reducing user errors), makes the user-system interface more efficient, and reduces system power consumption and improves battery life by allowing users to use the system more quickly and efficiently.

[0295] In some embodiments, before detecting a first input (e.g., 650e, 650g, 650u) directed to the camera user interface, the computer system (e.g., 600) communicates with one or more cameras, and the representation of the media (e.g., 630) is a representation (e.g., 630) of one or more objects in the physical environment (e.g., physical space) within the field of view of one or more cameras (e.g., live camera preview). In some embodiments, receiving a request to display a fourth representation of the media (e.g., a representation of the updated field of view of the cameras) includes detecting a change in the field of view of the cameras. In some embodiments, the fourth representation of the media includes the change in the field of view of the cameras. In some embodiments, when one or more objects (e.g., non-text objects) in the field of view are moving, the representation of the media is updated to indicate that one or more objects are moving. In some embodiments, the representation of the media is a live representation of the field of view of the cameras. By displaying a media representation that is a representation of one or more objects in physical space within the field of view of one or more cameras (e.g., a live camera preview), the user is given greater control over the computer system (e.g., changing the camera field of view of the system) to determine whether one or more objects in physical space can be captured without cluttering the user interface. Providing additional control of the system without cluttering the UI with additional displayed controls improves the usability of the system (e.g., by helping the user provide appropriate input when operating / interacting with the system and reducing user errors), makes the user-system interface more efficient, and in addition reduces the system's power consumption and improves battery life by allowing the user to use the system more quickly and efficiently.

[0296] In some embodiments, a representation of the media (e.g., 630) is a first representation of the media. In some embodiments, while displaying a first user interface object, the computer system detects a request (e.g., 750k) to display a fifth representation of the media (e.g., the same or different media as the first representation of the media). In some embodiments, a request to display a fifth representation of the media is detected when one or more changes in the field of view of one or more cameras communicating with the computer system are detected. In some embodiments, a request to display a fifth representation of the media is detected when a request to zoom out / zoom in and / or pan the representation of the media is detected. In some embodiments, a request to display a fifth representation of the media is detected when the computer system is moved. In some embodiments, upon detection of a request to display a fifth representation of the media (e.g., 750k) and in accordance with a determination that a particular set of criteria is not met (e.g., no individual text is detected in the fifth representation of the media, or individual text is detected but not sufficiently prominent), the computer system discontinues displaying the first user interface object (e.g., 680). In some embodiments, upon detection of a request to display a fifth representation of the media and in accordance with a determination that individual text is detected within the fifth representation of the media, the computer system continues to display the first user interface object. By discontinuing the display of the first user interface object when certain conditions are met (e.g., upon detection of a request to display a fifth representation of the media and in accordance with a determination that a particular set of criteria is not met), the computer system automatically provides the user with an indication of whether or not the representation of the media does not contain text detected by the computer system.By performing actions when a set of conditions is met without requiring further user input, the system's usability is enhanced, the user-system interface becomes more efficient (for example, by assisting the user in providing appropriate input when operating / interacting with the system and reducing user errors), and in addition, power consumption is reduced and battery life is improved by enabling users to use the system more quickly and efficiently.

[0297] In some embodiments, each criterion is met when it is determined that an individual text satisfies a predetermined prominence criterion (e.g., the text is of a size or location within the media representation that indicates the text is important and / or relevant) (determined to be relevant based on one or more techniques as described below in relation to Figures 7C, 7E-7J and 9) (e.g., the individual text indicates a first user interface object when it is on a sign, but does not indicate a first user object when it is detected on clothing).

[0298] In some embodiments, while displaying a media representation (e.g., 630) (and, in some embodiments, after detecting input corresponding to a selection of a first user interface object, and / or while the first user interface object is displayed as being in an active state, and / or while displaying multiple options for managing individual texts), and in accordance with a determination that individual texts (e.g., 642a-642b) contain a portion of text that has been determined to be an individual type of text (e.g., a phone number, an email address) (e.g., based on one or more regular expression patterns corresponding to different types of text), the computer system displays an indication (e.g., 638a-638b) (e.g., a data detector indication) that an individual type of text has been detected. In some embodiments, as part of displaying the indication that an individual type of text has been detected, the computer system highlights the portion of text (e.g., highlights, underlines, or puts parentheses around it). In some embodiments, the indication that an individual type of text has been detected is displayed adjacent to, around, or otherwise adjacent to the portion of text that belongs to an individual type of text. In some embodiments, if a computer system determines that an individual piece of text does not contain a portion of text belonging to a specific type of text (e.g., a phone number, an email address), it may not display an indication that a specific type of text has been detected (e.g., it may omit the display). By displaying an indication that a specific type of text has been detected within a media representation, the system provides the user with visual feedback on whether or not a media representation contains a certain type of text.By providing users with improved visual feedback, the usability of the computer system is enhanced, the user-system interface is made more efficient (for example, by assisting users in providing appropriate input when operating / interacting with the computer system and reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling users to use the computer system more quickly and efficiently.

[0299] In some embodiments, while displaying multiple options for managing individual text (e.g., 680), the computer system receives a third input (e.g., 650h) (e.g., tap input) directed to a portion of the camera user interface that does not contain individual text (e.g., a darkened or blurred portion of the media representation (e.g., a portion of the media representation that does not contain text) (and / or a darkened portion of the camera user interface)). In some embodiments, upon receiving the third input (e.g., 650h), the computer system ceases displaying the multiple options for managing individual text (e.g., 680). In some embodiments, upon receiving the third input, one or more interface objects (e.g., media capture affordances, camera setting affordances, camera mode affordances) are displayed (e.g., redisplayed) and / or displayed as active (e.g., not darkened) within the camera user interface (e.g., in response to user input on individual objects). By discontinuing the display of multiple options for managing individual texts in response to input directed towards a portion of the camera user interface, more control over the system is provided to the user without cluttering the user interface with additional user interface objects. Providing additional control over the system without cluttering the UI with additional displayed controls improves system usability (for example, by helping the user provide appropriate input when operating / interacting with the system and reducing user errors), makes the user-system interface more efficient, and in addition reduces system power consumption and improves battery life by allowing the user to use the system more quickly and efficiently.

[0300] In some embodiments, while simultaneously displaying a media representation (e.g., 630) and a media capture affordance (e.g., 610) (e.g., before displaying a first user interface object), if the computer system determines that the media representation (e.g., 630) includes a first machine-readable code (e.g., a linear barcode, a matrix barcode, or a QR code), it displays a first user interface object (e.g., 680) and a representation of a uniform resource locator (e.g., 668) corresponding to the first machine-readable code. Displaying the first user interface object and the representation of the uniform resource location improves security by informing the user of the location of the resource corresponding to the QR code before providing input to navigate to the resource. By providing improved security, the user interface is made more secure, unauthorized execution of secure operations is reduced, and in addition, power consumption of the computer system is reduced and battery life is improved by enabling the user to use the computer system more safely and efficiently. Displaying a first user interface object and a representation of a uniform resource location when certain conditions are met (for example, according to a determination that the media representation contains machine-readable code) notifies the user of the resources associated with the machine-readable code before the user selects the machine-readable code and provides the user with a uniform resource locator corresponding to the first machine-readable code. Performing an action when a set of conditions is met without requiring further user input improves the usability of the system, makes the user-system interface more efficient (for example, by helping the user provide appropriate input when operating / interacting with the system and reducing user errors), and reduces system power consumption and improves battery life by allowing the user to use the system more quickly and efficiently.

[0301] In some embodiments, while a media representation (e.g., 630) contains a second machine-readable code (and while the machine-readable code is selected), according to the determination that a first input (e.g., 650u) corresponds to a selection of a first user interface object, a plurality of options for managing individual texts (e.g., 672) includes one or more options for managing information corresponding to the second machine-readable code (e.g., uniform resource location). In some embodiments, while a media representation does not contain a machine-readable code (and / or while the machine-readable code is not selected), according to the determination that a first input corresponds to a selection of a first user interface object, a plurality of options for managing individual texts does not include one or more options for managing information. In some embodiments, one or more of the plurality of options for managing individual texts displayed when a machine-readable code is selected differ from one or more options for managing individual texts displayed when text without a machine-readable code is selected. Based on the determination that the first input corresponds to a selection in the first user interface, more control options (e.g., additional text management options) are provided to the user without cluttering the user interface by including one or more options for managing information corresponding to machine-readable code within multiple options. Providing additional system controls without cluttering the UI with additional displayed controls improves system usability (e.g., by helping the user provide appropriate input when operating / interacting with the system and reducing user errors), makes the user-system interface more efficient, and reduces system power consumption and improves battery life by allowing the user to use the system more quickly and efficiently.

[0302] In some embodiments, the camera user interface includes a selection of multiple camera setting affordances (e.g., 620a-620e) that modify the settings of one or more cameras (e.g., flash affordance, timer affordance, filter effect affordance, f-stop affordance, aspect ratio affordance, live photo affordance, etc.) (e.g., multiple user interface objects for accessing individual camera settings). In some embodiments, the camera user interface includes a selection of multiple camera mode affordances (e.g., 620) (e.g., multiple user interface objects for setting individual camera modes). In some embodiments, the selection of multiple camera setting affordances (e.g., 602a, 602b) is displayed simultaneously with media capture affordances (e.g., 610) and / or multiple camera mode affordances (e.g., 620). In some embodiments, each camera mode (e.g., video (e.g., 620b), photo (e.g., 620c), portrait (e.g., 620d), slow motion (e.g., 620a), and panorama (e.g., 620e)) (e.g., 620) has multiple settings (e.g., for portrait camera mode: studio lighting setting, contour lighting setting, and stage lighting setting) with multiple values ​​(e.g., light level for each setting) of modes (e.g., portrait mode) in which the camera (e.g., camera sensor) operates to capture media (including post-processing that is automatically performed after capture). In this way, for example, camera modes differ from modes that do not affect how the camera operates when capturing media or do not have multiple settings (e.g., flash mode, which has one setting with multiple values ​​(e.g., inactive, active, auto)).In some embodiments, camera modes allow the user to capture different types of media (...

Claims

1. It is a method, In a computer system that communicates with a display generation component and one or more input devices, The first representation of a media item containing a portion of text is displayed via the aforementioned display generation component, While displaying the first representation of a media item containing a portion of the aforementioned text, the system detects input via one or more input devices that corresponds to a request to display a second representation of the media item. In response to the detection of the input corresponding to the request to display the second representation of the media item, the display generation component displays the second representation of the media item, wherein the second representation of the media item includes the portion of the text in a certain location. While the second representation of the media item is being displayed, In accordance with the determination that the portion of text contained in the second representation of the media item satisfies a separate set of criteria, the display generation component displays a visual indication that highlights the portion of text corresponding to the portion of text contained in the second representation that was not displayed when the first representation of the media item was displayed, the visual indication being displayed in the location described above. Each set of the aforementioned criteria includes a criterion that is satisfied when it is determined that the amount of prominence in the portion of the text contained in the second representation of the media item exceeds a prominence threshold. The portion of the text included in the first representation of the media item is below the prominence threshold. method.

2. While the second representation of the media item is being displayed, The method according to claim 1, further comprising refraining from displaying the visual indication in accordance with the determination that a portion of the text contained in the second representation of the media item does not satisfy a particular set of criteria.

3. While the second representation of the media item is being displayed, an input corresponding to a request to display a third representation of the media item is detected via one or more input devices. In response to the detection of the input corresponding to the request to display the third representation of the media item, the third representation of the media item is displayed via the display generation component, While the third representation of the media item is being displayed, In accordance with the determination that a portion of the text contained in the third representation of the media item satisfies a separate set of the criteria, the display generation component displays a visual indication corresponding to the portion of the text contained in the third representation. The method according to claim 1, further comprising:

4. The method according to claim 1, wherein the first representation of the media item is a representation of the media item displayed at a first zoom level, and the second representation of the media item is a representation of the media item displayed at a second zoom level different from the first zoom level.

5. The method according to claim 1, wherein the first representation of the media item is a representation of the media item displayed with a first translation amount, and the second representation of the media item is a representation of the media item displayed with a second translation amount different from the first translation amount.

6. The method according to claim 1, wherein the input corresponding to the request to display the second representation of the media item is an input detected on the display generation component.

7. The second representation of the media item is displayed at a third zoom level, and the method is While the second representation of the media item is displayed at the third zoom level, the system detects inputs via one or more input devices that correspond to a request to change the zoom level of the second representation of the media item. In response to the detection of the input corresponding to the request to change the zoom level of the second representation of the media item, the display generation component displays the fourth representation of the media item at a fourth zoom level different from the third zoom level. While the fourth representation of the media item is displayed at the fourth zoom level, In accordance with the determination that the portion of text included in the fourth representation of the media item does not satisfy the individual set of criteria, the display of the visual indication is withheld. The method according to claim 1, further comprising:

8. While the second representation of the media item is being displayed, an input corresponding to a request to translate the second representation of the media item is detected via one or more input devices. In response to the detection of the input corresponding to the request to translate the second representation of the media item, a fifth representation of the media item is displayed, which includes a portion of the media item that was not included in the second representation of the media item. While the fifth representation of the media item is being displayed, In accordance with the determination that the second portion of the text included in the fifth representation of the media item does not satisfy the individual set of criteria, the display of the visual indication is withheld. The method according to claim 1, further comprising:

9. The method according to claim 1, wherein the amount of prominence exceeding the prominence threshold is based on individual portions of the text that occupy more than the threshold amount of the individual representation.

10. The method according to claim 1, wherein the amount of prominence exceeding the prominence threshold is based on the individual portion of text displayed at a specific location of the individual representation.

11. Displaying the second representation of the media item and, while displaying the visual indication, detecting inputs via one or more input devices that correspond to a request to change the second representation of the media item; In response to the detection of the input corresponding to a request to change the second representation of the media item, the twelfth representation of the media item is displayed. While the 12th representation of the media item is being displayed, In accordance with the determination that an individual portion of the text included in the 12th representation of the media item does not satisfy an individual set of the criteria, the display of the visual indication is withheld. The method according to claim 1, further comprising:

12. The second representation of the media item is displayed at a fifth zoom level, and the method is While the second representation of the media item is displayed at the fifth zoom level and the visual indication, an input corresponding to a request to zoom in on the second representation of the media item is detected via one or more input devices. In response to the detection of the input corresponding to the request to zoom in on the second representation of the media item, the display generation component displays a seventh representation of the media item, which includes a second portion of the text included in the second representation of the media item that is different from the portion of the text included in the second representation, at a sixth zoom level greater than the fifth zoom level. While the seventh representation of the media item is being displayed, In accordance with the determination that the third portion of the text satisfies the individual set of criteria, the display generation component shall display a visual indication corresponding to the third portion of the text that is different from the visual indication corresponding to the portion of the text, The method according to claim 1, further comprising:

13. The second representation of the media item is displayed at a seventh zoom level, and the method is While the second representation of the media item is displayed at the seventh zoom level, an input corresponding to a request to zoom out the second representation of the media item is detected via one or more input devices. In response to the detection of the input corresponding to the request to zoom out the second representation of the media item, the display generation component displays the eighth representation of the media item at an eighth zoom level smaller than the seventh zoom level. While the eighth representation of the media item is displayed at the eighth zoom level, The display of the visual indication is discontinued in accordance with the determination that a first individual portion of the text included in the eighth representation of the media item does not satisfy the individual set of criteria. The method according to claim 1, further comprising:

14. The method according to claim 1, wherein the visual indication surrounds the portion of the text included in the second expression.

15. The second representation of the media item is displayed at a ninth zoom level, and the method is While the second representation of the media item is displayed at the ninth zoom level and the visual indication, an input corresponding to a request to zoom in on the second representation of the media item is detected via one or more input devices. In response to the detection of the input corresponding to a request to zoom in on the second representation of the media item, the display generation component displays a ninth representation of the media item, which includes a separate portion of the text contained in the second representation of the media item that is different from the portion of the text contained in the second representation, at a tenth zoom level greater than the ninth zoom level. While the ninth representation of the media item is being displayed, To discontinue the display of the visual indication corresponding to a portion of the aforementioned text, In accordance with the determination that a second individual portion of the text satisfies a particular set of criteria, the display generation component shall display a visual indication corresponding to the second individual portion of the text that is different from the visual indication corresponding to the portion of the text, The method according to claim 1, further comprising:

16. The second representation of the media item includes a third portion of text that is not selectable, and the method is While the second representation of the media item is being displayed, an input corresponding to a request to display a tenth representation of the media item is detected via one or more input devices. The method according to claim 1, further comprising displaying the tenth representation of the media item, including the third portion of the text, in response to the detection of an input corresponding to a request to display the tenth representation of the media item, wherein the third portion of the text included in the tenth representation of the media item is selectable.

17. A computer program that causes a computer to perform the method described in any one of claims 1 to 16.

18. A computer system, A memory for storing the computer program described in claim 17, One or more processors capable of executing the computer program stored in the memory, A computer system comprising a computer system configured to communicate with a display generation component and one or more input devices.

19. A computer system configured to communicate with a display generation component and one or more input devices, A computer system comprising means for performing the method described in any one of claims 1 to 16.