User interfaces for altering visual media

A synthetic depth-of-field effect dynamically emphasizes subjects in visual media, addressing inefficiencies in existing technologies by reducing user input and conserving power in battery-operated devices.

JP2025108470APending Publication Date: 2025-07-23APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025060669
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-24
Filing Date
2025-04-01
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing technologies for modifying visual media on electronic devices are cumbersome and inefficient, requiring multiple key presses or keystrokes, wasting user time and device energy, particularly in battery-operated devices.

Method used

Implementing a synthetic depth-of-field effect that modifies visual content to emphasize a portion of the media, reducing cognitive burden and conserving power by applying the effect dynamically as the subject moves within the field of view.

Benefits of technology

The solution provides faster and more efficient methods for modifying visual content, reducing user input and conserving battery power while enhancing user experience and device efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108470000001_ABST
    Figure 2025108470000001_ABST
Patent Text Reader

Abstract

To provide electronic devices with faster, more efficient methods and interfaces for altering visual contents that include applying a synthetic depth-of-field effect to the visual contents so as to emphasize portions of media.SOLUTION: User interfaces comprise: capturing visual media (e.g., via a synthetic depth-of-field effect); playing back visual media (e.g., via a synthetic depth-of-field effect); editing visual media (e.g., that has a synthetic depth-of-field effect applied); and / or managing media capture.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Patent Application No. 17 / 484,321, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 24, 2021; U.S. Patent Application No. 17 / 484,307, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 24, 2021; U.S. Patent Application No. 17 / 484,279, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 24, 2021; U.S. Patent Application No. 17 / 483,684, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 23, 2021; U.S. Provisional Patent Application No. 63 / 244,213, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 14, 2021; U.S. Provisional Patent Application No. 63 / 243,724, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on September 13, 2021; U.S. Provisional Patent Application No. 63 / 197,460, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on June 6, 2021; and U.S. Provisional Patent Application No. 63 / 182,751, entitled "USER INTERFACES FOR ALTERING VISUAL MEDIA", filed on April 30, 2021. The entire contents of each of these applications are hereby incorporated by reference into this specification.

[0002] The present disclosure generally relates to computer user interfaces and related technologies, and more specifically, to user interfaces and technologies for altering visual media.

Background Art

[0003] Users of smartphones and other personal electronic devices capture, store, and edit media more frequently to safely protect and share memories with friends. Some existing technologies enable users to capture media such as images, audio, and / or video. Users can manage such media, for example, by capturing, storing, and editing the media.

Summary of the Invention

[0004] Some technologies for modifying visual information using computer systems and other electronic devices are generally cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may involve multiple key presses or keystrokes. Existing technologies take more time than necessary and waste the user's time and the device's energy. This latter consideration is particularly important in battery-operated devices.

[0005] Accordingly, the present technology provides electronic devices with faster and more efficient methods and interfaces for modifying visual content, including applying a synthetic depth-of-field effect to visual content to emphasize a portion of the media. Such methods and interfaces optionally complement or replace other methods for modifying visual content. Such methods and interfaces reduce the cognitive burden on the user and create a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the battery charging interval.

[0006] According to some embodiments, a method is described that is executed in a computer system that communicates with one or more cameras and one or more input devices. The method includes detecting, via one or more input devices, a request to capture a video representing the field of view of one or more cameras; in response to detecting the request to capture the video, capturing a video over a first capture duration, the video including a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of one or more cameras and a second subject within the field of view of one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of one or more cameras over the first capture duration; applying, to the plurality of frames of the video, a synthetic depth of field effect that modifies visual information captured by one or more cameras to emphasize the first subject over the second subject within the plurality of frames of the video, the synthetic depth of field effect changing over time as the first subject moves within the field of view of one or more cameras.

[0007] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras and one or more input devices, the one or more programs including detecting, via the one or more input devices, a request to capture a video representing a field of view of the one or more cameras; in response to detecting the request to capture the video, capturing a video over a first capture duration, the video including a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration; a synthetic subject depth-of-field effect that modifies visual information captured by the one or more cameras to emphasize the first subject relative to the second subject within the plurality of frames of the video, the synthetic subject depth-of-field effect varying over time as the first subject moves within the field of view of the one or more cameras; and applying the synthetic subject depth-of-field effect varying over time to the plurality of frames of the video.

[0008] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors that communicate with one or more cameras and one or more input devices, the one or more programs including detecting, via the one or more input devices, a request to capture a video representing a field of view of the one or more cameras; in response to detecting the request to capture the video, capturing a video over a first capture duration, the video including a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration; applying a synthetic defocus depth of field effect that modifies visual information captured by the one or more cameras to emphasize the first subject relative to the second subject within the plurality of frames of the video, the synthetic defocus depth of field effect changing over time as the first subject moves within the field of view of the one or more cameras, to the plurality of frames of the video.

[0009] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras and one or more input devices. The computer system includes one or more processors and one or more programs configured to be executed by the one or more processors, which detect a request to capture a video representing the field of view of the one or more cameras via the one or more input devices, and in response to detecting the request to capture the video, capture a video over a first capture duration, the video including a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration, and apply a synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the first subject relative to the second subject within the plurality of frames of the video, the synthetic depth of field effect changing over time as the first subject moves within the field of view of the one or more cameras, to the plurality of frames of the video.

[0010] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras and one or more input devices. The computer system includes means for detecting a request to capture a video representing the field of view of one or more cameras via one or more input devices, and in response to detecting the request to capture the video, capturing a video over a first capture duration, the video including a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of one or more cameras and a second subject within the field of view of one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of one or more cameras over the first capture duration, means for applying a synthetic defocus depth of field effect that changes the visual information captured by one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject within the plurality of frames of the video, the synthetic defocus depth of field effect changing over time as the first subject moves within the field of view of one or more cameras.

[0011] According to some embodiments, a computer program product is described. The computer program product is one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras and one or more input devices, the one or more programs detecting, via the one or more input devices, a request to capture a video representing the field of view of the one or more cameras, and in response to detecting the request to capture the video, capturing a video over a first capture duration, the video being a plurality of frames captured over the first capture duration and including a plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration, and applying a synthetic defocus depth effect that modifies the visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames of the video over the second subject in the plurality of frames of the video, the synthetic defocus depth effect varying over time as the first subject moves within the field of view of the one or more cameras, to the plurality of frames of the video.

[0012] According to some embodiments, a method is described that is executed in a computer system that communicates with one or more cameras, a display generation component, and one or more input devices. The method includes, via the display generation component, displaying a user interface that includes a representation of a video including a plurality of frames, the representation including a first subject and a second subject, and a first user interface object indicating that the first subject is emphasized by a synthetic subject depth of field effect that modifies visual information captured by one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject; detecting, via one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video while the user interface including the representation of the video and the first user interface object is being displayed; in response to detecting the gesture corresponding to the selection of the second subject in the representation of the video, modifying the synthetic subject depth of field effect to emphasize the second subject within the plurality of frames with respect to the first subject by modifying the visual information captured by one or more cameras; and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic subject depth of field effect that modifies the visual information captured by one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject.

[0013] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras, a display generation component, and one or more input devices, the one or more programs including: displaying a user interface including a representation of a video including a plurality of frames, the representation including a first subject and a second subject, and a first user interface object indicating that the first subject is emphasized by a synthetic subject depth-of-field effect that modifies visual information captured by one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject; detecting, via the one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video while the user interface including the representation of the video and the first user interface object is being displayed; in response to detecting a gesture corresponding to a selection of the second subject in the representation of the video, modifying the synthetic subject depth-of-field effect to emphasize the second subject within the plurality of frames with respect to the first subject by modifying the visual information captured by one or more cameras; and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic subject depth-of-field effect that modifies the visual information captured by one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject.

[0014] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras, a display generation component, and one or more input devices, the one or more programs including: displaying a user interface including a representation of a video including a plurality of frames, the representation including a first subject and a second subject, and a first user interface object indicating that the first subject is emphasized by a synthetic defocus depth of field effect that modifies visual information captured by one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject; detecting, via the one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video while the user interface including the representation of the video and the first user interface object is being displayed; in response to detecting the gesture corresponding to the selection of the second subject in the representation of the video, modifying the synthetic defocus depth of field effect to emphasize the second subject within the plurality of frames with respect to the first subject by modifying the visual information captured by one or more cameras; and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic defocus depth of field effect that modifies the visual information captured by one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject.

[0015] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras, a display generation component, and one or more input devices. The computer system includes one or more processors and one or more programs configured to be executed by the one or more processors. The programs include displaying a user interface including a first user interface object indicating that a first subject is emphasized by a synthetic depth-of-field effect that modifies visual information captured by one or more cameras to emphasize the first subject within a plurality of frames in a representation of a video including a first subject and a second subject, via the display generation component; detecting, via one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video while the user interface including the representation of the video and the first user interface object is being displayed; in response to detecting the gesture corresponding to the selection of the second subject in the representation of the video, modifying the synthetic depth-of-field effect to emphasize the second subject within a plurality of frames with respect to the first subject and modifying the visual information captured by one or more cameras; and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic depth-of-field effect that modifies the visual information captured by one or more cameras to emphasize the second subject within a plurality of frames with respect to the first subject.

[0016] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras, a display generation component, and one or more input devices. The computer system, via the display generation component, displays a user interface including a representation of a video including a plurality of frames, the representation including a first subject and a second subject, and a first user interface object indicating that the first subject is emphasized by a synthetic subject depth-of-field effect that modifies visual information captured by one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject. While displaying the representation of the video and the user interface including the first user interface object, means for detecting a gesture corresponding to a selection of the second subject in the representation of the video via one or more input devices, and in response to detecting a gesture corresponding to a selection of the second subject in the representation of the video, means for modifying the synthetic subject depth-of-field effect to emphasize the second subject within the plurality of frames with respect to the first subject by modifying the visual information captured by one or more cameras, and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic subject depth-of-field effect that modifies the visual information captured by one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject.

[0017] According to some embodiments, a computer program product is described. The computer program product includes one or more cameras, a display generation component, one or more input devices, one or more processors, and one or more programs configured to be executed by the one or more processors. The programs include displaying a user interface including a representation of a video including a plurality of frames, the representation including a first subject and a second subject, and a first user interface object indicating that the first subject is emphasized by a synthetic depth of field effect that modifies visual information captured by the one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject; detecting, via the one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video while the user interface including the representation of the video and the first user interface object is being displayed; in response to detecting the gesture corresponding to the selection of the second subject in the representation of the video, modifying the synthetic depth of field effect to emphasize the second subject within the plurality of frames with respect to the first subject by modifying the visual information captured by the one or more cameras; and displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject.

[0018] According to some embodiments, a method is described that is executed in a computer system that communicates with a display generation component. The method includes, via the display generation component, a representation of a video having a first duration, the representation including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements in the video, the plurality of changes in subject emphasis including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a video navigation user interface element for navigating through the video, the video including representations of the first time and the second time, the representation of the second time being visually distinguishable from other times during the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time being visually distinguishable from the representation of the second time, and displaying a user interface including simultaneously displaying the representation of the video and the video navigation user interface element.

[0019] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, and through the display generation component, a video having a first duration, a plurality of changes in subject emphasis within the video, where the changes in subject emphasis within the video include changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements within the video, a plurality of changes in subject emphasis within the video including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a representation of the video including the plurality of changes, and a video including a representation of the first time and a representation of the second time, where the representation of the second time is visually distinguishable from other times during the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time is visually distinguishable from the representation of the second time, and instructions for displaying a user interface including simultaneously displaying video navigation user interface elements for navigating through the video.

[0020] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, and through the display generation component, a video having a first duration, a plurality of changes in subject emphasis within the video, where the changes in subject emphasis within the video include changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements within the video, the plurality of changes in subject emphasis within the video including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a representation of the video including the plurality of changes, and a video including a representation of the first time and a representation of the second time, where the representation of the second time is visually distinguishable from other times during the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time is visually distinguishable from the representation of the second time, and instructions for displaying a user interface including simultaneously displaying video navigation user interface elements for navigating through the video.

[0021] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras and a display generation component. The computer system includes one or more processors and one or more programs configured to be executed by the one or more processors. The programs include instructions to display a user interface that simultaneously displays a representation of a video having a first duration, the video including a plurality of changes in subject emphasis, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements in the video, the plurality of changes in subject emphasis including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a video navigation user interface element for navigating through the video, the representation of the video including the representation of the first time and the representation of the second time, the representation of the second time being visually distinguishable from other times during the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time being visually distinguishable from the representation of the second time, and a memory storing the one or more programs.

[0022] According to some embodiments, a computer system is described. The computer system is configured to communicate with one or more cameras and a display generation component. The computer system, via the display generation component, presents a video having a first duration, the video including a plurality of changes in subject emphasis within the video, the changes in subject emphasis within the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements within the video, the plurality of changes in subject emphasis within the video including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, and presents a video navigation user interface element for navigating through the video, the video navigation user interface element including a representation of the first time and a representation of the second time, the representation of the second time being visually distinguishable from other times during the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time being visually distinguishable from the representation of the second time, and includes means for displaying, via the display generation component, a user interface that simultaneously displays the representation of the video and the video navigation user interface element.

[0023] According to some embodiments, a computer program product is described. The computer program product includes a display generation component, one or more processors, and one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for displaying a user interface that simultaneously displays a representation of a video having a first duration, the video including a plurality of changes in subject emphasis within the video, the changes in subject emphasis within the video including changes to the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements within the video, the plurality of changes in subject emphasis within the video including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a video navigation user interface element for navigating through the video, the video including a representation of the first time and a representation of the second time, the representation of the second time being visually distinguishable from other times during the first duration of the video that do not correspond to a change in subject emphasis, and the representation of the first time being visually distinguishable from the representation of the second time, and a memory storing the one or more programs.

[0024] According to some embodiments, a method is described that is executed in a computer system that communicates with a display generation component and a plurality of cameras including a first camera having first image capture parameters determined by the hardware of the first camera and a second camera having second image capture parameters determined by the hardware of the second camera, the second image capture parameters being different from the first image capture parameters. The method includes displaying, via the display generation component, a camera user interface that is a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and in response to detecting the decrease in the distance between the camera location and the focus location, transitioning from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance.

[0025] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is a program configured to be executed by one or more processors of a computer system that communicates with a plurality of cameras including a display generation component, a first camera having first image capture parameters determined by the hardware of the first camera, and a second camera having second image capture parameters determined by the hardware of the second camera, the second image capture parameters being different from the first image capture parameters, the program including: displaying, via the display generation component, a camera user interface including a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera having the first image capture parameters; detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and in response to detecting the decrease in the distance between the camera location and the focus location, transitioning from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance.

[0026] According to some embodiments, a non - transitory computer - readable storage medium is described. The non - transitory computer - readable storage medium is a program or programs configured to be executed by one or more processors of a computer system that communicates with a plurality of cameras including a display generation component, a first camera having first image capture parameters determined by the hardware of the first camera, and a second camera having second image capture parameters determined by the hardware of the second camera, where the second image capture parameters are different from the first image capture parameters. The program or programs cause the computer system to display, via the display generation component, a camera user interface including a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera having the first image capture parameters; detect a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and in response to detecting the decrease in the distance between the camera location and the focus location, transition from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance.

[0027] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a plurality of cameras including a first camera having first image capture parameters determined by the hardware of the first camera and a second camera having second image capture parameters determined by the hardware of the second camera, the second image capture parameters being different from the first image capture parameters. The computer system includes one or more processors and one or more programs configured to be executed by the one or more processors, which, via the display generation component, display a camera user interface that is a representation of one or more fields of view of one or more of the plurality of cameras, including a representation of a field of view displayed using visual information collected by the first camera together with the first image capture parameters; detect a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and in response to detecting the decrease in the distance between the camera location and the focus location, transition from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance.

[0028] According to some embodiments, a computer system is described. A computer system configured to communicate with a plurality of cameras including a display generation component, a first camera having first image capture parameters determined by the hardware of the first camera, and a second camera having second image capture parameters determined by the hardware of the second camera, the second image capture parameters being different from the first image capture parameters. The computer system includes means for displaying, via the display generation component, a camera user interface including a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; means for detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and means for transitioning from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view in accordance with a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance in response to detecting the decrease in the distance between the camera location and the focus location.

[0029] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, a first camera having first image capture parameters determined by the hardware of the first camera, and a plurality of cameras including a second camera having second image capture parameters determined by the hardware of the second camera, the second image capture parameters being different from the first image capture parameters. The one or more programs include, via the display generation component, displaying a camera user interface that is a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; and in response to detecting a decrease in the distance between the camera location and the focus location, transitioning from using the visual information collected by the first camera to display the representation of the field of view to using the visual information collected by the second camera to display the representation of the field of view according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance.

[0030] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component. The method includes, via the display generation component, playing a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; detecting a request to change subject emphasis at a second time in the video that is different from the first time after playing the portion of the video that includes the first subject emphasis change occurring at the first time; in response to detecting the request to change subject emphasis at the second time in the video, changing the subject emphasis in the video during a second period following the second time; and changing the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject for one or more elements in the video during the first period following the first time.

[0031] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, the one or more programs including, via the display generation component, playing a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; detecting a request to change subject emphasis at a second time in the video that is different from the first time after playing the portion of the video that includes the first subject emphasis change occurring at the first time; in response to detecting the request to change subject emphasis at the second time in the video, changing the subject emphasis in the video during a second period following the second time; and changing the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject for one or more elements in the video during the first period following the first time.

[0032] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium is one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, and through the display generation component, a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in a video during a first period following the first time, playing a portion of the video including the first subject emphasis change, detecting a request to change subject emphasis at a second time in the video that is different from the first time after playing the portion of the video including the first subject emphasis change occurring at the first time, in response to detecting the request to change subject emphasis at the second time in the video, changing the subject emphasis in the video during a second period following the second time, and changing the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject for one or more elements in the video during the first period following the first time.

[0033] According to some embodiments, a computer system configured to communicate with a display generation component is described. The computer system includes one or more processors and one or more programs configured to be executed by the one or more processors. The one or more programs include, via the display generation component, playing a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; detecting a request to change subject emphasis at a second time in the video that is different from the first time after playing the portion of the video that includes the first subject emphasis change occurring at the first time; in response to detecting the request to change subject emphasis at the second time in the video, changing subject emphasis in the video during a second period following the second time; and changing the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject for one or more elements in the video during the first period following the first time.

[0034] According to some embodiments, a computer system configured to communicate with a display generation component and one or more input devices is described. The computer system includes means for playing, via the display generation component, a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; means for detecting a request to change subject emphasis at a second time in the video that is different from the first time after playing the portion of the video that includes the first subject emphasis change occurring at the first time; means for changing subject emphasis in the video during a second period following the second time in response to detecting the request to change subject emphasis at the second time in the video; and means for changing the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject for one or more elements in the video during the first period following the first time.

[0035] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component. The one or more programs, via the display generation component, reproduce a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in appearance of visual information captured by one or more cameras to emphasize an individual subject with respect to one or more elements within the video during a first period following the first time; detect a request to change subject emphasis at a second time within the video that is different from the first time after reproducing the portion of the video that includes the first subject emphasis change occurring at the first time; in response to detecting the request to change subject emphasis at the second time within the video, change the subject emphasis within the video during a second period following the second time; and change the first subject emphasis change occurring at the first time, including changing the emphasis of an individual subject with respect to one or more elements within the video during the first period following the first time.

[0036] The executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors. The executable instructions for performing these functions are optionally included in a transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0037] Thus, faster and more efficient methods and interfaces for changing visual content are provided to a device, thereby increasing the effectiveness, efficiency, and user satisfaction of such a device. Such methods and interfaces may complement or replace other methods for changing visual content. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.

[0039]

Figure 1A

[0040]

Figure 1B

[0041]

Figure 2

[0042]

Figure 3

[0043]

Figure 4A

[0044]

Figure 4B

[0045]

Figure 5A

[0046]

Figure 5B

[0047]

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 6E

Figure 6F

Figure 6G

Figure 6H

Figure 6I

Figure 6J

Figure 6K

Figure 6L

Figure 6M

Figure 6N

Figure 6O

Figure 6P

Figure 6Q

Figure 6R

Figure 6R1

Figure 6S

Figure 6T

Figure 6U

Figure 6V

Figure 6W

Figure 6X

Figure 6Y

Figure 6Z

Figure 6AA

Figure 6AB

Figure 6AC

Figure 6AD

Figure 6AE

Figure 6AF

Figure 6AF1

Figure 6AG

Figure 6AH

Figure 6AI

Figure 6AJ

Figure 6AK

Figure 6AL

Figure 6AM

Figure 6AN

Figure 6AO

Figure 6AP

Figure 6AQ

Figure 6AR

Figure 6AS

Figure 6AT

Figure 6AU

Figure 6AV

Figure 6AW

Figure 6AX

Figure 6AY

Figure 6AZ

Figure 6BA

Figure 6BB

Figure 6BC

Figure 6BC1

Figure 6BC2

Figure 6BD

Figure 6BE

Figure 6BF

Figure 6BG

Figure 6BH

Figure 6BI

Figure 6BJ

[0048]

Figure 7

[0049]

Figure 8

[0050]

Figure 9

[0051]

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 10F

Figure 10G

Figure 10H

Figure 10I

[0052]

Figure 11

[0053]

Figure 12

[0054]

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0055] The following description sets forth exemplary methods, parameters, and the like. However, it should be recognized that such description is not intended as a limitation on the scope of the present disclosure, but rather as an illustration of exemplary embodiments.

[0056] There is a need for an electronic device that provides an efficient method and interface for changing visual content. For example, there is a need for an electronic device that enables a user to change visual content by applying a synthetic depth of field effect to a plurality of frames of media without the need to manually change and / or blur the frames of the media to simulate a depth of field effect. Such technology can reduce the cognitive burden on users who want to change the visual content of media, thereby increasing productivity. Further, such technology can reduce the use of the processor and the power of the battery that would otherwise be wasted on redundant user input.

[0057] The following FIGS. 1A-1B, 2, 3, 4A-4B, 5A-5B, and 12 provide an illustration of exemplary devices and systems for implementing techniques for managing and changing visual media.

[0058] FIGS. 6A-6BJ are user interfaces for changing visual media using a computer system according to some embodiments. FIG. 7 is a flowchart showing a method for changing visual content according to some embodiments. FIG. 8 is a flowchart showing a method for changing visual content according to some embodiments. FIG. 9 is a flowchart showing a method for changing visual content according to some embodiments. FIG. 13 is a flowchart showing a method for changing visual content according to some embodiments. The user interfaces in FIGS. 6A-6BJ are used to illustrate the processes described below, including the processes in FIGS. 7, 8, 9, and 13.

[0059] Figures 10A - 10I show exemplary user interfaces for managing media capture using a computer system, according to some embodiments. FIG. 11 is a flow diagram showing an exemplary method for managing media capture using a computer system, according to some embodiments. The user interfaces in FIGS. 10A - 10I are used to illustrate the processes described below, including the process in FIG. 11.

[0060] The processes described below enhance the operability of the device and streamline the user - device interface by various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, performing an operation without requiring further user input when a set of conditions is met, and / or other techniques (e.g., helping the user to make appropriate inputs when operating / interacting with the device and reducing user errors). These techniques also reduce power usage and improve the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0061] Furthermore, in the method described herein, conditioned on one or more conditions being met by one or more steps, it should be understood that the described method can be repeated in multiple iterations such that, over the course of the repetitions, all of the conditions conditioned on by the steps of the method are met in different repetitions of the method. For example, if a method requires performing a first step when a condition is met and a second step when the condition is not met, one of ordinary skill in the art would understand that the steps recited in the claims are repeated in a particular order until the condition is met and then ceases to be met. Thus, a method described in terms of one or more steps that depend on one or more met conditions can be rewritten as a method that is repeated until each condition described in the method is met. However, this is not required in claims for a system or computer-readable medium that includes instructions for performing conditional operations based on the fulfillment of corresponding one or more conditions, and thus can determine whether an eventuality is fulfilled without explicitly repeating the steps of the method until all conditions for the steps of the method being conditional are met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as necessary to ensure that all of the conditional steps are executed.

[0062] In the following description, terms such as "first", "second", etc. are used to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various embodiments described, a first touch can be referred to as a second touch, and similarly, a second touch can be referred to as a first touch. Both the first touch and the second touch are touches, but they are not the same touch.

[0063] The terms used in the description of the various embodiments described herein are for the purpose of describing particular embodiments only and are not intended to be limiting. In the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It should be understood that the terms "includes", "including", "comprises", and / or "comprising", when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0064] The term "if" is optionally interpreted to mean "when" or "upon", or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" are optionally interpreted to mean "upon determining" or "in response to determining", or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]", depending on the context.

[0065] Embodiments of electronic devices, user interfaces for such devices, and related processes for using such devices are described. In some embodiments, the device is a portable communication device such as a cellular phone that also includes other functions such as PDA functionality and / or music player functionality. Exemplary embodiments of portable multifunctional devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Optionally, other portable electronic devices such as laptop computers or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad) are also used. Also, in some embodiments, it should be understood that the device is not a portable communication device but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). In some embodiments, the electronic device is a computer system that communicates (e.g., via wired communication, via wireless communication) with a display generation component. The display generation component is configured to provide a visual output such as a display via a CRT display, a display via an LED display, or a display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. As used herein, "displaying" content includes transmitting data (e.g., image data or video data) via a wired or wireless connection to an integrated or external display generation component to visually generate the content (e.g., video data rendered or decoded by a display controller 156) for the purpose of visually generating the content.

[0066] In the following discussion, an electronic device including a display and a touch sensing surface will be described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick.

[0067] The device typically supports various applications such as one or more of a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0068] The various applications executed on the device optionally use at least one common physical user interface device such as a touch sensing surface. One or more functions of the touch sensing surface, as well as the corresponding information displayed on the device, are optionally adjusted and / or changed for each application and / or within an individual application. Thus, the common physical architecture of the device (such as the touch sensing surface) optionally supports various applications with a user interface that is intuitive and transparent to the user.

[0069] Attention is now directed to an embodiment of a portable device having a touch-sensing display. FIG. 1A is a block diagram showing a portable multifunctional device 100 having a touch-sensing display system 112 according to some embodiments. The touch-sensing display 112 may be referred to herein, for convenience, as a "touch screen" and may be known or referred to as a "touch-sensing display system". Device 100 includes a memory 102 (optionally including one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 (e.g., a touch-sensing surface such as the touch-sensing display system 112 of device 100) for detecting the intensity of a contact on device 100. Device 100 optionally includes one or more haptic output generators 167 for generating haptic output on device 100 (e.g., on a touch-sensing surface such as the touch-sensing display system 112 of device 100 or the touch pad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0070] As used in this specification and the claims, the term "intensity" of a contact on a touch-sensing surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensing surface, or a proxy for the force or pressure of a contact on the touch-sensing surface. The intensity of a contact has a range of values that includes at least four distinct values, and more typically, hundreds (e.g., at least 256) of distinct values. The intensity of a contact is optionally determined (or measured) using a variety of techniques and a variety of sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensing surface are optionally used to measure the force at various points on the touch-sensing surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine the estimated force of a contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensing surface. Alternatively, the size and / or change thereof of a contact area detected on the touch-sensing surface, the capacitance and / or change thereof of the touch-sensing surface proximate to the contact, and / or the resistance and / or change thereof of the touch-sensing surface proximate to the contact are optionally used as surrogates for the force or pressure of a contact on the touch-sensing surface. In some implementations, alternative measurements of the force or pressure of a contact are used directly to determine whether the intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the alternative measurement). In some implementations, a proxy measurement of the contact force or pressure is converted to an estimated value of the force or pressure, and the estimated value of the force or pressure is used to determine whether the intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). By using the intensity of a contact as an attribute of a user input, a user can access additional device functions that may otherwise be inaccessible to the user on a reduced-size device with a limited implementation area for displaying affordances (e.g., on a touch-sensing display), and / or receive a user input (e.g., via a touch-sensing display, a touch-sensing surface, or a physical / mechanical control such as a knob or button).

[0071] As used in this specification and the claims, the term "haptic output" refers to the physical displacement of the device relative to its previous position, the physical displacement of a component of the device (e.g., a touch-sensing surface) relative to another component of the device (e.g., the housing), or the displacement of a component relative to the center of mass of the device, which is to be detected by the user's sense of touch. For example, in a situation where the device or a component of the device is in contact with a touch-sensitive surface of the user (e.g., the finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or the component of the device. For example, the movement of a touch-sensing surface (e.g., a touch-sensing display or a trackpad) can be optionally interpreted by the user as a "down click" or "up click" of a physical actuator button. In some cases, even when there is no movement of the physical actuator button associated with the touch-sensing surface physically pressed (e.g., displaced) by the user's action, the user can feel a tactile sensation such as a "down click" or "up click". As another example, the movement of the touch-sensing surface can be optionally interpreted or perceived by the user as the "roughness" of the touch-sensing surface, even when there is no change in the smoothness of the touch-sensing surface. Such interpretation of the touch by the user depends on the user's individual sensory perception, but there are many sensory perceptions of touch that are common to the majority of users. Therefore, when the haptic output is described as corresponding to a specific sensory perception of the user (e.g., "up click", "down click", "roughness"), unless otherwise stated, the generated haptic output corresponds to the physical displacement of the device or a component of the device that generates the described sensory perception of a typical (or average) user.

[0072] Device 100 is merely an example of a portable multifunctional device, and it should be understood that Device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of those components. The various components shown in FIG. 1A are implemented in a combination of hardware, software, or both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0073] Memory 102 may optionally include high-speed random access memory and may also optionally include non-volatile memory such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 122 may optionally control access to memory 102 by other components of device 100.

[0074] Peripheral interface 118 can be used to couple input and output peripheral devices of the device to CPU 120 and memory 102. One or more processors 120 operate or execute various software programs (such as computer programs including instructions) and / or instruction sets stored in memory 102 to perform various functions for device 100 and process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0075] The RF (radio frequency) circuit 108 transmits and receives RF signals, also called electromagnetic signals. The RF circuit 108 converts electrical signals into electromagnetic signals or vice versa and communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, and the like. The RF circuit 108 optionally communicates wirelessly with networks such as the Internet, also called the World Wide Web (WWW), an intranet, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), as well as with other devices. The RF circuit 108 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as by a short-range communication radio. Wireless communication optionally includes, but is not limited to, Global System for Mobile Communications (GSM) for mobile communication, Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), Long Termevolution, LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE802.11a, IEEE802.11b, IEEE802.11g, IEEE802.11n, and / or IEEE802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols not yet developed as of the filing date of this specification. Any one of a plurality of communication standards, protocols, and technologies is used.

[0076] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts this audio data into an electrical signal, and transmits this electrical signal to the speaker 111. The speaker 111 converts the electrical signal into human audible sound waves. Also, the audio circuit 110 receives the electrical signal converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signal into audio data and transmits this audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., 212 of FIG. 2). The headset jack provides an interface between the audio circuit 110 and a detachable audio input / output peripheral device such as an output-only headset or a headset having both outputs (e.g., mono or stereo headphones) and inputs (e.g., a microphone).

[0077] The I / O subsystem 106 couples input / output peripheral devices on the device 100, such as the touch screen 112 and other input control devices 116, to the peripheral device interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input devices or control devices. The one or more input controllers 160 receive electrical signals from and transmit electrical signals to other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some embodiments, the input controller(s) 160 are optionally coupled to (or not coupled to any of) a pointer device such as a keyboard, an infrared port, a USB port, and a mouse. One or more buttons (e.g., 208 in FIG. 2) optionally include up / down buttons for volume control of the speaker 111 and / or the microphone 113. One or more buttons optionally include push buttons (e.g., 206 in FIG. 2). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wired communication or via wireless communication). In some embodiments, the one or more input devices include a touch sensing surface (e.g., a trackpad as part of a touch sensing display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175) for tracking user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system.In some embodiments, an air gesture is a gesture that is detected without the user touching (or independently of) an input element that is part of the device, and includes the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), the movement of the user's body relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of the user's other hand relative to one of the user's hands, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose with a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body at a predetermined speed or amount), and is based on the detected movement of a part of the user's body.

[0078] As described in U.S. Patent Application No. 11 / 322,549, filed December 23, 2005, "Unlocking a Device by Performing Gestures on an Unlock Image," U.S. Patent No. 7,657,849, which is hereby incorporated by reference in its entirety, a quick press of a push button optionally unlocks the touch screen 112 or optionally initiates a process of unlocking the device using gestures on the touch screen. A longer press of the push button (e.g., 206) optionally turns the power to the device 100 on or off. The functionality of one or more of the buttons can optionally be customized by the user. The touch screen 112 is used to implement virtual or soft buttons and one or more soft keyboards.

[0079] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from and / or transmits electrical signals to the touch screen 112. The touch screen 112 displays visual output to the user. This visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0080] The touch screen 112 has a touch-sensitive surface, sensor, or set of sensors that receives input from the user based on tactile and / or haptic contact. The touch screen 112 and the display controller 156 (along with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of the contact) on the touch screen 112 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.

[0081] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although in other embodiments other display technologies may also be used. The touch screen 112 and the display controller 156 optionally use, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112, of any of a plurality of touch sensing technologies currently known or later developed, to detect contact and any movement or interruption thereof. In an exemplary embodiment, a projected mutual capacitance sensing technology such as that found in the iPhone (registered trademark) and iPod Touch (registered trademark) from Apple Inc. of Cupertino, California is used.

[0082] The touch sensing display in some embodiments of the touch screen 112 optionally is similar to a multi-touch sensing touch pad described in U.S. Patent Nos. 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman), and / or U.S. Patent Application Publication No. 2002 / 0015024 (A1), each of which is hereby incorporated by reference in its entirety. However, the touch screen 112 displays visual output from the device 100, whereas the touch sensing touch pad does not provide visual output.

[0083] Touch sensing displays in some embodiments of touch screen 112 are described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, "Multipoint Touch Surface Controller"; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, "Multipoint Touchscreen"; (3) U.S. Patent Application No. 10 / 903,964, filed Jul. 30, 2004, "Gestures For Touch Sensitive Input Devices"; (4) U.S. Patent Application No. 11 / 048,264, filed Jan. 31, 2005, "Gestures For Touch Sensitive Input Devices"; (5) U.S. Patent Application No. 11 / 038,590, filed Jan. 18, 2005, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices"; (6) U.S. Patent Application No. 11 / 228,758, filed Sep. 16, 2005, "Virtual Input Device Placement On A Touch Screen User Interface"; (7) U.S. Patent Application No. 11 / 228,700, filed Sep. 16, 2005, "Operation Of A Computer With A Touch Screen Interface"; (8) U.S. Patent Application No. 11 / 228,737, filed Sep. 16, 2005, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard"; and (9) U.S. Patent Application No. 11 / 367,749, filed Mar. 3, 2006, "Multi-Functional Hand-Held Device". All of these applications are hereby incorporated by reference in their entirety.

[0084] The touch screen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user optionally touches the touch screen 112 using any suitable object or appendage such as a stylus, finger, etc. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, although this may be less accurate than stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts rough input by a finger into an accurate pointer / cursor position or command for performing the action desired by the user.

[0085] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touch pad for activating or deactivating certain functions. In some embodiments, the touch pad, unlike the touch screen, is a touch-sensitive area of the device that does not display a visual output. The touch pad is optionally a separate touch-sensitive surface from the touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.

[0086] The device 100 also includes a power system 162 that supplies power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharge system, a power outage detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power within a portable device.

[0087] Device 100 also optionally includes one or more optical sensors 164. FIG. 1A shows an optical sensor coupled to an optical sensor controller 158 within I / O subsystem 106. Optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Optical sensor 164 receives light from the environment projected through one or more lenses and converts that light into data representing an image. Optical sensor 164 cooperates with imaging module 143 (also referred to as a camera module) to optionally capture a still image or video. In some embodiments, the optical sensor is located on the back surface of device 100 opposite touch screen display 112 on the front of the device, and thus the touch screen display can be used as a viewfinder for the acquisition of still images and / or video. In some embodiments, the optical sensor is disposed on the front of the device such that the user's image is optionally acquired for a video conference while the user is viewing other video conference participants on the touch screen display. In some embodiments, the position of optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), and thus a single optical sensor 164 is used for both video conferencing and the acquisition of still images and / or video along with the touch screen display.

[0088] Device 100 optionally also includes one or more depth camera sensors 175. FIG. 1A shows a depth camera sensor coupled to a depth camera controller 169 within the I / O subsystem 106. The depth camera sensor 175 receives data from the environment and creates a three-dimensional model of an object (e.g., a face) within a scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in cooperation with the imaging module 143 (also referred to as the camera module), the depth camera sensor 175 is optionally used to determine depth maps of different portions of an image captured by the imaging module 143. In some embodiments, while a user is viewing other video conference participants on a touch screen display, a depth camera sensor is disposed on the front of device 100 to optionally acquire an image of the user with depth information for a video conference and also to capture a self-portrait image with depth map data. In some embodiments, the depth camera sensor 175 is disposed on the back of the device, or on both the back and the front of device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), such that the depth camera sensor 175 is used with the touch screen display for video conferencing as well as for acquiring still images and / or videos.

[0089] In some embodiments, a depth map (e.g., a depth map image) includes information (e.g., values) regarding the distance of objects within a scene from a viewpoint (e.g., a camera, a light sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines a position along the Z-axis of that viewpoint where its corresponding 2D pixel is located. In some embodiments, a depth map is composed of pixels, and each pixel is defined by a value (e.g., between 0 and 255). For example, the value "0" represents a pixel located at the farthest location within the "3D" scene, and the value "255" represents a pixel located closest to the viewpoint (e.g., a camera, a light sensor, a depth camera sensor) within that "3D" scene. In other embodiments, a depth map represents the distance between objects within a scene and a plane of the viewpoint. In some embodiments, a depth map includes information regarding the relative depth of various features of a target object (e.g., the relative depth of the eyes, nose, mouth, ears of a user's face) as seen from a depth camera. In some embodiments, a depth map includes information that enables a device to determine the outline of a target object in the z-direction.

[0090] Device 100 also optionally includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled to an intensity sensor controller 159 within I / O subsystem 106. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electrokinetic sensors, piezoelectric force sensors, optical force sensors, capacitive touch sensing surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of contact on a touch sensing surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with, or in proximity to, a touch sensing surface (e.g., touch sensing display system 112). In some embodiments, at least one contact intensity sensor is disposed on the back of device 100, which is opposite a touch screen display 112 disposed on the front of device 100.

[0091] Also, device 100 optionally includes one or more proximity sensors 166. FIG. 1A shows proximity sensor 166 coupled to peripheral device interface 118. Alternatively, proximity sensor 166 is optionally coupled to input controller 160 within I / O subsystem 106. Proximity sensor 166 functions as described, for example, in U.S. Patent Application Nos. 11 / 241,839, "Proximity Detector In Handheld Device", 11 / 240,788, "Proximity Detector In Handheld Device", 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output", 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices", and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals", which are hereby incorporated by reference in their entirety. In some embodiments, when a multifunctional device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables touch screen 112.

[0092] Device 100 also optionally includes one or more haptic output generators 167. FIG. 1A shows a haptic output generator coupled to a haptic feedback controller 161 within I / O subsystem 106. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components, and / or electromechanical devices that convert energy, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components that convert an electrical signal into a haptic output on the device), into linear movement. The contact intensity sensor 165 receives haptic feedback generation instructions from the haptic feedback module 133 and generates a haptic output on device 100 that can be sensed by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed with, or proximate to, a touch sensing surface (e.g., touch sensing display system 112) and optionally generates a haptic output by moving the touch sensing surface in a vertical direction (e.g., in / out of the surface of device 100) or in a horizontal direction (e.g., back and forth within the same plane as the surface of device 100). In some embodiments, at least one haptic output generator sensor is disposed on the back of device 100, which is opposite the touch screen display 112 disposed on the front of device 100.

[0093] Device 100 also optionally includes one or more accelerometers 168. FIG. 1A shows an accelerometer 168 coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to an input controller 160 within the I / O subsystem 106. The accelerometer 168 functions optionally as described in both of which are hereby incorporated by reference in their entirety, U.S. Patent Application Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices", and U.S. Patent Application Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer". In some embodiments, information is displayed on the touch screen display in a portrait or landscape display based on analysis of data received from one or more accelerometers. In addition to the accelerometer(s) 168, device 100 optionally includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for obtaining information regarding the location and orientation (e.g., portrait or landscape orientation) of device 100.

[0094] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application (or instruction set) 136. Further, in some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) stores a device / global internal state 157 as shown in FIGS. 1A and 3. The device / global internal state 157 includes an active application state indicating which application is active if there is a currently active application, a display state indicating which application, view, or other information occupies various regions of the touch screen display 112, a sensor state including information obtained from various sensors and input control devices 116 of the device, and one or more of location information regarding the location and / or orientation of the device.

[0095] The operating system 126 (e.g., an embedded operating system such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware components and software components.

[0096] The communication module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by the RF circuit 108 and / or the external port 124. The external ports 124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) are adapted to couple to other devices either directly or indirectly via a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or compatible with the 30-pin connector used on iPod (registered trademark) devices (a trademark of Apple Inc.).

[0097] The contact / motion module 130 optionally detects contact with the touch screen 112 and other touch sensing devices (e.g., a touch pad or a physical click wheel) (in cooperation with the display controller 156). The contact / motion module 130 performs various software components for performing various operations related to the detection of contact, such as determining whether contact has occurred (e.g., detecting a finger down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or an alternative to the force or pressure of the contact), determining whether there is movement of the contact, tracking movement across the touch sensing surface (e.g., detecting one or more events of dragging a finger), and determining whether the contact has ended (e.g., detecting a finger up event or an interruption of the contact). The contact / motion module 130 receives contact data from the touch sensing surface. Determining the movement of the contact point, represented by a series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These operations are optionally applied to a single contact (e.g., contact with one finger) or multiple simultaneous contacts (e.g., "multi-touch" / contact with multiple fingers). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touch pad.

[0098] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an action has been performed by the user (e.g., to determine whether the user has "clicked" on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds can be adjusted without changing the physical hardware of the device 100, rather than being determined by the activation thresholds of specific physical actuators). For example, the mouse "click" threshold of a trackpad or touch screen display can be set to any of a wide range of predefined thresholds without changing the trackpad or touch screen display hardware. Additionally, in some implementations, the user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or adjusting multiple intensity thresholds at once by system-level click "intensity" parameters).

[0099] The contact / motion module 130 optionally detects gesture inputs by the user. Different gestures on the touch-sensing surface have different contact patterns (e.g., the detected movement, timing, and / or intensity of the contact is different). Thus, gestures are optionally detected by detecting a specific contact pattern. For example, detecting a finger tap gesture includes detecting a finger down event followed by a finger up (lift-off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensing surface includes detecting a finger down event followed by one or more finger drag events and then followed by a finger up (lift-off) event.

[0100] The graphic module 132 includes various known software components for rendering and displaying graphics on the touch screen 112 or other display, including components for changing the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphic" includes, but is not limited to, any object that can be displayed to the user, including text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, and the like.

[0101] In some embodiments, the graphic module 132 stores data representing the graphics that will be used. Each graphic is optionally assigned a corresponding code. The graphic module 132 receives one or more codes specifying the graphics to be displayed, along with coordinate data and other graphic property data as needed, from an application or the like, and then generates the image data for the screen to be output to the display controller 156.

[0102] The tactile feedback module 133 includes various software components for generating the instructions used by the tactile output generator(s) 167 to generate tactile output at one or more locations on the device 100 in response to the user's interaction with the device 100.

[0103] The text input module 134 is optionally a component of the graphic module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).

[0104] The GPS module 135 determines the location of the device and provides this information for use within various applications (e.g., to the phone 138 for location-based dialing, to the camera 143 as picture / video metadata, and to applications that provide location-based services such as a weather widget, a local yellow pages widget, and a map / navigation widget).

[0105] The application 136 optionally includes one or more of the following modules (or sets of instructions) or subsets or supersets thereof. ● Contact module 137 (also sometimes referred to as an address book or contact list), ● Phone module 138, ● Video conferencing module 139, ● Email client module 140, ● Instant messaging (IM) module 141, ● Training support module 142, ● Camera module 143 for still images and / or videos, ● Image management module 144, ● Video player module, ● Music player module, ● Browser module 147, ● Calendar module 148, ● Optionally, a widget module 149 that includes one or more of a weather widget 149-1, a stock price widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, as well as a user-created widget 149-6, ● Widget creator module 150 for creating the user-created widget 149-6, ● Search module 151, ● A video and music player module 152 that integrates a video player module and a music player module, ● A memo module 153, ● A map module 154, and / or ● An online video module 155.

[0106] Examples of other applications 136 that are optionally stored in the memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-compatible applications, encryption, digital rights management, voice recognition, and voice replication.

[0107] Together with the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134, the contact module 137 is optionally used to manage an address book or contact list (e.g., stored in the application internal state 192 of the contact module 137 in the memory 102 or the memory 370), which includes adding a name(s) to the address book, deleting a name(s) from the address book, associating a phone number(s), an email address(es), an address(es), or other information with a name, associating an image with a name, classifying and sorting names, providing a phone number or email address to initiate and / or facilitate communication by phone 138, the video conferencing module 139, email 140, or IM 141, etc. It is used to manage the address book or contact list.

[0108] The telephone module 138 is used in cooperation with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134, optionally, for the input of a series of characters corresponding to a telephone number, access to one or more telephone numbers in the contact module 137, modification of the input telephone number, dialing of an individual telephone number, execution of a call, and disconnection and call hold at the end of a call. As described above, wireless communication optionally uses any of a plurality of communication standards, protocols, and technologies.

[0109] The videoconference module 139 includes executable instructions for starting, executing, and ending a videoconference between the user and one or more other participants according to the user's instructions in cooperation with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the optical sensor 164, the optical sensor controller 158, the contact / motion module 130, the graphic module 132, the text input module 134, the contact module 137, and the telephone module 138.

[0110] The email client module 140 includes executable instructions for creating, sending, receiving, and managing emails according to the user's instructions in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134. In cooperation with the image management module 144, the email client module 140 makes it very easy to create and send emails with still or moving images captured by the camera module 143.

[0111] The instant messaging module 141, in cooperation with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphic module 132, and text input module 134, includes executable instructions for input of a series of characters corresponding to an instant message, modification of previously input characters, (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for instant messages based on telephone communication, or XMPP, SIMPLE, or IMPS for instant messages based on the Internet) transmission of individual instant messages, reception of instant messages, and viewing of received instant messages. In some embodiments, the instant messages transmitted and / or received optionally include graphics, photos, audio files, video files, and / or other attached files as supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephone communication-based messages (e.g., messages transmitted using SMS or MMS) and Internet-based messages (e.g., messages transmitted using XMPP, SIMPLE, or IMPS).

[0112] In cooperation with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, GPS module 135, map module 154, and music player module, the training support module 142 includes executable instructions that create a training (e.g., having time, distance, and / or calorie burn goals), communicate with a training sensor (sports device), receive training sensor data, calibrate sensors used to monitor the training, select and play music for the training, and display, store, and transmit training data.

[0113] The camera module 143 includes executable instructions for capturing still images or videos (including video streams) and storing them in the memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from the memory 102, in cooperation with the touch screen 112, display controller 156, optical sensor(s) 164, optical sensor controller 158, contact / motion module 130, graphic module 132, and image management module 144.

[0114] The image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise operating on, labeling, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or videos, in cooperation with the touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, and camera module 143.

[0115] The browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching for, linking to, receiving, and displaying a web page or a part thereof, as well as attached files and other files linked to the web page, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134.

[0116] The calendar module 148 includes executable instructions for creating, displaying, modifying, and storing a calendar and data associated with the calendar (e.g., calendar items, to-do lists, etc.) according to user instructions, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the text input module 134, the email client module 140, and the browser module 147.

[0117] The widget module 149 cooperates with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the text input module 134, and the browser module 147, and optionally, mini-applications (e.g., weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) that are downloaded and used by the user, or mini-applications created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).

[0118] The widget creator module 150 cooperates with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the text input module 134, and the browser module 147, and is used by the user to optionally create a widget (e.g., make a user-specified portion of a web page into a widget).

[0119] The search module 151 cooperates with the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134, and includes executable instructions for searching for characters, music, sound, images, videos, and / or other files in the memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) according to the user's instructions.

[0120] The video and music player module 152, in cooperation with the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the audio circuit 110, the speaker 111, the RF circuit 108, and the browser module 147, includes executable instructions that enable a user to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing videos (e.g., on the touch screen 112 or on an external display connected via the external port 124). In some embodiments, the device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0121] The memo module 153, in cooperation with the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134, includes executable instructions for creating and managing memos, to-do lists, etc. according to the user's instructions.

[0122] The map module 154, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the text input module 134, the GPS module 135, and the browser module 147, is optionally used to receive, display, modify, and store maps and map-related data (e.g., driving routes, data on stores and other attractions near a particular location or its vicinity, and other location-based data) according to the user's instructions.

[0123] The online video module 155 cooperates with the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132, the audio circuit 110, the speaker 111, the RF circuit 108, the text input module 134, the email client module 140, and the browser module 147 to enable the user to access a specific online video, browse a specific online video, receive it (e.g., by streaming and / or downloading), play it (e.g., on the touch screen or on an external display connected via the external port 124), send an email having a link to a specific online video, and perform other management of online videos in one or more file formats such as H.264. In some embodiments, instead of the email client module 140, the instant messaging module 141 is used to send a link to a specific online video. For additional explanation of the online video application, reference is made to U.S. Provisional Patent Application No. 60 / 936,562, filed Jun. 20, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", and U.S. Patent Application No. 11 / 968,067, filed Dec. 31, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", the entire contents of which are incorporated herein by reference.

[0124] The modules and applications identified above each correspond to a set of executable instructions that perform one or more of the functions described above and the methods described in this application (e.g., the computer-executed methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as respective software programs (e.g., computer programs including instructions), procedures, or modules, and thus in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. For example, the video player module may optionally be combined with the music player module to form a single module (e.g., the video and music player module 152 of FIG. 1A). In some embodiments, the memory 102 optionally stores a subset of the modules and data structures identified above. Further, the memory 102 optionally stores additional modules and data structures not described above.

[0125] In some embodiments, the device 100 is a device in which the operation of a set of default functions in the device is performed only via a touch screen and / or a touch pad. By using the touch screen and / or the touch pad as the main input control device for the device 100 to operate, the number of physical input control devices (push buttons, dials, etc.) on the device 100 is optionally reduced.

[0126] The set of default functions that are executed only through the touch screen and / or touch pad optionally includes navigation between user interfaces. In some embodiments, the touch pad navigates the device 100 from any user interface displayed on the device 100 to the main menu, home menu, or root menu when touched by the user. In such embodiments, the "menu button" is implemented using the touch pad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touch pad.

[0127] FIG. 1B is a block diagram showing exemplary components for event processing according to some embodiments. In some embodiments, the memory 102 (FIG. 1A) or 370 (FIG. 3) includes an event sorting unit 170 (e.g., within the operating system 126) and an individual application 136-1 (e.g., any of the aforementioned applications 137-151, 155, 380-390).

[0128] The event sorting unit 170 receives event information and determines the application 136-1 that distributes the event information and the application view 191 of the application 136-1. The event sorting unit 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, the application 136-1 includes an application internal state 192 that indicates the current application view(s) displayed on the touch-sensitive display 112 when the application is active or running. In some embodiments, the device / global internal state 157 is used by the event sorting unit 170 to determine which application(s) is / are currently active, and the application internal state 192 is used by the event sorting unit 170 to determine the application view 191 to which the event information is distributed.

[0129] In some embodiments, the application internal state 192 includes additional information such as resume information to be used when application 136-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 136-1, a state queue that enables the user to return to a previous state or view of application 136-1, and a redo / undo queue of previous actions performed by the user, among one or more of these.

[0130] The event monitor 171 receives event information from the peripheral device interface 118. The event information includes information regarding sub-events (e.g., a user touch as part of a multi-touch gesture on the touch-sensitive display 112). The peripheral device interface 118 transmits information received from the I / O subsystem 106, or sensors such as the proximity sensor 166, accelerometer(s) 168, and / or microphone 113 (via the audio circuit 110). Information that the peripheral device interface 118 receives from the I / O subsystem 106 includes information from the touch-sensitive display 112 or a touch-sensitive surface.

[0131] In some embodiments, the event monitor 171 transmits requests to the peripheral device interface 118 at predetermined intervals. In response, the peripheral device interface 118 transmits event information. In other embodiments, the peripheral device interface 118 transmits event information only when there is an important event (e.g., receipt of an input that exceeds a predetermined noise threshold and / or exceeds a predetermined duration).

[0132] In some embodiments, the event sorter 170 also includes a hit view determination module 172 and / or an active event recognition unit determination module 173.

[0133] The hit view determination module 172 provides software procedures for determining where in one or more views a sub-event occurs when the touch-sensitive display 112 is displaying two or more views. A view is composed of control devices and other elements that a user can view on the display.

[0134] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, within which information is displayed and touch-based gestures occur. The application view (of an individual application) in which a touch is detected optionally corresponds to a program level within the program hierarchy or view hierarchy of the application. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events recognized as appropriate input is optionally determined based at least in part on the hit view of the initial touch that initiates a touch-based gesture.

[0135] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has a plurality of hierarchically structured views, the hit view determination module 172 identifies the hit view as the lowest-level view within the hierarchy in which the sub-event is to be processed. In most situations, the hit view is the lowest-level view in which the start sub-event (e.g., the first sub-event in a series of sub-events that form an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source as the touch or input source identified as the hit view.

[0136] The active event recognition unit determination module 173 determines which view(s) within the view hierarchy should receive a particular series of sub-events. In some embodiments, the active event recognition unit determination module 173 determines that only the hit view should receive a particular series of sub-events. In other embodiments, the active event recognition unit determination module 173 determines that all views including the physical location of the sub-events are views that are actively involved, and thus determines that all views that are actively involved should receive a particular series of sub-events. In other embodiments, even if a touch sub-event is completely limited to an area associated with one particular view, the upper-level views within the hierarchy remain views that are actively involved.

[0137] The event dispatcher module 174 dispatches event information to the event recognition unit (e.g., event recognition unit 180). In embodiments including the active event recognition unit determination module 173, the event dispatcher module 174 distributes event information to the event recognition unit determined by the active event recognition unit determination module 173. In some embodiments, the event dispatcher module 174 stores the event information retrieved by the individual event receiver 182 in the event queue.

[0138] In some embodiments, the operating system 126 includes the event sorter 170. Alternatively, the application 136-1 includes the event sorter 170. In still other embodiments, the event sorter 170 is a stand-alone module or part of another module stored in the memory 102 such as the touch / motion module 130.

[0139] In some embodiments, application 136-1 includes a plurality of event processing units 190 and one or more application views 191, each including instructions for processing touch events that occur within an individual view of the user interface of the application. Each application view 191 of application 136-1 includes one or more event recognition units 180. Typically, an individual application view 191 includes a plurality of event recognition units 180. In other embodiments, one or more of the event recognition units 180 are part of a separate module, such as a user interface kit, or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, an individual event processing unit 190 includes one or more of event data 179 received from data update unit 176, object update unit 177, GUI update unit 178, and / or event sorting unit 170. The event processing unit 190 optionally utilizes or invokes the data update unit 176, object update unit 177, or GUI update unit 178 to update the internal state 192 of the application. Alternatively, one or more of the application views 191 include one or more respective event processing units 190. Also, in some embodiments, one or more of the data update unit 176, object update unit 177, and GUI update unit 178 are included in an individual application view 191.

[0140] An individual event recognition unit 180 receives event information (e.g., event data 179) from event sorting unit 170 and identifies an event from the event information. The event recognition unit 180 includes an event receiving unit 182 and an event comparing unit 184. In some embodiments, the event recognition unit 180 also includes at least a subset of metadata 183 and event distribution instructions 188 (optionally including sub-event distribution instructions).

[0141] The event receiving unit 182 receives event information from the event sorting unit 170. The event information includes sub-events, for example, information about a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information such as the location of the sub-event. When the sub-event is related to the movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the posture of the device).

[0142] The event comparison unit 184 compares the event information with the definition of a defined event or sub-event, and based on the comparison, determines an event or sub-event, or determines or updates the state of an event or sub-event. In some embodiments, the event comparison unit 184 includes an event definition 186. The event definition 186 includes definitions of events (e.g., a predefined series of sub-events) such as event 1 (187-1) and event 2 (187-2). In some embodiments, the sub-events within the event (187-1 and / or 187-2) include, for example, a touch start, a touch end, a touch movement, a touch cancellation, and multiple touches. In one example, the definition of event 1 (187-1) is a double-tap on a displayed object. The double-tap includes, for example, a first touch (touch start) on the displayed object for a predetermined stage, a first lift-off (touch end) for the predetermined stage, a second touch (touch start) on the displayed object for the predetermined stage, and a second lift-off (touch end) for the predetermined stage. In another example, the definition of event 2 (187-2) is a drag on a displayed object. The drag includes, for example, a touch (or contact) on the displayed object for a predetermined stage, a movement of the touch across the touch-sensitive display 112, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event processing units 190.

[0143] In some embodiments, the event definition 186 includes a definition of an event for an individual user interface object. In some embodiments, the event comparison unit 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, within an application view in which three user interface objects are displayed on the touch-sensitive display 112, when a touch is detected on the touch-sensitive display 112, the event comparison unit 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each of the displayed objects is associated with an individual event processing unit 190, the event comparison unit determines which event processing unit 190 should be activated using the result of the hit test. For example, the event comparison unit 184 selects the event processing unit associated with the sub-event and object that triggered the hit test.

[0144] In some embodiments, the definition of an individual event 187 also includes a delay action that delays the delivery of event information until it is determined whether a series of sub-events corresponds to the event type of the event recognition unit.

[0145] If the individual event recognition unit 180 determines that a series of sub-events does not match any of the events of the event definition 186, the individual event recognition unit 180 enters a state of event impossible, event failure, or event end, and then ignores the next sub-event of the touch-based gesture. In this situation, if there is another event recognition unit that remains active for the hit view, that event recognition unit continues to track and process the sub-events of the ongoing touch-based gesture.

[0146] In some embodiments, an individual event recognition unit 180 includes metadata 183 having configurable properties, flags, and / or lists indicating how the event delivery system should actively participate in event recognition units that perform sub - event delivery. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists indicating how event recognition units interact with each other or how they can interact with each other. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists indicating whether sub - events are delivered to various levels in the view hierarchy or program hierarchy.

[0147] In some embodiments, an individual event recognition unit 180 activates an event processing unit 190 associated with an event when one or more specific sub - events of the event are recognized. In some embodiments, an individual event recognition unit 180 delivers event information associated with the event to the event processing unit 190. Activating the event processing unit 190 is separate from sending (and deferring sending) sub - events to individual hit views. In some embodiments, the event recognition unit 180 sets a flag associated with the recognized event, and the event processing unit 190 associated with that flag catches the flag and executes a predefined process.

[0148] In some embodiments, the event delivery instruction 188 includes a sub - event delivery instruction that delivers event information about sub - events without activating the event processing unit. Instead, the sub - event delivery instruction delivers event information to an event processing unit associated with a series of sub - events or to a view that is actively involved. The event processing unit associated with a series of sub - events or an actively involved view receives the event information and executes a predetermined process.

[0149] In some embodiments, the data update unit 176 creates and updates data used in application 136-1. For example, the data update unit 176 updates the phone numbers used in the contact module 137 or stores video files used in the video player module. In some embodiments, the object update unit 177 creates and updates objects used in application 136-1. For example, the object update unit 177 creates a new user interface object or updates the position of a user interface object. The GUI update unit 178 updates the GUI. For example, the GUI update unit 178 prepares display information and sends the display information to the graphic module 132 for display on the touch-sensitive display.

[0150] In some embodiments, the event processing unit(s) 190 includes or has access to the data update unit 176, the object update unit 177, and the GUI update unit 178. In some embodiments, the data update unit 176, the object update unit 177, and the GUI update unit 178 are included in a single module of the individual application 136-1 or the application view 191. In other embodiments, they are included in two or more software modules.

[0151] The foregoing description regarding event processing of a user's touch on the touch-sensitive display also applies to other forms of user input for operating the multifunctional device 100 using an input device, but it should be understood that not all of them are initiated on the touch screen. For example, the movement of a mouse and the pressing of a mouse button, optionally in conjunction with a single or multiple presses or holds of a keyboard, the movement of a contact such as a tap, drag, scroll on a touch pad, pen stylus input, the movement of the device, spoken commands, detected eye movement, biometric input, and / or any combination thereof are optionally utilized as input corresponding to sub-events that define events to be recognized.

[0152] Figure 2 shows a portable multifunctional device 100 having a touch screen 112, according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment, as well as in other embodiments described below, a user can select one or more of those graphics by performing a gesture on the graphics using, for example, one or more fingers 202 (not drawn to scale in the figure) or one or more styli 203 (not drawn to scale in the figure). In some embodiments, the selection of one or more graphics is performed when the user interrupts contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, from right to left, upward and / or downward), and / or rolling (from right to left, from left to right, upward and / or downward) of a finger in contact with the device 100. In some implementations or situations, an accidental contact with a graphic does not select that graphic. For example, if the gesture corresponding to selection is a tap, a swipe gesture that sweeps over an application icon does not optionally select the corresponding application.

[0153] Device 100 also optionally includes one or more physical buttons, such as a "home" button or a menu button 204. As described above, the menu button 204 is optionally used to navigate to any application 136 within a set of applications optionally running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touch screen 112.

[0154] In some embodiments, device 100 includes a touch screen 112, a menu button 204, a push button 206 for turning the device on / off and locking the device, volume adjustment button(s) 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and holding it in the pressed state for a predefined period, to lock the device by pressing the button and releasing it before a predefined time has elapsed, and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also accepts verbal input via a microphone 113 to activate or deactivate some functions. Device 100 optionally also includes one or more contact intensity sensors 165 for detecting the intensity of contact on the touch screen 112 and / or one or more haptic output generators 167 for generating haptic output to the user of device 100.

[0155] FIG. 3 is a block diagram of an exemplary multifunctional device having a display and a touch sensing surface, according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a child's learning toy), game system, or control device (e.g., a home or business controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more networks or other communication interfaces 360, memory 370, and one or more communication buses 320 interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communications between system components. Device 300 includes an input / output (I / O) interface 330 that includes a display 340, which is typically a touch screen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350, a touch pad 355, a tactile output generator 357 that generates tactile outputs on device 300 (e.g., similar to the tactile output generator(s) 167 described above with reference to FIG. 1A), and a sensor 359 (e.g., light, acceleration, proximity, touch sensing, and / or a contact intensity sensor similar to the contact intensity sensor(s) 165 described above with reference to FIG. 1A). Memory 370 includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices located remotely from the CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or a subset of, the programs, modules, and data structures stored in memory 102 of portable multifunctional device 100 (FIG. 1A). Further, memory 370 optionally stores additional programs, modules, and data structures that are not present in memory 102 of portable multifunctional device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk authoring module 388, and / or a spreadsheet module 390, whereas memory 102 of portable multifunctional device 100 (FIG. 1A) optionally does not store these modules.

[0156] Each of the elements identified above in FIG. 3 is optionally stored in one or more of the memory devices described above. Each of the modules identified above corresponds to a set of instructions for performing the functions described above. The modules or computer programs identified above (e.g., a set of instructions or including instructions) need not be implemented as separate software programs (e.g., computer programs including instructions), procedures, or modules, and thus, in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. In some embodiments, memory 370 optionally stores a subset of the modules and data structures identified above. Further, memory 370 optionally stores additional modules and data structures not described above.

[0157] Next, optionally direct attention to an embodiment of a user interface, for example, implemented on portable multifunctional device 100.

[0158] FIG. 4A shows an exemplary user interface of a menu of applications on a portable multifunctional device 100 according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, the user interface 400 includes the following elements, or subsets or supersets thereof. ● Signal strength indicator(s) 402 for wireless communication(s) such as cellular signal and Wi-Fi signal, ● Time 404, ● Bluetooth indicator 405, ● Battery status indicator 406, ● A tray 408 having icons of frequently used applications such as ○ An icon 416 of the phone module 138 labeled "Phone", optionally including an indicator 414 of the number of missed calls or voice mail messages, ○ An icon 418 of the email client module 140 labeled "Mail", optionally including an indicator 410 of the number of unread emails, ○ An icon 420 of the browser module 147 labeled "Browser", and ○ An icon 422 for the video and music player module 152, also referred to as the iPod (a trademark of Apple Inc.) module 152, labeled "iPod", and ● Icons of other applications such as ○ An icon 424 of the IM module 141 labeled "Message", ○ An icon 426 of the calendar module 148 labeled "Calendar", ○ An icon 428 of the image management module 144 labeled "Photos", ○ An icon 430 of the camera module 143 labeled "Camera", ○ An icon 432 of the online video module 155 labeled "Online Video", ○ The icon 434 of the stock price widget 149-2, labeled with "Stock Price", ○ The icon 436 of the map module 154, labeled with "Map", ○ The icon 438 of the weather widget 149-1, labeled with "Weather", ○ The icon 440 of the alarm clock widget 149-4, labeled with "Clock", ○ The icon 442 of the training support module 142, labeled with "Training Support", ○ The icon 444 of the memo module 153, labeled with "Memo", and ○ The icon 446 of the settings application or module, labeled with "Settings", which provides access to the settings of the device 100 and its various applications 136.

[0159] Note that the icon labels shown in FIG. 4A are merely exemplary. For example, the icon 422 of the video and music player module 152 is labeled with "Music" or "Music Player". Other labels may optionally be used for the various application icons. In some embodiments, the label for an individual application icon includes the name of the application corresponding to the individual application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to that particular application icon.

[0160] FIG. 4B shows an exemplary user interface on a device (e.g., device 300 of FIG. 3) having a touch sensing surface 451 (e.g., the tablet or touch pad 355 of FIG. 3) separate from the display 450 (e.g., touch screen display 112). The device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of the sensors 359) that detect the intensity of contact on the touch sensing surface 451, and / or one or more haptic output generators 357 that generate haptic output to the user of the device 300.

[0161] Some of the following examples are given with reference to inputs on a touch screen display 112 (where the touch sensing surface and the display are combined), but in some embodiments, the device detects inputs on a touch sensing surface separate from the display, as shown in FIG. 4B. In some embodiments, the touch sensing surface (e.g., 451 in FIG. 4B) has a primary axis (e.g., 452 in FIG. 4B) that corresponds to a primary axis (e.g., 453 in FIG. 4B) on the display (e.g., 450). According to these embodiments, the device detects contact (e.g., 460 and 462 in FIG. 4B) with the touch sensing surface 451 at locations (e.g., in FIG. 4B, 460 corresponds to 468 and 462 corresponds to 470) that correspond to each location on the display. In this way, user inputs (e.g., contacts 460 and 462 and their movements) detected by the device on the touch sensing surface (e.g., 451 in FIG. 4B) are used by the device to operate the user interface on the display (e.g., 450 in FIG. 4B) of the multifunctional device when the touch sensing surface is separate from the display. It should be understood that a similar method is optionally used for other user interfaces described herein.

[0162] In addition, while the following examples are given mainly with reference to finger inputs (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs are replaced by inputs from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be a mouse click (e.g., instead of a contact), followed by a mouse click with movement of the cursor along the path of the swipe (e.g., instead of movement of a contact). As another example, a tap gesture may optionally be a mouse click while the cursor is located on the location of the tap gesture (e.g., instead of detecting a contact and then ceasing to detect the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger contacts may optionally be used simultaneously.

[0163] FIG. 5A shows an exemplary personal electronic device 500. The device 500 includes a body 502. In some embodiments, the device 500 can include some or all of the functions described with respect to devices 100 and 300 (e.g., FIGS. 1A-4B). In some embodiments, the device 500 has a touch-sensitive display screen 504, hereinafter referred to as touch screen 504. Alternatively, or in addition to the touch screen 504, the device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, the touch screen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors that detect the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 504 (or touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of the device 500 can respond to the touch(es) based on its intensity, which means that touches of different intensities can call different user interface operations on the device 500.

[0164] Exemplary techniques for detecting and processing touch intensity are described, for example, in International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, published as International Publication No. WO 2013 / 169849, "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application", and International Patent Application No. PCT / US2013 / 069483, filed Nov. 11, 2013, published as International Publication No. WO 2014 / 105276, "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships", each of which is hereby incorporated by reference in its entirety.

[0165] In some embodiments, device 500 includes one or more input mechanisms 506 and 508. Input mechanisms 506 and 508, if included, may be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 includes one or more attachment mechanisms. Such attachment mechanisms, if included, can enable device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms enable a user to wear device 500.

[0166] FIG. 5B shows an exemplary personal electronic device 500. In some embodiments, device 500 can include some or all of the components described with respect to FIGS. 1A, 1B, and 3. Device 500 has a bus 512 that operably couples an I / O section 514 to one or more computer processors 516 and a memory 518. The I / O section 514 can be connected to a display 504, which can have a touch sensing component 522 and optionally an intensity sensor 524 (e.g., a contact intensity sensor). Additionally, the I / O section 514 can be connected to a communication unit 530 that receives application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 500 can include input mechanisms 506 and / or 508. Input mechanism 506 can optionally be, for example, a rotatable input device or a depressible and rotatable input device. In some examples, input mechanism 508 can optionally be a button.

[0167] In some examples, input mechanism 508 can optionally be a microphone. Personal electronic device 500 can optionally include various sensors such as a GPS sensor 532, an accelerometer 534, a direction sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which can be operably connected to the I / O section 514.

[0168] The memory 518 of the personal electronic device 500 can include one or more non-transitory computer-readable storage media for storing computer-executable instructions, which, when executed by one or more computer processors 516, can cause the computer processor to execute, for example, the techniques described below, including processes 700, 800, 900, 1100, and 1300 (FIGS. 7-9, FIG. 11, and FIG. 13). A computer-readable storage media can be any media that can tangibly contain or store computer-executable instructions used by or associated with an instruction execution system, apparatus, or device. In some embodiments, the storage media is a transitory computer-readable storage media. In some embodiments, the storage media is a non-transitory computer-readable storage media. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash, solid-state drives, and the like. The personal electronic device 500 is not limited to the components and configurations of FIG. 5B and can include other or additional components in a plurality of configurations.

[0169] As used herein, the term "affordance" refers to a user interaction graphical user interface object that is optionally displayed on the display screen of devices 100, 300, and / or 500 (FIGS. 1A, 3, and 5A-5B). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each optionally constitute an affordance.

[0170] As used herein, the term "focus selector" refers to an input element that indicates the current portion of the user interface with which the user is interacting. In some implementations that include a cursor or other location marker, the cursor acts as the "focus selector" such that while the cursor is positioned over a particular user interface element (e.g., a button, window, slider, or other user interface element), when an input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 of FIG. 3 or touch-sensitive surface 451 of FIG. 4B), the particular user interface element is adjusted according to the detected input. In some implementations that include a touch screen display (e.g., touch-sensitive display system 112 of FIG. 1A or touch screen 112 of FIG. 4A) that enables direct interaction with user interface elements on the touch screen display, the detected contact on the touch screen acts as the "focus selector" such that when an input (e.g., a press input by contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted according to the detected input. In some implementations, the focus is moved from one region of the user interface to another region of the user interface without movement of the corresponding cursor or movement of the contact on the touch screen display (e.g., by using the tab key or arrow keys to move the focus from one button to another button), and in these implementations, the focus selector moves in accordance with the movement of the focus between various regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or a contact on the touch screen display) that is controlled by the user to communicate (e.g., by indicating to the device the user interface element with which the user intends to interact through it) the user's intended interaction with the user interface.For example, the location of a focus selector (e.g., a cursor, a contact, or a selection box) over an individual button while a press input is detected on a touch sensing surface (e.g., a touch pad or a touch screen) indicates that the user intends to activate that individual button (as opposed to other user interface elements shown on the device's display).

[0171] As used in this specification and the claims, the term "characteristic strength" of a contact refers to the characteristics of that contact based on one or more strengths of the contact. In some embodiments, the characteristic strength is based on a plurality of strength samples. The characteristic strength is optionally based on a set of strength samples collected during a predetermined time (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined number of strength samples, i.e., a predetermined event (e.g., after detecting the contact, before detecting the lift-off of the contact, before or after detecting the start of movement of the contact, before detecting the end of the contact, before or after detecting an increase in the strength of the contact, and / or before or after detecting a decrease in the strength of the contact). The characteristic strength of a contact is optionally based on one or more of the maximum value of the strength of the contact, the mean value of the strength of the contact, the average value of the strength of the contact, the top 10 percentile value of the strength of the contact, the median value of the strength of the contact, the top 90 percent value of the strength of the contact, etc. In some embodiments, the duration of the contact is used when determining the characteristic strength (e.g., when the characteristic strength is the average of the strength of the contact over time). In some embodiments, the characteristic strength is compared to a set of one or more strength thresholds to determine whether an action has been performed by a user. For example, the set of one or more strength thresholds optionally includes a first strength threshold and a second strength threshold. In this example, a contact having a characteristic strength that does not exceed the first threshold results in a first action, a contact having a characteristic strength that exceeds the first strength threshold but does not exceed the second strength threshold results in a second action, and a contact having a characteristic strength that exceeds the second threshold results in a third action. In some embodiments, the comparison between the characteristic strength and one or more thresholds is not used to determine whether to perform a first action or a second action, but is used to determine whether to perform one or more actions (e.g., whether to perform an individual action or whether to refrain from performing an individual action).

[0172] In some embodiments, for the purpose of determining characteristic intensity, a portion of a gesture is identified. For example, the touch sensing surface optionally receives a continuous swipe contact that transitions from a start location to reach an end location, where the intensity of the contact increases at that location. In this example, the characteristic intensity of the contact at the end location is optionally based on only a portion of the continuous swipe contact (e.g., only the portion of the swipe contact at the end location) rather than the entire swipe contact. In some embodiments, optionally, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of a non - weighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the width of the swipe contact intensity for the purpose of determining the characteristic intensity.

[0173] The intensity of a contact on the touch sensing surface is optionally characterized relative to one or more intensity thresholds such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold typically corresponds to the intensity at which the device performs an operation associated with clicking a button of a physical mouse or a trackpad. In some embodiments, the deep press intensity threshold typically corresponds to the intensity at which the device performs an operation different from the operation associated with clicking a button of a physical mouse or a trackpad. In some embodiments, when a contact having a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold below which the contact is not detected) is detected, the device moves the focus selector in accordance with the movement of the contact on the touch sensing surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise specified, these intensity thresholds are consistent among various sets of user interface values.

[0174] An increase in the characteristic intensity of contact from an intensity below a light pressing intensity threshold to an intensity between the light pressing intensity threshold and a deep pressing intensity threshold may be referred to as an input of "light pressing". An increase in the characteristic intensity of contact from an intensity below the deep pressing intensity threshold to an intensity above the deep pressing intensity threshold may be referred to as an input of "deep pressing". An increase in the characteristic intensity of contact from an intensity below the contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light pressing intensity threshold may be referred to as the detection of contact on the touch surface. A decrease in the characteristic intensity of contact from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold may be referred to as the detection of lift-off of contact from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0175] In some embodiments described herein, one or more operations are performed in response to detecting a gesture that includes an individual pressing input, or in response to detecting an individual pressing input performed by an individual contact (or multiple contacts), and the individual pressing input is detected at least in part based on detecting an increase in the intensity of a contact (or multiple contacts) above a pressing input intensity threshold. In some embodiments, the individual operation is performed in response to detecting an increase in the intensity of an individual contact above the pressing input intensity threshold (e.g., the "downstroke" of an individual pressing input). In some embodiments, the pressing input includes an increase in the intensity of an individual contact above the pressing input intensity threshold and a subsequent decrease in the intensity of the contact below the pressing input intensity threshold, and the individual operation is performed in response to detecting a subsequent decrease in the intensity of the individual contact below the pressing input threshold (e.g., the "upstroke" of an individual pressing input).

[0176] In some embodiments, the device employs intensity hysteresis to avoid spurious inputs sometimes referred to as "jitter", and the device defines or selects a hysteresis intensity threshold having a predefined relationship to the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some other reasonable percentage of the press input intensity threshold). Thus, in some embodiments, a press input includes an increase in the intensity of an individual contact above the press input intensity threshold, and a subsequent decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and an individual operation is performed in response to detecting a subsequent decrease in the intensity of an individual contact below the hysteresis intensity threshold (e.g., the "upstroke" of an individual press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of a contact from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, and optionally, a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and an individual operation is performed in response to detecting the press input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, depending on the situation).

[0177] For ease of explanation, the description of an operation performed in response to a press input associated with a press input intensity threshold, or a gesture including a press input, is optionally triggered in response to detecting any of an increase in the intensity of a contact above the press input intensity threshold, an increase in the intensity of a contact from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, a decrease in the intensity of a contact below the press input intensity threshold, and / or a decrease in the intensity of a contact below the hysteresis intensity threshold corresponding to the press input intensity threshold. Further, in examples where an operation is described as being performed in response to detecting a decrease in the intensity of a contact below the press input intensity threshold, the operation is optionally performed in response to detecting a decrease in the intensity of a contact corresponding to and below a lower hysteresis intensity threshold corresponding to the press input intensity threshold.

[0178] Next, attention is directed to embodiments of a user interface (“UI”) and related processes implemented on an electronic device such as the portable multifunctional device 100, device 300, or device 500.

[0179] Figures 6A - 6BJ illustrate exemplary user interfaces for modifying visual content in a media, according to some embodiments. The user interfaces in these figures are used to explain processes described below, including the processes in FIGS. 7, 8, and 9. The examples in FIGS. 6A - 6BJ are described with respect to touch input on a touch - sensitive surface, but it should be understood that taps, long - presses, press - and - holds, swipes, and other touch gestures can be replaced with other inputs directed to the relevant user interface elements. For example, a tap can be replaced with a mouse click, a swipe can be replaced with a click - and - drag, a double - tap can be replaced with a double - click, and / or a long - press (and / or press - and - hold) can be replaced with a right - click or a click while holding a modifier key. Similarly, air gestures such as pinching two fingers together or touching a finger to a hand can replace a tap, pinching two fingers together followed by a movement can replace a touch - and - drag, a double - pinch can replace a double - tap, and a long - pinch can replace a long - tap or a tap - and - hold. In some embodiments, the location within the user interface to which an input is directed is determined based on a direct touch (e.g., a tap, double - tap, long - press, press - and - hold, or swipe on a user interface element), but the location to which an input is directed may also be determined based on other indications of the user's intent, such as the location of a displayed cursor or the location to which the user's line of sight is directed.

[0180] FIG. 6A shows a computer system 600 (e.g., an electronic device) that displays a camera user interface including a live preview 630 that optionally extends from the top of the display of the computer system 600 to the bottom of the display of the computer system 600. In some embodiments, the computer system 600 optionally includes one or more features of device 100, device 300, or device 500. In some embodiments, the computer system 600 is a tablet, phone, laptop, desktop, and / or camera.

[0181] The live preview 630 is a representation of the field of view (the "FOV") of one or more cameras of the computer system 600. In some embodiments, the live preview 630 is a representation of a partial FOV. In some embodiments, the live preview 630 is based on an image detected by one or more camera sensors. In some embodiments, the computer system 600 uses multiple camera sensors to capture images and combines them to display the live preview 630. In some embodiments, the computer system 600 uses a single camera sensor to capture an image and display the live preview 630.

[0182] The camera user interface of FIG. 6A includes an indicator region 602 and a control region 606, which are positioned relative to the live preview 630 such that indicators and controls can be displayed simultaneously with the live preview 630. The camera display region 604 does not substantially overlap with the indicators and / or controls. As shown in FIG. 6A, the camera user interface also includes a visual boundary 608 that indicates the boundary between the indicator region 602 and the camera display region 604 and the boundary between the camera display region 604 and the control region 606.

[0183] As shown in FIG. 6A, the indicator area 602 includes indicators such as a flash indicator 602a, a mode setting indicator 602b, and an animation image indicator 602c. The flash indicator 602a indicates whether the flash mode is on (e.g., activated), off (e.g., deactivated), or in another mode (e.g., auto mode). In FIG. 6A, the flash indicator 602a indicates that the flash mode is off, and thus, the flash operation is not used when the computer system 600 is capturing media. Further, the mode - setting indicator 602b, when selected, causes the computer system 600 to replace the camera mode control 620 with a camera setting control for setting a plurality of settings of the currently selected camera mode (e.g., the photo camera mode of FIG. 6A). The animation image indicator 602c indicates whether the camera is configured to capture a single image or a plurality of images (e.g., in response to detecting a request to capture media). In some embodiments, the indicator area 602 includes an overlay that is superimposed on the live preview 630 and optionally colored (e.g., gray, translucent).

[0184] As shown in FIG. 6A, the camera display area 604 includes a live preview 630 and a zoom control (e.g., an affordance) 622. The zoom control 622 includes a 0.5× zoom control 622a, a 1× zoom control 622b, and a 2× zoom control 622c. As shown in FIG. 6A, the 1x zoom control 622b is enlarged compared to the other zoom controls, which indicates that the 1x zoom control 622b is selected and the computer system 600 is displaying the live preview 630 at the "1x" zoom level. In some embodiments, the computer system 600 displays the 1x zoom control 622b as being selected by displaying it in a color different from the other zoom controls 622.

[0185] As shown in FIG. 6A, the control region 606 includes representations of a camera mode control 620, a shutter control 610, a camera switcher control 614, and a media collection 612. In FIG. 6A, camera mode controls 620a-620e including a panorama mode control 620a, a portrait mode control 620b, a photo mode control 620c, a video mode control 620d, and a cinematic video mode control 620e are shown. As shown in FIG. 6A, the photo mode control 620c is selected, which is indicated by the bolded photo mode control 620c. When the photo mode control 620c is selected, the computer system 600 starts (e.g., and / or captures) the capture of photo media (e.g., still images) in response to detecting an input directed by the computer system 600 to the shutter control 610. The photo media captured by the computer system 600 represents a live preview 630 that is displayed when the input is directed to the shutter control 610. In some embodiments, in response to detecting an input directed to the panorama mode control 620a, the computer system 600 starts the capture of panorama media (e.g., panorama photos). In some embodiments, in response to detecting an input directed to the portrait mode control 620b, the computer system 600 starts the capture of portrait media (e.g., still photos, still photos with applied blur). In some embodiments, in response to detecting an input directed to the video mode control 620d, the computer system 600 starts the capture of video media (e.g., video). In some embodiments, indicators and / or controls displayed on the camera user interface are based on the selected mode (e.g., and / or a mode configured to operate based on the camera mode selected by the computer system 600).

[0186] In FIG. 6A, when activated, shutter control 610 causes computer system 600 to capture media (e.g., a photograph when shutter control 610 is activated in FIG. 6A) using one or more camera sensors based on the current state of live preview 630 and the current state of the camera application (e.g., which camera mode is selected). The captured media is stored locally in computer system 600 and / or transmitted to a remote server for storage. When activated, camera switcher control 614 causes computer system 600 to switch, such as by switching between a rear camera sensor and a front camera sensor, to show the view of different cameras within live preview 630. The representation of media collection 612 shown in FIG. 6A is a representation of media (e.g., images, videos) most recently captured by computer system 600. In some embodiments, in response to detecting an input to media collection 612, computer system 600 displays a user interface similar to the user interface shown in FIG. 7 (described below). In some embodiments, indicator region 602 includes an overlay that is superimposed on live preview 630 and optionally colored (e.g., gray, translucent).

[0187] As described above, FIGS. 6A-6BJ illustrate exemplary user interfaces for modifying visual content according to some embodiments. In particular, FIGS. 6A-6AC illustrate exemplary embodiments in which a synthetic (e.g., simulated, computer-generated) depth-of-field effect is applied to the visual content of the media currently being captured. The synthetic depth-of-field effect is applied automatically (e.g., without responding to one or more inputs) and / or in response to user input. When the synthetic depth-of-field effect is applied automatically, the computer system 600 makes one or more decisions based on a set of criteria for determining how the synthetic depth-of-field effect is to be applied and applies the synthetic depth-of-field effect (e.g., without detecting an input for applying the synthetic depth-of-field effect). When the synthetic depth-of-field effect is applied in response to user input, the computer system 600 detects the input and applies the synthetic depth-of-field effect based on the type of the detected input.

[0188] As shown in FIG. 6A, the computer system 600 displays a live preview 630 that includes John 632 and Jane 634. As shown by the live preview 630, John 632 is located closer to one or more rear cameras of the computer system 600 than Jane 634. The live preview 630 of FIG. 6A is displayed without the synthetic depth-of-field effect being applied. However, it should be understood that the live preview 630 of FIG. 6A is displayed by the natural depth-of-field effect.

[0189] As used herein, the natural depth of field is different from the synthetic depth of field effect. The natural depth of field effect is generated based on the distance between the subject (e.g., person, animal, object) in the scene and one or more cameras, and the size and focal length of the aperture of one or more cameras that capture the scene. Therefore, the natural depth of field effect is directly limited by the physical specifications (e.g., focal length, aperture size) of one or more cameras used to capture the scene. However, the synthetic depth of field effect is a computer-generated depth of field effect (e.g., via software) and is not strictly limited by the physical specifications of one or more cameras and / or the distance between the subject in the scene and one or more cameras.

[0190] Therefore, applying the synthetic depth of field effect can have distinct advantages over applying only the natural depth of field effect to the medium. For example, applying the synthetic depth of field effect can be more advantageous than applying only the natural depth of field effect because the synthetic depth of field effect can be applied and adjusted in more ways (e.g., in real time) during the capture of the medium (whereas adjusting the natural depth of field effect is limited by the physical specifications of one or more cameras). Further, the synthetic depth of field effect provides an advantage in that the hardware of the computer system 600 (e.g., one or more cameras) does not need to be switched in order to apply a particular depth of field effect (e.g., and / or to replace a depth of field effect having one type of tracking with a depth of field effect having another type of tracking during a portion of a video). In some embodiments, the type of tracking with respect to the depth of field effect includes emphasizing a particular subject with respect to one or more other subjects in the medium (e.g., during the duration of the medium, only during a portion of the duration of the medium), emphasizing a subject at a particular location in the medium with respect to other subjects in the medium, and the like.

[0191] As shown in FIGS. 6A - 6BJ, the synthetic depth of field effect of the scenes (e.g., 630, 640, and / or 660) being displayed by computer system 600 is indicated by shading (e.g., white, gray, black). The portions of the scene indicated by darker shading have a greater amount of synthetic blur (e.g., synthetic depth of field effect) than the portions of the scene having lighter shading. It should be understood that the shading shown in FIGS. 6A - 6BJ does not represent an exact / accurate representation of the synthetic depth of field effect applied to the scenes shown in these figures. However, the shading shown in FIGS. 6A - 6BJ is provided to illustrate how the synthetic depth of field effect is applied and / or modified to the subjects within the scene automatically and / or in response to user input. As shown in FIG. 6A, the live preview 630 is unshaded (e.g., white), indicating that the live preview 630 has only the blur caused by the natural depth of field effect. In FIG. 6A, computer system 600 detects a right - swipe input 650a1 on the live preview 630 and / or a tap input 620e on the cinematic video mode control 650a2.

[0192] In FIG. 6B, in response to detecting a rightward swipe input 650a1 and / or a tap input 650a2, computer system 600 moves camera mode control 620 to the right so that cinematic video mode control 620e is displayed at the center of the camera user interface. In FIG. 6B, computer system 600 displays cinematic video mode control 620e as being selected (e.g., in bold) and stops displaying photo mode control 620a as being selected. Further, in response to detecting a rightward swipe input 650a, computer system 600 transitions from a state configured to operate in photo camera mode to cinematic video camera mode. In some embodiments, while cinematic video mode control 620e is displayed as being selected, computer system 600 detects a left swipe input, and in response to detecting the left swipe input (e.g., in a direction opposite to rightward swipe input 650a1), computer system 600 moves camera mode control to the left so that photo mode control 620c is displayed as being selected.

[0193] While the computer system 600 is operating in the cinematic video camera mode, the computer system 600 applies a synthetic depth of field effect. In some embodiments, one camera mode employs a synthetic depth of field effect (e.g., the cinematic video camera mode), while other camera modes do not employ a synthetic depth of field effect (e.g., the photo mode, the portrait mode, the video mode). In some embodiments, the synthetic depth of field can be manually enabled or disabled for any given camera mode. In FIG. 6B, the applied synthetic depth of field effect emphasizes John 632 relative to Jane 634 (e.g., makes John more prominent than Jane due to less blur), which can be seen via the live preview 630 showing that John 632 and the area around John 632 are brighter and shaded than Jane 634 and the area around Jane 634. In particular, John 632 is not shaded in the live preview 630, indicating that John 632 is being displayed with only the natural blur (if any) generated by the natural depth of field effect of one or more cameras of the computer system 600. Further, John 632 not being shaded in the live preview 630 indicates that the synthetic depth of field effect is not applying synthetic blur to John 632. On the other hand, since the computer system 600 is applying synthetic blur to Jane 634 via the synthetic depth of field effect applied in FIG. 6B, Jane 634 is displayed shaded (e.g., darker than John 632). In some embodiments, the natural blur is less visually prominent (or has less blur) than a portion of the blur displayed when the synthetic depth of field effect is applied.

[0194] As shown in FIG. 6B, the computer system 600 displays a primary subject indicator 672a around the head of John 632 and a secondary subject indicator 674b around the head of Jane 634. The primary subject indicator 672a is displayed around the head of John 632 because John 632 is emphasized via the applied synthetic depth-of-field effect. Since Jane 634 is not emphasized via the applied synthetic depth-of-field effect, the secondary subject indicator 674b is displayed around the head of Jane 634. Thus, in FIG. 6B, the computer system 600 displays different indicators to distinguish the subject emphasized by the synthetic depth-of-field effect from the subjects not emphasized by the synthetic depth-of-field effect. In some embodiments, since the computer system 600 has sufficient visual content to track and / or focus on (and / or apply the synthetic depth-of-field effect to emphasize) Jane 632, the secondary subject indicator 674b is displayed around the head of Jane 634. In some embodiments, if the computer system 600 does not have sufficient visual content to track and / or focus on Jane 632, the secondary subject indicator is not displayed around the head of Jane 634 (and / or the secondary subject indicator corresponding to Jane 634 is not displayed).

[0195] As shown in FIG. 6B, different portions of the scene shown in the live preview 630 have different levels of blur applied. For example, the trees and grass in the live preview 630 of FIG. 6B are shown less detailed than the trees and grass in the live preview 630 of FIG. 6A, indicating that the background, foreground, and / or different portions of the scene are also blurred (e.g., not just the subjects within the scene). Further, the portions of the background of the scene within the live preview 630 are shown more blurred (e.g., with darker shadows) than the subjects (e.g., John 632 and Jane 634) within the live preview 630 after the synthetic depth-of-field effect is applied.

[0196] In addition to applying the compositing write boundary depth effect, in response to detecting a rightward swipe input 650a1 and / or a tap input 650a2, the computer system 600 enlarges the live preview 630 such that the live preview 630 in FIG. 6B occupies more of the computer system 600 than the live preview 630 in FIG. 6A. In response to detecting the rightward swipe input 650a1 and / or the tap input 650a2, the computer system 600 continues to display the flash indicator 602a and stops displaying the setting mode indicator 602b and the animation image indicator 602c in the indicator region 602 of FIG. 6A. As shown in FIG. 6B, the computer system 600 displays an elapsed time indicator 602d at the position where the mode setting indicator 602b was previously displayed in FIG. 6A. Further, the computer system 600 displays a depth indicator 602e instead of the animation image indicator 602c. In some embodiments, in response to receiving an input directed to the depth indicator 602e, the computer system 600 displays controls for adjusting the blur effect applied to the captured media (e.g., as described below with respect to FIGS. 6AD-6AH). In some embodiments, the computer system 600 updates the live preview 630 when the controls for adjusting the blur effect (e.g., using one or more techniques as described below in connection with FIGS. 6AD-6AF) are changed.

[0197] As shown in FIG. 6B, in response to detecting a rightward swipe input 650a1 and / or a tap input 650a2, the computer system 600 also stops displaying the 0.5x zoom control 622a and the 2x zoom control 622c and maintains the display of the 1x zoom control 622b. In some embodiments, the computer system 600 continues to display the 1x zoom control 622b because a determination is made that the composite depth of field effect is applied only when the computer system 600 is displaying a particular zoom level (e.g., 1x) and / or a range of zoom levels (e.g., 0.8x zoom to 1.7x zoom). In some embodiments, the computer system 600 continues to display the 1x zoom control 622b because a set of cameras (e.g., a wide-angle camera (e.g., having an f / 1.6 aperture (e.g., and / or an f / 1.4 to f / 8.0 aperture) and a 60° to 120° field of view)) is used to capture cinematic video media at the 1x zoom level (and / or a range of zoom values including the 1x zoom level). In some embodiments, the computer system 600 stops displaying the zoom control 622a and the 2x zoom control 622c because a particular set of cameras (e.g., an ultra-wide-angle camera (e.g., having an f / 2.4 aperture (e.g., and / or an f / 1.4 to f / 8.0 aperture) and a field of view greater than 120°), a telephoto camera (e.g., having an f / 2.0 aperture (e.g., and / or an f / 1.4 to f / 8.0 aperture) and a 30° to 60° field of view and / or a field of view less than 60°)) does not capture cinematic media at the 0.5x and / or 2x zoom levels. In some embodiments, the computer system 600 determines that using a particular camera set when applying the composite depth of field effect is undesirable and / or not optimal (e.g., due to the physical specifications of the particular camera set). In FIG. 6B, the computer system 600 detects a rotation 650b1 and a tap input 650b2 directed at the shutter control 610.

[0198] As shown in FIG. 6C, in response to detecting rotation 650b1, computer system 600 shifts the camera user interface from a portrait orientation to a landscape orientation. In particular, FIG. 6C shows two computer systems. Computer system 600 is disposed on the right side of FIG. 6C, and computer system 690 is disposed on the left side of FIG. 6C. Both computer system 600 and computer system 690 are shown with their respective user interfaces in landscape orientation. Computer system 600 of FIG. 6C captures video and displays stop control 616 in response to tap input 650b2. In particular, computer system 600 of FIG. 6C is shown such that the frame of the video being captured (e.g., live preview 630) is at a capture duration of 1 second (as indicated, for example, by elapsed time indicator 602d) and / or 1 second has elapsed since tap input 650b2 was received. Computer system 690 is provided to show how the computer system displays the frame of the video being captured by computer system 600 in FIG. 6C during video playback (e.g., after the complete video has been captured by computer system 600). One reason for providing computer system 690 is to show the differences and / or similarities between how the frame of the video is shown while the video is being captured and how the frame of the video is shown after the video has been captured and played back. In some embodiments, computer system 600 and computer system 690 are the same system (e.g., at different times). In some embodiments, computer system 600 and computer system 690 are different systems (e.g., a file representing the video captured by computer system 600 is transferred to computer system 690 after the video has been captured).

[0199] As shown in FIG. 6C, computer system 690 shows a media playback user interface that includes a previously captured media representation 640 and an elapsed time indicator 646. As suggested above, the previously captured media representation 640 is a frame that is displayed during playback of video captured by computer system 600 (e.g., a frame captured and shown via live preview 630). Thus, as shown in FIG. 6C, live preview 630 and the previously captured media representation 640 represent the same frame of video captured by computer system 600, but are shown at different temporal instances (e.g., during video capture vs. during video playback). Thus, the previously captured media representation 640 is shown during a 1-second capture duration (and / or 1-second mark) of the video (as indicated, for example, by elapsed time indicator 646). Thus, elapsed time indicator 602d and elapsed time indicator 646 are shown for the same elapsed time (e.g., 1 second) of the video.

[0200] FIG. 6C also includes a graph 680 that includes activity trackers 680a, 680b, and 680c. Displayed within activity tracker 680a is John's activity level 680a1 (e.g., the activity level of John 632). Displayed within activity tracker 680b is Jane's activity level 680b1 (e.g., the activity level of Jane 634). John's activity level 680a1 and Jane's activity level 680b2 are activity levels that the computer system 600 detected and registered to correspond to the activity levels of John 632 and Jane 634 in real time. Further, John's activity level 680a1 does not represent John 632's absolute activity level, and Jane's activity level 680b2 does not represent Jane 634's absolute activity level. Rather, John's activity level 680a1 represents John 632's relative activity compared to Jane 634's activity level, and Jane's activity level 680a1 represents Jane 634's relative activity compared to John 632's activity level. Further, the activity levels shown in FIG. 6E represent activity levels detected / processed by the computer system 600 in real time, which can lag actual characteristics (e.g., the physical / visual characteristics of the subject to determine whether the subject is speaking, moving, gazing in a particular direction, hidden by one or more other objects in the scene, etc.) used to determine the activity level of the subject within the scene. As shown in FIG. 6C, activity tracker 680c does not include an activity level because the dog 638 was not captured by the computer system 600 (e.g., not displayed in the live preview 630) prior to the one second elapsed time indicated by the elapsed time indicator 602d.Referring to FIG. 6W, when a dog 638 (e.g., the dog 638 shown in the live preview 630 of FIG. 6W) is captured by the computer system 600, the activity tracker 680c (e.g., FIG. 6C) includes the activity level 680c1 of the dog (e.g., the activity level of the dog 638). The activity level displayed on the graph 680 represents the activity level of the subject at a specific time (e.g., 0:00 to 0:45) within the video being captured by the computer system 600. As shown in FIG. 6C, John's activity level 680a1 is higher than Jane's activity level 680b1 (as indicated, for example, by John's activity level 680a1 occupying more area than Jane's activity level 680b1). In FIG. 6C, John 632 is closer to one or more cameras of the computer system 600 (e.g., capturing the scene shown in the live preview 630), and John 632 is currently speaking (as indicated, for example, by John 632's mouth being higher), so John's activity level 680a1 is higher. Further, Jane 634 is farther away from one or more cameras of the computer system 600, and Jane 634 is not speaking (as indicated, for example, by Jane 634's mouth being closed), so Jane's activity level 680b1 is lower.

[0201] In FIG. 6C, in response to detecting the tap input 650b2, the computer system 600 starts capturing video and makes a determination that John 632 meets a set of automatic selection criteria (e.g., based on John's activity level). Specifically, since John 632 had a higher activity level than Jane 634 during the duration of the video being captured (as indicated by John's activity level 680a1 being higher than Jane's activity level 680b1, for example, between 0 seconds and 1 second), John 632 meets the set of automatic selection criteria. As shown in FIG. 6C, since a determination is made that John 632 meets the set of automatic selection criteria, the computer system 600 applies a synthetic depth of field effect to the frames of the video being captured with a 1 - second capture duration. As shown by the live preview 630 in FIG. 6C, the applied synthetic depth of field effect emphasizes John 632 relative to Jane 634 such that John 632 is displayed less blurred than Jane 634 (as indicated, for example, by John 632 having a brighter shadow than Jane 634). Further, since John 632 is emphasized by the synthetic depth of field effect, the computer system 600 displays a primary subject indicator 672a around John 632's head, and since Jane 634 is not emphasized by the synthetic depth of field effect, the computer system 600 displays a secondary subject indicator 674b around Jane 634's head.

[0202] As shown in FIG. 6C, graph 680 is provided to show which subject is being emphasized by the synthetic depth of field effect at a particular point in time. As shown in FIG. 6C, graph 680 includes a media capture line 680d1 and a media playback line 680d2. Media capture line 680d1 indicates the subject for which the synthetic depth of field effect is being emphasized at a particular time during video capture (e.g., by computer system 600). Further, media playback line 680d2 indicates the subject for which the synthetic depth of field effect is being emphasized at a particular time during video playback (e.g., by computer system 690). When media capture line 680d1 is at (or near) the centerline of an individual activity tracker (e.g., media capture line 680d1 is on the centerline of John's activity tracker 680a in FIG. 6C), computer system 600 applies the synthetic depth of field effect to emphasize the individual subject relative to other subjects within the FOV at a particular time. Similarly, when media playback line 680d2 is at (or near) the centerline of an individual activity tracker (e.g., media playback line 680d2 is on the centerline of John's activity tracker 680a in FIG. 6C), computer system 600 applies the synthetic depth of field effect to emphasize the individual subject over other subjects within the FOV at a particular time. Thus, computer system 600, which displays live preview 630 having a synthetic depth of field effect that emphasizes John 632 relative to Jane 634, is indicated by media capture line 680d1 that is at the center of John's activity tracker 680a. And computer system 690, which displays previously captured media representation 640 having a synthetic depth of field effect that emphasizes John 632 relative to Jane 634, is indicated by media playback line 680d2 that is at the center of John's activity tracker 680a.At a particular time on graph 680 where media capture line 680d1 or media playback line 680d2 is not at the center of an individual media tracker (e.g., graph 680 of FIG. 6F from 2 seconds to 3 seconds within the media), the computer system is transitioning the synthetic depth of field effect so that a new subject is emphasized more than an individual subject within the media.

[0203] Figures 6D - 6G illustrate an exemplary embodiment in which computer system 600 automatically changes the synthetic depth of field effect to emphasize Jane 634 over John 632. As shown in Figure 6D, computer system 600 displays the scene shown in live preview 630 (e.g., representing a video frame) at 2 seconds during the capture of the video (as indicated by, e.g., elapsed time indicator 602d). Live preview 630 shows John 632's eyes as viewed from one or more of the cameras in Figure 6D, which is a change from John 632's eyes in live preview 630 of Figure 6C. Thus, John 632's line of sight has changed from being directed towards one or more of the cameras of computer system 600 in Figure 6C to being directed away from one or more of the cameras of computer system 600 in Figure 6D. A subject's line of sight directed towards one or more of the cameras of computer system 600 can increase the subject's activity level, which increases the probability that the subject meets the automatic selection criteria. However, a subject's line of sight being directed away from one or more of the cameras of computer system 600 can potentially decrease the subject's activity, which decreases the likelihood that the subject meets the automatic selection criteria. Thus, in Figure 6D, John 632's activity level has begun to decrease along with the probability that John 632 continues to meet the set of automatic selection criteria. In addition to the change in line of sight, John 632 has stopped talking in Figure 6D and Jane 634 has started talking in Figure 6D. However, since computer system 600 detects the subject's activity level in real - time (e.g., while the video is being captured) and more information (e.g., data, visual content) is required to make this determination, computer system 600 has not made a determination that Jane 634 meets the set of automatic selection criteria.As shown in FIG. 6D, since the computer system 600 has not made a determination that Jane 634 meets the set of automatic selection criteria within the video time frame (e.g., the computer system 600 still depends on the determination made regarding John who meets the set of automatic selection criteria described above in FIG. 6C), the computer system 600 continues to apply the synthetic depth-of-field effect to emphasize John 632 over Jane 634. In particular, to show that the computer system 600 did not detect a relative change in the activity levels of John 632 and Jane 634, the activity level 680a1 of John continues to be greater than the activity level 680b2 of Jane in the graph 680 of FIG. 6D.

[0204] In contrast to computer system 600 of FIG. 6D, computer system 690 of FIG. 6D is playing back video previously captured by computer system 600. Thus, computer system 690 has sufficient information to make a determination that Jane 634 meets a set of automatic selection criteria. This is because, at least, computer system 690 has more (or all) of the information corresponding to the captured video. Thus, since computer system 690 can access the information within the previously captured video, computer system 690 can make a determination as to whether a subject meets a set of automatic criteria within a particular time frame of the video. In FIG. 6D, computer system 690 makes a determination that Jane 634 meets the automatic selection criteria within the time frame of the video, and based on this determination, automatically applies a synthetic depth-of-field effect to emphasize Jane 634 relative to John 632. However, as shown in FIGS. 6D-6G, computer system 690 displays the animation of the previously captured media representation 640 to smoothly transition to emphasizing Jane 634 relative to John 632 by emphasizing John 632 relative to Jane 634 (e.g., instead of a more abrupt transition). As part of the animation, computer system 690 gradually displays a more blurred John 632 and a more blurred Jane 634 such that Jane 634 is emphasized relative to John 632 in FIG. 6G (e.g., the difference in blur is approximately the same as when John 632 was emphasized relative to Jane 634 in FIG. 6B).

[0205] As shown in FIG. 6E, computer system 600 displays the scene shown in live preview 630 at 3 seconds during the capture of the video (as indicated, for example, by elapsed time indicator 602d). Live preview 630 continues to show John 632's eyes as seen from one or more of the cameras in FIG. 6E (which has not been changed, for example, from the live preview 630 in FIG. 6D). In FIG. 6E, computer system 600 has not made a determination that Jane 634 meets a set of automatic selection criteria because computer system 600 requires more information (e.g., data, content) to make this determination. As shown in FIG. 6D, since no determination has been made that Jane 634 meets a set of automatic selection criteria, computer system 600 continues to apply a synthetic depth-of-field effect to emphasize John 632 over Jane 634.

[0206] As shown in FIG. 6F, computer system 600 displays the scene shown in live preview 630 during the capture video. In FIG. 6F, elapsed time indicator 602d indicates 3 seconds, but the live preview 630 in FIG. 6F is displayed after the live preview 630 in FIG. 6E was displayed. In FIG. 6F, computer system 600 makes a determination that Jane 634 meets a set of automatic selection criteria (e.g., because computer system 600 has sufficient information in FIG. 6F). Based on this determination, computer system 600 automatically changes the synthetic depth-of-field effect to emphasize Jane 634 over John 632 and, in FIGS. 6F-6G, displays an animation of John 632 with more blur and Jane 634 with less blur.

[0207] In particular, the animations displayed by computer system 600 in FIGS. 6F - 6G include transitions that are more abrupt and less smooth compared to the transitions including the animations by computer system 690 in FIGS. 6E - 6G. This is because, at least, the computer system 690 determines that a set of automatic selection criteria is met and that a change in the synthetic depth - of - field effect to emphasize Jane 634 over John 632 needs to occur by 4 seconds of video playback / capture (for example, because the live preview 630 of computer system 600 is updated to show the full change in the synthetic depth - of - field effect in FIG. 6G) before the computer system 600 can make this decision. In FIG. 6G, the media capture line 680d1 and the media playback line 680d2 of graph 680 provide context for the comparison of the animations displayed by computer systems 600 and 690. The media capture line 680d1 moves from John's activity tracker 680a to the female activity tracker 680b at a later time than the media playback line 680d2. Further, the media capture line 680d1 ramps down faster (for example, the shorter and more abrupt animation in FIGS. 6F - 6G displayed by computer system 600) than the media playback line 680d2 (for example, the longer and smoother animation in FIGS. 6E - 6G displayed by computer system 600).

[0208] As shown in FIG. 6G, computer system 600 and computer system 690 apply a synthetic depth-of-field effect (e.g., when the shadow of live preview 630 coincides with the shadow of previously captured media representation 640) to emphasize Jane 634 over John 632. As shown in FIG. 6G, while applying the synthetic depth-of-field effect to emphasize Jane 634 over John 632, computer system 600 stops displaying primary subject indicator 672a around John 632's head and secondary subject indicator 674b around Jane 632's head, and displays primary subject indicator 672b around Jane 634's head and secondary subject indicator 674a around John 634's head. Primary subject indicator 672b indicates that Jane 634 is currently being emphasized by the synthetic depth-of-field effect, and secondary subject indicator 674b indicates that John 632 is not being emphasized by the synthetic depth-of-field effect. As shown in FIGS. 6F-6G, primary subject indicator 672a in FIG. 6F and primary subject indicator 672b in FIG. 6G have the same visual appearance (e.g., focus brackets, same shape, and / or same object). Similarly, secondary subject indicator 674a in FIG. 6G and secondary subject indicator 674b in FIG. 6F have the same visual appearance (e.g., rectangle, same shape, and / or same object). However, the primary subject indicator and the secondary subject indicator do not have the same visual appearance (e.g., 672a-672b compared to 674a-674b in FIGS. 6F-6G). In some embodiments, computer system 600 stops displaying primary subject indicator 672a around John 632's head and secondary subject indicator 674b around Jane 634's head during the animation of the transition of the change in the application of the synthetic depth-of-field effect, and / or displays primary subject indicator 672b around Jane 634's head and secondary subject indicator 674a around John 632's head.

[0209] In some embodiments, computer system 600 and computer system 690 are illustrated in connection with FIGS. 6AD - 6AG and display respective animations that are different from the animations discussed above. In some embodiments, computer system 600 determines that an automatic change to the synthetic depth - of - field effect should be made (e.g., computer system 600 makes this determination at 4 seconds during video capture). In some embodiments, when it is determined that an automatic change to the synthetic depth - of - field effect should be made, computer system 600 automatically displays an animation of the change to the synthetic depth - of - field effect (e.g., an animation that is played between 4 and 5 times during video capture). In some embodiments, the live preview 630 is updated to show completion of the change to the synthetic depth - of - field effect at some time after the determination is made (e.g., at 5 seconds during video capture) such that the displayed animation is fully complete. In some embodiments, computer system 690 determines that an automatic change to the synthetic depth - of - field effect should be made at the time (e.g., 4 seconds) that computer system 600 made this determination while capturing a live video (e.g., computer system 690 makes this determination at 3 seconds during video playback). In some embodiments, when computer system 690 determines that an automatic change to the synthetic depth - of - field effect should be made, computer system 690 displays an animation of the change to the synthetic depth - of - field effect (e.g., an animation that is displayed between 3 and 4 seconds during video playback). In some embodiments, the animation of the change to the synthetic depth - of - field effect displayed by computer system 690 is fully complete, and as a result, the previously captured media representation 640 is updated to show completion of the change to the synthetic depth - of - field effect at the time (e.g., 4 seconds) that computer system 600 made its determination while capturing a live video.In some embodiments, the animation displayed by computer system 690 is the same length as the animation displayed by computer system 600 (e.g., both animations are 1 to 5 seconds). In some embodiments, the animation displayed by computer system 690 completes at a time corresponding to a time in the video earlier than the time at which the animation displayed by computer system 600 completely finishes.

[0210] Figures 6H - 6K illustrate exemplary embodiments in which computer system 600 automatically changes the synthetic depth - of - field effect to emphasize John 632 relative to Jane 634. As shown in Figure 6H, at 6 seconds during the capture of the video (as indicated, for example, by elapsed time indicator 602d), computer system 600 displays the scene shown in live preview 630 (e.g., representing a frame of the video). The live preview 630 in Figure 6H shows that John 632's head has moved (e.g., horizontally), indicating that John 632 is moving within the field of view of one or more cameras. An increase in the movement of a subject within the field of view of one or more cameras can increase the activity level of the subject, which can increase the probability that the subject meets the automatic selection criteria. Conversely, a decrease in the movement of a subject within the field of view of one or more cameras can decrease the activity level of the subject, which can decrease the probability that the subject meets the automatic selection criteria. Further, Jane 634 is not talking (as indicated, for example, by Jane 634's mouth being closed in Figure 6H). As shown in Figure 6H, since computer system 600 does not make a determination that Jane 634 meets a set of automatic selection criteria because it does not have sufficient information (for the same reasons described above in relation to Figure 6D), computer system 600 continues to apply the synthetic depth - of - field effect to emphasize Jane 634 relative to John 632.

[0211] In contrast to the computer system 600 of FIG. 6H, the computer system 690 makes a determination that Jane 634 meets a set of automatic selection criteria within a particular time frame of the video (e.g., for the same reasons described above in connection with FIGS. 6D - 6G), and based on this determination, automatically changes the synthetic depth - of - field effect to emphasize John 632 with respect to Jane 634. As shown in FIGS. 6H - 6K, the computer system 690 displays an animation of a previously captured media representation 640 that smoothly transitions from emphasizing Jane 634 with respect to John 632 to emphasizing John 632 with respect to Jane 634. As part of the animation, the computer system 690 gradually displays Jane 634 more blurred and John 632 less blurred such that (e.g., using one or more of the same techniques described above in connection with FIGS. 6D - 6G) John 632 is emphasized with respect to Jane 634 in FIG. 6K.

[0212] As shown in FIG. 6I, the computer system 600 displays the scene shown in the live preview 630 at 7 seconds during the capture of the video (e.g., as indicated by the elapsed - time indicator 602d). The live preview 630 continues to show that John 632 is moving within the FOV (e.g., the head of John 632 is in a different position in FIG. 6I than in FIG. 6H). In FIG. 6I, the computer system 600 has not made a determination that John 632 meets a set of automatic selection criteria because more information is required to make this determination. As shown in FIG. 6I, since no determination has been made that John 632 meets a set of automatic selection criteria (e.g., depending on the determination made in FIG. 6F), the computer system 600 continues to apply the synthetic depth - of - field effect to emphasize Jane 634 with respect to John 632.

[0213] As shown in FIG. 6J, computer system 600 displays the scene shown in live preview 630 during the capture video, and computer system 600 continues to indicate that John 632 is moving within the FOV. After the live preview 630 of FIG. 6I is displayed while the elapsed time indicator 602d indicates 7 seconds, the live preview 630 of FIG. 6J is displayed. In FIG. 6J, computer system 600 makes a determination that John 632 meets a set of automatic selection criteria (for the same reasons as described above in connection with FIGS. 6F-6G, for example). Based on this determination, computer system 600 automatically changes the synthetic depth-of-field effect to emphasize John 632 with respect to Jane 634, and (using one or more techniques, for the same reasons as described above in connection with FIGS. 6F-6G, for example) displays an animation of blur in which John 632 is displayed decreasing and blur in which John 632 is displayed increasing. As shown in FIG. 6G, while applying the synthetic depth-of-field effect to emphasize John 632 with respect to Jane 634, computer system 600 also displays a primary subject indicator 672a around the head of John 632 and a secondary subject indicator 674b around the head of Jane 634 (using one or more techniques, for the same reasons as described above in connection with FIGS. 6F-6G, for example). The media capture line 680d1 and the media playback line 680d2 of the graph 680 in FIGS. 6G-6J are also updated and displayed for the same reasons as described above in connection with FIGS. 6F-6G.

[0214] Figures 6L - 6M illustrate an exemplary embodiment in which the computer system 600 does not change the synthetic depth of field effect previously applied. As shown in Figure 6L, the computer system 600 displays the scene shown in the live preview 630 at 10 seconds during the capture of the video (as indicated, for example, by the elapsed time indicator 602d), where John 632 is wiping his face with a towel 642. As shown in Figure 6L, the towel 642 covers (and / or obscures) John 632's face. In some embodiments, the towel 642 covers John 632's face such that the computer system 600 cannot detect John 632's face within the field of view of one or more cameras (e.g., using one or more face detection techniques). As shown in Figure 6M, the computer system 600 displays the scene shown in the live preview 630 at 11 seconds, and the live preview 630 indicates that John 632 has removed the towel 642 from his face in Figure 6L. Thus, in Figure 6M, John 632's face is no longer covered.

[0215] In FIGS. 6L - 6M, computer system 600 and computer system 690 make individual determinations that John 632's face is covered and / or obscured (e.g., and / or the individual computer system is unable to detect John 632's face) for a period less than a predetermined period (e.g., 2 - 60 seconds). In FIGS. 6L - 6M, for these individual determinations, computer system 600 and computer system 690 continue to individually apply the previously applied synthetic depth - of - field effect (e.g., to emphasize John 632 relative to Jane 634 in FIGS. 6H - 6K) regardless of whether the towel 642 obscures John 632's face. As shown in FIG. 6L, even when the towel 642 is hiding John 632's face, John 632 is emphasized relative to Jane 634 in both the live preview 630 and the previously captured media representation 640. As shown in FIGS. 6L - 6M, since computer system 600 continues to apply the synthetic depth - of - field effect that was previously applied before John 632 covered his face with the towel 642 in FIG. 6L, computer system 600 continues to display the primary subject indicator 672a and the secondary subject indicator 674a. In some embodiments, the determinations made by computer system 690 in FIGS. 6L - 6M are made earlier (e.g., for the same reasons as described above in connection with FIGS. 6D - 6G) with respect to the elapsed time of the video than the determinations made by computer system 600.

[0216] Figures 6N through 6T illustrate an exemplary embodiment in which computer system 600 modifies a synthetic depth of field effect in response to a first type of user input (e.g., a user-specified change). As shown in FIG. 6O, at 12 seconds during the capture of a video (as indicated, for example, by elapsed time indicator 602d), computer system 600 displays the scene shown in live preview 630 (e.g., representing a video frame). In FIG. 6N, computer system 600 continues to apply a synthetic depth of field effect that emphasizes John 632 over Jane 634 (as indicated, for example, by the shading of live preview 630 in FIG. 6N) to the content being captured by one or more cameras of computer system 600. In FIG. 6O, computer system 600 detects a single tap input 650o on Jane 634.

[0217] In FIG. 6P, in response to detecting a single tap input 650o, computer system 600 changes the synthetic depth-of-field effect to emphasize Jane 634 over John 632 (as shown, for example, by the shading of the live preview 630 in FIG. 6P). In response to detecting the single tap input 650o, computer system 600 makes an immediate change to the synthetic depth-of-field effect and does not display a transition animation indicating that the synthetic depth-of-field effect (as shown, for example, by the live preview 630 in FIG. 6P being displayed at 12 seconds during video capture) is being changed. Thus, the live preview 630 is updated to reflect a user-specified change to the synthetic depth-of-field effect (e.g., a change that occurs in response to detecting an input), as opposed to being updated to reflect an automatic change to the synthetic depth-of-field effect. When a user-specified change to the synthetic depth-of-field effect occurs, the live preview 630 is immediately updated (e.g., and / or a change to the application of the synthetic depth-of-field immediately occurs). However, when an automatic change to the synthetic depth-of-field effect occurs, the live preview 630 is updated more gradually (e.g., an animation of the transition between the current synthetic depth-of-field effect and the new synthetic depth-of-field effect is displayed, as described in connection with FIGS. 6D-6K). Additionally, graph 680 shows this as well. In graph 680, media capture line 680d1 is drawn at a right angle at 12 seconds to reflect how an immediate change in a user-specified change to the synthetic depth-of-field effect occurred (e.g., in response to the single tap input 650o), and media capture line 680d1 is drawn as a curve between 3 seconds and 10 seconds and at 12 seconds to reflect how a smoother automatic change to the synthetic depth-of-field effect occurred.

[0218] Returning to FIGS. 6N - 6P, computer system 690 displays previously captured media representation 640 with an animation of a user - specified change (e.g., occurring in response to detecting single - tap input 650o) in the synthetic depth - of - field effect (e.g., during playback of the captured video). As shown in FIGS. 6N - 6P, since computer system 690 has information indicating that a user - specified change has occurred (for the same reasons as described above in relation to FIGS. 6D - 6K, for example), when displaying the previously captured media representation 640 with the user - specified change in the synthetic depth - of - field effect, a smoother transition is provided. Thus, in FIG. 6N, unlike live preview 630, previously captured media representation 640 has begun to show the change in the synthetic depth - of - field effect, while live preview 630 has not. In particular, in FIG. 6O, computer system 690 of the previously captured media representation 640 represents the change in the synthetic depth - of - field effect in its final state. In FIG. 6O, computer system 690 completes the change in the synthetic depth - of - field effect (e.g., the blur of the previously captured media representation 640 in FIG. 6O looks the same as live preview 630 in FIG. 6P) to emphasize Jane 634 over John 632 in the frame in which single - tap input 650o was received. Thus, computer system 690 can display the user - specified change in the frame corresponding to when the input that caused the user - specified change was received. Further, the comparison between media capture line 680d1 and media playback line 680d2 shows how a user - specified change affects visual content during playback of a video (e.g., via live preview 630 and previously captured media representation 640), as opposed to during media capture.As shown by graph 680, in response to detecting single tap input 650o, media playback line 680d2 exhibits a smoother and / or longer transition (e.g., generating a right angle in 12 seconds) than media capture line 680d1 to change the synthetic depth of field effect.

[0219] Referring to FIG. 6Q, live preview 630 (and previously captured media representation 640) is displayed with a user-specified synthetic depth of field effect change initiated via single tap input 650o, even if in FIG. 6Q John's activity level 680a1 is greater than Jane's activity level 680b1. When a user-specified changed synthetic depth of field effect occurs, computer system 600 uses a modified set of auto-selection criteria. The modified set of auto-criteria is different from the set of criteria (e.g., those that occurred before a request for a user-specified to change the synthetic depth of field effect was received, before single tap input 650o was detected) used to perform the automatic change of the synthetic depth of field effect described above in FIGS. 6B - 6K. In some embodiments, the modified set of auto-selection criteria has a higher threshold for automatically changing the synthetic depth of field effect than the set of criteria used to perform the automatic change of the synthetic depth of field effect described above in FIGS. 6B - 6K. In some embodiments, John 632 has to speak louder, move more, approach the camera, look straight into the camera, etc. for a longer period of time so that computer system 600 automatically changes the synthetic depth of field effect to emphasize John 632 over Jane 634. In some embodiments, after changing the application of the synthetic depth of field effect in response to detecting single tap input 650o, computer system 600 does not change the application of the synthetic depth of field effect for a predetermined period regardless of the activity level of the subject (e.g., unless the face of the subject is not detected for a predetermined period).

[0220] As shown in FIG. 6Q, Jane 634 has started walking out of the field of view of one or more cameras (e.g., walking out of the scene as shown by live preview 630 in FIG. 6Q). Looking at FIGS. 6P - 6Q, Jane 634 is emphasized relative to John 632 in live preview 630 (and previously captured media representation 640), but Jane 634 is moving within the field of view of one or more cameras. This indicates that the synthetic depth of field effect applied to emphasize a subject relative to other subjects follows and / or tracks the emphasized subject. Further, subject indicators (such as those shown by primary subject indicator 672b in FIGS. 6P - 6Q) move with each of the respective subjects surrounded by the individual subject indicators. In some embodiments, in response to detecting an input at a location in live preview 630 that is not on a subject, the applied synthetic depth of field effect does not follow and / or track the subject.

[0221] In FIG. 6R, Jane 634 is outside the field of view of one or more cameras (e.g., has walked out of the scene). In FIG. 6R, a determination is made that John 632 meets a modified set of automatic selection criteria (e.g., as indicated by Jane's activity level 680b1, Jane 634 is outside the frame and / or the computer system 600 is not detecting activity from Jane 634). As shown in FIG. 6R, the computer system 600 automatically changes the synthetic depth of field effect to emphasize John 632 (e.g., John 632 is shown with only natural blur (e.g., no shadows), and other portions of live preview 630 include a certain amount of synthetic blur (e.g., shadows)). The computer system 600 automatically changes the synthetic depth of field effect to emphasize John 632 relative to other portions of live preview 630 because a determination is made that John 632 meets a modified set of automatic selection criteria and / or because Jane has not had an activity level for a predetermined period (e.g., 1 second).

[0222] Figure 6R1 shows an exemplary embodiment of the position of Jane 634 relative to John 632 within the FOV of computer system 600. In Figure 6R1, a live preview 630 is shown at the 17 - second mark using one or more of the similar techniques described above in relation to Figure 6R. In Figure 6R1, a boundary 601 indicates the size of the FOV, and one or more cameras of computer system 600 can capture visual content inside the boundary 601 (e.g., within region 603 that includes live preview 630). As shown in Figure 6R1, Jane 634 is within region 603. Thus, Jane 634 is being captured by one or more cameras, but Jane 634 is not positioned within region 603 enough to be captured by one or more cameras such that Jane 634 is displayed in live preview 630. As shown in Figure 6R1, when Jane 634 is within region 603 but located outside the content within the FOV used to display live preview 630, computer system 600 continues to track Jane 634 for a predetermined period (e.g., 0.1 - 5 seconds). In some embodiments, while Jane 634 is within region 603 but located outside the content within the FOV used to display live preview 630 (as shown in Figure 6R1), if computer system 600 determines that it cannot capture Jane 634 within the visual content corresponding to live preview 630, it stops tracking Jane 634 after a predetermined period. In some embodiments, a neural network (e.g., discussed in Figure 12) continues to track Jane after a certain period, and computer system 600 can provide one or more representations of Jane 634 (e.g., old representations and / or previously captured representations of Jane 634) for a second predetermined period. In some embodiments, after the second predetermined period, computer system 600 automatically switches to highlighting and / or tracking another subject and / or focal plane within the visual content captured in the FOV corresponding to live preview 630.In some embodiments, when Jane 634 is located outside region 603 (e.g., outside boundary 601), computer system 600 does not track Jane 634 (e.g., and / or does not store the corresponding identifier). In some embodiments, when Jane 634 is located within region 603 and inside the content within the FOV used to display the live preview, computer system 600 tracks Jane 634 regardless of the predetermined period. In some embodiments, computer system 600 has information regarding a user who is within region 603 but outside the content within the FOV used to display the live preview (e.g., the time period during which Jane 632 was inside and / or outside the FOV for the content used to display live preview 630, and / or whether Jane 634 was moving towards and / or away from the content used to display live preview 630 while Jane 634 was within region 603), and automatically switches to emphasize and / or track another subject (e.g., "John" and / or the plane of focus) within the visual content captured within the FOV corresponding to live preview 630. This enables computer system 600 to more quickly switch the emphasis to a subject entering the portion of the FOV used to display the live preview. Because computer system 600 (and, optionally, the neural network that makes the automatic emphasis determination) has more time to track the subject and observe the behavior of the subject occurring outside the FOV used to display the live preview but within region 603, compared to a situation where computer system 600 does not have the opportunity to observe the behavior of the subject before the subject enters the portion of the FOV used to display the live preview, to determine the relative importance of the subject compared to other subjects that can be emphasized.

[0223] As shown in FIG. 6S, Jane 634 has come back into the field of view of one or more cameras (e.g., standing within the scene as shown by the live preview 630 in FIG. 6S). In FIG. 6S, the live preview 630 continues to be displayed with a synthetic depth-of-field effect that emphasizes John 632 relative to Jane 634, which is due to the single tap input 650o in FIG. 6O being of the first type. In particular, since the single tap input 650 in FIG. 6O is of the first type, the computer system 600 treats a change to the synthetic depth-of-field effect that emphasizes Jane 634 relative to John 632 as a temporary user-specified change to the application of the synthetic depth-of-field effect. When a temporary user-specified change to the application synthetic depth-of-field effect occurs, the computer system 600 does not automatically reapply the temporary change to the synthetic depth-of-field effect after an automatic change to the synthetic depth-of-field effect has occurred (e.g., regardless of how long Jane 634 has gone outside the visual content within the FOV corresponding to the live preview 630). Thus, since the single tap input 650o in FIG. 6O is of the first type and an automatic change to the synthetic depth-of-field effect (e.g., the change described in FIG. 6P) has occurred after the single tap input 650o was detected, the computer system 600 continues to apply the synthetic depth-of-field effect to emphasize John 632 relative to other portions of the live preview 630.

[0224] As shown in FIG. 6T, although 4 seconds have elapsed since the live preview 630 of FIG. 6S was displayed (e.g., as shown by 602d in FIGS. 6S - 6T), it continues to be displayed with a synthetic depth - of - field effect that emphasizes John 632 with respect to Jane 634. In FIG. 6T, the computer system 600 continues to apply a synthetic depth - of - field effect that emphasizes John 632 with respect to Jane 634 because the single - tap input 650o of FIG. 6O is a first - type input and an automatic change to the synthetic depth - of - field effect (e.g., the change described in FIG. 6P) occurred after the single - tap input 650o was detected.

[0225] FIGS. 6U - 6Y are exemplary embodiments in which the computer system 600 changes the synthetic depth - of - field effect in response to a second - type of user input (e.g., a user - specified change). As shown in FIG. 6U, although 10 seconds have elapsed since the live preview 630 of FIG. 6S was displayed (e.g., as shown by 602d in FIGS. 6S - 6T), the live preview 630 of the computer system 600 continues to be displayed with a synthetic depth - of - field effect that emphasizes John 632 with respect to Jane 634. In FIG. 6U, for the same reasons as described above in connection with FIGS. 6S - 6T, the live preview 630 is displayed with a synthetic depth - of - field effect that emphasizes John 632 with respect to Jane 634. In FIG. 6U, the computer system 600 detects a double - tap input 650u.

[0226] As shown in FIG. 6V, in response to detecting a double-tap input 650u, the computer system 600 immediately changes the synthetic depth-of-field effect to emphasize Jane 634 more than John 632 (as shown, for example, by the shading of the live preview 630 in FIG. 6V). In response to detecting the double-tap input 650u, the computer system 600 immediately makes a change to the synthetic depth-of-field effect and does not display a transition animation showing the change in the synthetic depth-of-field effect (for the same reasons as described above in connection with FIG. 6P, as shown by 680d1 of 30 seconds).

[0227] In FIG. 6V, computer system 600 displays primary subject indicator 678b around the head of Jane 634 and secondary subject indicator 674a around the head of John 632. In particular, since each individual indicator is displayed in response to detecting a different type of input, primary subject indicator 678b is different from primary subject indicator 672b, which was displayed in response to detecting single tap input 650o. In particular, since a determination has been made that a second type of input (e.g., double tap input 650u in FIG. 6U) has been detected, primary subject indicator 678b is displayed in FIG. 6V, and since a determination has been made that a first type of input (e.g., single tap input 650o in FIG. 6O) has been detected, primary subject indicator 672b is displayed in FIG. 6P. Further, computer system 600 displays different subject indicators since a different type of tracking is applied when a second type of input is received than when a first type of input is received. As described above in connection with FIGS. 6O-6P, computer system 600 makes a temporary change to the composite depth of field effect applied when a first type of input (e.g., single tap input 650o in FIG. 6O) is received. As described above in connection with FIGS. 6O-6P, computer system 600 does not automatically reapply the temporary change to the composite depth of field effect after the automatic change to the composite depth of field effect has been made. However, when a second type of input (e.g., double tap input 650u in FIG. 6U) is received, computer system 600 makes a user-specified change to the composite depth of field effect applied. When computer system 600 makes a user-specified change to the applied composite depth of field effect, computer system 600 automatically reapplies the user-specified change to the composite depth of field effect after the automatic change to the composite depth of field effect has been made (e.g., as further described below in connection with FIG. 6Y).As shown in FIG. 6V, since the computer system 600 has determined that the double-tap input 650v is a second type of input, the computer system 600 displays a tracking indicator 694a (e.g., "AF Tracking Lock"). The tracking indicator 694a indicates that the autofocus setting (e.g., and / or the currently applied synthetic depth of field) is not automatically changed by the computer system 600. The tracking indicator 694a is displayed within the camera user interface simultaneously with the live preview 630 of FIG. 6V.

[0228] Returning to FIGS. 6T - 6V, the computer system 690 displays the previously captured media representation 640 along with an animation of a user-specified change (e.g., occurring in response to detecting the double-tap input 650u) in the synthetic depth of field effect (e.g., during playback of the captured video). As shown in FIGS. 6T - 6V, since the computer system 690 has information indicating that a user-specified change is made (for reasons similar to those described above in relation to FIGS. 6N - 6P), when displaying the previously captured media representation 640 along with the user-specified change in the synthetic depth of field effect, it provides a smoother transition (e.g., compared to when displaying the live preview 630 of FIGS. 6T - 6V).

[0229] As shown by the live preview 630 of FIG. 6V, Jane 634 begins to walk out of the field of view of one or more cameras (e.g., walks out of the scene as shown by the live preview 630 of FIG. 6Q), and the synthetic depth of field effect moves with Jane 634 (e.g., for the same reasons described in connection with FIGS. 6P - 6Q as shown in FIGS. 6U - 6T). In FIG. 6W, Jane 634 is not within the field of view of one or more cameras (e.g., has walked out of the scene). In FIG. 6W, a decision is made that John 632 meets a modified set of automatic selection criteria (e.g., because Jane 634 is outside the FOV, Jane 634's face cannot be detected by the computer system 600, and / or the computer system 600 is not detecting activity from Jane 634 as indicated by Jane's activity level 680). As shown in FIG. 6W, the computer system 600 automatically changes the synthetic depth of field effect to emphasize John 632 (e.g., John 632 is displayed with only natural blur (e.g., no shadows) relative to the dog 638 that has entered the field of view of one or more cameras). The computer system 600 automatically changes the synthetic depth of field effect to emphasize John 632 relative to the dog 638 (e.g., using the same techniques for the same reasons disclosed above in connection with FIG. 6W). As shown in FIG. 6W, since the computer system 600 applies the synthetic depth of field effect to emphasize John 632 relative to the dog 638, the primary subject indicator 672a is displayed around John 632's head, and the secondary subject indicator 674c is displayed around the dog 638's head.

[0230] As shown in FIG. 6X, since a determination has been made that the dog 638 meets a set of automatic selection criteria (e.g., as indicated by the dog's activity level 680c1 exceeding John's activity level 680a1 at approximately 34 seconds on graph 680), the computer system 600 changes the synthetic depth of field effect to emphasize the dog 638 relative to John 632. Here, the dog 638 meets the set of automatic selection criteria because Jane 634 is not within the field of view of one or more cameras, and does not meet the modified set of criteria. Further, since a determination has been made that the dog 638 meets the set of automatic selection criteria, the computer system 600 displays the primary subject indicator 672c around the head of the dog 638 and the secondary subject indicator 674a around the head of John 632.

[0231] As shown in FIG. 6Y, Jane 634 has come back into the field of view of one or more cameras (e.g., standing within the scene shown by the live preview 630 in FIG. 6Y). In FIG. 6Y, the computer system 600 changes the synthetic depth-of-field effect to emphasize Jane 634 relative to other subjects (e.g., John 632, dog 638) within the field of view of one or more cameras. In particular, since the computer system 600 applied the change to the synthetic depth-of-field effect in response to detecting a user-specified change to the synthetic depth-of-field effect as a double-tap input 650u, the computer system 600 changes the synthetic depth-of-field effect to emphasize Jane 634 relative to other subjects. That is, the computer system 600 changes the synthetic depth-of-field effect to emphasize Jane 634 relative to other subjects in FIG. 6Y, regardless of whether an automatic change to the synthetic depth-of-field effect was applied after a permanent change to the synthetic depth-of-field effect was made (e.g., in response to detecting the double-tap input 650u). As shown in FIG. 6Y, since the synthetic depth-of-field effect is applied to emphasize Jane 634 relative to other subjects, the computer system 600 displays a primary subject indicator 678b around the head of Jane 634 and secondary subject indicators 674a and 674c around the heads of John 632 and dog 638, respectively. In some embodiments, in FIG. 6Y, the computer system 600 applies the synthetic depth-of-field effect to emphasize Jane 634 relative to other subjects based on a determination that Jane 634 has been inside and / or inside the region 603 of FIG. 6R1 for less than a predetermined period (e.g., 0.5 seconds to 5 seconds). In some embodiments, based on a determination that Jane 634 has been outside and / or inside the region 603 of FIG. 6R1 for more than a predetermined period, the computer system 600 does not apply the synthetic depth-of-field effect to emphasize Jane 634 relative to other subjects.

[0232] Figures 6Z through 6AB are exemplary embodiments in which computer system 600 modifies the synthetic depth-of-field effect in response to a third type of user input (e.g., a user-specified change). As shown in Figure 6Z, live preview 630 is displayed with a synthetic depth-of-field effect that emphasizes Jane 634 relative to other subjects within the media. In Figure 6Z, computer system 600 detects a press-and-hold input 650z on dog 638. In some embodiments, press-and-hold input 650z is detected at a different location on live preview 630 (e.g., a location not occupied by John 632, Jane 634, and dog 638, a location not corresponding to the location of the subject, etc.).

[0233] In FIG. 6AA, in response to detecting a press-and-hold input 650z on dog 638, computer system 600 changes the synthetic depth-of-field effect (e.g., because the press-and-hold input is a third type of input different from the first and second types of inputs) to emphasize the focal plane of the field of view of one or more cameras. The focal plane to be emphasized includes the location, object, and / or subject corresponding to the location, object, and / or subject where the press-and-hold input 650z was detected. Since dog 638 is located within the focal plane, dog 638 is emphasized relative to other subjects in the live preview 630 (e.g., as shown by dog 638 having no shadow). Further, since John 632 is closer to the focal plane that is more emphasized than Jane 634 (e.g., as shown by the shadow in the live preview 630), John 632 is displayed with less blur than Jane 634. In response to detecting the press-and-hold input 650z, computer system 600 displays a focus indicator 676 at the location corresponding to the location where the press-and-hold input 650z was detected. Further, in response to detecting the press-and-hold input 650z, computer system 600 displays secondary subject indicators 674a and 674b around the heads of John 632 and Jane 634, respectively. In FIG. 6AA, the focus indicator 676 is displayed to indicate that the focal plane is being emphasized by the synthetic depth-of-field effect. In some embodiments, the focus indicator 676 is displayed because dog 638 is in the focal plane and is currently being emphasized. However, in some embodiments, a secondary subject indicator 674c is displayed around the head of dog 638.

[0234] In FIG. 6AB, the live preview 630 shows John 632, Jane 634, and dog 638 moving away from the currently emphasized focal plane (as indicated, for example, by the focus indicator 676). As shown in FIG. 6AB, since John 632, Jane 634, and dog 638 are not within the currently emphasized focal plane, they are displayed with a blurred composite amount. In some embodiments, one or more portions of the live preview 630 that are within the focal plane are emphasized (for example, while the focal plane is emphasized in response to detecting the press-and-hold input 650z). In FIG. 6AB, the computer system 600 detects a tap input 650ab on the stop control 616.

[0235] Figures 6AC - 6AQ show exemplary embodiments in which the video captured in FIGS. 6B - 6AB is displayed and edited (e.g., in response to detecting a tap input 650b2). In FIG. 6AC, in response to detecting a tap input 650ab, the computer system 600 stops capturing the video and saves the captured video (e.g., the one captured in FIGS. 6B - 6AB). As shown in FIG. 6C, in response to detecting a tap input 650ab, the computer system 600 updates the media collection 624 to display a representation of the captured video (captured in FIGS. 6B - 6AB). In some embodiments, the computer system 600 detects one or more inputs and navigates to the cinematic video editing user interface shown in FIG. 6AD. In some embodiments, the one or more inputs include inputs directed to the media collection 624. In some embodiments, in response to detecting an input on the media collection 624, a representation of the captured video is displayed and controls for editing the captured video are displayed. In some embodiments, the one or more inputs include inputs on controls for editing the captured video. In some embodiments, in response to detecting an input directed to a control for editing the captured video, the computer system 600 displays the cinematic video editing user interface of FIG. 6AD.

[0236] FIG. 6AD shows a computer system 600 that displays a cinematic video editing user interface including a control region 662, a media representation 660, media navigation elements 664, and media editing mode control 684. The control region 662 is disposed on top of the media representation 660 and includes a completion control 662a, a redo control 662b1, an undo control 662b2, a cinematic video control 662c, a synthetic depth of field effect (SDOFE) control 662d, a depth indicator control 662e, a mute control 662f, and a cancel control 662g. In some embodiments, in response to detecting an input directed to the completion control 662a, the computer system 600 saves the representation of the media edited while the cinematic video editing user interface is being displayed. In some embodiments, the computer system 600 displays the completion control 662a as non-selectable when no changes and / or modifications are being made to the media (e.g., the media represented by the media representation 660). In some embodiments, the computer system 600 displays the completion control 662a as selectable when at least one change and / or modification has been made to the media using the cinematic video editing user interface. In some embodiments, when the completion control 662a is non-selectable, the computer system 600 does not save the representation of the media in response to detecting an input directed to the completion control 662a. In some embodiments, in response to detecting an input directed to the redo control 662b1, the computer system 600 reverses the most recent excessive operation. In some embodiments, in response to detecting an input directed to the undo control 662b2, the computer system 600 cancels the most recent edit made to the media (and, in some embodiments, queues all edits).In some embodiments, in response to detecting an input directed to the cinematic video control 662c, the computer system 600 performs one or more operations as described below in connection with FIGS. 6AP - 6AQ. In some embodiments, the SDOFE control 662d indicates that the computer system 600 is displaying and / or is currently configured to display media frames via the media representation 660, and that a synthetic depth of field effect is manually applied to the frames (e.g., a user - specified change to the synthetic depth of field effect as described above in connection with FIGS. 6O - 6AB). In some embodiments, the SDOFE control 662d indicates that the computer system 600 is displaying and / or is currently configured to display media frames via the media representation 660, and that a synthetic depth of field effect is automatically applied to the frames (e.g., an automatic change to the synthetic depth of field effect as described above in connection with FIGS. 6B - 6N). In some embodiments, in response to detecting an input directed to the SDOFE control 662d, the computer system 600 stops displaying the media using the user - specified change to the synthetic depth of field effect in the media while continuing to display the media using the automatic change to the synthetic depth of field effect. In some embodiments, in response to detecting an input directed to the SDOFE 662d, the computer system 600 modifies the media representation 660 such that one or more user - specified changes to the synthetic depth of field effect are not applied to one or more frames of the media while maintaining the application of the automatic change to the synthetic depth of field effect (as further described below in connection with FIGS. 6AZ - 6BC). In some embodiments, in response to detecting an input directed to the SDOFE control 662d, the computer system 600 modifies the media representation 660 such that one or more automatic changes to the synthetic depth of field effect are not applied to one or more frames of the media while maintaining the user - specified change to the application of the synthetic depth of field effect (e.g., the user - specified change as described in connection with FIGS. 6O - 6AB).In some embodiments, in response to detecting an input directed to the depth indicator control 662e, the computer system 600 performs one or more operations as described above in connection with FIGS. 6AD - 6AG. In some embodiments, in response to detecting an input directed to the mute control 662f, the computer system 600 switches the setting (e.g., on / off) that configures the computer system 600 to output sound while playing media. In some embodiments, in response to detecting an input directed to the cancel control 662g, the computer system 600 displays a confirmation screen for canceling one or more edits made to the media.

[0237] As shown in FIG. 6AD, media representation 660 is a representation of a frame of the video (the “captured video”) captured in FIGS. 6B - 6AB. In FIG. 6AD, media representation 660 is the first frame of the video and is captured before the live preview 630 in FIG. 6B is captured (e.g., the live preview 630 is captured at 0:00). In particular, since media representation 660 is displayed using a synthetic depth of field effect applied to emphasize John 632 relative to Jane 634 (for example, for the same reasons as described above in connection with FIG. 6B), media representation 660 includes a primary subject indicator 672a around the head of John 632 and a secondary subject indicator 674b around the head of Jane 634. Thus, computer system 600 displays subject indicators (e.g., primary subject indicators and / or secondary subject indicators) during the capture of the video (e.g., live preview 630) and while displaying a representation of a previously captured video (e.g., media representation 660). As shown herein, computer system 600 displays subject indicators while the media is not being played (e.g., media representation 660 in FIG. 6B) and during the playback of the media (e.g., media representation 660 in FIG. 6AK described below). In some embodiments, computer system 600 does not display subject indicators (and / or any subject indicators) while the media is not being played and during the playback of the media (e.g., a previously captured media representation 640).

[0238] As shown in FIG. 6AD, the media editing mode control 684 includes a cinematic video mode editing control 684a, a visual characteristic editing mode control 684b, a filter editing mode control 684c, and an aspect ratio editing mode control 684d. As shown in FIG. 6AD, the cinematic video mode editing control 684a is shown as being selected (e.g., as indicated by a selection indicator 684a1 displayed below the cinematic video mode editing control 684a in FIG. 6AD), which indicates that the cinematic video editing user interface is being displayed. In some embodiments, in response to detecting an input directed to the filter editing mode control 684c or the aspect ratio editing mode control 684d, the computer system 600 displays one or more controls corresponding to the control (e.g., the control to which the input is directed) selected for editing one or more frames of the video. In some embodiments, in response to detecting an input directed to the filter editing mode control 684c or the aspect ratio editing mode control 684d, the display of one or more user interface objects displayed within the cinematic video editing media user interface is stopped.

[0239] As shown in FIG. 6AD, the media navigation element 664 includes a scrubber region 664a, an effect region 664b, and playback controls 668a. The scrubber region 664a includes multiple representations of frames within the captured video, a playback head 664a1, a start crop control 664a2, and an end crop control 664a3. As shown in FIG. 6AD, the playback head 664a1 is displayed at a location corresponding to the start of the representation of the first frame of the captured video (e.g., the leftmost frame within the scrubber region 664a). Since the playback head 664a1 is displayed at a location corresponding to the start of the representation of the first frame (e.g., 0 seconds of the captured video), the media representation 660 of FIG. 6A is a representation of the first frame of the captured video (e.g., the time within the video corresponding to the location of the playback head 664a1). The start crop control 664a2 and the end crop control 664a3 indicate the portion of the captured video that is cropped and saved in response to the computer system 600 receiving a request to save the edited media (e.g., selection of the completion control 662a). In particular, the portion of the video to be cropped is the portion of the captured video that is between the start crop control 664a2 and the end crop control 664a3 (and / or from the time within the video corresponding to the location of the start crop control 664a2 within the scrubber region 664a to the time within the captured video corresponding to the location of the end crop control 664a3 within the scrubber region 664a).

[0240] As shown in FIG. 6AD, the effect region 664b includes a time bar 664b1 and change indicators 686a, 686b, 688c, 686d, 688e, 686f, 686g, and 688h (the "change indicators"). The time bar 664b1 has a plurality of tick marks, and each tick mark corresponds to a time in the captured video. The tick marks displayed on the time bar 664b1 cover at least a portion of the entire length of the captured video. In FIG. 6AD, each change indicator is displayed near (e.g., on and / or adjacent to) a tick mark on the time bar 664b1 corresponding to a time in the captured video where the computer system 600 changed the application of the synthetic depth-of-field effect applied to the visual content of the captured video. In FIG. 6AD, the effect region 664b is copied onto the graph 680 (the "effect region 664b - enlarged") to show how the change indicators correspond to changes in the application of the synthetic depth-of-field effect applied to the visual content of the video. In some embodiments, one or more change indicators are displayed at a start, end, or intermediate (average) position (e.g., relative to the tick marks of the time bar 664b1) when the actual application of the synthetic depth-of-field effect applied to the visual content is changed (e.g., while the video is being captured and / or after the video has been captured). In some embodiments, each of the change indicators is displayed below an individual representation of a frame within the scrubber region 664a corresponding to the time when the synthetic depth-of-field effect was applied to the content representing the individual frame. In some embodiments, the individual representation of a frame within the scrubber region is displayed with the synthetic depth-of-field effect applied during the time the individual frame within the scrubber region was captured (e.g., such that the frames within the scrubber region include blur). In some embodiments, the representation of the frame does not include blur and / or indicates the synthetic depth-of-field effect being applied.

[0241] In particular, the change indicators 686a, 686b, 686d, 686f, and 686g (the "automatic change indicators") indicate that the changes in the application of the synthetic depth of field effect were automatically performed by the computer system 600. Table 1 (Change Indicator Correspondence Table) is provided below to quickly summarize the connection of each of the change indicators in FIG. 6AD to the captured video.

Table 1

[0242] As shown in FIG. 6AD, the automatic change indicator is represented by a change indicator indicated using X, and the user-specified change indicator is represented by a change indicator indicated using O. Since the automatic change indicator has a different visual appearance from the user-specified change indicator, the automatic change indicator is represented differently from the user-specified change indicator. Further, each of the user-specified change indicators is displayed together with a transition indicator (e.g., 688c1, 688e1, and / or 688h1) that extends from the user-specified change to the next change (e.g., a change immediately to the right of the user-specified change and / or a change to the right end of the effect region 664b). In some embodiments, the transition indicator represents an individual period between media to which the user-specified change is applied and frames of the media that occur during the individual period. In some embodiments, one or more other techniques (e.g., using different colors, sizes, changes, text, locations, etc.) can be used to distinguish the automatic change indicator from the user-specified change indicator. In some embodiments, the user-specified change indicator is displayed and the automatic change indicator is not displayed, and / or vice versa. In some embodiments, the computer system 600 includes selectable options (e.g., SDOFE control 662d) for stopping the display of the automatic change indicator and / or the user-specified change indicator while maintaining the display of the user-specified change indicator and / or vice versa. In some embodiments, the user-specified change indicator that occurs during video capture is displayed differently (e.g., with a different visual appearance) from the user-specified change indicator that occurs after the video is captured (e.g., during video editing). In FIG. 6AD, the computer system 600 detects a tap input 650ad on the depth indicator control 662e.

[0243] As shown in FIG. 6AE, in response to detecting a tap input 650ad, computer system 600 displays a depth control 682 to the left of media editing mode control 684 (e.g., upward in portrait orientation when computer system 600 is in portrait orientation). Depth control 682 is a slider that is displayed with a depth control value 682a (e.g., displayed within depth indicator control 662e of FIG. 6AD). In some embodiments, in response to detecting a tap input 650ad, computer system 600 stops displaying the scrubber region 664a and the effect region 664b (e.g., the scrubber region 664a and the effect region 664b are not displayed while the depth control 682 is not displayed and / or are displayed while the depth control 682 is displayed). In FIG. 6AE, computer system 600 detects a rightward swipe input 650ae on depth control 682.

[0244] In FIG. 6AF, in response to detecting a rightward swipe input 650ae, computer system 600 changes the depth control value 682a from a 4.5 f-stop value to a 1.4 f-stop value, which increases the blur applied to portions of the media representation 660 that are not included in (e.g., not focused on) John 632, which is currently being emphasized (e.g., focused on) by the synthetic depth-of-field effect applied to the frame corresponding to the media representation 660 of FIG. 6AF. In FIG. 6AF, John 632 is not displayed with an additional amount of blur (e.g., not darker compared to John 632 in FIG. 6AE) in response to detecting the rightward swipe input 650ae, but Jane 634 as well as the background and foreground portions of the media representation 660 are displayed with an additional amount of blur (e.g., become darker compared to how each individual portion was blurred in FIG. 6AE). Thus, the adjustment to the depth control 682 causes the synthetic depth-of-field effect applied to be adjusted. In some embodiments, the adjustment to the depth control 682 causes an adjustment only to the representation of the frames of the captured video that are being displayed via the media representation 660 when the adjustment is being performed. In some embodiments, the adjustment to the depth control 682 causes an adjustment to the frames of the captured video (e.g., all and / or most of the frames), regardless of whether the synthetic depth-of-field effect was applied to the frames of the captured video (e.g., a global change) or not. In some embodiments, the adjustment to the depth control 682 causes an adjustment to the frames of the captured video for which the same application of the synthetic depth-of-field effect was applied (e.g., the frames of the video in which John 632 is being emphasized by the synthetic depth-of-field effect in FIG. 6AF, and / or the frames of the video that correspond to and / or occur after the change in the synthetic depth-of-field effect of the media representation 660 of FIG. 6AF but before a different change in the synthetic depth-of-field effect (e.g., between 0 seconds and 3 seconds in FIG. 6AF)).In FIG. 6AF, computer system 600 detects a tap input 650af1 on depth control 682 and / or a leftward swipe input 650af2 on depth control 682.

[0245] As shown in FIG. 6AF1, in response to detecting tap input 650af1, computer system 600 stops displaying depth control 682 and continues to display media representation 660 having the same amount of blur as it had prior to the detection of tap input 650af1. Further, computer system 600 updates the display of depth indicator control 682 to include the value (e.g., 1.4) previously set by depth control 662e (e.g., in response to detecting a rightward swipe input 650ae). In some embodiments, computer system 600 updates the display of depth indicator control 662e to include the value (e.g., 1.4) selected in response to detecting a rightward swipe input 650ae.

[0246] As shown in FIG. 6AG, in response to detecting a leftward swipe input 650af2, computer system 600 changes the depth control value 682a from a 1.4f stop value to a 4.5f stop value, reducing the blur applied to unfocused portions of media representation 660 (e.g., indicated by lighter shading as compared to FIG. 6AF). In some embodiments, the techniques described herein related to depth control 682 also function on depth indicator 602e (e.g., prior to / during capture of media, as discussed above in relation to FIG. 6B). In FIG. 6AG, computer system 600 detects a tap input 650ag on depth indicator control 662e. As shown in FIG. 6AH, in response to detecting tap input 650ag, computer system 600 stops displaying depth control 682 and continues to display media representation 660 with the same amount of blur it had prior to the detection of tap input 650ag. Further, computer system 600 updates the display of depth indicator control 682 to include the value (e.g., 4.5) previously set for depth control 662e (e.g., in response to detecting leftward swipe input 650af2). As shown in FIG. 6AH, computer system 600 detects a tap input 650ah on media playback control 668a. In response to detecting tap input 650ah, computer system 600 begins playback of the captured video.

[0247] Figures 6AI through 6AO illustrate exemplary embodiments in which user-specified changes are made in the captured video. In Figure 6AI, computer system 600 is playing the captured video, as indicated by the paused play control 668b of Figure 6AH being displayed and the media play control 668a being displayed to stop. As shown in Figure 6AI, playback head 664a1 is displayed at a location corresponding to the frame displayed during the 7-second duration of the captured video (indicated by elapsed time indicator 664c displayed above playback head 664a1), and media representation 660 is updated to be the representation of the frame displayed during the 7-second duration of the captured video. In particular, media representation 660 corresponds to the live preview 630 of Figure 6K (e.g., represents the same frame), and an automatic change to the synthetic depth of field effect has been applied to emphasize John 632 over Jane 634. Thus, the media representation 660 of Figure 6AI includes a primary subject indicator 672a around the head of John 632 and a secondary subject indicator 674b around the head of Jane 634 to reflect the applied synthetic depth of field effect. In Figure 6AI, computer system 600 detects a single tap input 650ai on Jane 634 at the 7-second mark during media playback.

[0248] In FIG. 6AJ, in response to detecting a single tap input 650ai, computer system 600 changes the synthetic depth of field effect to emphasize Jane 634 over John 632. As shown in FIG. 6AJ, the synthetic depth of field effect is applied to the representation of the frame of the video displayed at the 8 second mark within the captured video (e.g., as indicated by elapsed time indicator 664c). FIG. 6AJ shows the representation of the frame of the video that occurred after the single tap input 650ai was detected, but computer system 600 changes the synthetic depth of field effect for all of the frames of the edited media that were applied between the 5 second mark (e.g., when the single tap input 650ai was detected) and the 12 second mark (e.g., when the next change to the synthetic depth of field effect occurs within the captured video, as indicated by user-specified change representation 688c) within the captured video. The edited media playback line 680d3 of graph 680 also indicates when and how the synthetic depth of field effect was changed in response to detecting the single tap input 650ai. As shown by graph 680, the edited media playback line 680d3 is separated from the media playback line 680d2 and indicates that computer system 600 changed the application of the synthetic depth of field effect in response to detecting the single tap input 650ai and when the change occurred. In particular, because computer system 600 replaces the automatic change indicator 686b of FIG. 6AI with the user-specified change indicator 688i in response to detecting the single tap input 650ai, the edited media playback line 680d3 shifts to be positioned over the activity tracker 680b (e.g., "Jane's tracker") between the 5 second mark and the 12 second mark.

[0249] As shown in FIG. 6AJ, in response to detecting a single tap input 650ai, computer system 600 stops displaying the automatic change indicator 686b of FIG. 6AI and displays a user-specified change indicator 688i (e.g., together with the transition indicator 688i1) at the location where the automatic change indicator 688b was displayed. Thus, in some embodiments, user-specified changes during media editing can supersede automatic and / or user-specified changes that occurred during media capture and / or during media editing. In some embodiments, computer system 600 detects individual inputs on the representation of a frame on a video that do not correspond to an individual time within a video in which a change in the synthetic depth of field effect occurred, and in response to detecting the individual inputs, computer system 600 displays an additional user-specified change indicator. In some embodiments, computer system 600 displays an additional user-specified change indicator while continuing to display other change indicators. In some embodiments, in response to detecting an individual input, computer system 600 changes (e.g., based on the input) the application of the synthetic view to a plurality of frames of the video starting from an individual time within the video. In some embodiments, in response to detecting a single tap input 650ai, computer system 600 displays an animation of a transition indicator 688i1 that gradually fills (e.g., gradually increases in size by expanding from the right end of the transition indicator) from the position of the user-specified change indicator 688i to the next change indicator (e.g., the user-specified change indicator 688c). In FIG. 6AJ, computer system 600 detects a tap input 650aj on the pause playback control 668b. In response to detecting the tap input 650aj, computer system 600 pauses the playback of the media.

[0250] As shown in FIG. 6AK, media representation 660 is displayed with a representation of a frame corresponding to the 10 - second mark of the video (as indicated, for example, by playback head 664a1 and elapsed time indicator 664c). Further, playback control 668a is displayed at the location where pause - playback control 668b was previously displayed in FIG. 6AJ. In FIG. 6AK, media representation 660 is a representation of the same frame within the captured media to which live preview 630 of FIG. 6AL corresponds. In particular, media representation 660 of FIG. 6AK is different from live preview 630 of FIG. 6AL, which is due to the fact that media representation 660 is a frame having a synthetic depth - of - field effect applied to emphasize Jane 634 with respect to John 632, and live preview 630 is a frame having a synthetic depth - of - field effect applied to emphasize John 632 with respect to Jane 634. When computer system 600 changes the application of the depth - of - field effect due to an input detected on a frame of the video (e.g., a representation of a frame of the video), computer system 600 also changes the application of the depth - of - field effect applied to frames of the video that occur after the frame of the video on which the input was received. In FIG. 6AK, computer system 600 detects a tap input 650ak on user - specified change indicator 688h.

[0251] As shown in FIG. 6AL, in response to detecting a tap input 650ak, computer system 600 displays playback head 664a1 over a user-specified change indicator 688h. By having playback head 664a1 over user-specified change indicator 688h, playback head 664a1 is displayed at a location corresponding to the time at which a user-specified change (e.g., the user-specified change represented by user-specified change indicator 688h) occurred in the captured video. In response to detecting a tap input 650ak, computer system 600 updates media representation 660 to be the representation of the frame that was displayed when the user-specified change occurred (e.g., as shown by media representation 660 of FIG. 6AL which is the live preview 630 of FIG. 6Z in which a synthetic depth-of-field effect has been applied to emphasize the focal plane and / or live preview 630 of FIG. 6AA). In FIG. 6AL, computer system 600 detects a double-tap input 650al.

[0252] As shown in FIG. 6AM, in response to detecting a double-tap input 650al, computer system 600 changes the synthetic depth-of-field effect to emphasize John 632 with respect to Jane 634. Further, computer system 600 displays a primary subject indicator 678a around John 632's head and secondary subject indicators 674b - 674c around Jane 634's and the dog 638's heads, respectively. Since the double-tap input 650al is a double-tap input, computer system 600 applies the synthetic depth-of-field effect to emphasize John 632 with respect to Jane 634 and does not automatically change the applied synthetic depth-of-field effect as long as computer system 600 can detect John 632 (e.g., John 632's face) in the visual content of the captured video (using, for example, one or more of the techniques described above in relation to the detection of double-tap input 650u). In particular, computer system 600 performs the same actions in response to detecting the same type of input (e.g., changing the synthetic depth-of-field effect in the same way and displaying the same type of indicators), regardless of whether computer system 600 is capturing and / or editing media (e.g., performing the same actions as above in response to detecting a single-tap input 650o, 650ai, in response to detecting a double-tap input 650u, 650al, in response to detecting a press-and-hold input). As shown by graph 680, the edited media playback line 680d3 is separated from the media playback line 680d2 after the 4-second mark to indicate that computer system 600 changed the application of the synthetic depth-of-field effect in response to detecting the double-tap input 650al and when a change occurred.Specifically, after the 42 - second mark (e.g., the frame of the media while the double - tap input 650al is detected), in the edited media, to indicate that John 632 is being emphasized and tracked (not the selected focal plane), the edited media playback line 680d3 is changed such that it is on the activity tracker 680a (e.g., the "John tracker"). In some embodiments, in response to detecting the double - tap input 650al, the computer system 600 replaces the user - specified change indicator 688h with a new user - specified change indicator.

[0253] FIG. 6AN shows a computer system 600 that displays a media representation 660 including a captured video representation that occurs after the previously captured media representation 660 of FIG. 6AM. As shown in FIG. 6AN, the computer system 600 applies a synthetic depth - of - field effect to emphasize John 632 relative to Jane 634 in the media representation shown by the media representation 660 (e.g., the media representation 660 is different from the live preview 630 of FIG. 6AB for the same reasons described above in relation to FIG. 6AK).

[0254] Figures 6AO to 6AP illustrate exemplary embodiments in which options for removing changes in the application of the synthetic depth of field effect are displayed. In Figure 6AN, computer system 600 detects a tap input 650an on user-specified change indicator 688h. As shown in Figure 6AO, in response to detecting tap input 650an, computer system 600 displays a delete option 688h2 adjacent to user-specified change indicator 688h and does not emphasize (e.g., grays out) scrubber region 664a and effect region 664b. Here, computer system 600 does not emphasize (e.g., grays out) scrubber region 664a and effect region 664b to indicate that other portions (e.g., portions not including delete option 669h1) are unavailable, inactive, and / or do not respond to user input. Computer system 600 makes other portions unavailable, inactive, and / or not respond to user input to avoid the possibility that the user causes the computer system to perform an unintended operation when the user attempts to select delete option 688h2. In some embodiments, in response to detecting an input at a location not corresponding to delete option 688h2, computer system 600 re-emphasizes scrubber region 664a and effect region 664b and / or stops displaying delete option 688h2. In Figure 6AO, computer system 600 detects a tap input 650ao on delete option 688h2. As shown in Figure 6AP, in response to detecting tap input 650ao, computer system 600 changes the application of the synthetic depth of field effect from emphasizing John 632 with respect to Jane 634 and re-emphasizes scrubber region 664a and effect region 664b (e.g., makes scrubber region 664a and effect region 664b active). When computer system 600 changes the application of the synthetic depth of field effect from emphasizing John 632 with respect to Jane 634, computer system 600 returns to the application of the synthetic depth of field effect that would have been applied if the removed user-specified change had not been made.Thus, in FIG. 6AP, in response to detecting a double - tap input 650u (using, for example, one or more techniques as described above in connection with FIGS. 6U - 6Y), since a permanent change in the application of the synthetic depth - of - field effect has been applied, the computer system 600 updates the media representation 660 to emphasize Jane 634 over John 632. As shown by graph 680, the edited media playback line 680d3 is changed to indicate that the computer system 600 has changed the application of the synthetic depth - of - field effect in response to detecting a tap input 650an and when the change was made. In FIG. 6AP, the computer system 600 detects a tap input 650ap1 on the cinematic video control 662c.

[0255] As shown in FIG. 6AQ, in response to detecting a tap input 650ap1, computer system 600 displays the cinematic video control 662c in a non-active state and stops applying the synthetic depth-of-field effect to the captured video within the media editing user interface (e.g., as shown by the media representation 660 without shadows). In some embodiments, in response to detecting a tap input 650ap1, computer system 600 displays the change indicator as non-selectable (e.g., grayed out) or stops displaying one or more of the change indicators. In some embodiments, in response to detecting an input directed to the cinematic video control 662c in FIG. 6AQ, computer system 600 reapplies the synthetic depth-of-field effect to the captured video within the media editing user interface. In some embodiments, in response to detecting a tap input on the completion control 662a, computer system 600 saves a version of the captured video to which the synthetic depth-of-field effect is not applied (e.g., a version of the captured video having only natural blur for one or more and / or all of the frames within the video). In some embodiments, in response to detecting a tap input 650ap1, computer system 600 stops displaying the effect region 664b within the region 664d. In some embodiments, computer system 600 moves the scrubber region 664a downward, and a portion of the scrubber region 664a is moved downward within the region 664d. In some embodiments, computer system 600 enlarges the size of the media representation 660 and / or the scrubber region 664a in response to detecting a tap input 650ap1. In some embodiments, in response to detecting a tap input 650ap1, computer system 600 de-emphasizes the effect region 664b and / or displays the effect region 664b as non-active.

[0256] FIG. 6AR shows an exemplary embodiment in which the playback head 664a1 is dragged across the scrubber region 664a so that the playback head 664a1 snaps to a location corresponding to the change indicator. As shown in FIG. 6AR, a rightward swipe input 650ar is detected at location 654a, and the computer system 600 determines that the location 654a is not within a first predetermined distance from a location corresponding to a user-specified change indicator 654c (a "change indicator location") (e.g., it is determined that the playback head 664a1 is not displayed at the change indicator location), so the playback head 664a1 is displayed as being at location 654a. When the rightward swipe input 650ar is detected at location 654b, the computer system 600 displays the playback head 664a1 at the change location (e.g., above the user-specified change indicator 688c), which is in front of location 654b because it has been determined that the location 654b is within a first predetermined distance from the change indicator location (e.g., it has been determined that the playback head 664a1 is not displayed at the change indicator location). As shown in FIG. 6AR, when the playback head 664a1 is displayed at the change location, the computer system emits an output 656 (e.g., a haptic output (e.g., vibration), a sound). When the rightward swipe input 650ar is detected at location 654c, the computer system 600 continues to display the playback head 664a1 at the change location because it has been determined that the location 654c is not within a second predetermined distance from the change indicator location (e.g., it has been determined that the playback head 664a1 is displayed at the change indicator location). When the rightward swipe input 650ar is detected at location 654d, the computer system 600 displays the playback head 664a1 at location 654d because it has been determined that the location 654d is within a second predetermined distance from the change indicator location (e.g., it has been determined that the playback head 664a1 is displayed at the change indicator location).Thus, in some embodiments, the playback head snaps to the location associated with the change indicator when the playback head is near the change indicator. Further, in some embodiments, the playback head remains at the location associated with the change indicator until the playback head is a certain distance away from the change indicator. In some embodiments, the first predetermined distance and / or the second predetermined distance is a non-zero distance and / or a distance greater than a certain number of tick marks (e.g., 2 - 5 tick marks) away from the changing location.

[0257] Figures 6AS - 6AU illustrate exemplary embodiments in which computer system 600 transitions from being configured to operate in a cinematic video camera mode to being configured to operate in a portrait camera mode. As shown in Figure 6AS, computer system 600 is configured to operate in a cinematic video camera mode (as indicated by, e.g., cinematic video mode control 620e in an active state), and while configured to operate in the cinematic video camera mode, computer system 600 displays a camera user interface using one or more of the techniques described above in connection with Figure 6B. In particular, as shown in Figure 6AS, computer system 600 applies a synthetic depth of field effect to the visual content captured by one or more cameras of computer system 600 to emphasize John 632 over Jane 634 (as indicated, e.g., by the shading of live preview 630 in Figure 6AS). As shown in Figure 6AS, computer system 600 displays a primary subject indicator 672a around John 632's head and a secondary subject indicator 674b around Jane 634's head. In Figure 6AS, computer system 600 detects a left - swipe input 650 on camera mode control 620.

[0258] As shown in FIG. 6AT, in response to detecting a leftward swipe input 650as, computer system 600 moves camera mode control 620 to the left such that portrait mode control 620b is displayed at the center of the camera user interface. In FIG. 6AT, computer system 600 displays portrait mode control 620b as being selected (e.g., in bold) and stops displaying cinematic video mode control 620e (e.g., indicating that cinematic video mode control 620e is not selected). Further, in response to detecting a leftward swipe input 650as, computer system 600 transitions from a state configured to operate in cinematic video camera mode to portrait camera mode. As shown in FIG. 6AT, in response to detecting a leftward swipe input 650as, computer system 600 compresses live preview 630, and the live preview 630 of FIG. 6AT is smaller than the live preview 630 of FIG. 6AS and has a different aspect ratio. In addition to shortening the live preview 630, computer system 600 is updated to include lighting effect control 618. Lighting effect control 618 indicates that a natural light effect is being applied to live preview 630 (e.g., as indicated by the displayed natural light control 618a and natural light indicator 618a1). In some embodiments, when a natural light effect is applied to live preview 630, a blur effect and / or lighting effect is used / applied when capturing media. In some embodiments, adjustments to lighting effect control 618 are also reflected in live preview 630.

[0259] As shown in FIG. 6AT, the computer system 600 does not display any subject indicators (e.g., primary subject indicator 672a, secondary subject indicator 674b) to indicate whether an individual subject is emphasized or not. While operating in portrait camera mode, the computer system 600 does not apply a synthetic depth of field effect to emphasize one subject over another. However, while operating in portrait camera mode, the computer system 600 applies a blur effect and / or an illumination effect based on the selected natural light control 618a (e.g., indicated by the shadow in the live preview 630 of FIG. 6AT). In FIG. 6AT, the computer system 600 detects a press-and-hold input 650at on the live preview 630.

[0260] As shown in FIG. 6AU, in response to detecting the press-and-hold input 650at, the computer system 600 displays a focus and exposure control 696 that includes an exposure control indicator 696a1. While displaying the focus and exposure control 696, the computer system 600 also displays a focus setting indicator 602c (“AE / AF lock”) within an indicator region 694c, which indicates that the computer system 600 does not allow automatic exposure settings and autofocus settings to be changed automatically. In FIG. 6AU, in response to detecting the press-and-hold input 650at, the computer system 600 blurs a portion of the display such that the computer system 600 focuses on the location corresponding to the location where the press-and-hold input 650at was received and blurs other portions of the region. In some embodiments, in response to detecting a swipe input on the live preview 630, the computer system 600 adjusts the exposure settings based on the size and direction of the swipe input.

[0261] In response to detecting a press-and-hold input, computer system 600 is configured to focus on a particular location within the FOV, regardless of whether the computer system 600 is operating in a cinematic camera mode (e.g., as described above in connection with the detection of press-and-hold input 650z in FIGS. 6Z - 6AA) or a portrait camera mode (e.g., as described above in connection with the leftward swipe input 650as in FIGS. 6AS - 6AU). Further, the visual appearance of the focus and exposure control 696 in FIG. 6AU appears similar to the focus indicator 676 in FIG. 6AA. However, the focus and exposure control 696 includes an exposure control indicator 696a1, which the focus indicator 676 does not. Further, the exposure control indicator 696a1 in FIG. 6AU also differs from the focus control indicator 694b. The exposure control indicator 696a1 indicates that the computer system 600 has locked the focus setting (e.g., the blur effect applied in FIG. 6AU) and the exposure setting, while the focus control indicator 694b indicates only that the computer system 600 has locked the focus setting (e.g., the synthetic depth of field effect applied in FIG. 6AA). Thus, while the computer system 600 is operating in the cinematic video camera mode, the computer system 600 displays controls (e.g., as described above in connection with FIGS. 6Z - 6AA) that indicate that the computer system 600 is configured to focus on a particular location and that enable adjustment and / or locking of the exposure setting used by the computer system 600 to capture media. Further, while the computer system 600 is operating in the portrait camera mode, the computer system 600 displays controls that indicate that the computer system 600 is configured to focus on a particular location and that enable adjustment and / or locking of the exposure setting used by the computer system 600 to capture media (e.g., as described above in connection with FIGS. 6AS - 6AU).

[0262] Figures 6AV - 6AY illustrate exemplary embodiments where automatic changes for applying a synthetic depth of field effect are removed during media editing. Returning to Figure 6AP, the computer system 600 detects one or more inputs including a tap input 650ap2 on the cancel control 662g (e.g., as an alternative to the detection of the tap input 650ap1 as described above in connection with Figure 6AP). Referring to Figure 6AV, in response to detecting one or more inputs including the tap input 650ap2, the computer system 600 discards previous changes made to the media (e.g., changes to the application of one or more synthetic depth of field effects as described above in connection with Figures 6AD - 6AP). In other words, the computer system 600 resets the media to the state it was in before the media was edited in Figures 6AD - 6AP and / or after the media was captured. Thus, in Figure 6AV, the computer system 600 redisplayed the cinematic video editing user interface of Figure 6AD, including, among other things, the change indicators 686a, 686b, 688c, 686d, 688e, 686f, 686g, and 688h (automatic and user - specified synthetic depth of field changes as described above in connection with Figures 6A - 6AC). In Figure 6AV, the computer system 600 detects a tap input 650av on the automatic change indicator 686b.

[0263] As shown in FIG. 6AW, in response to detecting a tap input 650av, computer system 600 updates media representation 660 to a representation of a frame of media that occurs at the 7 - second mark in the media (e.g., a frame of media corresponding to the occurrence of an automatic change to synthetic depth - of - field shown by automatic change indicator 686b). As shown by media representation 660 in FIG. 6AW, computer system 600 automatically applies a synthetic depth - of - field effect to emphasize John 632 over Jane 634 at the 7 - second mark in the media. In FIG. 6AW, computer system 600 detects a tap input 650aw (or a press - and - hold input) on automatic change indicator 686b. As shown in FIG. 6AX, in response to detecting tap input 650aw, computer system 600 displays a delete option 686b2 adjacent to automatic change indicator 686b and de - emphasizes (e.g., grays out) scrubber region 664a and effect region 664b (using one or more similar techniques as described above in connection with FIGS. 6AN - 6AO). In FIG. 6AX, computer system 600 detects a tap input 650ax on delete option 686b2.

[0264] As shown in FIG. 6AY, in response to detecting the tap input 650ax, the computer system 600 removes the automatic change indicator 686b in FIG. 6AX and the automatic change to the synthetic depth of field effect applied to the 7-second mark in the media. As part of removing the automatic change to the synthetic depth of field effect, the computer system 600 updates the media representation 660 to show Jane 634 highlighted with respect to John 632 at the 7-second mark in the media. Here, since the automatic depth of field effect corresponding to the automatic change indicator 686a (e.g., the most recent synthetic depth of field effect applied before the 7-second mark) is currently applied to the frame of the media that occurs at the 7-second mark in the media, Jane 634 is highlighted with respect to John 632. Further, it should also be understood that the automatic synthetic depth of field effect corresponding to the automatic change indicator 686a is applied to other frames of the media captured between the time corresponding to the automatic change indicator 686a (e.g., 4 seconds) and the time corresponding to the user-specified change indicator 688c (e.g., 12 seconds). Thus, when the automatic change indicator 686b is removed, the computer system 600 applies the synthetic depth of field effect corresponding to the automatic change indicator 686a to the frame of the media to which the synthetic depth of field effect corresponding to the automatic change indicator 686b was previously applied. As shown by the graph 680 in FIG. 6AY, the edited media playback line 680d3 is separated from the media playback line 680d2 between the 6-second mark and the 10-second mark to show the change to the synthetic depth of field effect that occurred in response to detecting the tap input 650ax (e.g., the edited media playback line 680d3 is on the activity tracker 680b "Jane's tracker" between the 6-second mark and the 10-second mark in FIG. 6AY, which is different from the position of the edited media playback line 680d3 in the corresponding time frame in FIG. 6AX).

[0265] Figures 6AZ - 6BC illustrate an exemplary embodiment in which computer system 600 detects one or more inputs on SDOFE control 662d. In FIG. 6AY, computer system 600 detects a tap input 650ay on user - specified change indicator 688h. As shown in FIG. 6AZ, computer system 600 moves playback head 664a1 from the 7 - second mark to the 42 - second mark to the right and updates media representation 660 to show the frame of the media corresponding to the 42 - second mark (e.g., the frame corresponding to user - specified change indicator 688h). As shown in FIG. 6AZ, media representation 660 has a synthetic depth - of - field effect applied to emphasize the focal plane (e.g., as described above in connection with FIGS. 6Z - 6AB). In FIG. 6AZ, since dog 638 is located within the focal plane (shown by, for example, focus indicator 676), dog 638 is emphasized relative to other subjects within media representation 660 (shown, for example, by dog 638 having no shadow within media representation 660). Further, since John 632 is closer to the focal plane that is more emphasized than Jane 634 (shown by, for example, the shadow of media representation 660), John 632 is displayed with less blur than Jane 634. In FIG. 6AZ, computer system 600 detects a tap input 650az on SDOFE control 662d.

[0266] As shown in FIG. 6BA, in response to detecting the tap input 650az, the computer system 600 stops applying the change in the depth of field effect corresponding to the user-specified changes (e.g., user-specified change indicators 688c, 688e, and 688h in FIG. 6AZ) in the edited media. Further, since the computer system 600 is configured not to apply the previously applied user-specified synthetic depth of field effect changes (e.g., in response to detecting the tap input 650az), in response to detecting the tap input 650az, the computer system 600 stops displaying the user-specified change indicators 688c, 688e, and 688h, as well as the transition indicators 688c1, 688e1, and 688h1. In particular, the computer system 600 removes the user-specified change indicators 688c and 688e without replacing them with another change indicator. However, at the 42-second mark, the computer system 600 replaces the user-specified change indicator 688h in FIG. 6AZ with the automatic change indicator 686ba in FIG. 6BA. Thus, based on the determination that an automatic change to the synthetic depth of field effect should be made, when the computer system 600 removes the user-specified change to the synthetic depth of field effect, it can insert an automatic change to the synthetic depth of field effect (e.g., using one or more techniques described below in connection with FIG. 12). Here, this individual determination (e.g., the determination that an automatic change to the synthetic depth of field effect should be made) was made because the activity level 680a1 ("John's activity level") increased relative to the activity levels 680b1 ("Jane's activity level") and 680c1 (the dog's activity level) at the 42-second mark. Thus, as shown by the media representation 660, based on this individual determination and since the user-specified change is no longer applied at the 42-second mark, the computer system 600 automatically applies the synthetic depth of field effect to emphasize John 632 relative to Jane 634 and the dog 638 at the 42-second mark in the video.In some embodiments, this individual determination is made while (e.g., as described below with respect to FIG. 12) capturing media (e.g., and / or before user-specified changes are removed). In some embodiments, this individual determination can be made available to be applied (or reapplied) when user-specified changes are removed, and is saved during the capture of the media (e.g., as described below in connection with FIG. 12). In some embodiments, user-specified changes can disable the saved automatic changes to the synthetic depth-of-field effect (e.g., as described below in connection with FIG. 12). In some embodiments, this individual determination is made after user-specified changes are removed. In FIG. 6BA, computer system 600 detects a leftward swipe gesture 650ba on playback head 664a1.

[0267] As shown in FIG. 6BB, in response to detecting a leftward swipe gesture 650ba, computer system 600 moves playback head 664a1 leftward from the location corresponding to 42 seconds in the media to the location corresponding to 34 seconds in the media. As shown in FIG. 6BB, in response to detecting a leftward swipe gesture 650ba, computer system 600 updates media representation 660 to show the frame of the media corresponding to 34 seconds in the media. At the 34 - second mark, computer system 600 applies a compositing effect depth that emphasizes John 632 with respect to wagon 628 (as described above in connection with, for example, FIG. 6W). In some embodiments, in response to detecting input 650bb1 on SDOFE control 662d, computer system 600 reapplies a user - specified depth - of - field change to the media's representation and redisplay user - specified change indicators 688c, 688e, and 688h and transition indicators 688c1, 688e1, and 688h1 (e.g., the edited media and the cinematic video editing user interface return to the state shown in FIG. 6AZ and / or the state prior to the detection of tap input 650az). In FIG. 6BB, computer system 600 detects input 650bb2 on wagon 628.

[0268] As shown in FIG. 6BC, in response to detecting the input 650bb2 and based on the determination that the input 650bb2 is a press-and-hold input, the computer system 600 changes the synthetic depth-of-field effect to emphasize the focal plane at the location of the press-and-hold input 650bb2 (starting from the 42-second mark in the media). Further, the computer system 600 displays a user-specified change indicator 688j and a transition indicator 688j1 at a location within the effect region 664b corresponding to the 42-second mark in the media. As shown in FIG. 6BC, in response to detecting the input 650bb2 and based on the determination that the input 650bb2 is a press-and-hold input, the computer system 600 also displays a focus setting indicator 694bc ("AF Lock - 5M") that includes a display of the distance (e.g., "5M") between the computer system 600 and the currently selected focal plane (e.g., the focal plane selected by the input 650bb2). After applying the synthetic depth-of-field effect that emphasizes the focal plane in FIG. 6BC, the media representation 660 shows the wagon 628 that is emphasized with respect to John 632 and Jane 634. Here, since the wagon 628 is located on the focal plane where the wagon 628 is emphasized, it is emphasized with respect to John 632 and Jane 634 in the media representation 660. In particular, since it has been determined that no automatic change to the synthetic depth-of-field effect corresponding to the automatic change indicator 686g is required, the computer system 600 stops displaying the automatic change indicators 686g and 686ba in FIG. 6BB. Returning to FIG. 6W, since it has been determined that Jane 634 (e.g., the currently emphasized subject) is outside the field of view of one or more cameras of the computer system 600, an automatic change to the synthetic depth-of-field effect corresponding to the automatic change indicator 686g has been made. However, Jane 634 is no longer emphasized immediately before the time corresponding to the automatic change indicator 686g by the synthetic depth-of-field effect.Therefore, in FIG. 6BC, since Jane 634 is no longer emphasized, the computer system 600 removes the automatic change to the synthetic depth-of-field effect that was made because the currently emphasized subject (e.g., Jane 634) could not be detected within the field of view of one or more cameras of the computer system 600. The computer system 600 also removes the automatic change indicator 686ba for similar reasons (e.g., since the user specified that the focal plane be emphasized, the computer system determines that there is no need to implement changes to emphasize subjects in the media via the application of the synthetic depth-of-field effect). Thus, as shown in FIGS. 6BB - 6BC, the computer system 600 can remove changes to the s...

Claims

Claim 1 A method comprising: In a computer system communicating with one or more cameras and one or more input devices, Detecting a request to capture a video representing the field of view of the one or more cameras via the one or more input devices; In response to detecting the request to capture the video, Capturing a video over a first capture duration, the video comprising a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration; Applying to the plurality of frames of the video a synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the first subject relative to the second subject within the plurality of frames of the video, the synthetic depth of field effect varying over time as the first subject moves within the field of view of the one or more cameras. Claim 2 Applying the synthetic depth of field effect to the plurality of frames of the video comprises: Displaying a first set of frames of the plurality of frames, the first set of frames including displaying the second subject at a first distance from the one or more cameras and having a first blur amount; Displaying a second set of frames of the plurality of frames, the second set of frames including displaying the second subject at a second distance from the one or more cameras and having a second blur amount different from the first blur amount, wherein the first distance is different from the second distance. The method according to claim 1. Claim 3 When applying the synthetic depth of field effect, the first subject is displayed with a third blur amount and the second subject is displayed with a fourth blur amount greater than the third blur amount. The method according to claim 1 or 2. Claim 4 Applying the synthetic depth-of-field effect to the plurality of frames of the video comprises: Applying a fifth amount of blur to a first portion of a third frame among the plurality of frames; Applying a sixth amount of blur greater than the fifth amount of blur to a second portion of the third frame among the plurality of frames, the method according to any one of claims 1 to 3. **Claim 5** Applying the synthetic depth-of-field effect to the plurality of frames of the video comprises: Blurring a portion of a fourth frame among the plurality of frames that does not include a subject within the field of view of the one or more cameras, the method according to any one of claims 1 to 4. **Claim 6** Applying the synthetic depth-of-field effect to the plurality of frames of the video comprises blurring a foreground of a fifth frame of the plurality of frames with respect to the first subject and a background of the fifth frame with respect to the subject, the method according to any one of claims 1 to 5. **Claim 7** The video includes a second plurality of frames captured over a second capture duration, The second plurality of frames represents the first subject within the field of view of the one or more cameras and a third subject within the field of view of the one or more cameras, The method comprises: Detecting an instruction that the third subject in the second plurality of frames should be emphasized with respect to the first subject in the second plurality of frames while the video is being captured over the first capture duration; In response to detecting the instruction, applying a second synthetic depth-of-field effect to the second plurality of frames of the video to change the visual information captured by the one or more cameras to emphasize the third subject in the second plurality of frames with respect to the first subject in the second plurality of frames, further comprising the method according to any one of claims 1 to 6. The method according to any one of claims 1 to 6. **Claim 8** The method according to claim 7, wherein the computer system automatically detects the instruction when the third subject in the second plurality of frames meets a set of automatic selection criteria. **Claim 9** The method according to claim 7 or 8, wherein the set of automatic selection criteria includes criteria that are satisfied based on the movement of the third subject within the field of view of the one or more cameras.

10. The method according to any one of claims 7 to 9, wherein the set of automatic selection criteria includes criteria that are satisfied when it is determined that the face of the third subject is detected within the field of view of the one or more cameras.

11. The method according to any one of claims 7 to 10, wherein the set of automatic selection criteria includes criteria that are satisfied based on the voice corresponding to the third subject.

12. The method according to any one of claims 7 to 11, wherein the set of automatic selection criteria includes criteria that are satisfied based on the distance between the third subject and the one or more cameras in one or more of the second plurality of frames.

13. The method according to any one of claims 7 to 12, wherein the set of automatic selection criteria includes criteria that are satisfied based on the line of sight of the third subject.

14. The method according to any one of claims 7 to 13, wherein the set of automatic selection criteria includes criteria that are satisfied based on the position of the appendage organs of the third subject.

15. The method according to any one of claims 7 to 14, wherein the set of automatic selection criteria includes criteria that are satisfied based on one or more changes in features detected in the captured video.

16. Detecting a first gesture via the one or more input devices while capturing the video over the first capture duration; Modifying the set of automatic selection criteria in response to detecting the first gesture; The method according to any one of claims 7 to 15, further comprising.

17. The method according to claim 7, wherein the computer system detects the instruction when a second gesture is detected via the one or more input devices.

18. In response to detecting the instruction, while capturing the video Displaying a first animation including a first transition from a display of one or more representations of a plurality of frames having the synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames with respect to the applied second subject, to a display of one or more representations of a second plurality of frames having a second synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the third subject in the second plurality of frames of the video with respect to the first subject in the second plurality of frames of the video. The method according to any one of claims 7 to 17, further comprising. **Claim 19** While playing the video at a time after the capture of the video has ended, displaying a second animation corresponding to the first animation, the second animation starting in the playback of the video at a time corresponding to a time point in the video that occurred before the time point in the video at which the instruction was detected. The method according to claim 18, further comprising. **Claim 20** The second synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the third subject in the second plurality of frames of the video with respect to the first subject in the second plurality of frames of the video is a synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize a selected focal plane in the video, and the transition characteristics for displaying the first animation are based on a difference between the selected focal plane in the video and a previous focal plane in the video. The method according to claim 18 or 19. **Claim 21** According to a determination that the distance between the selected focal plane and the previous focal plane is a first distance, the speed of the animation is a first speed, According to a determination that the distance between the selected focal plane and the previous focal plane is a second distance shorter than the first distance, the speed of the animation is a second speed faster than the first speed. The method according to claim 20. **Claim 22** Applying the synthetic depth of field effect includes maintaining focus at a location corresponding to the first subject while the first subject is at least partially occluded. The method according to any one of claims 1 to 21.

23. Displaying a first user interface object indicating that the first subject is emphasized while applying the synthetic depth of field effect. The method according to any one of claims 1 to 22, further comprising.

24. The method according to claim 23, wherein the first user interface object indicating that the first subject is emphasized is displayed while the video is being captured.

25. The method according to claim 23, wherein the first user interface object indicating that the first subject is emphasized is displayed after the capture of the video is completed.

26. While applying the synthetic depth of field effect, displaying a second user interface object corresponding to the second subject, the appearance of which is different from that of the user interface object indicating the first subject to which the synthetic depth of field effect is applied. The method according to any one of claims 1 to 25, further comprising.

27. The method according to any one of claims 1 to 26, wherein the first subject is a person, an animal, or an object.

28. Before detecting the request to capture the video. While the computer system is configured to operate in a first capture mode, detecting a third gesture. In response to detecting the third gesture, configuring the computer system to operate in a cinematic video capture mode different from the first capture mode. The method according to any one of claims 1 to 27, further comprising.

29. While the computer system is configured to operate in the first capture mode, a first representation of the field of view of the one or more cameras is displayed. While the computer system is configured to operate in the cinematic video capture mode, a second representation of the field of view of the one or more cameras is displayed, the second representation having less blur than the first representation. The method according to claim 28.

30. While the computer system is configured to operate in the cinematic video capture mode, detecting a fourth gesture in a direction different from the third gesture; In response to detecting the fourth gesture, configuring the computer system to operate in a still image capture mode; The method according to claim 28 or 29, further comprising:

31. Before detecting the request to capture the video, While the computer system is configured to operate in a second capture mode, detecting a fifth gesture; In response to detecting the fifth gesture, configuring the computer system to operate in a portrait capture mode; The method according to any one of claims 1 to 30, further comprising:

32. Applying the synthetic depth of field effect to the plurality of frames of the video includes adjusting the magnitude of the synthetic depth of field effect applied to the video. The method according to any one of claims 1 to 31.

33. The computer system communicates with a display generation component, and the method further includes: After adjusting the magnitude of the synthetic depth of field effect applied to the video, displaying a representation of the magnitude of the synthetic depth of field effect applied to the video. The method according to claim 32.

34. After applying the synthetic depth of field effect to the plurality of frames of the video, detecting a second request to apply the synthetic depth of field effect to a second plurality of frames of the captured video; In response to detecting the second request, Applying the synthetic depth of field effect to the second plurality of frames of the video captured by the first type of tracking according to a determination that the second request was detected based on the detection of the first type of gesture. In accordance with the determination that the second request has been detected based on the detection of the second type of gesture, applying the synthetic depth of field effect to the second plurality of frames of the video captured by a second type of tracking, which is different from the first type of tracking, of the second type of gesture; The method according to any one of claims 1 to 33, further comprising.

35. In response to detecting the second request, In accordance with the determination that the second request has been detected based on the detection of a third type of gesture, applying the synthetic depth of field effect to the second plurality of frames of the video captured by a third type of tracking, which is different from the first type of tracking and the second type of tracking, of the third type of gesture; The method according to claim 34, further comprising.

36. The method according to claim 34, wherein the second request is one of a single tap gesture, a multi - tap gesture, and a press - and - hold gesture.

37. The method according to any one of claims 34 to 36, wherein the second request is based on a gesture not directed at one or more subjects within the plurality of frames.

38. The first subject within the plurality of frames of the video is at a third distance from the one or more cameras, The second subject within the plurality of frames of the video is at a fourth distance from the one or more cameras, which is closer to the one or more cameras than the third distance. The method according to any one of claims 1 to 37.

39. Capturing the video over the first capture duration, At a first time during the first capture duration, adjusting one or more settings of a first camera among the one or more cameras to focus on a first focal plane corresponding to the first subject; At a second time during the first capture duration, detecting a change in the distance between the first subject and the first camera while the first camera is aligned with the first focal plane; In response to detecting the change in the distance between the first subject and the first camera, adjusting the one or more settings of the first camera to focus on a second focal plane different from the first focal plane corresponding to the first subject; After capturing the video over the first capture duration, detecting an instruction that, in the first plurality of frames corresponding to the second time among the first plurality of frames with respect to the first subject in the second plurality of frames, a second subject should be emphasized; In response to detecting the instruction that the second subject should be emphasized in the first plurality of frames with respect to the first subject in the second plurality of frames, applying, to the plurality of frames of the video, an individual synthetic depth of field effect that modifies visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames of the video with respect to the first frame of the plurality of frames of the video while the second focal plane has not been changed. The method according to any one of claims 1 to 38, comprising:

40. Applying, to the plurality of frames of the video, the synthetic depth of field effect that modifies visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames of the video with respect to the second subject in the plurality of frames of the video; Using an object detection algorithm to identify, in the plurality of frames of the video, one or more objects and one or more characteristics of the one or more objects; Providing the one or more identified objects and the one or more identified characteristics of the one or more identified objects to a neural network; Based on the one or more identified objects and the one or more identified characteristics of the one or more identified objects, obtaining an output from the neural network that identifies the first subject from among the one or more objects for application of the synthetic depth of field effect. The method according to any one of claims 1 to 39, comprising:

41. The method according to claim 40, wherein the neural network is trained using training data including user preference data that identifies which object in a video within a set of the captured videos the user selected for emphasis at a plurality of times within the set of the captured videos.

42. After applying the synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames of the video with respect to the second subject in the plurality of frames of the video, while the neural network continues to identify the first subject from among the one or more objects for individual application of an individual synthetic depth-of-field effect, detecting a request to emphasize the second subject in the plurality of frames of the video; In response to detecting the request to emphasize different subjects in the plurality of frames of the video, applying a different synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames of the video with respect to the first subject in the plurality of frames of the video, wherein after applying the different synthetic depth-of-field effect, the synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames of the video with respect to the second subject in the plurality of frames of the video is stored as a change to a default depth-of-field effect; The method according to claim 40 or 41, further comprising.

43. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 42.

44. A computer system configured to communicate with one or more cameras and one or more input devices, One or more processors, A computer system comprising: a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to perform the method according to any one of claims 1 to 42.

45. A computer system configured to communicate with one or more cameras and one or more input devices, means for performing the method according to any one of claims 1 to 42 comprising a computer system.

46. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 42.

47. One or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras and one or more input devices, detecting a request to capture a video representing the field of view of the one or more cameras via the one or more input devices; in response to detecting the request to capture the video, capturing a video over a first capture duration, the video being a plurality of frames captured over the first capture duration, the plurality of frames including a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration; applying to the plurality of frames of the video a synthetic depth-of-field effect that modifies visual information captured by the one or more cameras to emphasize the first subject over the second subject within the plurality of frames of the video, the synthetic depth-of-field effect changing over time as the first subject moves within the field of view of the one or more cameras, the one or more programs storing instructions for performing the above.

48. A computer system configured to communicate with one or more cameras and one or more input devices, one or more processors, one or more programs configured to be executed by the one or more processors, detecting a request to capture a video representing the field of view of the one or more cameras via the one or more input devices; in response to detecting the request to capture the video, capturing a video over a first capture duration, the video being a plurality of frames captured over the first capture duration, the plurality of frames including a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration; applying, to the plurality of frames of the video, a synthetic defocus depth effect that modifies visual information captured by the one or more cameras to emphasize the first subject relative to the second subject within the plurality of frames of the video, the synthetic defocus depth effect varying over time as the first subject moves within the field of view of the one or more cameras, a memory storing one or more programs including instructions for performing the above; Claim 49 A computer system configured to communicate with one or more cameras and one or more input devices, means for detecting a request to capture a video representing the field of view of the one or more cameras via the one or more input devices; in response to detecting the request to capture the video, Means for capturing a video over a first capture duration, the video comprising a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration. Means for applying a synthetic depth of field effect that modifies visual information captured by the one or more cameras to emphasize the first subject within the plurality of frames of the video relative to the second subject within the plurality of frames of the video, the synthetic depth of field effect varying over time as the first subject moves within the field of view of the one or more cameras, to the plurality of frames of the video. A computer system comprising: **Claim 50** One or more programs configured to be executed by one or more processors of a computer system that communicates with one or more cameras and one or more input devices, Detecting a request to capture a video representing the field of view of the one or more cameras via the one or more input devices; In response to detecting the request to capture the video, Capturing a video over a first capture duration, the video comprising a plurality of frames captured over the first capture duration, the plurality of frames representing a first subject within the field of view of the one or more cameras and a second subject within the field of view of the one or more cameras, wherein in the plurality of frames, the first subject is moving relative to the field of view of the one or more cameras over the first capture duration. A computer program product comprising one or more programs for performing: changing visual information captured by the one or more cameras with a synthetic depth of field effect that emphasizes the first subject in the plurality of frames of the video with respect to the second subject in the plurality of frames of the video; and applying the synthetic depth of field effect that changes over time as the first subject moves within the field of view of the one or more cameras to the plurality of frames of the video. [

51. ] A method comprising: In a computer system communicating with one or more cameras, a display generation component, and one or more input devices, Via the display generation component, a user interface comprising: A representation of a video comprising a plurality of frames, the representation including a first subject and a second subject; and Displaying a user interface including a first user interface object indicating that the first subject is emphasized by a synthetic depth of field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames with respect to the second subject; While displaying the user interface including the representation of the video and the first user interface object, detecting, via the one or more input devices, a gesture corresponding to selection of the second subject in the representation of the video; In response to detecting the gesture corresponding to selection of the second subject in the representation of the video, Changing the synthetic depth of field effect to change the visual information captured by the one or more cameras and emphasizing the second subject in the plurality of frames with respect to the first subject; and Displaying a second user interface object indicating that the second subject is emphasized by the changed synthetic depth of field effect that changes the visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames with respect to the first subject. [

52. ] The method according to claim 51, wherein the first user interface object and the second user interface object have the same visual appearance. **Claim 53** Before detecting the gesture corresponding to the selection of the second subject, displaying, via the display generation component, a third user interface object indicating that the second subject is not emphasized. The method according to claim 51 or 52, further comprising: **Claim 54** The method according to claim 53, wherein the first user interface object has a visual appearance different from that of the third user interface object. **Claim 55** The representation of the video includes a third subject, and the method further comprises: Before detecting the gesture corresponding to the selection of the second subject, displaying, via the display generation component, a fourth user interface object indicating that the second subject is not emphasized and a fifth user interface object indicating that the third subject is not emphasized. The method according to any one of claims 51 to 54. **Claim 56** The method according to claim 55, wherein the fourth user interface object and the fifth user interface object have different visual appearances. **Claim 57** In response to detecting the gesture corresponding to the selection of the second subject, stopping the display of the first user interface object. The method according to any one of claims 51 to 56, further comprising: **Claim 58** In response to detecting the gesture corresponding to the selection of the second subject, displaying a sixth user interface object indicating that the first subject is not emphasized. The method according to any one of claims 51 to 57, further comprising: **Claim 59** The method according to any one of claims 51 to 58, wherein the gesture corresponding to the selection of the second subject is detected while the one or more cameras are capturing the visual information. **Claim 60** The method according to any one of claims 51 to 59, wherein the gesture corresponding to the selection of the second subject is detected during playback of the video after capture of the video has ended. **Claim 61** The method according to any one of claims 51 to 60, wherein the gesture corresponding to the selection of the second subject is a first single-tap gesture.

62. The method according to any one of claims 51 to 61, wherein the gesture corresponding to the selection of the second subject is a first multi-tap gesture.

63. The method according to any one of claims 51 to 61, wherein the gesture corresponding to the selection of the second subject is a first press-and-hold gesture.

64. Changing the synthetic depth-of-field effect to emphasize the second subject within the plurality of frames with respect to the first subject and changing the visual information captured by the one or more cameras, changing the visual information captured by the one or more cameras to emphasize the second subject until a first criterion is met, according to a determination that the gesture corresponding to the selection of the second subject is a first type of gesture; changing the visual information captured by the one or more cameras to emphasize the second subject until a second criterion different from the first criterion is met, according to a determination that the gesture corresponding to the selection of the second subject is a second type of gesture different from the first type of gesture; The method according to any one of claims 51 to 63, comprising:

65. wherein the first type of gesture is a second single-tap gesture, and the second type of gesture is a second multi-tap gesture, The method according to claim 64.

66. detecting a gesture of the first type of gesture directed at the second subject while the visual information captured by the one or more cameras is being changed to emphasize the second subject until a first criterion is met; in response to detecting the gesture of the first type of gesture directed at the second subject, changing the visual information captured by the one or more cameras to emphasize the second subject until a second criterion is met; The method according to claim 64 or 65, further comprising:

67. Changing the synthetic depth of field effect to emphasize the second subject in the plurality of frames with respect to the first subject and changing the visual information captured by the one or more cameras. The method according to any one of claims 51 to 66, including changing the visual information captured by the one or more cameras to emphasize the second subject by applying the synthetic depth of field effect to a fixed focal plane in the plurality of frames according to a determination that the gesture corresponding to the selection of the second subject is a third type of gesture.

68. Displaying an indication of the distance to the fixed focal plane according to a determination that the gesture corresponding to the selection of the second subject is the third type of gesture. The method according to claim 67, further comprising.

69. While displaying the second user interface object and not displaying the first user interface object, Displaying the first user interface object and stopping the display of the second user interface object according to a determination that the first subject in the plurality of frames meets a set of automatic selection criteria. The method according to any one of claims 51 to 68, further comprising.

70. According to a determination that the gesture corresponding to the selection of the second subject is a fourth type of gesture, the set of automatic selection criteria is a first set of automatic selection criteria, According to a determination that the gesture corresponding to the selection of the second subject is a fifth type of gesture different from the fourth type of gesture, the set of automatic selection criteria is a second set of automatic selection criteria different from the first set of automatic selection criteria. The method according to claim 69.

71. Before detecting the gesture corresponding to the selection of the second subject, the set of automatic selection criteria includes criteria that are satisfied when an individual subject in the representation of the media meets a first selection confidence threshold. In response to detecting the gesture corresponding to the selection of the second subject, the set of automatic selection criteria includes criteria that are satisfied when the individual subject in the representation of the media meets a second selection confidence threshold higher than the first selection confidence threshold. The method according to claim 69 or 70. **Claim 72** The synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject changes over time as the second subject moves within the field of view of the one or more cameras, the method according to any one of claims 51 to 71. **Claim 73** The user interface includes video navigation user interface elements, and the method further displays a user interface object indicating that a user-specified change has occurred at a certain time within the video in the video navigation user interface element in response to detecting the gesture corresponding to the selection of the second subject while the video navigation user interface element is being displayed, the method according to any one of claims 51 to 72. **Claim 74** The user interface object indicating that the user-specified change has occurred includes a fourth visual appearance according to a determination that the gesture corresponding to the selection of the second subject is a sixth type of gesture, and a fifth visual appearance different from the fourth visual appearance according to a determination that the gesture corresponding to the selection of the second subject is a seventh type of gesture different from the sixth type of gesture, the method according to claim 73. **Claim 75** Displaying the second user interface object includes displaying the second user interface object having a sixth visual appearance according to a determination that the gesture corresponding to the selection of the second subject is an eighth type of gesture, and displaying the second user interface object having a seventh visual appearance different from the sixth visual appearance according to a determination that the gesture corresponding to the selection of the second subject is a ninth type of gesture different from the eighth type of gesture, the method according to any one of claims 51 to 74. **Claim 76** The user interface is a media capture user interface, and the method further After detecting the gesture corresponding to the selection of the second subject, while the user interface is being displayed, detecting one or more gestures via the one or more input devices; In response to detecting the one or more gestures, a media editing user interface, A second representation of the video including a third plurality of frames, the second representation including the first subject and the second subject; A sixth user interface object indicating that the first subject is emphasized by a synthetic subject depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject, and displaying the media editing user interface; While the media editing user interface is being displayed, detecting a second gesture corresponding to the selection of the second subject in the second representation of the video via the one or more input devices; In response to detecting the second gesture corresponding to the selection of the second subject in the second representation of the video, Changing the synthetic subject depth-of-field effect to change the visual information captured by the one or more cameras and emphasizing the second subject with respect to the first subject within the third plurality of frames; Displaying a seventh user interface object indicating that the second subject is emphasized by the changed synthetic subject depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the second subject within the third plurality of frames with respect to the first subject, the method according to any one of claims 51 to 75. **Claim 77** After detecting the gesture corresponding to the selection of the second subject and changing the synthetic subject depth-of-field effect to change the visual information captured by the one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject, detecting a first gesture directed to the representation of the media; In response to detecting the first gesture directed to the representation of the media, to modify the changed synthetic depth-of-field effect to change the visual information captured by the one or more cameras. The method according to any one of claims 51 to 76, further comprising.

78. The method according to any one of claims 51 to 77, wherein the user interface includes a selectable user interface object for changing the synthetic depth-of-field effect, which changes the synthetic depth-of-field effect when selected.

79. The user interface includes a selectable user interface object for controlling the video capture mode. The selectable user interface object for controlling the video capture mode is displayed together with a status indication indicating that the video capture mode is in an active state. The method is as follows. While displaying the user interface including the representation of the video, the first user interface object and the selectable user interface object for controlling the video capture mode are displayed together with the status indication indicating that the video capture mode is in an active state, and applying the synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames relative to the second subject. While applying the synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames relative to the second subject, detecting a gesture directed to the selectable user interface object for controlling the video capture mode. In response to detecting the gesture directed to the selectable user interface object for controlling the video capture mode, stopping the application of the synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames relative to the second subject. Further comprising. The method according to any one of claims 51 to 78.

80. Before detecting the gesture directed to the selectable user interface object for controlling the video capture mode, the representation is displayed with a first amount of blur, and the method further In response to detecting the gesture directed to the selectable user interface object for controlling the video capture mode, displaying the representation of the video with a second amount of blur smaller than the first amount of blur via the display generation component, the method according to claim 79.

81. Configuring the focus setting of one or more cameras to focus on the second subject in the representation of the video in response to detecting the gesture corresponding to the selection of the second subject, wherein the computer system is not configured to automatically change the focus setting of the one or more cameras for at least a predetermined period, the configuring; Detecting a second gesture directed to the representation of the video while the computer system is configured to focus on the second subject in the representation of the video; Enabling the computer system to automatically change the focus setting of the one or more cameras for at least the predetermined period in response to detecting the second gesture directed to the representation of the video; The method according to any one of claims 51 to 80, further comprising.

82. The representation of the video includes a representation of a subset of content from a first portion of the field of view of one or more cameras, The field of view of the one or more cameras extends beyond the first portion of the field of view to a second portion of the field of view of the one or more cameras not included in the representation, A decision regarding which subject to emphasize is based on information from the second portion of the field of view of the one or more cameras in the video. The method according to any one of claims 51 to 81.

83. The method according to claim 82, wherein the decision regarding which subject to emphasize includes automatically selecting the individual subject before the individual subject to be emphasized is visible within the first portion of the field of view.

84. The decision regarding which subject to emphasize is detecting that the individual subject moves outside the first portion of the field of view while the individual subject is being emphasized; in response to detecting that the individual subject has moved outside the first portion of the field of view, automatically selecting a different subject to be emphasized according to a determination that the individual subject has moved outside the second portion of the field of view; refraining from selecting a different subject to be emphasized for at least a predetermined time period according to a determination that the first subject remains within the second portion of the field of view, the method according to claim 82 or 83.

85. One or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras, a display generation component, and one or more input devices, the one or more programs storing instructions for executing the method according to any one of claims 51 to 84, a non-transitory computer-readable storage medium.

86. A computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs storing instructions for executing the method according to any one of claims 51 to 84, a computer system.

87. A computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices, means for executing the method according to any one of claims 51 to 84 a computer system.

88. One or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras, a display generation component, and one or more input devices, the one or more programs including instructions for executing the method according to any one of claims 51 to 84, a computer program product.

89. One or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras, a display generation component, and one or more input devices, Via the display generation component, a user interface comprising: A representation of a video including a plurality of frames, the representation including a first subject and a second subject; and A first user interface object indicating that the first subject is emphasized by a synthetic depth of field effect that modifies visual information captured by the one or more cameras to emphasize the first subject within the plurality of frames with respect to the second subject. While displaying the user interface including the representation of the video and the first user interface object, detecting, via the one or more input devices, a gesture corresponding to a selection of the second subject in the representation of the video; In response to detecting the gesture corresponding to the selection of the second subject in the representation of the video, Changing the synthetic depth of field effect to modify the visual information captured by the one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject; and Displaying a second user interface object indicating that the second subject is emphasized by the modified synthetic depth of field effect that modifies the visual information captured by the one or more cameras to emphasize the second subject within the plurality of frames with respect to the first subject. **Claim 90** A computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices, the computer system comprising: One or more processors; One or more programs configured to be executed by the one or more processors, the one or more programs comprising: Via the display generation component, a user interface comprising: A representation of a video including a plurality of frames, the representation including a first subject and a second subject; A first user interface object indicating that the first subject is emphasized by a synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames with respect to the second subject; and, displaying a user interface including the same; while displaying the user interface including the representation of the video and the first user interface object, detecting, via the one or more input devices, a gesture corresponding to selection of the second subject in the representation of the video; in response to detecting the gesture corresponding to selection of the second subject in the representation of the video; changing the synthetic depth-of-field effect to change the visual information captured by the one or more cameras, and emphasizing the second subject in the plurality of frames with respect to the first subject; displaying a second user interface object indicating that the second subject is emphasized by the changed synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames with respect to the first subject; a memory storing one or more programs including instructions for performing the same; a computer system comprising the same. A computer system configured to communicate with one or more cameras, a display generation component, and one or more input devices, means for displaying, via the display generation component, a user interface comprising: a representation of a video including a plurality of frames, the representation including a first subject and a second subject; a first user interface object indicating that the first subject is emphasized by a synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames with respect to the second subject; While displaying the user interface including the representation of the video and the first user interface object, means for detecting a gesture corresponding to selection of the second subject in the representation of the video via the one or more input devices; In response to detecting the gesture corresponding to selection of the second subject in the representation of the video; Changing the synthetic depth-of-field effect to change the visual information captured by the one or more cameras, and emphasizing the second subject in the plurality of frames with respect to the first subject; Means for displaying a second user interface object indicating that the second subject is emphasized by the changed synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames with respect to the first subject; A computer system comprising:

92. One or more programs configured to be executed by one or more processors of a computer system communicating with one or more cameras, a display generation component, and one or more input devices, Via the display generation component, a user interface, A representation of a video including a plurality of frames, the representation including a first subject and a second subject, Displaying a user interface including: a first user interface object indicating that the first subject is emphasized by a synthetic depth-of-field effect that changes visual information captured by the one or more cameras to emphasize the first subject in the plurality of frames with respect to the second subject; While displaying the user interface including the representation of the video and the first user interface object, detecting, via the one or more input devices, a gesture corresponding to selection of the second subject in the representation of the video; In response to detecting the gesture corresponding to selection of the second subject in the representation of the video; Changing the synthetic depth of field effect to change the visual information captured by the one or more cameras and emphasizing the second subject in the plurality of frames relative to the first subject; A computer program product comprising one or more programs including instructions for displaying a second user interface object indicating that the second subject is emphasized by the changed synthetic depth of field effect that changes the visual information captured by the one or more cameras to emphasize the second subject in the plurality of frames relative to the first subject. **Claim 93** A method, in a computer system communicating with a display generation component, via the display generation component, a user interface comprising a video having a first duration, a representation of the video including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject relative to one or more elements in the video, the plurality of changes including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time; and simultaneously displaying a video navigation user interface element for navigating through the video including the representation of the first time and the representation of the second time, wherein the representation of the second time is visually distinguishable from other times in the first duration of the video that do not correspond to the change in subject emphasis, and the representation of the first time is visually distinguishable from the representation of the second time. **Claim 94** wherein the automatic change in subject emphasis is a first synthetic depth of field effect that changes the visual information captured by one or more cameras to emphasize a first subject in the video relative to a second subject in the video; The user-specified change in subject emphasis is a second synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize a third subject in the video relative to a fourth subject in the video. The method according to claim 93. **Claim 95** The video navigation user interface element for navigating through the video, The method according to claim 93 or 94, wherein the graphical user interface object indicating that the automatic change occurred at the first time is not included. **Claim 96** The video navigation user interface element for navigating through the video, At a first location on the video navigation user interface element, a first graphical user interface object indicating that the automatic change occurred at the first time in the video, the first graphical user interface object having a first visual appearance; and At a second location on the video navigation user interface element different from the first location, a second graphical user interface object indicating that the user-specified change occurred at a second time different from the first time in the video, the second graphical user interface object having a second visual appearance different from the first visual appearance, the method according to any one of claims 93 to 95. **Claim 97** The video navigation user interface element for navigating through the video includes graphical user interface objects indicating that individual changes occurred at individual times in the video that occurred before the second time in the video at individual locations on the video navigation user interface element, and the method further includes Displaying a visual indication extending from the individual location on the video navigation user interface element to the second location on the video navigation user interface element according to a determination that the individual change that occurred at the individual time in the video is an individual user-specified change, the method according to claim 96. **Claim 98** The method according to claim 96 or 97, wherein the second graphical user interface object is displayed in or adjacent to the representation at the second time.

99. The method according to any one of claims 93 to 98, wherein the change in the user-specified in subject emphasis is caused in response to a gesture detected while the video is being captured.

100. Detecting a gesture directed at the representation at the second time while the representation at the second time is being displayed; In response to detecting the gesture directed at the representation at the second time, displaying a second representation at the second time during the first duration of the video; The method according to any one of claims 93 to 99, further comprising.

101. Detecting a gesture directed at the video navigation user interface element while the video navigation user interface element is being displayed; Navigating through the representation of the video in response to detecting the gesture directed at the video navigation user interface element; The method according to any one of claims 93 to 100, further comprising.

102. Before detecting the gesture directed at the video navigation user interface, the video navigation user interface element includes a first playback head element at a first playback head location, The representation of the video is the representation of the video at a time corresponding to the first playback head location, The method includes In response to detecting the input indicated in the first user interface and detecting the gesture directed at the video navigation user interface element, Moving the first playback head from the first playback head location to a second playback head location; While stopping displaying the representation of the video at the time corresponding to the first playback head location, displaying the representation of the video at the time corresponding to the second playback head location, further comprising The method according to claim 101.

103. While detecting the gesture directed to the video navigation user interface element, moving a selectable indicator, displaying the selectable indicator that moves according to the detected speed of the gesture directed to the video navigation user interface element according to a determination that the selectable indicator is not within a threshold distance from the representation at the second time; displaying the selectable indicator at the representation at the second time according to a determination that the selectable indicator is within a threshold distance from the representation at the second time, moving the selectable indicator, The method according to any one of claims 101 or 102, further comprising:

104. Providing a haptic output corresponding to snapping to the second time according to a determination that the selectable indicator is within a threshold distance from the representation at the second time, The method according to claim 103.

105. The method according to claim 103 or 104, wherein the selectable indicator is the first playback head.

106. The method according to claim 103 or 104, wherein the selectable indicator is a trim indicator.

107. The representation of the video is a representation at a third time during the first duration, including a fifth subject and a sixth subject, displaying the representation of the video, displaying a first user interface object indicating that the fifth subject is emphasized by a synthetic depth of field effect that changes the visual information captured by the one or more cameras to emphasize the fifth subject in the representation of the video with respect to the sixth subject; The method according to any one of claims 93 to 106.

108. The fifth subject in the plurality of frames is displayed with a first visual characteristic, The sixth subject in the plurality of frames is displayed with a second visual characteristic different from the first visual characteristic, The method according to claim 107.

109. While displaying the representation of the video and the first user interface object, detecting a gesture corresponding to selection of the sixth subject in the representation of the video, In response to detecting the gesture corresponding to the selection of the sixth subject in the representation of the video, changing the synthetic depth-of-field effect to change the visual information captured by the one or more cameras to emphasize the sixth subject with respect to the fifth subject in the representation of the video; and The method according to claim 107 or 108, further comprising. **Claim 110** In response to detecting the gesture corresponding to the selection of the sixth subject in the representation of the video, displaying a seventh graphical user interface object indicating that the sixth subject is emphasized by the changed synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the sixth subject with respect to the fifth subject in the representation of the video; The method according to claim 109, further comprising. **Claim 111** The video navigation user interface element for navigating through the video includes, at a seventh location on the video navigation user interface element, the seventh graphical user interface object; at an eighth location on the video navigation user interface element, an eighth graphical object indicating that a synthetic depth-of-field change occurred at an eighth time within the video; and a portion between the seventh location and the eighth location, Before detecting the gesture corresponding to the selection of the sixth subject in the representation of the video, the portion of the video navigation user interface element between the seventh location and the eighth location is displayed in a first visual state, The method In response to detecting the gesture corresponding to the selection of the sixth subject in the representation of the video, further includes displaying an animation in which the portion of the video navigation user interface element between the seventh location and the eighth location changes from the first visual state to a second visual state different from the first visual state. The method according to claim 110. **Claim 112** In response to detecting the gesture corresponding to the selection of the sixth subject in the representation of the video, display, within the video navigation user interface element, a second representation of the third time, the second representation of the third time being a user-specified change in subject emphasis. The method according to any one of claims 109 to 111, further comprising. **Claim 113** The representation of the third time includes a seventh subject, and the method further comprises detecting a gesture corresponding to the selection of the seventh subject in the representation of the video while the representation of the video and the first user interface object are being displayed; and in response to detecting the gesture corresponding to the selection of the seventh subject in the representation of the video, changing the synthetic depth-of-field effect to change the visual information captured by the one or more cameras to emphasize the seventh subject relative to the fifth subject in the representation of the video; and displaying a third user interface object indicating that the seventh subject is emphasized by the changed synthetic depth-of-field effect that changes the visual information captured by the one or more cameras to emphasize the seventh subject relative to the fifth subject in the representation of the video. The method according to any one of claims 107 to 112 includes. **Claim 114** The video navigation user interface element for navigating through the video includes, at a third location on the video navigation user interface element, a third graphical user interface object indicating that the user-specified change occurred at the second time within the video, and the method further comprises detecting a gesture directed at the third graphical user interface object while the third graphical user interface object is being displayed; and in response to detecting the gesture directed at the third graphical user interface object, displaying an option for removing the user-specified change that occurred at the second time within the video. The method according to any one of claims 93 to 113 includes.

115. The video navigation user interface element for navigating through the video includes at a fourth location on the video navigation user interface element, a fourth graphical user interface object indicating that the user-specified change occurred at the second time within the video, after the representation of the second time, a plurality of representations including the one subject emphasized with respect to one or more elements within the video are displayed, the method according to any one of claims 93 to 114.

116. The representation of the video is a third representation of the second time, The third representation of the second time has a third visual appearance according to a determination that the user-specified change is a first type of user-specified change, and has a fourth visual appearance different from the third visual appearance according to a determination that the user-specified change is a second type of user-specified change different from the first type of user-specified change, the method according to any one of claims 93 to 115.

117. while displaying the video navigation user interface element, detecting a gesture directed to a sixth location on the video navigation user interface element; and in response to detecting the gesture directed to the sixth location on the video navigation user interface element, displaying a progress indicator representing the time in the playback of the video corresponding to the sixth location, the method according to any one of claims 93 to 116, further comprising.

118. The user interface includes a selectable user interface object for controlling a video editing mode, the selectable user interface object for controlling the video editing mode is displayed with a status indication indicating that the video editing mode is active, the video navigation user interface element for navigating through the video includes at a seventh location on the video navigation user interface element, a sixth graphical user interface object indicating that the user-specified change occurred at the second time within the video, the sixth graphical user interface object is displayed in a selectable state, the method while displaying the selectable user interface object for controlling the video editing mode together with the status indication indicating that the video editing mode is in the active state, detecting a gesture directed at the selectable user interface object for controlling the video editing mode; in response to detecting the gesture directed at the selectable user interface object for controlling the video editing mode, withholding the display of the sixth graphical user interface object in the selectable state; The method according to any one of claims 93 to 117. **Claim 119** Before detecting the gesture directed at the selectable user interface object for controlling the video editing mode, the video navigation user interface element for navigating through the video is displayed with a first amount of visual emphasis, and the method further in response to detecting the gesture directed at the selectable user interface object for controlling the video editing mode, displaying the video navigation user interface element for controlling the video editing mode with a second amount of visual emphasis less than the first amount of visual emphasis. The method according to claim 118. **Claim 120** One or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 93 to 119. A non-transitory computer-readable storage medium storing the program. **Claim 121** A computer system configured to communicate with a display generation component, one or more processors; one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 93 to 119. A computer system comprising a memory storing the program. **Claim 122** A computer system configured to communicate with a display generation component, means for executing the method according to any one of claims 93 to 119 A computer system comprising **Claim 123** One or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs comprising instructions for executing the method according to any one of claims 93 to 119, A computer program product. **Claim 124** A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs comprising: Via the display generation component, a user interface, A video having a first duration, a representation of the video including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject for one or more elements in the video, the plurality of changes including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, A representation of the video and, Instructions for displaying a user interface including simultaneously displaying a video navigation user interface element for navigating through the video including the representation of the first time and the representation of the second time, The representation of the second time is visually distinguishable from other times in the first duration of the video that do not correspond to the change in subject emphasis, The representation of the first time is visually distinguishable from the representation of the second time, a non-transitory computer-readable storage medium. **Claim 125** A computer system configured to communicate with a display generation component, One or more processors and, One or more programs configured to be executed by the one or more processors, Via the display generation component, a user interface, A video having a first duration, the representation of the video including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements in the video, the plurality of changes including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, the representation of the video, A memory storing one or more programs including instructions for displaying a user interface including simultaneously displaying a video navigation user interface element for navigating through the video including the representation of the first time and the representation of the second time, The representation of the second time is visually distinguishable from other times in the first duration of the video that do not correspond to the change in subject emphasis, A computer system in which the representation of the first time is visually distinguishable from the representation of the second time.

126. A computer system configured to communicate with a display generation component, A user interface via the display generation component, A user interface via the display generation component, A video having a first duration, the representation of the video including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements in the video, the plurality of changes including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, the representation of the video, Means for displaying a user interface including simultaneously displaying a video navigation user interface element for navigating through the video including the representation of the first time and the representation of the second time, The representation of the second time is visually distinguishable from other times in the first duration of the video that do not correspond to the change in subject emphasis, A computer system in which the representation of the first time is visually distinguishable from the representation of the second time. **Claim 127** One or more programs configured to be executed by one or more processors of a computer system that communicate with a display generation component, a user interface via the display generation component, a video having a first duration, the representation of the video including a plurality of changes in subject emphasis in the video, the changes in subject emphasis in the video including changes in the appearance of visual information captured by one or more cameras to emphasize one subject with respect to one or more elements in the video, the plurality of changes including an automatic change in subject emphasis at a first time during the first duration and a user-specified change in subject emphasis at a second time during the first duration different from the first time, a representation of the video, and a video navigation user interface element for navigating through the video including the representation of the first time and the representation of the second time, the one or more programs including instructions for displaying a user interface including simultaneously displaying. the representation of the second time is visually distinguishable from other times in the first duration of the video that do not correspond to changes in subject emphasis, A computer program product in which the representation of the first time is visually distinguishable from the representation of the second time. **Claim 128** A method, in a computer system that communicates with a display generation component and a plurality of cameras including a first camera having first image capture parameters determined by the hardware of the first camera and a second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, displaying, via the display generation component, a camera user interface that is a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view including the visual information collected by the first camera using the first image capture parameters. Detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the visual field using the visual information collected by the first camera; In response to detecting the decrease in the distance between the camera location and the focus location; Shifting from displaying the representation of the visual field using the visual information collected by the first camera to displaying the representation of the visual field using the visual information collected by the second camera according to a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance, the method comprising:

129. The method according to claim 128, wherein the predetermined threshold distance is based on the first image capture parameter.

130. Detecting a request to capture media while displaying the representation of the visual field using the visual information collected by the first camera; In response to detecting the request to capture media; Using second visual information collected by the first camera according to a determination that a current distance between the camera location and the focus location is closer than a second predetermined threshold distance; and Capturing media using second visual information collected by the second camera according to a determination that the current distance between the camera location and the focus location is not closer than the second predetermined threshold distance, the method according to claim 128 or 129 further comprising: The method according to claim 128 or 129, further comprising:

131. In response to detecting the decrease in the distance between the camera location and the focus location; Refraining from shifting from using the visual information collected by the first camera to display the representation of the visual field to using the visual information collected by the second camera according to a determination that the decreased distance between the camera location and the focus location is not closer than the predetermined threshold distance, the method according to any one of claims 128 to 130 further comprising: The method according to any one of claims 128 to 130, further comprising:

132. The method according to any one of claims 128 to 131, wherein the decrease in the distance between the camera location and the focus location is detected based on the movement of the computer system.

133. The method according to any one of claims 128 to 132, wherein the decrease in the distance between the camera location and the focus location is detected based on a new focus being selected.

134. Detecting an increase in the distance between the camera location and the focus location while displaying the representation of the visual field using the visual information collected by the second camera; In response to detecting the increase in the distance between the camera location and the focus location; According to a determination that the increased distance between the camera location and the focus location is not closer than a third predetermined threshold distance, shifting from displaying the representation of the visual field using the visual information collected by the second camera to displaying the representation of the visual field using the visual information collected by the first camera; The method according to any one of claims 128 to 133, further comprising.

135. The representation of the visual field is displayed at an effective zoom level before the decrease in the distance between the camera location and the focus location is detected; Shifting from displaying the representation of the visual field using the visual information collected by the first camera to displaying the representation of the visual field using the visual information collected by the second camera includes continuing to display the representation of the visual field at the effective zoom level; The method according to any one of claims 128 to 134.

136. Shifting from displaying the representation of the visual field using the visual information collected by the first camera to displaying the representation of the visual field using the visual information collected by the second camera includes changing the appearance of the representation of the visual field;

137. The first camera is disposed at a first position on the computer system; The second camera is disposed at a second position on the computer system; Displaying the representation of the visual field using the visual information collected by the first camera, and then transitioning to displaying the representation of the visual field using the visual information collected by the second camera, while reducing the alignment between the visual fields of the first camera and the second camera in one or more portions of the representation of the visual field that are further away from the predetermined portion, and increasing the alignment between the visual fields of the first camera and the second camera near the predetermined portion more than the amount of translation near the predetermined portion of the camera user interface, including displaying the representation of the visual field shifted as described above. The method according to any one of claims 128 to 136.

138. The plurality of cameras includes a third camera having third image capture parameters different from the first image capture parameters and the second image capture parameters determined by the hardware of the third camera, and the method further includes Before displaying the representation of the visual field using the visual information collected by the first camera together with the first image capture parameters, displaying the representation of the visual field using the visual information collected by the third camera together with the third image capture parameters; While displaying the representation of the visual field using the visual information collected by the third camera, detecting a second decrease in the distance between the camera location corresponding to at least one of the plurality of cameras and the focus location corresponding to the focus; In response to detecting the second decrease in the distance between the camera location and the focus location; According to the determination that the second decreased distance between the camera location and the focus location is closer than a fourth predetermined distance, transitioning from displaying the representation of the visual field using the visual information collected by the third camera to displaying the representation of the visual field using the visual information collected by the first camera, the method according to any one of claims 128 to 137.

139. According to the determination that the amount of light within the field of view of one or more of the plurality of cameras exceeds a threshold amount of light, the predetermined threshold distance is a first threshold distance, According to the determination that the amount of light within the field of view of one or more of the plurality of cameras does not exceed the threshold amount of light, the predetermined threshold distance is a second threshold distance different from the first threshold distance, The method according to any one of claims 128 to 138.

140. The method according to any one of claims 128 to 139, wherein the first camera has a first fixed focal length and the second camera has a second fixed focal length different from the first fixed focal length.

141. The first camera has a first minimum focal length, The second camera has a second minimum focal length, The first minimum focal length is longer than the second minimum focal length, The method according to any one of claims 128 to 140.

142. The first camera has a first minimum zoom level, The second camera has a second minimum zoom level, The first minimum zoom level is different from the second minimum zoom level, The method according to any one of claims 128 to 141.

143. The first camera has a first maximum zoom level, The second camera has a second maximum zoom level, The first maximum zoom level is different from the second maximum zoom level, The method according to any one of claims 128 to 142.

144. One or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component and a plurality of cameras including the first camera having first image capture parameters determined by the hardware of the first camera and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, the one or more programs storing instructions for performing the method according to any one of claims 128 to 143, a non-transitory computer-readable storage medium.

145. A computer system configured to communicate with a plurality of cameras including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, one or more processors, a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 128 to 143. A computer system comprising. **Claim 146** A computer system configured to communicate with a plurality of cameras including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, means for performing the method according to any one of claims 128 to 143 A computer system comprising. **Claim 147** One or more programs configured to be executed by one or more processors of a computer system including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, the one or more programs including instructions for performing the method according to any one of claims 128 to 143. A computer program product comprising. **Claim 148** One or more programs configured to be executed by one or more processors of a computer system that communicates with a plurality of cameras including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, wherein: displaying, via the display generation component, a camera user interface that is a representation of the field of view of one or more of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; in response to detecting the decrease in the distance between the camera location and the focus location; in accordance with a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance, transitioning from displaying the representation of the field of view using the visual information collected by the first camera to displaying the representation of the field of view using the visual information collected by the second camera. A non-transitory computer-readable storage medium storing one or more programs including instructions for performing the above. **Claim 149** A computer system configured to communicate with a plurality of cameras including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, the computer system comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; wherein the one or more programs are: displaying, via the display generation component, a camera user interface including a representation of one or more fields of view of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; detecting a decrease in a distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; in response to detecting the decrease in the distance between the camera location and the focus location; in accordance with a determination that the decreased distance between the camera location and the focus location is closer than a predetermined threshold distance, transitioning from displaying the representation of the field of view using the visual information collected by the first camera to displaying the representation of the field of view using the visual information collected by the second camera; a computer system comprising instructions for performing the foregoing.

150. A computer system configured to communicate with a display generation component, a first camera having first image capture parameters determined by hardware of the first camera, and a second camera having second image capture parameters different from the first image capture parameters determined by hardware of the second camera, the computer system comprising: means for displaying, via the display generation component, a camera user interface including a representation of one or more fields of view of the plurality of cameras, the representation of the field of view being displayed using visual information collected by the first camera together with the first image capture parameters; means for detecting a decrease in a distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the field of view using the visual information collected by the first camera; in response to detecting the decrease in the distance between the camera location and the focus location; Means for transitioning from displaying the representation of the visual field using the visual information collected by the first camera in accordance with a determination that the reduced distance between the camera location and the focus location is closer than a predetermined threshold distance to displaying the representation of the visual field using the visual information collected by the second camera, a computer system comprising the same. **Claim 151** One or more programs configured to be executed by one or more processors of a computer system that communicates with a plurality of cameras including a display generation component, the first camera having first image capture parameters determined by the hardware of the first camera, and the second camera having second image capture parameters different from the first image capture parameters determined by the hardware of the second camera, Displaying, via the display generation component, a camera user interface that is a representation of the visual field of one or more of the plurality of cameras, the representation of the visual field including visual information collected by the first camera along with the first image capture parameters; Detecting a decrease in the distance between a camera location corresponding to at least one of the plurality of cameras and a focus location corresponding to a focus while displaying the representation of the visual field using the visual information collected by the first camera; In response to detecting the decrease in the distance between the camera location and the focus location; One or more programs comprising instructions for making a determination that the reduced distance between the camera location and the focus location is closer than a predetermined threshold distance, and transitioning from displaying the representation of the visual field using the visual information collected by the first camera to displaying the representation of the visual field using the visual information collected by the second camera. A computer program product. **Claim 152** A method, In a computer system communicating with a display generation component Play a part of the video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize individual subjects for one or more elements within the video during a first period following the first time, via the display generation component. After playing the portion of the video that includes the first subject emphasis change occurring at the first time, detect a request to change subject emphasis at a second time within the video that is different from the first time. In response to detecting the request to change subject emphasis at the second time within the video. Change the subject emphasis within the video during a second period following the second time. Changing the first subject emphasis change occurring at the first time, including changing the emphasis of the individual subjects for the one or more elements within the video during the first period following the first time.

153. Before detecting the request to change subject emphasis at the second time, the video includes a second subject emphasis change that occurred at the second time. Changing the subject emphasis within the video during the second period following the second time includes removing the second subject emphasis change that occurred at the second time. The method according to claim 152.

154. Before detecting the request to change subject emphasis at the second time, display a first graphical user interface object indicating that the second subject emphasis that occurred at the second time is changed, and detecting the request to change subject emphasis that occurred at the second time. While displaying the first graphical user interface object, detect an input directed to the first graphical user interface object. In response to detecting the input directed to the first graphical user interface object, display an option to remove the second subject emphasis change that occurred at the second time. Detecting an input directed to the option for removing the emphasis change of the second subject that occurred at the second time while displaying the option for removing the emphasis change of the second subject that occurred at the second time, and displaying, including: In response to detecting the input directed to the option for removing the emphasis change of the second subject that occurred at the second time, changing the subject emphasis in the video during the second period following the second time by removing the emphasis change of the second subject that occurred at the second time. The method according to claim 153, further comprising:

155. Before detecting the input directed to the first graphical user interface object, the first graphical user interface object is displayed simultaneously with a video navigation user interface element having a first amount of visual emphasis. In response to detecting the input directed to the first graphical user interface object, the option for removing the emphasis change of the second subject that occurred at the second time is displayed simultaneously with the video navigation user interface element having a second amount of visual emphasis that is less than the first amount of visual emphasis. The method according to claim 154.

156. Before detecting the request to change the subject emphasis at the second time, the video does not include the change in subject emphasis that occurred at the second time. Changing the subject emphasis in the video during the second period following the second time includes adding a third subject emphasis change that occurred at the second time. The method according to any one of claims 152 to 155.

157. Detecting the request to change the subject emphasis that occurred at the second time includes detecting a first type of input directed to a first representation of the video, the first type of input being a first input for selecting a first fixed focal plane in the video. Changing the subject emphasis in the video during the second period following the second time includes applying a synthetic depth of field effect to the first fixed focal plane in the first plurality of frames of the video corresponding to the second period. The method according to any one of claims 152 to 156.

158. Detecting the request to change the subject emphasis that occurred at the second time includes detecting a second type of input directed to a second representation of the video, the second type of input being for selecting a first subject to be focused on within the video, Changing the subject emphasis within the video during the second period following the second time includes applying a synthetic depth of field effect to emphasize the first subject with respect to a second subject within a second plurality of frames of the video corresponding to the second period, The method according to any one of claims 152 to 157.

159. Detecting the request to change the subject emphasis that occurred at the second time includes detecting a third type of input directed to a third representation of the video, the third type of input being a second input for selecting a second fixed focal plane within the video, The method is Further including, in response to detecting the request to change the subject emphasis at the second time within the video, displaying an indication of the distance to the second fixed focal plane. The method according to any one of claims 152 to 158.

160. The first subject emphasis change that occurred at the first time is a first type of subject emphasis change, Changing the first subject emphasis change that occurred at the first time includes adding a fourth subject emphasis change that is a fourth type of subject emphasis change different from the first type of subject emphasis change at the first time. The method according to any one of claims 152 to 159.

161. The method according to any one of claims 152 to 160, wherein the first time corresponds to a first subset of the video in which an emphasized subject that was visible in a second portion of the video preceding the first time becomes invisible.

162. Changing the first subject emphasis change that occurred at the first time includes removing the first subject emphasis change that occurred at the first time. The method according to any one of claims 152 to 159 and 161.

163. The method according to any one of claims 152 to 162, wherein the first subject emphasis change that occurred at the first time is an automatic change in subject emphasis.

164. Before detecting the request to change the subject emphasis at the second time in the video that is different from the first time, the video includes a fifth subject emphasis change that occurs at a third time, The method is, In response to detecting the request to change the subject emphasis at the second time in the video, a set of emphasis change criteria, including criteria that are satisfied when the fifth subject emphasis change made at the third time is a user-specified change in subject emphasis, is satisfied. Further including refraining from changing the fifth subject emphasis change that occurred at the third time according to the determination, The method according to any one of claims 152 to 163.

165. The method according to any one of claims 152 to 164, wherein the second time occurs after the first time in the video.

166. The method according to any one of claims 152 to 164, wherein the second time occurs before the first time in the video.

167. The video includes a fifth subject emphasis change that occurs at a fourth time, The method is, Displaying a first selectable user interface object; Detecting a first input directed to the first selectable user interface object while the first selectable user interface object is being displayed and while the video includes the fifth subject emphasis change that occurs at the fourth time; In response to detecting the first input directed to the first selectable user interface object and according to the determination that the fifth subject emphasis change that occurs at the fourth time is a user-specified change in subject emphasis, removing the fifth subject emphasis change that occurs at the fourth time from the video. Further including, The method according to any one of claims 152 to 166.

168. In response to detecting the input directed to the first selectable user interface object and according to the determination that the fifth subject emphasis change that occurs at the fourth time is an automatic change in subject emphasis, refraining from removing the change in the fifth subject emphasis that occurs at the fourth time from the video. Further including, the method according to claim 167.

169. While displaying the first selectable user interface object and while the fifth subject emphasis change occurring at the fourth time is removed from the video, detecting a second input directed to the first selectable user interface object; In response to detecting the second input directed to the first selectable user interface object, adding the fifth subject emphasis change occurring at the fourth time to the video; The method according to claim 167 or 168, further comprising.

170. While the fifth subject emphasis change occurring at the fourth time is removed from the video and the first selectable user interface object is displayed in an inactive state, detecting a request to add one or more user-specified changes in subject emphasis; In response to detecting the request to add one or more user-specified changes in subject emphasis, displaying the first selectable user interface object in an active state different from the inactive state without adding the fifth subject emphasis change occurring at the fourth time to the video; The method according to any one of claims 167 to 169, further comprising.

171. While the video includes the first subject emphasis change occurring at the first time, Displaying a second graphical user interface object that indicates the first subject emphasis change occurring at the first time with a first visual appearance according to a determination that the first subject emphasis change is a user-specified change in subject emphasis; Displaying the second graphical user interface object having a second visual appearance different from the first visual appearance according to a determination that the first subject emphasis change is an automatic change in subject emphasis; The method according to any one of claims 152 to 170, further comprising.

172. The subject emphasis at the second time in the video is a third type of subject emphasis, The method is, After playing the portion of the video that includes the first subject emphasis change at the first time, detecting a second request to change the subject emphasis at the second time; In response to detecting the second request to change subject emphasis at the second time, and in accordance with the determination that the second request to change subject emphasis at the second time is a request to change the subject emphasis in the video at the second time to the third type of subject emphasis, refraining from changing the subject emphasis in the video during the second period following the second time. The method according to any one of claims 152 to 171. **Claim 173** One or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs storing instructions for executing the method according to any one of claims 152 to 172. **Claim 174** A computer system configured to communicate with a display generation component, One or more processors; A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for executing the method according to any one of claims 152 to 172. **Claim 175** A computer system configured to communicate with a display generation component, Means for executing the method according to any one of claims 152 to 172 Comprising a computer system. **Claim 176** One or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs comprising instructions for executing the method according to any one of claims 152 to 172. **Claim 177** A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs being Play a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize individual subjects for one or more elements within the video during a first period following the first time, via the display generation component; After playing the portion of the video that includes the first subject emphasis change occurring at the first time, detect a request to change subject emphasis at a second time within the video that is different from the first time; In response to detecting the request to change subject emphasis at the second time within the video, Change the subject emphasis within the video during a second period following the second time; Change the first subject emphasis change occurring at the first time, including changing the emphasis of the individual subjects for the one or more elements within the video during the first period following the first time. A non-transitory computer-readable storage medium including instructions for performing the above.

178. A computer system configured to communicate with a display generation component, comprising: One or more processors; A memory storing one or more programs configured to be executed by the one or more processors, Wherein the one or more programs include: Play a portion of a video that includes a first subject emphasis change occurring at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize individual subjects for one or more elements within the video during a first period following the first time, via the display generation component; After playing the portion of the video that includes the first subject emphasis change occurring at the first time, detect a request to change subject emphasis at a second time within the video that is different from the first time; In response to detecting the request to change subject emphasis at the second time within the video, Change the subject emphasis within the video during a second period following the second time; Changing the first subject emphasis change that occurs at the first time, including changing the emphasis of the individual subject on the one or more elements in the video during the first period following the first time. A computer system including instructions for performing the above.

179. A computer system configured to communicate with a display generation component and one or more input devices, Means for playing a portion of a video including a first subject emphasis change that occurs at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; Means for detecting a request to change subject emphasis at a second time in the video that is different from the first time, after playing the portion of the video including the first subject emphasis change that occurs at the first time; In response to detecting the request to change subject emphasis at the second time in the video, Changing the subject emphasis in the video during a second period following the second time; Means for changing the first subject emphasis change that occurs at the first time, including changing the emphasis of the individual subject on the one or more elements in the video during the first period following the first time. A computer system comprising the above.

180. One or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, Playing a portion of a video including a first subject emphasis change that occurs at a first time, the first subject emphasis change including a change in the appearance of visual information captured by one or more cameras to emphasize an individual subject for one or more elements in the video during a first period following the first time; Detecting a request to change subject emphasis at a second time in the video that is different from the first time, after playing the portion of the video including the first subject emphasis change that occurs at the first time; In response to detecting the request to change subject emphasis at the second time in the video, changing the subject emphasis in the video during a second period following the second time; changing the first subject emphasis change occurring at the first time, including changing the emphasis of the individual subject on the one or more elements in the video during the first period following the first time. A computer program product comprising one or more programs including instructions for performing the above.

Citation Information

Patent Citations

  • Camera module, camera, camera control method and control program

    JP2014106274A

  • Dual Aperture Zoom Digital Camera

    JP2016527734A

  • Lens controller and control method thereof

    JP2018036508A

  • Intelligent photography method and device, intelligent terminal

    JP2020504969A

  • Multiple lenses system, operation method and electronic device employing the same

    US20190109979A1