DEVICE, METHOD, AND GRAPHICAL USER INTERFACE FOR DEPTH-BASED ANNOTATION - Patent application

The computer system addresses inefficiencies in augmenting media by using depth data for intuitive annotation and virtual object placement, improving user interaction and conserving power in battery-operated devices.

JP7774679B2Active Publication Date: 2025-11-21APPLE INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024113573
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-17
Filing Date
2024-07-16
Publication Date
2025-11-21
Estimated Expiration
2039-09-20

AI Technical Summary

Technical Problem

Existing methods for augmenting media with virtual objects and annotations are cumbersome, inefficient, and place a cognitive burden on users, particularly in battery-operated devices, due to the need for multiple inputs and separate provision of augmentations, which wastes energy and reduces user experience.

Method used

A computer system with improved methods and interfaces that allow for intuitive annotation and virtual object placement using depth data to maintain spatial relationships, enabling efficient interaction through touch-sensitive surfaces and cameras, and supporting shared annotation sessions.

Benefits of technology

Enhances user interaction by reducing the number of user inputs, conserving power, and extending battery life, while allowing for seamless annotation and virtual object placement in various contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007774679000001
    Figure 0007774679000001
  • Figure 0007774679000002
    Figure 0007774679000002
  • Figure 0007774679000003
    Figure 0007774679000003
Patent Text Reader

Abstract

To provide an electronic device for displaying an annotation in a spatial position in an image corresponding to a spatial position in a physical environment captured in the image.SOLUTION: A computer system displays a representation of a field of view of at least one camera that is updated with changes in the field of view. In response to a request to add an annotation, the representation of the field of view of the camera is replaced with a still image of the field of view of the camera. An annotation is received on a portion of the still image that corresponds to a portion of a physical environment captured in the still image. The still image is replaced with the representation of the field of view of the camera. An indication of a current spatial relationship of the camera relative to the portion of the physical environment is displayed or not displayed based on a determination of whether the portion of the physical environment captured in the still image is currently within the field of view of the camera.SELECTED DRAWING: Figure 5G
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to electronic devices that display images of physical environments, including, but not limited to, electronic devices that display annotations at spatial locations within the images that correspond to spatial locations within the physical environment captured in the images. [Background technology]

[0002] The development of computer systems for augmented media has progressed significantly in recent years. Examples of augmented media include augmented reality environments, which include at least some virtual elements that replace or augment the physical world, and augmented stored media, which include at least some virtual elements that replace or augment stored media, such as image and video content. Input devices, such as touch-sensitive surfaces of computer systems and other electronic computing devices, are used to augment the media. Exemplary touch-sensitive surfaces include touchpads, touch-sensitive remote controls, and touchscreen displays. These surfaces are used to manipulate user interfaces and objects therein on the display. Exemplary user interface objects include digital images, video, text, icons, and control elements (such as buttons and other graphics).

[0003] However, methods and interfaces for augmenting media are cumbersome, inefficient, and limited. For example, augmentations, such as user-input annotations that have a fixed spatial location relative to a portion of the physical environment, can be difficult for a user to locate if the user's device's current camera view does not correspond to that portion of the physical environment. Searching for the augmentations places a significant cognitive burden on the user, detracting from the experience of using the augmented media. Additionally, providing augmented input for stored media (e.g., previously captured video) is time-intensive when augmented input must be provided separately for various portions of the stored media. Additionally, these methods take longer than necessary, thereby wasting energy. The latter problem is particularly acute in battery-operated devices. Summary of the Invention

[0004] Therefore, there is a need for a computer system with improved methods and interfaces for enhancing media data. Such methods and interfaces optionally complement or replace conventional methods for enhancing media data. Such methods and interfaces reduce the number, range, and / or type of inputs from a user, creating a more efficient human-machine interface. For battery-operated devices, such methods and interfaces conserve power and extend the time between battery charges.

[0005] The above-mentioned deficiencies and other problems associated with interfaces for augmenting media data with virtual objects and / or annotation input are reduced or eliminated by the disclosed computer system. In some embodiments, the computer system comprises a desktop computer. In some embodiments, the computer system is portable (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system comprises a personal electronic device (e.g., a wearable electronic device such as a watch). In some embodiments, the computer system has (and / or is in communication with) a touchpad. In some embodiments, the computer system has (and / or is in communication with) a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI, in part, through contacts and gestures with a stylus and / or fingers on a touch-sensitive surface. In some embodiments, the functionality optionally includes game playing, image editing, drawing, presenting, word processing, spreadsheet creation, telephony, video conferencing, email, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note taking, and / or digital video playback, and executable instructions to perform those functionality are optionally contained on a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0006] According to some embodiments, a method is performed on a computer system having a display generation element, one or more input devices, and one or more cameras. The method includes displaying, via the display generation element, a first user interface region including a representation of the field of view of one or more cameras that is updated with changes in the field of view of the one or more cameras over time. The method further includes receiving, while displaying the first user interface region including the representation of the field of view of the one or more cameras, a first request to add an annotation to the displayed representation of the field of view of the one or more cameras, via the one or more input devices. In response to the first request to add an annotation to the displayed representation of the field of view of the one or more cameras, the method further includes replacing the display of the representation of the field of view of the one or more cameras in the first user interface region with a still image of the field of view of the one or more cameras captured at a time corresponding to receiving the first request to add the annotation. The method further includes receiving, while displaying the still image in the first user interface region, a first annotation on a first portion of the still image, the first portion of the still image corresponding to a first portion of the physical environment captured in the still image, via the one or more input devices. The method further includes receiving, via the one or more input devices, a first request to redisplay representations of the fields of view of the one or more cameras in the first user interface region while displaying the first annotation on the first portion of the still image in the first user interface region. The method further includes replacing the display of the still image with the representations of the fields of view of the one or more cameras in the first user interface region in response to receiving the first request to redisplay the representations of the fields of view of the one or more cameras in the first user interface region. The method further includes displaying, in accordance with a determination that the first portion of the physical environment captured in the still image is currently outside the fields of view of the one or more cameras, an indication of a current spatial relationship of the one or more cameras with respect to the first portion of the physical environment captured in the still image, and withholding display of the indication in accordance with a determination that the first portion of the physical environment captured in the still image is currently within the fields of view of the one or more cameras.

[0007] In some embodiments, a method is implemented in a computer system having a display generation element and one or more input devices. The method includes displaying, via the display generation element, a user interface on a display, the user interface including a video playback area. The method further includes receiving, via the one or more input devices, a request to add an annotation to the video playback while displaying a playback of a first portion of the video in the video playback area. The method further includes, in response to receiving the request to add the annotation, pausing playback of the video at a first location within the video and displaying a still image corresponding to the first paused location of the video. The method further includes, while displaying the still image, receiving, via the one or more input devices, an annotation on a first portion of the physical environment captured in the still image. After receiving the annotation, the method further includes displaying, within the video playback area, a second portion of the video corresponding to a second location in the video that is different from the first location in the video, wherein the first portion of the physical environment is captured in the second portion of the video and the annotation is displayed in the second portion of the video.

[0008] In some embodiments, a method is implemented in a computer system having a display generation element and one or more input devices. The method includes displaying, via the display generation element, a first previously captured media object including one or more first images, the first previously captured media object being stored with first depth data corresponding to a first physical environment captured in each of the one or more first images. The method further includes receiving a first user request via the one or more input devices while displaying the first previously captured media object and adding a first virtual object to the first previously captured media object. In response to the first user request to add the first virtual object to the first previously captured media object, the method further includes displaying a first virtual object over at least a portion of each image in the first previously captured media object, the first virtual object being displayed at at least a first position or orientation determined using the first depth data corresponding to each image in the first previously captured media object.

[0009] According to some embodiments, a method is performed on a computer system having a display generating element, a first set of one or more input devices, and a first set of one or more cameras. The method includes sending a request to a remote device to initiate a shared annotation session with a second device including a second display generating element, a second set of one or more input devices, and a second set of one or more cameras. The method further includes receiving, in response to sending the request to initiate the shared annotation session with the second device, an indication of acceptance of the request to initiate the shared annotation session. The method further includes, in response to receiving the indication of acceptance of the request to initiate the shared annotation session, displaying, via the first display generating element, a first prompt that causes the first device to move toward the second device. The method further includes, after displaying the first prompt, displaying a representation of the field of view of the first camera set in a shared annotation session with the second device in accordance with a determination that a connectivity criterion for the first device and the second device is satisfied, the connectivity criterion requiring that at least a portion of the field of view of the first device and a portion of the field of view of the second device correspond to the same portion of a physical environment surrounding the first and second devices. The method further includes, during the shared annotation session, displaying one or more first virtual annotations corresponding to annotation input by the first device targeted to the respective positions in the physical environment via the first display generation element if the respective positions are included within the field of view of the first camera set, and displaying one or more second virtual annotations corresponding to annotation input by the second device targeted to the respective positions in the physical environment via the first display generation element.

[0010] According to some embodiments, an electronic device includes a display generating element, optionally one or more input devices, optional one or more touch-sensitive surface, optional one or more cameras, optional one or more sensors for detecting intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, optional one or more posture sensors for detecting changes in posture, one or more processors, and a memory having stored thereon one or more programs, the one or more programs configured to be executed by the one or more processors, and the one or more programs including instructions for performing or causing to be performed any of the operations of the methods described herein. According to some embodiments, a computer-readable storage medium has stored thereon instructions that, when executed by an electronic device comprising a display generating element, optionally one or more input devices, optional one or more touch-sensitive surface, optional one or more cameras, optional one or more sensors that detect intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, and optional one or more posture sensors, cause the device to perform or cause the performance of any of the operations of the methods described herein. According to some embodiments, a graphical user interface of an electronic device comprising a display generating element, optionally one or more input devices, optionally one or more touch-sensitive surface, optional one or more cameras, one or more sensors for detecting intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, optional one or more posture sensors, memory, and one or more processors executing one or more programs stored in the memory, includes one or more of the elements displayed in any of the methods described herein being updated to include input as described in any of the methods described herein.According to some embodiments, an electronic device includes a display generating element, optionally one or more input devices, optional one or more touch-sensitive surface, optional one or more cameras, one or more sensors that detect the intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, optional one or more posture sensors that detect changes in posture, and means for performing or causing the performance of any of the operations of the methods described herein. According to some embodiments, an information processing apparatus for use in an electronic device including a display generating element, optional one or more input devices, optional one or more touch-sensitive surface, optional one or more cameras, one or more sensors that detect the intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, and optional one or more posture sensors that detect changes in posture includes means for performing or causing the performance of any of the operations of the methods described herein.

[0011] Thus, electronic devices comprising a display generating element, optionally one or more input devices, optional one or more touch-sensitive surfaces, optional one or more cameras, optional one or more sensors for detecting the intensity of contact with the touch-sensitive surface, optional one or more audio output generators, optional one or more device orientation sensors, optional one or more tactile output generators, and optional one or more posture sensors are provided with improved methods and interfaces for displaying virtual objects in a variety of contexts, thereby increasing the effectiveness, efficiency, and user satisfaction of these devices. These methods and interfaces can complement or replace conventional methods for displaying virtual objects in a variety of contexts. [Brief explanation of the drawings]

[0012] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:

[0013] [Figure 1A] FIG. 1 is a block diagram illustrating a portable multifunction device with a touch-sensitive display in accordance with some embodiments.

[0014] [Figure 1B] FIG. 1 is a block diagram illustrating example components for event handling, according to some embodiments.

[0015] [Figure 1C] FIG. 1 is a block diagram illustrating a tactile output module, according to some embodiments.

[0016] [Figure 2] FIG. 1 illustrates a portable multifunction device with a touch screen according to some embodiments.

[0017] [Figure 3] FIG. 1 is a block diagram illustrating an exemplary multifunction device with a display and a touch-sensitive surface in accordance with some embodiments.

[0018] [Figure 4A] 1 illustrates an exemplary user interface for a menu of applications on a portable multifunction device, according to some embodiments.

[0019] [Figure 4B] 1A-1C illustrate exemplary user interfaces for a multifunction device with a touch-sensitive surface separate from a display in accordance with some embodiments.

[0020] [Figure 4C] FIG. 10 illustrates an example of a dynamic intensity threshold, according to some embodiments. [Figure 4D] FIG. 10 illustrates an example of a dynamic intensity threshold, according to some embodiments. [Figure 4E] FIG. 10 illustrates an example of a dynamic intensity threshold, according to some embodiments.

[0021] [Figure 5A] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5B] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5C] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5D] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5E] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5F] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5G] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5H] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5I] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5J] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5K] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5L] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5M]10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5N] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5O] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5P] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5Q] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5R] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5S] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5T] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5U] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5V] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5W] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5X] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5Y] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5Z] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AA]10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AB] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AC] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AD] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AE] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments. [Figure 5AF] 10 illustrates an example of a user interface for relocating an annotation according to some embodiments.

[0022] [Figure 6A] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6B] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6C] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6D] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6E] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6F]1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6G] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6H] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6I] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6J] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6K] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6L] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6M] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 6N] 1 illustrates an exemplary user interface for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments.

[0023] [Figure 7A]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7B] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7C] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7D] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7E] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7F] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7G] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7H] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7I] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7J] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7K]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7L] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7M] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7N] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7O] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7P] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7Q] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7R] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7S] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7T] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7U]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7V] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7W] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7X] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7Y] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7Z] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AA] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AB] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AC] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AD] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AE]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AF] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AG] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AH] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AI] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AJ] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AK] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AL] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AM] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AN] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AO]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AP] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AQ] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AR] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AS] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AT] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AU] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AV] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AW] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AX] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AY]1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7AZ] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BA] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BB] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BC] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BD] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BE] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments. [Figure 7BF] 1 illustrates an exemplary user interface for adding a virtual object to a previously captured media object according to some embodiments.

[0024] [Figure 8A] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8B] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8C]1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8D] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8E] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8F] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8G] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8H] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8I] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8J] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8K] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8L] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8M]1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8N] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8O] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8P] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8Q] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8R] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8S] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8T] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8U] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8V] 1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments. [Figure 8W]1 illustrates an exemplary user interface for illustrating an exemplary user interface for initiating a shared annotation session according to some embodiments.

[0025] [Figure 9A] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments. [Figure 9B] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments. [Figure 9C] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments. [Figure 9D] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments. [Figure 9E] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments. [Figure 9F] FIG. 1 is a flow diagram of a process for relocating an annotation according to some embodiments.

[0026] [Figure 10A] FIG. 1 is a flow diagram of a process for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. [Figure 10B] FIG. 1 is a flow diagram of a process for receiving annotations on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments.

[0027] [Figure 11A] FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments. [Figure 11B] FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments. [Figure 11C]FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments. [Figure 11D] FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments. [Figure 11E] FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments. [Figure 11F] FIG. 1 is a flow diagram of a process for adding a virtual object to a previously captured media object according to some embodiments.

[0028] [Figure 12A] FIG. 1 is a flow diagram of a process for initiating a shared annotation session according to some embodiments. [Figure 12B] FIG. 1 is a flow diagram of a process for initiating a shared annotation session according to some embodiments. [Figure 12C] FIG. 1 is a flow diagram of a process for initiating a shared annotation session according to some embodiments. [Figure 12D] FIG. 1 is a flow diagram of a process for initiating a shared annotation session according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0029] Conventional methods of augmenting media often require multiple separate inputs (e.g., individual annotations of multiple frames and / or placement of augmentations relative to objects in the media) to achieve an intended result (e.g., annotating a portion of a stored video or live video feed and / or displaying a virtual object at a location corresponding to the surface of a physical object in the stored media). Embodiments herein provide an intuitive way for users to augment media, such as stored content, still images, and / or live video captured by one or more cameras of a device (e.g., using depth data stored and / or captured along with the image data to position the augmentations and maintain a constant spatial relationship between the augmentations and portions of the physical environment within the camera's field of view).

[0030] The systems, methods, and GUIs described herein improve user interface interaction with augmented media in several ways: for example, they make it easier to rearrange annotations, annotate video, add virtual objects to previously captured media, and initiate shared annotation sessions.

[0031] Below, Figures 1A-1C, 2, and 3 provide a description of an exemplary device. Figures 4A-4B, 5A-5AF, 6A-6N, 7A-7BF, and 8A-8W show exemplary user interfaces displaying virtual objects in various contexts. Figures 9A-9F show a process for repositioning annotations. Figures 10A-10B show a process for receiving annotations on a portion of the physical environment captured in a still image corresponding to a paused position in a video. Figures 11A-11F show a process for adding a virtual object to a previously captured media object. Figures 12A-12D show a process for initiating a shared annotation session. The user interfaces of Figures 5A-5AF, 6A-6N, 7A-7BF, and 8A-8W are used to illustrate the processes of Figures 9A-9F, 10A-10B, 11A-11F, and 12A-12D. Exemplary Devices

[0032] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments being described. However, it will be apparent to those skilled in the art that the various embodiments described may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0033] In this specification, terms such as "first," "second," etc. are used to describe various elements in some examples, but it will be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first contact can be referred to as a second contact, and similarly, a second contact can be referred to as a first contact, without departing from the scope of the various embodiments described. Although a first contact and a second contact are both contacts, they are not the same contact unless the context clearly dictates otherwise.

[0034] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. In the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0035] As used herein, the term "if" is optionally interpreted to mean "when," "upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" are optionally interpreted to mean "upon determining" or "in response to determining," or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]," depending on the context.

[0036] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Other portable electronic devices, such as laptops or tablet computers having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad), are optionally used. It should also be understood that in some embodiments, the device is not a portable communication device, but rather a desktop computer having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad).

[0037] In the following discussion, electronic devices are described that include a display and a touch-sensitive surface. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick.

[0038] The device typically supports a variety of applications such as one or more of a note-taking application, a drawing application, a presentation application, a word processing application, a website creation application, a disc authoring application, a spreadsheet application, a gaming application, a telephony application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0039] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more features of the touch-sensitive surface and corresponding information displayed on the device are optionally adjusted and / or changed for each application and / or within each application. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally supports various applications with user interfaces that are intuitive and transparent to the user.

[0040] Attention now turns to embodiments of portable devices with touch-sensitive displays. FIG. 1A is a block diagram illustrating portable multifunction device 100 having touch-sensitive display system 112, according to some embodiments. Touch-sensitive display system 112 may conveniently be referred to as a "touch screen" or simply a touch-sensitive display. Device 100 includes memory 102 (optionally including one or more computer-readable storage media), memory controller 122, one or more processing units (CPUs) 120, peripherals interface 118, RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, input / output (I / O) subsystem 106, other input or control devices 116, and external port 124. Device 100 optionally includes one or more light sensors 164. Device 100 optionally includes one or more intensity sensors 165 that detect the intensity of a contact on device 100 (e.g., a touch-sensitive surface, such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 that generate tactile output on device 100 (e.g., generate tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0041] It should be understood that device 100 is only one example of a portable multifunction device, and that device 100 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of its components. The various components shown in FIG. 1A are implemented in hardware, software, firmware, or a combination thereof, including one or more signal processing circuits and / or application specific integrated circuits.

[0042] Memory 102 optionally includes high-speed random access memory, and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Access to memory 102 by other components of device 100, such as CPU(s) 120 and peripherals interface 118, is optionally controlled by memory controller 122.

[0043] A peripheral interface 118 may be used to couple input and output peripherals of the device with the CPU(s) 120 and memory 102. The one or more processors 120 operate or execute various software programs and / or instruction sets stored in memory 102 to perform various functions and process data for the device 100.

[0044] In some embodiments, peripheral interface 118, CPU(s) 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0045] RF (radio frequency) circuitry 108 transmits and receives RF signals, also called electromagnetic signals. RF circuitry 108 converts electrical signals to electromagnetic signals and communicates with communication networks and other communication devices via electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. RF circuitry 108 optionally communicates via wireless communication with networks, such as the Internet, also known as the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and with other devices. Radio options include Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), long term evolution (LTE), near field communication (NFC), wideband code division multiple access (W-CDMA), and code division multiple access (CDMA).Wireless technologies include, but are not limited to, standard IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n, voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol (XMPP)), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), and Instant Messaging and Presence Services (IMP). The present invention may use any of a number of communication standards, protocols, and technologies, including, but not limited to, Intermediate Message Service (IMPS), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this document.

[0046] Audio circuit 110, speaker 111, and microphone 113 provide an audio interface between a user and device 100. Audio circuit 110 receives audio data from peripherals interface 118, converts the audio data into electrical signals, and transmits the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves audible to humans. Audio circuit 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuit 110 converts the electrical signals into audio data and transmits the audio data to peripherals interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to memory 102 and / or RF circuit 108 by peripherals interface 118. In some embodiments, audio circuit 110 also includes a headset jack (e.g., 212, FIG. 2 ). The headset jack provides an interface between audio circuitry 110 and a detachable audio input / output peripheral, such as an output-only headphone or a headset with both an output (e.g., mono or binaural headphones) and an input (e.g., a microphone).

[0047] I / O subsystem 106 couples input / output peripherals on device 100, such as touch-sensitive display system 112 and other input or control devices 116, with peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, light sensor controller 158, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. One or more input controllers 160 receive electrical signals from and send electrical signals to other input or control devices 116. Other input or control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, input controller(s) 160 are optionally coupled to any (or none) of a keyboard, infrared port, USB port, stylus, and / or pointer device such as a mouse. The one or more buttons (e.g., 208, FIG. 2) optionally include up / down buttons for volume control of the speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., 206, FIG. 2).

[0048] Touch-sensitive display system 112 provides an input and output interface between the device and a user. Display controller 156 receives electrical signals from and / or sends electrical signals to touch-sensitive display system 112. Touch-sensitive display system 112 displays visual output to the user. This visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects. As used herein, the term "affordance" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object that is configured to respond to input directed towards the graphical user interface object). Examples of user-interactive graphical user interface objects include, but are not limited to, a button, a slider, an icon, a selectable menu item, a switch, a hyperlink, or other user interface control.

[0049] Touch-sensitive display system 112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touch-sensitive display system 112 and display controller 156 (along with any associated modules and / or instruction sets in memory 102) detect contacts (and any movement or disruption of contact) on touch-sensitive display system 112 and translate the detected contacts into interactions with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touch-sensitive display system 112. In some embodiments, the point of contact between touch-sensitive display system 112 and the user corresponds to the user's finger or stylus.

[0050] Touch-sensitive display system 112 optionally uses liquid crystal display (LCD), light emitting polymer display (LPD), or light emitting diode (LED) technology, although other display technologies are used in other embodiments. Touch-sensitive display system 112 and display controller 156 optionally use any of a number of now-known or later-developed touch-sensing technologies to detect contact and any movement or disruption thereof, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with touch-sensitive display system 112. In some embodiments, projected mutual capacitance sensing technology is used, such as that found in the iPhone®, iPod Touch®, and iPad® from Apple Inc. of Cupertino, California.

[0051] Touch-sensitive display system 112 optionally has a video resolution greater than 100 dpi. In some embodiments, the touchscreen video resolution exceeds 400 dpi (e.g., 500 dpi, 800 dpi, or higher). A user optionally contacts touch-sensitive display system 112 using any suitable object or accessory, such as a stylus, finger, or the like. In some embodiments, the user interface is designed to work with finger-based contacts and gestures, which may be less precise than stylus-based input due to the larger contact area of ​​a finger on a touchscreen than that of a stylus. In some embodiments, the device translates coarse finger input into precise pointer / cursor positions or commands to perform actions desired by the user.

[0052] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad (not shown) for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is optionally a touch-sensitive surface separate from touch-sensitive display system 112 or an extension of the touch-sensitive surface formed by the touchscreen.

[0053] Device 100 also includes a power system 162 that provides power to the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, power failure detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with generating, managing, and distributing electrical power within a portable device.

[0054] Device 100 also optionally includes one or more light sensors 164. FIG. 1A shows a light sensor coupled to light sensor controller 158 in I / O subsystem 106. Light sensor(s) 164 optionally include a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Light sensor(s) 164 receive light from the environment, projected through one or more lenses, and convert the light into data representing an image. In conjunction with imaging module 143 (also called a camera module), light sensor(s) 164 optionally capture still images and / or video. In some embodiments, the light sensor is located on the back of device 100, opposite touch-sensitive display system 112 on the front of the device, so that the touchscreen can be used as a viewfinder for still and / or video image acquisition. In some embodiments, another light sensor is placed on the front of the device so that an image of the user is captured (e.g., for a selfie, for a video conference while the user is viewing other video conference participants on the touchscreen, etc.).

[0055] Device 100 also optionally includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled with intensity sensor controller 159 in I / O subsystem 106. Contact intensity sensor(s) 165 optionally include one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor(s) 165 receive contact intensity information (e.g., pressure information or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with or proximate to the touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the back of device 100, opposite touchscreen display system 112, which is located on the front of device 100.

[0056] Device 100 also optionally includes one or more proximity sensors 166. Figure 1A shows proximity sensor 166 coupled with peripherals interface 118. Alternatively, proximity sensor 166 is coupled with input controller 160 in I / O subsystem 106. In some embodiments, the proximity sensor turns off and disables touch-sensitive display system 112 when the multifunction device is placed near a user's ear (e.g., when the user is making a phone call).

[0057] Device 100 also optionally includes one or more tactile output generators 167. FIG. 1A shows tactile output generators coupled to haptic feedback controller 161 in I / O subsystem 106. In some embodiments, tactile output generator(s) 167 include one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear movement, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile output on the device). Tactile output generator(s) 167 receive tactile feedback generation instructions from haptic feedback module 133 and generate tactile outputs on device 100 that can be sensed by a user of device 100. In some embodiments, at least one tactile output generator is juxtaposed with or proximate to a touch-sensitive surface (e.g., touch-sensitive display system 112) and generates tactile output, optionally by moving the touch-sensitive surface vertically (e.g., in / out of the surface of device 100) or horizontally (e.g., back and forth in the same plane as the surface of device 100). In some embodiments, at least one tactile output generator sensor is located on the back of device 100, opposite touch-sensitive display system 112, which is located on the front of device 100.

[0058] Device 100 also optionally includes one or more accelerometers 168. FIG. 1A shows accelerometer 168 coupled to peripherals interface 118. Alternatively, accelerometer 168 is optionally coupled to input controller 160 in I / O subsystem 106. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from the one or more accelerometers. In addition to accelerometer(s) 168, device 100 optionally includes a magnetometer (not shown) and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the location and orientation (e.g., portrait or landscape) of device 100.

[0059] In some embodiments, software components stored in memory 102 include an operating system 126, a communications module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a haptic feedback module (or instruction set) 133, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and applications (or instruction sets) 136. Additionally, in some embodiments, as shown in Figures 1A and 3, memory 102 stores device / global internal state 157. Device / global internal state 157 includes one or more of: active application state, which indicates which applications, if any, are currently active; display state, which indicates what applications, views, or other information are occupying various areas of touch-sensitive display system 112; sensor state, which includes information obtained from the device's various sensors and other input or control devices 116; and position and / or location information regarding the device's position and / or orientation.

[0060] Operating system 126 (e.g., an embedded operating system such as iOS, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks) includes various software components and / or drivers for controlling and managing overall system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware and software components.

[0061] Communications module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by RF circuitry 108 and / or external port 124. External port 124 (e.g., Universal Serial Bus (USB), FIREWIRE®, etc.) is adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector identical to, similar to, and / or compatible with the 30-pin connector used in some iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. In some embodiments, the external port is a Lightning connector identical to, similar to, and / or compatible with the Lightning connector used in some iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California.

[0062] Contact / motion module 130 optionally detects contact with touch-sensitive display system 112 (in cooperation with display controller 156) and with other touch-sensitive devices (e.g., a touchpad or physical click wheel). Contact / motion module 130 includes software components for performing various operations related to detecting contact (e.g., by a finger or stylus), such as determining if contact has occurred (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining if there is movement of the contact and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining if the contact has ended (e.g., detecting a finger-up event or an interruption of the contact). Contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point, as represented by the series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These actions are optionally applied to a single contact (e.g., a single finger contact or a stylus contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, contact / motion module 130 and display controller 156 detect contacts on a touchpad.

[0063] Contact / motion module 130 optionally detects gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different movements, timing, and / or strength of the detected contacts). Thus, gestures are optionally detected by detecting particular contact patterns. For example, detecting a finger tap gesture includes detecting a finger down event, followed by detecting a finger up (lift-off) event at the same location (or substantially the same location) as the finger down event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger down event, followed by detecting one or more finger drag events, and then detecting a finger up (lift-off) event. Similarly, taps, swipes, drags, and other gestures are optionally detected with respect to a stylus by detecting particular contact patterns with respect to the stylus.

[0064] In some embodiments, detecting a finger tap gesture depends on the length of time between detecting a finger-down event and detecting a finger-up event, but is not related to the intensity of the finger contact between detecting the finger-down event and detecting the finger-up event. In some embodiments, a tap gesture is detected according to determining that the length of time between the finger-down event and the finger-up event is less than a predetermined value (e.g., less than 0.1, 0.2, 0.3, 0.4, or 0.5 seconds), regardless of whether the intensity of the finger contact during the tap meets a given intensity threshold (greater than a nominal contact-detection intensity threshold), such as a light or deep pressure intensity threshold. Thus, a finger tap gesture can satisfy certain input criteria that do not require the characteristic intensity of the contact to meet a given intensity threshold for the particular input criteria to be met. For clarity, finger contacts in a tap gesture generally need to meet a nominal contact-detection intensity threshold below which the contact is not detected in order to detect a finger-down event. A similar analysis applies to detecting tap gestures or other contacts with a stylus. In cases where the device is capable of detecting contact of a finger or stylus hovering over the touch-sensitive surface, the nominal contact-detection intensity threshold optionally does not correspond to physical contact between the finger or stylus and the touch-sensitive surface.

[0065] In a similar manner, the same concepts apply to other types of gestures. For example, swipe gestures, pinch gestures, de-pinch gestures, and / or long press gestures are optionally detected based on meeting criteria that are either unrelated to the intensity of the contacts included in the gesture or that do not require the contacts performing the gesture to reach an intensity threshold in order to be recognized. For example, swipe gestures are detected based on the amount of movement of one or more contacts, pinch gestures are detected based on the movement of two or more contacts toward each other, de-pinch gestures are detected based on the movement of two or more contacts away from each other, and long press gestures are detected based on the duration of contacts on the touch-sensitive surface that is less than a threshold amount of movement. Thus, a statement that a particular gesture recognition criterion does not require the intensity of a contact(s) to meet a corresponding intensity threshold in order for the particular gesture recognition criterion to be met means that the particular gesture recognition criterion can be met when the contact(s) in the gesture do not reach the corresponding intensity threshold, and can also be met in situations where one or more of the contacts in the gesture reach or exceed the corresponding intensity threshold. In some embodiments, a tap gesture is detected based on a determination that a finger-down event and a finger-up event are detected within a predetermined time period, regardless of whether the contacts are above or below the respective intensity thresholds during the predetermined time period, and a swipe gesture is detected based on a determination that a movement of the contact is greater than a predetermined magnitude, even if the contacts exceed the respective intensity thresholds at the end of the movement of the contacts. Even in implementations in which gesture detection is affected by the intensity of the contact performing the gesture (e.g., the device detects long presses more quickly when the intensity of the contact exceeds an intensity threshold, or the device is slow to detect tap inputs when the intensity of the contact is higher), detection of those gestures does not require the contact to reach a particular intensity threshold, as long as the criteria for recognizing the gesture can be met in situations in which the contact does not reach the particular intensity threshold (e.g., even if the amount of time required to recognize the gesture varies).

[0066] The contact intensity threshold, duration threshold, and movement threshold may, in some circumstances, be combined in various different combinations to create heuristics for distinguishing between two or more different gestures directed at the same input element or region, thereby enabling multiple different interactions with the same input element to provide a richer set of user interactions and responses. A statement that a particular set of gesture recognition criteria does not require the intensity of the contact(s) to meet the respective intensity threshold for that particular gesture recognition criterion to be met does not preclude simultaneously evaluating other intensity-dependent gesture recognition criteria to identify other gestures whose criteria are met when the gesture includes a contact having an intensity exceeding the respective intensity threshold. For example, in some circumstances, a first gesture recognition criterion for a first gesture that does not require the intensity of the contact(s) to meet the corresponding intensity threshold for the first gesture recognition criterion to be met competes with a second gesture recognition criterion for a second gesture that relies on the contact(s) reaching the corresponding intensity threshold. In such a competition, a gesture is optionally not recognized as satisfying the first gesture recognition criteria for the first gesture if the second gesture recognition criteria for the second gesture are satisfied first. For example, if the contact reaches the corresponding intensity threshold before moving a predetermined amount of movement, a deep press gesture is detected rather than a swipe gesture. Conversely, if the contact moves a predetermined amount of movement before reaching the corresponding intensity threshold, a swipe gesture is detected rather than a deep press gesture. Even in such a situation, the first gesture recognition criteria for the first gesture still do not require the intensity of the contact(s) to meet the corresponding intensity threshold for the first gesture recognition criteria to be satisfied, because if the contact remains below the corresponding intensity threshold until the end of the gesture (e.g., a swipe gesture with a contact that does not increase in intensity above the corresponding intensity threshold), the gesture would be recognized by the first gesture recognition criteria as a swipe gesture.In this way, certain gesture recognition criteria that do not require the intensity of the contact(s) to meet a corresponding intensity threshold for the particular gesture recognition criterion to be satisfied (A) still depend on the intensity of the contact with respect to the intensity threshold (e.g., for a tap gesture) in some circumstances, and / or (B) in some circumstances, the particular gesture recognition criterion (e.g., for a long press gesture) will not function if a competing set of intensity-dependent gesture recognition criteria (e.g., for a deep press gesture) recognizes an input as corresponding to an intensity-dependent gesture before the particular gesture recognition criterion recognizes the gesture corresponding to the input (e.g., for a long press gesture that competes with a deep press gesture for recognition).

[0067] Graphics module 132 includes various known software components for rendering and displaying graphics on touch-sensitive display system 112 or other displays, including components for modifying the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual characteristics) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (such as user interface objects including soft keys), digital images, video, and animations.

[0068] In some embodiments, graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 132 receives one or more codes specifying the graphics to be displayed, including coordinate data and other graphic characteristic data, as needed, from an application or the like, and then generates screen image data to output to display controller 156.

[0069] The haptic feedback module 133 includes various software components that generate instructions (e.g., instructions used by the haptic feedback controller 161) that use the tactile output generator(s) 167 to create tactile outputs at one or more locations on the device 100 in response to user interaction with the device 100.

[0070] Text input module 134 is optionally a component of graphics module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).

[0071] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for use in location-based calling, to the camera 143 as photo / video metadata, and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and maps / navigation widgets).

[0072] Application 136 optionally includes the following modules (or instruction sets), or a subset or superset thereof: • a contacts module 137 (sometimes called an address book or contact list); ●Telephone module 138, ●Videoconferencing module 139, ● an email client module 140; ● Instant messaging (IM) module 141; ●Training support module 142, a camera module 143 for still and / or video images; ● Image management module 144; ● Browser module 147, ● Calendar module 148, a widget module 149, optionally including one or more of a weather widget 149-1, a stock price widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widgets 149-6; a widget creation module 150 for creating user-created widgets 149-6; ● Search module 151, ● A video and music player module 152, optionally consisting of a video player module and a music player module; ● Memo module 153, Map module 154, and / or ●Online video module 155.

[0073] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA® enabled applications, encryption, digital rights management, voice recognition, and voice duplication.

[0074] Contacts module 137, along with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, includes executable instructions (e.g., stored in memory 102 or in the application internal state 192 of contacts module 137 in memory 370) for managing an address book or contact list, including adding name(s) to the address book, removing name(s) from the address book, associating phone number(s), email address(es), physical address(es), or other information with names, associating images with names, categorizing and sorting names, providing phone numbers and / or email addresses to initiate and / or facilitate communication by telephone 138, video conference 139, email 140, or IM 141, etc.

[0075] In cooperation with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, telephone module 138 includes executable instructions for entering a series of characters corresponding to a telephone number, accessing one or more telephone numbers in address book 137, modifying an entered telephone number, dialing each telephone number, conducting a conversation, and disconnecting or hanging up when the conversation is completed. As noted above, wireless communication optionally uses any of a number of communication standards, protocols, and technologies.

[0076] In cooperation with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, light sensor(s) 164, light sensor controller 158, contact module 130, graphics module 132, text input module 134, contact list 137, and telephone module 138, videoconferencing module 139 includes executable instructions for initiating, conducting, and terminating videoconferences between a user and one or more other participants according to the user's commands.

[0077] In cooperation with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, email client module 140 contains executable instructions for creating, sending, receiving, and managing emails in response to user commands. In cooperation with image management module 144, email client module 140 greatly facilitates the creation and sending of emails with still or video images captured by camera module 143.

[0078] In cooperation with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for entering a series of characters corresponding to an instant message, modifying previously entered characters, sending each instant message (e.g., using Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for telephone-based instant messaging, or using XMPP, SIMPLE, Apple Push Notification Service (APNs), or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, sent and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments, such as those supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both telephone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, APNs, or IMPS).

[0079] In cooperation with the RF circuitry 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, the GPS module 135, the map module 154, and the video and music player module 152, the training support module 142 includes executable instructions to create workouts (e.g., with time, distance, and / or calorie burn goals), communicate with training sensors (in the sports device and the smartwatch), receive training sensor data, calibrate sensors used to monitor workouts, select and play music for workouts, and display, store, and transmit workout data.

[0080] Camera module 143, along with touch-sensitive display system 112, display controller 156, light sensor(s) 164, light sensor controller 158, contact module 130, graphics module 132, and image management module 144, includes executable instructions to capture still images or video (including video streams) and store them in memory 102, change characteristics of the still images or video, and / or delete the still images or video from memory 102.

[0081] Image management module 144, along with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and camera module 143, includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slide show or album), and storing still and / or video images.

[0082] Browser module 147, along with RF circuitry 108, touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, and text input module 134, contains executable instructions for browsing the Internet according to user commands, including retrieving, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.

[0083] Calendar module 148, along with RF circuitry 108, touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, includes executable instructions that create, display, modify, and store calendars and data associated with calendars (e.g., calendar entries, to-do lists, etc.) according to user commands.

[0084] Widget modules 149, along with RF circuitry 108, touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, text input module 134, and browser module 147, are optionally mini-applications downloaded and used by users (e.g., weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5), or mini-applications created by users (e.g., user-created widget 149-6). In some embodiments, widgets include Hypertext Markup Language (HTML) files, Cascading Style Sheets (CSS) files, and JavaScript files. In some embodiments, widgets include Extensible Markup Language (XML) files and JavaScript files (e.g., Yahoo! Widgets).

[0085] In cooperation with RF circuitry 108, touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, text input module 134, and browser module 147, widget creation module 150 contains executable instructions for creating widgets (e.g., turning user-specified portions of a web page into widgets).

[0086] In cooperation with touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, and text input module 134, search module 151 includes executable instructions to search memory 102 for text, music, sound, images, video, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user commands.

[0087] In cooperation with touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that enable a user to download and play recorded music or other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions to display, present, or otherwise play videos (e.g., on touch-sensitive display system 112 or on an external display connected wirelessly or via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (a trademark of Apple Inc.).

[0088] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, notes module 153 contains executable instructions for creating and managing notes, to-do lists, and the like according to user commands.

[0089] In conjunction with RF circuitry 108, touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 may be used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data about stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.

[0090] In cooperation with touch-sensitive display system 112, display system controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, online video module 155 contains executable instructions that enable a user to access, view, receive (e.g., by streaming and / or downloading), and play (e.g., on touchscreen 112 or on an external display connected wirelessly or via external port 124) online videos in one or more file formats, such as H.264, and send and otherwise manage emails with links to particular online videos. In some embodiments, instant messaging module 141 is used to send links to particular online videos, rather than email client module 140.

[0091] Each of the above-identified modules and applications corresponds to executable instruction sets that perform one or more of the functions described above, as well as methods described in the present application (e.g., computer-implemented methods and other information processing methods described herein). The modules (i.e., instruction sets) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of the modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 102 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 102 optionally stores additional modules and data structures not described above.

[0092] In some embodiments, device 100 is a device in which operation of a predetermined set of functions on the device is performed solely via a touchscreen and / or touchpad. Using the touchscreen and / or touchpad as the primary input control device for operation of device 100 optionally reduces the number of physical input control devices (push buttons, dials, etc.) on device 100.

[0093] The set of default functions performed only through the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 to a main menu, home menu, or root menu from any user interface displayed on device 100. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touchpad.

[0094] 1B is a block diagram illustrating exemplary components for event processing, according to some embodiments. In some embodiments, memory 102 (in FIG. 1A) or 370 (in FIG. 3) includes event sorter 170 (e.g., within operating system 126) and a respective application 136-1 (e.g., any of applications 136, 137-155, 380-390 described above).

[0095] Event sorter 170 receives the event information and determines which application 136-1 to deliver the event information to and application view 191 for application 136-1. Event sorter 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192 that indicates the current application view(s) that are displayed on touch-sensitive display system 112 when the application is active or running. In some embodiments, device / global internal state 157 is used by event sorter 170 to determine which application(s) are currently active, and application internal state 192 is used by event sorter 170 to determine which application(s) are currently active, and application internal state 192 is used by event sorter 170 to determine which application view(s) to deliver the event information to.

[0096] In some embodiments, application internal state 192 includes additional information such as one or more of resume information to be used when application 136-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 136-1, state cues that allow the user to return to a previous state or view of application 136-1, and redo / undo cues for previous actions taken by the user.

[0097] Event monitor 171 receives event information from peripherals interface 118. The event information includes information about sub-events (e.g., a user's touch on touch-sensitive display system 112 as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, accelerometer(s) 168, and / or microphone 113 (via audio circuitry 110). The information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display system 112 or a touch-sensitive surface.

[0098] In some embodiments, event monitor 171 sends requests to peripherals interface 118 at predetermined intervals. In response, peripherals interface 118 transmits event information. In other embodiments, peripherals interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or for longer than a predetermined period of time).

[0099] In some embodiments, event sorter 170 also includes a hit view determination module 172 and / or an active event recognizer determination module 173 .

[0100] Hit view determination module 172 provides software procedures for determining where in one or more views a sub-event occurred when touch-sensitive display system 112 displays more than one view. A view consists of the controls and other elements that a user can see on the display.

[0101] Another aspect of a user interface associated with an application is the set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the respective application) in which the touch is detected optionally corresponds to a programmatic level within the application's programmatic or view hierarchy. For example, the lowest-level view in which the touch is detected is optionally referred to as the hit view, and the set of events that are recognized as appropriate inputs is optionally determined based at least in part on the hit view of the initial touch that initiates the touch gesture.

[0102] Hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized hierarchically, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy that should process the sub-events. In most situations, the hit view is the lowest-level view in which the initiating sub-event occurs (i.e., the first sub-event in a series of sub-events that form an event or potential event). Once a hit view is identified by the hit view determination module, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.

[0103] Active event recognizer determination module 173 determines which view(s) in the view hierarchy should receive the particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive the particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that contain the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the particular sequence of sub-events. In other embodiments, even if the touch sub-event is completely confined to the area associated with one particular view, views higher in the hierarchy still remain actively participating views.

[0104] Event dispatcher module 174 dispatches event information to event recognizers (e.g., event recognizer 180). In embodiments that include active event recognizer determination module 173, event dispatcher module 174 delivers the event information to the event recognizers determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores event information obtained by each event receiver module 182 in an event queue.

[0105] In some embodiments, operating system 126 includes event sorter 170. Alternatively, application 136-1 includes event sorter 170. In still other embodiments, event sorter 170 is a stand-alone module or is part of another module stored in memory 102, such as contact / motion module 130.

[0106] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each containing instructions for processing touch events that occur within a respective view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, each application view 191 includes multiple event recognizers 180. In other embodiments, one or more of event recognizers 180 are part of a separate module, such as a User Interface Kit (not shown) or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, each event handler 190 includes one or more of data updaters 176, object updaters 177, GUI updaters 178, and / or event data 179 received from event sorter 170. Event handler 190 optionally utilizes or calls data updaters 176, object updaters 177, or GUI updaters 178 to update application internal state 192. Instead, one or more of the application views 191 include one or more respective event handlers 190. Also, in some embodiments, one or more of the data updater 176, the object updater 177, and the GUI updater 178 are included in each application view 191.

[0107] Each event recognizer 180 receives event information (e.g., event data 179) from event sorter 170 and identifies an event from the event information. Event recognizer 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event recognizer 180 also includes at least a subset of metadata 183 and event delivery instructions 188 (optionally including sub-event delivery instructions).

[0108] Event receiver 182 receives event information from event sorter 170. The event information includes information about a sub-event, e.g., a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information, such as the position of the sub-event. When the sub-event involves a movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also called the device's attitude).

[0109] The event comparator 184 compares the event information with predefined event or sub-event definitions and determines the event or sub-event, or determines or updates the state of the event or sub-event, based on the comparison. In some embodiments, the event comparator 184 includes an event definition 186. The event definition 186 includes definitions of events (e.g., a predefined sequence of sub-events), such as Event 1 (187-1) and Event 2 (187-2). In some embodiments, sub-events in Event 187 include, for example, a touch start, a touch end, a touch movement, a touch cessation, and multiple touches. In one example, the definition for Event 1 (187-1) is a double tap on a displayed object. The double tap includes, for example, a first touch (touch start) for a predetermined stage on the displayed object, a first lift-off (touch end) for a predetermined stage on the displayed object, a second touch (touch start) for a predetermined stage on the displayed object, and a second lift-off (touch end) for a predetermined stage. In another example, a definition of event 2 (187-2) is a drag on a displayed object. Drag includes, for example, a touch (or contact) of a predetermined magnitude on a displayed object, a movement of the touch across the touch-sensitive display system 112, and a lift-off of the touch (end of the touch). In some embodiments, the event also includes information about one or more associated event handlers 190.

[0110] In some embodiments, event definition 187 includes a definition of the event for each user interface object. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view in which three user interface objects are displayed on touch-sensitive display system 112, when a touch is detected on touch-sensitive display system 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a respective event handler 190, event comparator 184 uses the results of the hit test to determine which event handler 190 to activate. For example, event comparator 184 selects the event handler associated with the sub-event and object that triggers the hit test.

[0111] In some embodiments, each event 187 definition also includes a delay action that delays delivery of the event information until it is determined whether a set of sub-events corresponds to the event recognizer's event type.

[0112] If the respective event recognizer 180 determines that the sequence of sub-events does not match any of the events in the event definition 186, the respective event recognizer 180 enters an event-disabled, event-failed, or event-ended state and thereafter ignores the next sub-event of the touch-based gesture. In this situation, any other event recognizers that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.

[0113] In some embodiments, each event recognizer 180 includes metadata 183 with configurable properties, flags, and / or lists that indicate to actively participating event recognizers how the event delivery system should perform sub-event delivery. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact with each other or how event recognizers are enabled to interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how sub-events are delivered to various levels in a view or programmatic hierarchy.

[0114] In some embodiments, each event recognizer 180 activates an event handler 190 associated with an event when one or more specific sub-events of the event are recognized. In some embodiments, each event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is separate from sending (and postponing sending) sub-events to the respective hit view. In some embodiments, the event recognizer 180 pops a flag associated with the recognized event, and the event handler 190 associated with the flag captures the flag and performs a predetermined process.

[0115] In some embodiments, the event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with a set of sub-events or to an actively participating view. The event handler associated with the set of sub-events or the actively participating view receives the event information and performs predetermined processing.

[0116] In some embodiments, data updater 176 creates and updates data used by application 136-1. For example, data updater 176 updates phone numbers used by contacts module 137 or stores video files used by video and music player module 152. In some embodiments, object updater 177 creates and updates objects used by application 136-1. For example, object updater 177 creates new user interface objects or updates the positions of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on the touch-sensitive display.

[0117] In some embodiments, event handler(s) 190 include or have access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the respective application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.

[0118] It should be understood that the foregoing description of event processing of a user's touch on a touch-sensitive display also applies to other forms of user input for operating multifunction device 100 using input devices, not all of which are initiated on the touchscreen. For example, mouse movements and mouse button presses, contact movements such as tapping, dragging, scrolling on a touchpad, optionally coordinated with single or multiple keyboard presses or holds, pen stylus input, device movement, verbal commands, detected eye movement, biometric input, and / or any combination thereof, optionally utilize as inputs corresponding to sub-events that define the recognized event.

[0119] 1C is a block diagram illustrating a tactile output module, according to some embodiments. In some embodiments, I / O subsystem 106 (e.g., haptic feedback controller 161 (FIG. 1A) and / or other input controller(s) 160 (FIG. 1A) includes at least some of the example components shown in FIG. 1C. In some embodiments, peripherals interface 118 includes at least some of the example components shown in FIG. 1C.

[0120] In some embodiments, the tactile output module includes a haptic feedback module 133. In some embodiments, the haptic feedback module 133 aggregates and combines tactile output in response to user interface feedback from software applications on the electronic device (e.g., user input corresponding to a displayed user interface, as well as feedback responsive to alerts and other notifications indicating the performance of an action or the occurrence of an event in the user interface of the electronic device). The haptic feedback module 133 includes one or more of a waveform module 123 (which provides the waveforms used to generate the tactile output), a mixer 125 (which mixes waveforms, such as waveforms in different channels), a compressor 127 (which reduces or compresses the dynamic range of the waveforms), a low-pass filter 129 (which filters high-frequency signal components in the waveforms), and a thermal controller 131 (which adjusts the waveform according to thermal conditions). In some embodiments, the haptic feedback controller 161 ( FIG. 1A ) includes the haptic feedback module 133. In some embodiments, a separate unit of haptic feedback module 133 (or a separate implementation of haptic feedback module 133) is also included in an audio controller (e.g., audio circuit 110 of FIG. 1A) and used to generate audio signals. In some embodiments, a single haptic feedback module 133 is used to generate waveforms for audio signals and tactile outputs.

[0121] In some embodiments, haptic feedback module 133 also includes trigger module 121 (e.g., a software application, operating system, or other software module that determines a tactile output to be generated and initiates the process of generating the corresponding tactile output). In some embodiments, trigger module 121 generates a trigger signal that initiates the generation of a waveform (e.g., by waveform module 123). For example, trigger module 121 generates the trigger signal based on pre-set timing criteria. In some embodiments, trigger module 121 receives trigger signals from outside haptic feedback module 133 (e.g., in some embodiments, haptic feedback module 133 receives trigger signals from hardware input processing module 146 located outside haptic feedback module 133) based on the activation of a user interface element (e.g., an application icon or affordance within an application) or a hardware input device (e.g., an intensity-sensitive input surface such as a home button or an intensity-sensitive touchscreen), and relays those trigger signals to other components within haptic feedback module 133 (e.g., waveform module 123) or software applications that trigger actions (e.g., by trigger module 121). In some embodiments, trigger module 121 also receives tactile feedback generation instructions (e.g., from haptic feedback module 133, FIGS. 1A and 3). In some embodiments, trigger module 121 generates a trigger signal in response to haptic feedback module 133 (or trigger module 121 within haptic feedback module 133) receiving a tactile feedback instruction (e.g., from haptic feedback module 133, FIGS. 1A and 3).

[0122] Waveform module 123 receives a trigger signal as an input (e.g., from trigger module 121) and, in response to the received trigger signal, provides a waveform for generation of one or more tactile outputs (e.g., a waveform selected from a predetermined set of waveforms designed for use by waveform module 123, such as the waveforms described in more detail below with reference to Figures 4F-4G).

[0123] Mixer 125 receives waveforms (e.g., from waveform module 123) as inputs and mixes these waveforms together. For example, when mixer 125 receives two or more waveforms (e.g., a first waveform on a first channel and a second waveform on a second channel that at least partially overlaps with the first waveform), mixer 125 outputs a combined waveform that corresponds to the sum of the two or more waveforms. In some embodiments, mixer 125 also modifies one or more of the two or more waveforms to emphasize a particular waveform relative to the rest of the two or more waveforms (e.g., by increasing the scale of a particular waveform and / or decreasing the scale of the rest of the waveforms). In some situations, mixer 125 selects one or more waveforms to remove from the combined waveform (e.g., when there are waveforms from four or more sources requested to be output simultaneously by tactile output generator 167, the waveform from the oldest source is dropped).

[0124] Compressor 127 receives as input a waveform (e.g., a composite waveform from mixer 125) and modifies the waveform. In some embodiments, compressor 127 reduces the waveform (e.g., according to the physical specifications of tactile output generator 167 (FIG. 1A) or 357 (FIG. 3)) so that the tactile output corresponding to the waveform is reduced. In some embodiments, compressor 127 limits the waveform, such as by enforcing a predetermined maximum amplitude for the waveform. For example, compressor 127 reduces the amplitude of portions of the waveform that exceed a predetermined amplitude threshold, while maintaining the amplitude of portions of the waveform that do not exceed the predetermined amplitude threshold. In some embodiments, compressor 127 reduces the dynamic range of the waveform. In some embodiments, compressor 127 dynamically reduces the dynamic range of the waveform so that the combined waveform remains within the performance specifications of tactile output generator 167 (e.g., force and / or displacement limits of a moving mass).

[0125] Low-pass filter 129 receives waveforms (e.g., compressed waveforms from compressor 127) as inputs and filters (e.g., smooths) these waveforms (e.g., removes or reduces high-frequency signal components within the waveforms). For example, in some instances, when a tactile output is generated according to a compressed waveform, compressor 127 includes extraneous signals (e.g., high-frequency signal components) within the compressed waveform that interfere with the generation of the tactile output and / or exceed the performance specifications of tactile output generator 167. Low-pass filter 129 reduces or removes such extraneous signals within the waveform.

[0126] Thermal controller 131 receives waveforms (e.g., filtered waveforms from low-pass filter 129) as inputs and adjusts these waveforms according to the thermal state of device 100 (e.g., based on an internal temperature detected within device 100, such as the temperature of haptic feedback controller 161, and / or an external temperature detected by device 100). For example, in some cases, the output of haptic feedback controller 161 varies with temperature (e.g., haptic feedback controller 161 generates a first tactile output when haptic feedback controller 161 is at a first temperature and a second tactile output when haptic feedback controller 161 is at a second temperature different from the first temperature, in response to receiving the same waveform). For example, the magnitude (or amplitude) of the tactile output may vary with temperature. To reduce the effects of temperature fluctuations, the waveform is modified (e.g., the amplitude of the waveform is increased or decreased based on temperature).

[0127] In some embodiments, haptic feedback module 133 (e.g., trigger module 121) is coupled to hardware input processing module 146. In some embodiments, other input controller(s) 160 of FIG. 1A includes hardware input processing module 146. In some embodiments, hardware input processing module 146 receives input from hardware input device 145 (e.g., other input or control device 116 of FIG. 1A, such as a home button or an intensity-sensitive input surface such as an intensity-sensitive touchscreen). In some embodiments, hardware input device 145 is any of touch-sensitive display system 112 (FIG. 1A), keyboard / mouse 350 (FIG. 3), touchpad 355 (FIG. 3), one of other input or control devices 116 (FIG. 1A), or an input device such as an intensity-sensitive home button. In some embodiments, hardware input device 145 consists of an intensity-sensitive home button and is not touch-sensitive display system 112 (FIG. 1A), keyboard / mouse 350 (FIG. 3), or touchpad 355 (FIG. 3). In some embodiments, in response to input from hardware input device 145 (e.g., an intensity-sensitive home button or a touchscreen), hardware input processing module 146 provides one or more trigger signals to haptic feedback module 133 to indicate that a user input that meets predetermined input criteria has been detected, such as an input corresponding to a home button "click" (e.g., a "down click" or an "up click"). In some embodiments, haptic feedback module 133, in response to an input corresponding to a home button "click," provides a waveform corresponding to a home button "click" to simulate the haptic feedback of pressing a physical home button.

[0128] In some embodiments, the tactile output module includes a haptic feedback controller 161 (haptic feedback controller 161 in FIG. 1A ) that controls generation of the tactile output. In some embodiments, the haptic feedback controller 161 is coupled to a plurality of tactile output generators, selects one or more of the plurality of tactile output generators, and sends waveforms to the selected one or more tactile output generators that generate the tactile output. In some embodiments, the haptic feedback controller 161 aggregates tactile output requests corresponding to activation of the hardware input device 145 and tactile output requests corresponding to software events (e.g., tactile output requests from the haptic feedback module 133), modifies one or more of the two or more waveforms, emphasizes a particular waveform relative to the rest of the two or more waveforms (e.g., by increasing the scale of a particular waveform and / or decreasing the scale of the rest of the waveforms, such as to prioritize tactile output corresponding to activation of the hardware input device 145 over tactile output corresponding to the software event).

[0129] In some embodiments, as shown in FIG. 1C , the output of haptic feedback controller 161 is coupled to an audio circuit of device 100 (e.g., audio circuit 110, FIG. 1A ) and provides an audio signal to the audio circuit of device 100. In some embodiments, haptic feedback controller 161 provides both a waveform used to generate the tactile output and an audio signal used to provide the audio output in conjunction with generating the tactile output. In some embodiments, haptic feedback controller 161 modifies the audio signal and / or waveform (used to generate the tactile output) so that the audio output and the tactile output are synchronized (e.g., by delaying the audio signal and / or waveform). In some embodiments, haptic feedback controller 161 includes a digital-to-analog converter used to convert the digital waveform to an analog signal, which is received by amplifier 163 and / or tactile output generator 167.

[0130] In some embodiments, tactile output module includes amplifier 163. In some embodiments, amplifier 163 receives a waveform (e.g., from haptic feedback controller 161) and amplifies the waveform before sending the amplified waveform to tactile output generator 167 (e.g., either tactile output generator 167 (FIG. 1A) or 357 (FIG. 3)). For example, amplifier 163 amplifies the received waveform to a signal level according to the physical specifications of tactile output generator 167 (e.g., the voltage and / or current required by tactile output generator 167 to generate a tactile output, such that the signal sent to tactile output generator 167 generates a tactile output that corresponds to the waveform received from tactile feedback controller 161), and sends the amplified waveform to tactile output generator 167. In response, tactile output generator 167 generates a tactile output (e.g., by shifting the movable mass back and forth in one or more dimensions relative to the movable mass's neutral position).

[0131] In some embodiments, tactile output module includes sensor 169 coupled to tactile output generator 167. Sensor 169 detects the state or state change (e.g., mechanical position, physical displacement, and / or movement) of tactile output generator 167 or one or more components of tactile output generator 167 (e.g., one or more moving parts, such as a membrane, used to generate the tactile output). In some embodiments, sensor 169 is a magnetic field sensor (e.g., a Hall Effect sensor) or other displacement and / or movement sensor. In some embodiments, sensor 169 provides information (e.g., the position, displacement, and / or movement of one or more parts within tactile output generator 167) to tactile feedback controller 161, and tactile feedback controller 161 adjusts the waveform output from tactile feedback controller 161 (e.g., the waveform sent to tactile output generator 167, optionally via amplifier 163) according to the information provided by sensor 169 regarding the state of tactile output generator 167.

[0132] FIG. 2 illustrates portable multifunction device 100 having a touchscreen (e.g., touch-sensitive display system 112, FIG. 1A ) according to some embodiments. The touchscreen optionally displays one or more graphics within user interface (UI) 200. In these embodiments, as well as those described below, a user is enabled to select one or more of the graphics by making a gesture on the graphics, for example, with one or more fingers 202 (not drawn to scale) or one or more styluses 203 (not drawn to scale). In some embodiments, selection of one or more graphics is performed when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward and / or downward), and / or rolling of a finger in contact with device 100 (right to left, left to right, upward and / or downward). In some implementations or situations, accidental contact with a graphic does not select the graphic, for example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon optionally does not select the corresponding application.

[0133] Device 100 also optionally includes one or more physical buttons, such as a "home" or menu button 204. As mentioned above, menu button 204 is optionally used to navigate to any application 136 in a set of applications optionally running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on a touchscreen display.

[0134] In some embodiments, device 100 includes a touchscreen display, a menu button 204 (sometimes referred to as a home button 204), a push button 206 for powering the device on / off and locking the device, volume control button(s) 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and an external docking / charging port 124. Push button 206 is optionally used to power the device on / off by pressing and holding the button down for a predetermined period of time, to lock the device by pressing and releasing the button before the predetermined time has elapsed, and / or to unlock the device or initiate the unlocking process. In some embodiments, device 100 also accepts verbal input through microphone 113 to activate or deactivate some features. Device 100 also optionally includes one or more contact intensity sensors 165 that detect the intensity of a contact on touch-sensitive display system 112 and / or one or more tactile output generators 167 that generate a tactile output for a user of device 100.

[0135] FIG. 3 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or commercial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communication between system components. Device 300 includes input / output (I / O) interface 330, which includes display 340, which is typically a touchscreen display. I / O interface 330 optionally also includes a keyboard and / or mouse (or other pointing device) 350, as well as a touchpad 355, a tactile output generator 357 for generating tactile outputs on device 300 (e.g., similar to tactile output generator(s) 167 described above with reference to FIG. 1A ), sensors 359 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors similar to contact intensity sensor(s) 165 described above with reference to FIG. 1A ). Memory 370 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or a subset of, programs, modules, and data structures stored in memory 102 of portable multifunction device 100 (FIG. 1A). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores drawing module 380, presentation module 382, ​​word processing module 384, website creation module 386, disc authoring module 388, and / or spreadsheet module 390, while memory 102 of portable multifunction device 100 (FIG. 1A) optionally does not store those modules.

[0136] 3 is optionally stored in one or more of the memory devices mentioned above. Each of the above-identified modules corresponds to an instruction set that performs the functions described above. The above-identified modules or programs (i.e., instruction sets) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of those modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 370 optionally stores additional modules and data structures not described above.

[0137] Attention is now directed to embodiments of a user interface (“UI”) that is optionally implemented on portable multifunction device 100.

[0138] 4A shows an exemplary user interface 400 for a menu of applications on portable multifunction device 100, according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof: signal strength indicator(s) for wireless communication(s), such as cellular and Wi-Fi signals; ●Time, ●Bluetooth (registered trademark) indicator, ● Battery status indicator, Tray 408 with icons of frequently used applications, such as: an icon 416 for the phone module 138, labeled "Phone," that optionally includes an indicator 414 of the number of missed calls or voicemail messages; icon 418 of the email client module 140, labeled "Mail," optionally including an indicator 410 of the number of unread emails; ○ An icon 420 for the Browser module 147, labeled "Browser"; and ○ An icon 422 for the video and music player module 152 labeled "Music", and Icons of other applications, such as: ○ Icon 424 of IM module 141, labeled "Messages"; icon 426 of the calendar module 148, labeled "Calendar"; ○ Icon 428 of the image management module 144, labeled "Photos" ○ An icon 430 of the camera module 143, labeled "camera"; ○ Icon 432 of the online video module 155, labeled "Online Video"; Icon 434 of Stock Price Widget 149-2, labeled "Stock Price" ○ Icon 436 of the map module 154, labeled "Map"; Icon 438 of weather widget 149-1, labeled "Weather" ○ Icon 440 of alarm clock widget 149-4, labeled "Clock" ○ Icon 442 of Training Support Module 142, labeled "Training Support"; icon 444 of the Notes module 153, labeled "Notes"; and ○ An icon 446 for a settings application or module that provides access to settings for the device 100 and its various applications 136.

[0139] 4A are merely examples. For example, other labels are optionally used for various application icons. In some embodiments, the label for each application icon includes the name of the application corresponding to the respective application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to that particular application icon.

[0140] FIG. 4B shows an exemplary user interface on a device (e.g., device 300, FIG. 3) that has a touch-sensitive surface 451 (e.g., tablet or touchpad 355, FIG. 3) that is separate from a display 450. While many of the following examples are given with reference to input on touchscreen display 112 (where the touch-sensitive surface and display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, as shown in FIG. 4B . In some embodiments, the touch-sensitive surface (e.g., 451 in FIG. 4B ) has a major axis (e.g., 452 in FIG. 4B ) that corresponds to a major axis (e.g., 453 in FIG. 4B ) on the display (e.g., 450). According to these embodiments, the device detects contact with touch-sensitive surface 451 (e.g., 460 and 462 in FIG. 4B ) at locations that correspond to respective locations on the display (e.g., in FIG. 4B , 460 corresponds to 468, and 462 corresponds to 470). In this manner, when the touch-sensitive surface is separate from the display, user input (e.g., contacts 460 and 462, and their movement) detected by the device on the touch-sensitive surface (e.g., 451 in FIG. 4B ) is used by the device to operate a user interface on the display (e.g., 450 in FIG. 4B ) of the multifunction device. It should be understood that similar methods are optionally used for the other user interfaces described herein.

[0141] Additionally, while the following examples are given primarily with reference to finger input (e.g., finger touches, finger tap gestures, finger swipe gestures, etc.), it should be understood that in some embodiments, one or more of those finger inputs are replaced with input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a touch) followed by movement of a cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click (e.g., instead of detecting a touch and subsequently terminating contact detection) while the cursor is located over the location of the tap gesture. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice are optionally used simultaneously, or a mouse and finger touches are optionally used simultaneously.

[0142] As used herein, the term “focus selector” refers to an input element that indicates the current portion of a user interface with which a user is interacting. In some implementations involving a cursor or other location marker, the cursor functions as a “focus selector” such that when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 in FIG. 3 or touch-sensitive surface 451 in FIG. 4B ) while the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations involving a touchscreen display (e.g., touch-sensitive display system 112 in FIG. 1A or touchscreen in FIG. 4A ) that enables direct interaction with user interface elements on the touchscreen display, a contact detected on the touchscreen functions as a “focus selector” such that when input (e.g., a press input by contact) is detected on the touchscreen display at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one region of the user interface to another region of the user interface (e.g., by using the tab key or arrow keys to move focus from one button to another) without a corresponding cursor movement or contact movement on the touchscreen display. In these implementations, the focus selector moves to follow the movement of focus between various regions of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is generally a user interface element (or a contact on a touchscreen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user intends to interact with).For example, while a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of a focus selector (e.g., a cursor, touch, or selection box) over a corresponding button indicates that the user intends to activate that corresponding button (rather than other user interface elements shown on the device's display).

[0143] As used herein and in the claims, the term “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact or a stylus contact) on the touch-sensitive surface, or a proxy for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values ​​that includes at least four distinct values ​​and more typically includes hundreds (e.g., at least 256) distinct values. The intensity of a contact is optionally determined (or measured) using various techniques and various sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted average or sum) to determine an estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in the contact area detected on the touch-sensitive surface, the capacitance and / or change in the capacitance of the touch-sensitive surface proximate the contact, and / or the resistance and / or change in the capacitance of the touch-sensitive surface proximate the contact are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the surrogate measure of the force or pressure of the contact is used directly to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measure). In some implementations, the surrogate measure for the force or pressure of the contact is converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure).Using the intensity of contact as an attribute of user input enables a user to access additional device functionality that may not otherwise be easily accessible by the user on devices of reduced size with limited assets for displaying affordances (e.g., on a touch-sensitive display) and / or receiving user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as knobs or buttons).

[0144] In some embodiments, contact / motion module 130 uses one or more sets of intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and may be adjusted without modifying the physical hardware of device 100). For example, the mouse “click” threshold of a trackpad or touchscreen display can be set to any of a wide range of pre-defined thresholds without modifying the trackpad or touchscreen display hardware. Furthermore, in some implementations, a user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once with a system-level click “intensity” parameter).

[0145] As used herein and in the claims, the term "characteristic intensity" of a contact refers to a characteristic of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on a plurality of intensity samples. The characteristic intensity is optionally based on a predetermined number of intensity samples, i.e., a set of intensity samples collected during a predetermined time (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detecting the contact, before detecting lift-off of the contact, before or after detecting the start of contact movement, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The characteristic intensity of the contact is optionally based on one or more of the maximum intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the top 10% of the intensity of the contact, half the maximum intensity of the contact, 90% of the maximum intensity of the contact, a value generated by low-pass filtering the intensity of the contact starting over a predefined period or at a predefined time, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action is performed by the user. For example, the set of one or more intensity thresholds may include a first intensity threshold and a second intensity threshold. In this example, a contact having a characteristic intensity not exceeding the first threshold results in a first action being performed, a contact having a characteristic intensity above the first intensity threshold but not above the second intensity threshold results in a second action being performed, and a contact having a characteristic intensity above the second intensity threshold results in a third action being performed. In some embodiments, the comparison between the characteristic intensity and one or more intensity thresholds is not used to determine whether to perform a first operation or a second operation, but rather is used to determine whether to perform one or more operations (e.g., whether to perform a respective option or refrain from performing a respective operation).

[0146] In some embodiments, a portion of the gesture is identified for purposes of determining the characteristic intensity. For example, a touch-sensitive surface may receive a continuous swipe contact (e.g., a drag gesture) that transitions from a start position to an end position, at which point the intensity of the contact increases. In this example, the characteristic intensity of the contact at the end position may be based on only a portion of the continuous swipe contact (e.g., only the portion of the swipe contact at the end position) rather than the entire swipe contact. In some embodiments, a smoothing algorithm may be applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the swipe contact intensity for purposes of determining the characteristic intensity.

[0147] The user interface diagrams described herein optionally include various intensity diagrams showing the current intensity of a contact on the touch-sensitive surface relative to one or more intensity thresholds (e.g., a touch-detection intensity threshold ITO, a light press intensity threshold ITL, a deep press intensity threshold ITD (e.g., at least initially higher than ITL), and / or one or more other intensity thresholds (e.g., an intensity threshold ITH lower than ITL). This intensity diagram is typically not part of the displayed user interface, but is provided to aid in interpretation of the diagram. In some embodiments, the light press intensity threshold is set to the intensity at which the device performs the action normally associated with clicking a button on a physical mouse or a trackpad. In some embodiments, the deep pressure intensity threshold corresponds to an intensity at which the device performs an action different from the action typically associated with clicking a physical mouse button or trackpad. In some embodiments, when a contact is detected having a characteristic intensity below the shallow pressure intensity threshold (e.g., above a nominal contact-detection intensity threshold IT0 below which contact is no longer detected), the device moves the focus selector according to the movement of the contact on the touch-sensitive surface without performing any action associated with the shallow or deep pressure intensity threshold. Generally, unless otherwise specified, these intensity thresholds are consistent across various sets of values ​​for a user interface.

[0148] In some embodiments, the device's response to an input detected by the device depends on criteria based on the intensity of the contact during the input. For example, for some "light press" inputs, the intensity of the contact during the input that exceeds a first intensity threshold triggers a first response. In some embodiments, the device's response to an input detected by the device depends on criteria including both the intensity of the contact during the input and time-based criteria. For example, for some "deep press" inputs, the intensity of the contact during the input that exceeds a second intensity threshold that is greater than the first intensity threshold for light presses triggers a second response only if a delay time has elapsed between meeting the first intensity threshold and meeting the second intensity threshold. This delay time is typically less than 200 ms (milliseconds) (e.g., 40 ms, 100 ms, or 120 ms, depending on the magnitude of the second intensity threshold, with the delay time increasing as the second intensity threshold increases). This delay time helps avoid accidental recognition of deep press inputs. As another example, for some "deep pressure" inputs, there is a period of reduced sensitivity that occurs after the first intensity threshold is met. During the period of reduced sensitivity, the second intensity threshold is increased. This temporary increase in the second intensity threshold also helps to avoid accidental deep pressure inputs. For other deep pressure inputs, the response to the detection of a deep pressure input does not depend on a time-based criterion.

[0149] In some embodiments, one or more of the input intensity threshold and / or corresponding output vary based on one or more factors, such as user settings, contact movement, input timing, running application, rate at which intensity is applied, number of simultaneous inputs, user history, environmental factors (e.g., ambient noise), position of focus selector, etc. Exemplary factors are described in U.S. Patent Application Publication Nos. 14 / 399,606 and 14 / 624,296, which are incorporated herein by reference in their entireties.

[0150] For example, FIG. 4C illustrates a dynamic intensity threshold 480 that varies over time based in part on the intensity of touch input 476 over time. Dynamic intensity threshold 480 is the sum of two components: a first component 474 that decays over time after a predetermined delay time p1 from when touch input 476 is first detected, and a second component 478 that tracks the intensity of touch input 476 over time. The initial high intensity threshold of first component 474 reduces accidental triggering of a "deep press" response while still allowing an immediate "deep press" response if touch input 476 provides sufficient intensity. Second component 478 reduces unintentional triggering of a "deep press" response due to gradual intensity variations of the touch input. In some embodiments, a "deep press" response is triggered when touch input 476 meets dynamic intensity threshold 480 (e.g., at point 481 in FIG. 4C ).

[0151] FIG. 4D illustrates another dynamic intensity threshold 486 (e.g., intensity threshold ID). FIG. 4D also illustrates two other intensity thresholds: a first intensity threshold IH and a second intensity threshold IL. In FIG. 4D, a touch input 484 satisfies the first intensity threshold IH and the second intensity threshold IL before time p2, but no response is provided until delay time p2 has elapsed at time 482. Also in FIG. 4D, the dynamic intensity threshold 486 decays over time, beginning at time 488 after a predetermined delay time p1 has elapsed from time 482 (when the response associated with the second intensity threshold IL is triggered). This type of dynamic intensity threshold reduces the accidental triggering of a response associated with the dynamic intensity threshold ID immediately after or simultaneously with triggering a response associated with a lower intensity threshold, such as the first intensity threshold IH or the second intensity threshold IL.

[0152] 4E illustrates yet another dynamic intensity threshold 492 (e.g., intensity threshold ID). In FIG. 4E, a response associated with intensity threshold ID is triggered after a delay time p2 from when touch input 490 is first detected. At the same time, dynamic intensity threshold 492 decays after a predetermined delay time p1 from when touch input 490 is first detected. Thus, a decrease in the intensity of touch input 490 after triggering a response associated with intensity threshold ID, followed by an increase in the intensity of touch input 490 without releasing touch input 490, can trigger a response associated with intensity threshold ID (e.g., at time 494) even when the intensity of touch input 490 falls below another intensity threshold, e.g., intensity threshold ID.

[0153] An increase in the characteristic intensity of a contact from an intensity below the shallow pressure intensity threshold ITL to an intensity between the shallow pressure intensity threshold ITL and the deep pressure intensity threshold ITD may be referred to as a “shallow press” input. An increase in the characteristic intensity of a contact from an intensity below the deep pressure intensity threshold ITD to an intensity above the deep pressure intensity threshold ITD may be referred to as a “deep press” input. An increase in the characteristic intensity of a contact from an intensity below the contact detection intensity threshold IT0 to an intensity between the contact detection intensity threshold ITL and the shallow pressure intensity threshold ITL may be referred to as detecting a contact on the touch surface. A decrease in the characteristic intensity of a contact from an intensity above the contact detection intensity threshold IT0 to an intensity below the contact detection intensity threshold IT0 may be referred to as detecting a lift-off of the contact from the touch surface. In some embodiments, IT0 is zero. In some embodiments, IT0 is greater than zero. In some examples, a shaded circle or ellipse is used to represent the intensity of a contact on the touch-sensitive surface. In some examples, an unshaded circle or ellipse is used to represent each contact on the touch-sensitive surface without specifying the intensity of each contact.

[0154] In some embodiments described herein, one or more actions are performed including upon detecting a gesture including the respective pressure input, or including upon detecting a respective pressure input performed on the respective contact(s), where the respective pressure inputs are detected based at least in part on detecting an increase in intensity of the contact(s) above a pressure input intensity threshold. In some embodiments, the respective actions are performed including upon detecting an increase in intensity of the respective contact(s) above a pressure input intensity threshold (e.g., the respective actions are performed on the “downstroke” of the respective pressure input). In some embodiments, the pressure inputs include an increase in intensity of the respective contact(s) above a pressure input intensity threshold and a subsequent decrease in intensity of the contact(s) below the pressure input intensity threshold, and the respective actions are performed including upon detecting a subsequent decrease in intensity of the respective contact(s) below the pressure input threshold (e.g., the respective actions are performed on the “upstroke” of the respective pressure input).

[0155] In some embodiments, the device employs intensity hysteresis to avoid accidental input, sometimes referred to as “jitter,” and the device defines or selects a hysteresis intensity threshold that has a predetermined relationship to the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units below the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Thus, in some embodiments, the press input includes an increase in the intensity of each contact above the press input intensity threshold and a subsequent decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and a respective action is performed in response to detecting the subsequent decrease in intensity of each contact below the hysteresis intensity threshold (e.g., the respective action is performed on the “upstroke” of the respective press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, and optionally a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and a respective action is performed in response to detecting the press input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, as the case may be).

[0156] For ease of explanation, descriptions of actions performed in response to a press input associated with a press input intensity threshold or in response to a gesture including a press input are, optionally, triggered in response to detecting an increase in the intensity of the contact above the press input intensity threshold, an increase in the intensity of the contact from an intensity below a hysteresis intensity threshold to an intensity above the press input intensity threshold, a decrease in the intensity of the contact below the press input intensity threshold, or a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to the press input intensity threshold. Further, in examples where an action is described as being performed in response to detecting a decrease in the intensity of the contact below a press input intensity threshold, the action is optionally performed in response to detecting a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to and lower than the press input intensity threshold. As noted above, in some embodiments, the triggering of these responses also depends on time-based criteria being met (e.g., a delay time elapsed between the first intensity threshold being met and the second intensity threshold being met).

[0157] As used herein and in the claims, the term “tactile output” refers to a physical displacement of a device relative to its previous position, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, that will be detected by a user with the user's sense of touch. For example, in a situation where a device or a component of a device is in contact with a touch-sensitive surface of a user (e.g., the fingers, palm, or other part of the user's hand), the tactile output produced by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of the touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a “downclick” or “upclick” of a physical actuator button. In some cases, the user feels a tactile sensation such as a “downclick” or “upclick” even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's action. As another example, movement of a touch-sensitive surface is optionally interpreted or perceived by a user as "roughness" of the touch-sensitive surface, even when there is no change in the smoothness of the touch-sensitive surface. Such user interpretation of touch depends on the user's personal sensory perception, but there are many sensory perceptions of touch that are common to a majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "upclick," "downclick," "roughness"), unless otherwise specified, the generated tactile output corresponds to a physical displacement of the device, or a component of the device, that produces the described sensory perception for a typical (or average) user.Using tactile output to provide haptic feedback to the user improves usability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), as well as reducing power usage and improving the battery life of the device by enabling the user to use the device more quickly and efficiently. User Interface and Related Processes

[0158] Attention is now directed to embodiments of user interfaces (“UI”) and associated processes that may be implemented on an electronic device such as portable multifunction device 100 or device 300, which includes a display, a touch-sensitive surface, (optionally) one or more tactile output generators for generating tactile output, and (optionally) one or more sensors for detecting the intensity of contact with the touch-sensitive surface.

[0159] 5A-5AF show example user interfaces for relocating an annotation, according to some embodiments. The user interfaces in those figures are used to illustrate processes described below, including those in FIGS. 9A-9F, 10A-10B, 11A-11F, and 12A-12D. For ease of explanation, some of the embodiments are discussed with reference to operations performed on a device having touch-sensitive display system 112. In such embodiments, the focus selector is optionally a respective finger or stylus contact, a representative point corresponding to a finger or stylus contact (e.g., the centroid of each contact or a point associated with each contact), or the centroid of two or more contacts detected on touch-sensitive display system 112. However, similar operations, along with a focus selector, are optionally performed on a device having display 450 and a separate touch-sensitive surface 451 in response to detecting contacts on touch-sensitive surface 451 while displaying the user interface shown in the figures on display 450.

[0160] 5A shows an annotation user interface 5003 displayed on touchscreen 112 of device 100. The user interface displays the field of view of the camera of device 100 as it captures the physical environment 5000 of device 100. A table 5001a and a mug 5002a are located within physical environment 5000. The camera view includes a visual representation 5001b of the physical table 5001a as displayed in user interface 5003. User interface 5003 includes controls 5004 for toggling between still and video modes, controls 5006 for controlling camera flash settings, controls 5008 for accessing camera settings, and a mode indicator 5010 that indicates the current mode.

[0161] From Figure 5A to Figure 5B, device 100 is moved relative to the physical environment such that the portion of table 5001b visible within the camera's field of view changes and visual representation 5002b of physical mug 5002a becomes visible within the camera's field of view.

[0162] 5C , stylus 5012 is touched on touchscreen display 112 at a point within annotation user interface 5003 indicated by contact 5014. In response to detecting contact 5014, a still image of the camera's field of view is captured and displayed within user interface 5003. The state of mode control 5004 and mode indicator 5010 is changed to indicate that the active mode of the user interface has changed from video mode to still image mode. The transition to still image mode, which occurs in response to stylus touchdown (and / or in response to other types of input, as described further below), allows annotation input to be received relative to the view of the physical environment captured by camera 143 without being affected by changes in the spatial relationship between device 100 and physical environment 5000 caused by movement of device 100.

[0163] 5C to 5D, device 100 is moved relative to physical environment 5000. The image displayed by touchscreen display 112 does not change from FIG. 5C to FIG. 5D because the active mode of the user interface has changed from video mode to still image mode.

[0164] From Figure 5E to Figure 5G, contact 5014 moves along the path indicated by arrow 5016, creating a first annotation 5018 on a portion of the still image that includes visual representation 5002b of physical mug 5002a.

[0165] In Figure 5H, stylus 5012 provides input at a location corresponding to mode control 5004, as indicated by contact 5019. In Figure 51, in response to the input by stylus 5012, the state of mode control 5004 and mode indicator 5010 has changed to indicate that the active mode of annotation user interface 5003 has changed from still image mode to video mode.

[0166] 5I-5L, the still image displayed in FIG. 5I gradually transitions from a full-size view of the still image (as shown in user interface 5003 of FIG. 5I) to a miniature view 5020 (as shown in user interface 5003 of FIGS. 5J-5K) up to indicator dot 5022 (as shown in user interface 5003 of FIG. 5L). In FIG. 5J, a miniature view 5020 of the still image shown in FIG. 5I is shown that covers the camera's current field of view. The size of miniature view 5020 is reduced from FIG. 5J to FIG. 5K (e.g., to provide an indication of the correspondence between the still image displayed by device 100 of FIG. 5I and indicator dot 5022 displayed over the video corresponding to the device camera's field of view of FIG. 5L).

[0167] 5L-5M, the position of device 100 is changed relative to physical environment 5000. As device 100 moves, the field of view of device 100's camera changes, causing indicator 5022 in user interface 5003 to change position. Movement of indicator 5022 in user interface 5003 provides an indication of the virtual spatial location of annotation 5018 relative to the current location of device 100. In this way, the user is provided visual feedback that movement of device 100 in the direction indicated by indicator 5022 is required to redisplay annotation 5018.

[0168] 5M-5N, the position of device 100 continues to change relative to physical environment 5000. As a result of device 100's movement, the field of view of device 100's camera is updated so that the portion of physical environment 5000 captured in the annotated still image of FIG. 5I is visible within the camera's field of view. Annotation 5018 is displayed in the video (e.g., within each image frame containing a visual representation of physical mug 5002a) at a location corresponding to the location in the still image from which annotation 5018 was received (e.g., at a location corresponding to the location of visual representation 5002b of physical mug 5002a (e.g., as shown in the still image of FIG. 5I)). In some embodiments, as device 100 moves closer to physical mug 5002a, visual representation 5002b of physical mug 5002a appears larger in the video image, and annotation 5018 also appears larger in accordance with the changed size of visual representation 5002b. In some embodiments, as device 100 moves around physical mug 5002a, visual representation 5002b of physical mug 5002a is updated to reflect different viewing angles of physical mug 5002a, and the appearance of annotations 5018 in the video image is also updated according to the changing viewing angle of the physical mug (e.g., as viewed from different angles).

[0169] 5O, stylus 5012 is touched down on touchscreen display 112 at a point in annotation user interface 5003 indicated by contact 5024. In response to detecting contact 5024, a second still image of the camera's field of view is captured and displayed in user interface 5003.

[0170] 5P through 5R, while a second still image is displayed, input by stylus 5012 is received at a location indicated by contact 5026. Movement of contact 5026 creates a second annotation 5028 in a portion of the still image that includes visual representation 5002b of physical mug 5002a. In FIG. 5S, contact 5012 lifts off touchscreen 112.

[0171] In Figure 5T, input is detected at a location on touchscreen display 112 corresponding to control 5004 for toggling between still image mode and video mode, as indicated by contact 5030. In Figure 5U, in response to the input, the state of mode control 5004 and mode indicator 5010 has changed to indicate that the active mode of annotation user interface 5003 has changed from still image mode to video mode. Because the portion of physical environment 5000 captured in the annotated still image of Figure 5T is already visible within the field of view of the camera displayed in the video mode of user interface 5003 of Figure 5U, annotations 5018 and 5028 are displayed at locations in the video corresponding to the corresponding locations in the still image where the annotations were received (e.g., locations corresponding to the corresponding locations of visual representation 5002b of physical mug 5002a shown in Figures 5G and 5R).

[0172] 5U through 5V, the position of device 100 is changed relative to physical environment 5000 so that the field of view of the device camera displayed in user interface 5003 does not include the portion of physical environment 5000 that includes mug 5002a. Indicator dot 5022 corresponding to annotation 5018 and indicator dot 5032 corresponding to annotation 5028 are displayed (e.g., indicator dots 5022 and 5032 are displayed in positions that indicate the off-screen virtual spatial locations of annotations 5018 and 5028, respectively, relative to physical environment 5000 (e.g., the sides of physical mug 5002a shown in FIGS. 5G and 5R ).

[0173] 5V through 5W, the position of device 100 continues to change relative to physical environment 5000. As device 100 moves downward, the field of view of device 100's camera changes and indicators 5022 and 5032 move upward in user interface 5003 (e.g., indicating the virtual spatial positions of annotations 5018 and 5028 relative to device 100's current location). In FIG. 5W, stylus 5012 is touched down on touchscreen display 112 at a point in annotation user interface 5003 indicated by contact 5034. In response to detecting contact 5034, a third still image of the camera's field of view is captured and displayed in user interface 5003.

[0174] From Figure 5W to Figure 5X, while the third still image is displayed, contact 5034 moves along touch screen display 112 to create third annotation 5036 in a portion of the third still image including the lower right surface of visual representation 5001b of physical table 5001a.

[0175] In Figure 5Y, stylus 5012 provides input at a location corresponding to mode control 5004, as indicated by contact 5038. In Figure 5Z, in response to the input by stylus 5012, the state of mode control 5004 and mode indicator 5010 has changed to indicate that the active mode of annotation user interface 5003 has changed from still image mode to video mode. Because the lower right surface of table 5001a is visible within the field of view of the camera displayed in the video mode of user interface 5003 of Figure 5Z, annotation 5036 is displayed at a location in the video (e.g., in the image frame including the position of the table surface shown in Figure 5X) that corresponds to the location in the still image from which annotation 5036 was received (e.g., a location corresponding to the lower right surface of visual representation 5001b of physical table 5001a). The portion of physical environment 5000 visible within the camera's field of view displayed in the video mode of user interface 5003 does not include the portion of the physical environment corresponding to the spatial locations of annotations 5018 and 5028 (e.g., mug 5002a is not visible within the camera's field of view), and indicator dots 5022 and 5032 are displayed in positions relative to physical environment 5000 that indicate the off-screen virtual spatial locations of annotations 5018 and 5028.

[0176] 5Z through 5AA, the position of device 100 changes relative to physical environment 5000. As device 100 moves upward, the field of view of the camera on device 100 changes and indicators 5022 and 5032 move downward in user interface 5003. Because the lower right face of table 5001a is no longer visible in the field of view of the camera displayed in the video mode of the user interface, indicator dot 5040 is displayed at a position in the video that corresponds to the position in the still image where annotation 5036 was received.

[0177] 5AA to 5AB, the position of device 100 changes relative to physical environment 5000. As device 100 continues to move upward, the field of view of the camera of device 100 changes and indicators 5022, 5032, and 5040 move downward in user interface 5003 to indicate the off-screen virtual spatial locations of annotations 5018, 5028, and 5036 relative to physical environment 5000.

[0178] 5AB to 5AC, the position of device 100 changes, causing a change in the field of view of the camera of device 100 displayed in user interface 5003. The positions of indicators 5022, 5032, and 5040 are updated based on the off-screen virtual spatial positions of annotations 5018, 5028, and 5036.

[0179] 5AC through 5AD, the position of device 100 is changed relative to physical environment 5000 such that the field of view of the device camera displayed in user interface 5003 includes partial physical environment 5000, including mug 5002a. Annotations 5018 and 5028 are displayed in user interface 5003, and indicator dots 5022 and 5032 are no longer displayed.

[0180] 5AD through 5AE, the position of device 100 is changed relative to physical environment 5000 so that the field of view of the device camera displayed in user interface 5003 includes partial physical environment 5000 that includes the lower right surface of table 5001 a. Annotation 5036 is displayed in user interface 5003, and indicator dot 5040 is no longer displayed.

[0181] From Figures 5AE to 5AF, device 100 is moved around the periphery and over table 5001a such that the position and perspective of annotations 5018, 5028 and 5036 change.

[0182] 6A-6N show exemplary user interfaces for receiving annotations on a portion of a physical environment captured in a still image corresponding to a paused position of a video, according to some embodiments. The user interfaces in those figures are used to illustrate processes described below, including those in FIGS. 9A-9F, 10A-10B, 11A-11F, and 12A-12D. For ease of explanation, some of the embodiments are discussed with reference to operations performed on a device having touch-sensitive display system 112. In such embodiments, the focus selector is optionally a respective finger or stylus contact, a representative point corresponding to a finger or stylus contact (e.g., the face centroid of each contact or a point associated with each contact), or the face centroid of two or more contacts detected on touch-sensitive display system 112. However, similar operations, along with a focus selector, are optionally performed on a device having display 450 and a separate touch-sensitive surface 451 in response to detecting contacts on touch-sensitive surface 451 while displaying the user interface shown in the figures on display 450.

[0183] 6A shows a user interface 6000 that includes a video playback area 6002. In some embodiments, the user interface 6000 is accessed via a list of media content objects (e.g., within an image and / or video browsing application). In some embodiments, the user interface 6000 also includes a timeline 6004 (e.g., a set of sample frames 6006 that correspond to successive segments of a video). The timeline 6004 includes a current position indicator 6008 that indicates the position on the timeline 6004 that corresponds to the frame displayed in the video playback area 6002 through the video displayed in the video playback area 6002. In some embodiments, the user interface includes a markup control 6010 (e.g., for entering a markup mode for marking the video displayed in the video playback area 6002), a rotate control 6012 (e.g., for rotating the video displayed in the video playback area 6002), an edit control 6014 (e.g., for editing the video displayed in the video playback area 6002), a cancel control 6020 (e.g., for canceling a current operation), a rotate control 6022 (e.g., for rotating the video displayed in the video playback area 6002), and a play / pause toggle control 6016 (e.g., for playing and pausing the video displayed in the video playback area 6002). Contact 6018 (e.g., input by a user's finger) with the touchscreen display 112 is detected at a location corresponding to the play / pause toggle control 6016.

[0184] In FIG. 6B, in response to input detected at a location corresponding to play / pause toggle control 6016, video playback area 6002 begins playing a video.

[0185] 6C , as the video continues to play, input is detected at a location corresponding to markup control 6010, as indicated by contact 6024. In response to input by contact 6024 at a location corresponding to markup control 6010, playback of the video is paused, a still image corresponding to the paused location of the video is displayed, a markup mode is initiated in which input received at a location corresponding to video playback region 6002 marks the video, and the state of markup control 6010 is changed to display the text "Done" (e.g., to indicate that input into selection control 6010 terminates markup mode).

[0186] In Figures 6D-6F, annotation input is detected at a location within video playback region 6002, as indicated by contact 6026. As the contact moves along the path shown in Figures 6D-6F, annotation 6030 is received. In some embodiments, the annotation is received at a location corresponding to an object in the video (e.g., kite 6028). In Figure 6G, while markup mode is active and the text "Done" is displayed, input by contact 6031 is detected at a location corresponding to markup control 6010. In response to input by contact 6031, the markup session ends (and / or the time since the last input was received increases above a threshold amount of time), and play / pause toggle control 6016 reappears, as shown in Figure 6H.

[0187] In FIG. 6I, input is detected at a location corresponding to play / pause toggle control 6016, as indicated by contact 6032.

[0188] 6I-6K, in response to an input detected at a location corresponding to play / pause toggle control 6016, the video resumes playing in video playback area 6002. As shown in FIG. 6J, sample frames of timeline 6004 include markings at locations within their respective images corresponding to kite 6028. For example, sample frame 6006 includes marking 6034 at a location corresponding to kite 6036. Annotation 6030 applied to kite object 6028 in video playback area 6002 has been applied to sample frames in timeline 6004 that include frames of video occurring before the point in time in the video at which annotation 6030 was received (e.g., frames to the left of scrub control 6008 in timeline 6004, such as sample frame 6006) and frames of video occurring after the point in time in the video at which annotation 6030 was received (e.g., frames to the right of scrub control 6008 in timeline 6004). In each sample frame of timeline 6004 in which kite object 6036 (corresponding to kite object 6028 to which annotation 6030 was applied in video playback area 6002) is visible, a marking (e.g., marking 6034 corresponding to annotation 6030) is shown in a position corresponding to the changing position of kite object 6036 in the sample frame. The marking is displayed at a size and orientation that is scaled to correspond to the change in size and orientation of kite object 6036 in the sample frame. As the video displayed in video playback area 6002 plays forward beyond the frame at which the video was stopped to receive annotation input later in the video, annotation 6030 is displayed at a size and orientation that is scaled to correspond to the change in size and orientation of kite object 6026, as shown in FIG. 6K. In this manner, annotations received in video playback area 6002 are applied to the object (e.g., so that the annotation tracks the object) as it moves, changes size (e.g., due to a changing distance from the camera), and changes orientation (e.g., in any direction in three-dimensional space).

[0189] In Figure 6L, input is detected at a location corresponding to timeline 6004, as indicated by contact 6038. As contact 6038 moves along the path indicated by arrows 6040 and 6044, the video displayed in video playback area 6002 rewinds, as shown in Figures 6L-6N. For example, as contact 6038 moves across timeline 6004, time indication 6042 associated with the currently displayed frame in video playback area 6002 decreases. When the video displayed in video playback area 6002 rewinds to the frame before the frame at which the video was stopped to receive the annotation input, annotation 6030 is displayed with a scaled size and orientation corresponding to the change in size and orientation of kite object 6026.

[0190] 7A-7BF show exemplary user interfaces for adding a virtual object to a previously captured media object, according to some embodiments. The user interfaces in those figures are used to illustrate processes described below, including the processes in FIGS. 9A-9F, 10A-10B, 11A-11F, and 12A-12D. For ease of explanation, some of the embodiments are discussed with reference to operations performed on a device having touch-sensitive display system 112. In such embodiments, the focus selector is optionally a respective finger or stylus contact, a representative point corresponding to a finger or stylus contact (e.g., the face center of each contact or a point associated with each contact), or the face center of two or more contacts detected on touch-sensitive display system 112. However, similar operations, along with a focus selector, are optionally performed on a device having display 450 and a separate touch-sensitive surface 451 in response to detecting contacts on touch-sensitive surface 451 while displaying the user interface shown in the figures on display 450.

[0191] 7A shows a user interface 7000 displayed by touchscreen display 112 of device 100, including a media object display area 7002 and a navigation area 7004. In some embodiments, user interface 7000 is accessed via a list of media content objects (e.g., in an image and / or video viewing application). Previously captured images are displayed in the media object display area. Information corresponding to the previously captured image (e.g., the location where the image was captured, "Cupertino") is displayed in information area 7003. Navigation area 7004 includes a previous media object control 7006 (e.g., for navigating to a previous media object), a current media object indicator 7008 (e.g., indicating the location of the previously captured image (represented by an enlarged dot) relative to other media objects (represented by non-enlarged dots) stored by device 100), and a next media object control 7010 (e.g., for navigating to a next media object). User interface 7000 includes controls 7012-7024 for adding various virtual objects to previously captured images displayed within media object display area 7002, as discussed further below with respect to Figures 7B-7BG. As described in accordance with various embodiments herein, the manner in which virtual objects are displayed relative to physical objects in previously captured images provides a user with an indication that depth data is stored for the previously captured images and that the virtual objects can interact with various surfaces in the various images. Interaction of the virtual objects with physical objects in the physical environment captured in the images provides a user with an indication of the presence of detectable surfaces in the images.

[0192] 7B-7L show how a virtual ball object interfaces with the surface of a physical object depicted in a previously captured image displayed in media object display area 7002. For example, the captured image is stored along with depth data that is used to determine the location of surfaces (e.g., horizontal and / or vertical planes) corresponding to the physical object within the physical environment captured in the image.

[0193] 7B , input by contact 7026 (e.g., tap input) is received at a location corresponding to ball control 7012 for adding a virtual ball to a previously captured image displayed within media object display area 7002. In FIG. 7C , in response to detecting input selecting ball control 7012, the mode of user interface 7000 is changed to a ball generation mode, as indicated by the "Ball" label in information area 7003 and as indicated by the changed visual state of ball control 7012. Input (e.g., tap input) is received at a location indicated by contact 7028. In response to detecting the contact, a virtual ball 7030 is added to a previously captured image displayed within media object display area 7002. For example, adding virtual ball 7030 to a previously captured image includes launching virtual ball 7030 with an upward trajectory along the path indicated by dotted line 7032 from the point where contact 7028 is detected (e.g., virtual ball 7030, under the influence of pseudo gravity, bounces off the surface of chair object 7036 in the previously captured image, bounces off the surface of table object 7038 in the previously captured image, falls onto floor surface 7040 in the previously captured image, and rolls along floor surface 7040).

[0194] 7D , input (e.g., a tap input) is received at the location indicated by contact 7042. In response to detecting the contact, a virtual ball 7044 is added to the previously captured image displayed within media object display area 7002, and virtual ball 7044 travels along path 7046 (e.g., bounces off floor 7040 and lands on floor 7040).

[0195] In FIG. 7E , several additional virtual balls have been added to the previously captured image displayed within media object display area 7002, some of which have landed and anchored on the surfaces of chair object 7036 and table object 7038. Input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to the subsequent media object control 7010, indicated by contact 7048. In response to the input, the display of the first previously captured image shown in FIG. 7E is replaced by a second previously captured image, as shown in FIG. 7F . The virtual balls added to the previously captured image displayed within media object display area 7002 are added to the second previously captured image, as shown in FIG. 7F (e.g., the virtual balls are animated as if under the influence of pseudo-gravity from floor surface 7040 in the first previously captured image of FIG. 7E to the top of the physical environment in the second previously captured image, as shown in FIGS. 7F-7I ). As the virtual ball falls, it settles on surfaces in the physical environment captured in the second, previously captured image, such as the surfaces of ramp 7052 and table 7054 and the arm grooves of people 7056 and 7058. In some embodiments, the virtual object is compared to depth data corresponding to physical objects in the previously captured image to determine the placement of the virtual object relative to the physical objects (e.g., to determine whether the physical object occludes the virtual object, or vice versa). For example, in FIG. 7G , virtual ball 7045 is partially occluded by physical table 7054.

[0196] In FIG. 7J , an input (e.g., a tap input) is detected at a location on the touchscreen display 112 corresponding to the subsequent media object control 7010, indicated by contact 7050. In response to the input, the display of the second previously captured image shown in FIG. 7J is replaced by a third previously captured image, as shown in FIG. 7K. The virtual ball added to the second previously captured image displayed within the media object display area 7002 is added to the first previously captured image, as shown in FIG. 7K (e.g., the virtual ball is animated as if under the influence of pseudo-gravity from a surface in the second previously captured image of FIG. 7J to the top of the physical environment in the third previously captured image, as shown in FIGS. 7K-7L). As the virtual ball falls, it sinks onto a surface in the physical environment captured in the third previously captured image, e.g., the surface of a sofa 7060.

[0197] In Figure 7L, input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to text insertion control 7014 indicated by contact 7062. In response to the input, the virtual balls (e.g., balls 7034 and 7044) disappear and text object 7064 is added to the third previously captured image, as shown in Figure 7M.

[0198] 7M-7T show how the virtual text object 7064 interfaces with the surface of a physical object depicted in a previously captured image displayed in the media object display area 7002.

[0199] In Figure 7M, a tap input by contact 7066 is detected at a location on touch screen 112 corresponding to text object 7064. From Figures 7M to 7P, contact 7066 moves along a path indicated by arrows 7068, 7070, and 7072. As contact 7066 moves, text object 7064 is "dragged" by contact 7066 such that the movement of text object 7064 corresponds to the movement of contact 7066. As text object 7064 is dragged to a position corresponding to sofa 7060, text object 7064 interacts with the detected surface of sofa 7060 by "marching" across the arms of sofa 7060, as shown in Figures 7N-7O. For example, in FIG. 7O, when text object 7064 is dragged to an area of ​​the third previously captured image corresponding to sofa 7060, a first portion of text object 7064 is adjacent to a first surface of sofa 7060 (e.g., above the arms of the sofa) and a second portion of text object 7064 is adjacent to a second surface of sofa 7060 (e.g., above the seat of sofa 7060).

[0200] In Figure 7Q, input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to text object 7064, as indicated by contact 7062. In response to the input, a text edit mode is initiated for text object 7064, as indicated by the display of cursor 7080 and keyboard 7078 in Figure 7R. In Figure 7S, the input provided via keyboard 7078 has changed the text of text object 7064 from the word "text" to the word "chills."

[0201] In Figure 7T, input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to decal insertion control 7016, as indicated by contact 7064. In response to the input, text object 7064 disappears and decal object 7084 is added to the third previously captured image, as shown in Figure 7U.

[0202] 7U-7Y show how the virtual decal object 7084 interfaces with the surface of a physical object depicted in a previously captured image displayed in the media object display area 7002.

[0203] In Figure 7U, a tap input by contact 7086 is detected at a location on touch screen 112 corresponding to decal object 7084. From Figure 7U to Figure 7X, contact 7086 moves along the path indicated by arrows 7088, 7090, and 7092. As contact 7086 moves, decal object 7084 is "dragged" by contact 7086 such that the movement of decal object 7084 corresponds to the movement of contact 7086. When decal object 7084 is dragged onto the surface of sofa 7060, decal object 7084 conforms to the detected horizontal and vertical surfaces of sofa 7060 and floor 7094, as shown in Figures 7U-7X. For example, in Figure 7V, when decal object 7084 is dragged onto sofa 7060, a first portion of decal object 7084 is adjacent to a first surface of sofa 7060 (flat on the seat of the sofa), and a second portion of decal object 7084 is adjacent to a second surface of sofa 7060 (e.g., dragged in front of sofa 7060). In Figure 7X, when decal object 7084 is dragged onto floor 7094, decal object 7084 is partially occluded by sofa 7060.

[0204] In Figure 7Y, an input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to emoji insertion control 7018 as indicated by contact 7096. In response to the input, decal object 7084 disappears and emoji object 7098 is added to the third previously captured image, as shown in Figure 7Z.

[0205] 7Z-7AE show how the virtual emoji object 7098 interfaces with the surface of a physical object depicted in a previously captured image displayed in the media object display area 7002.

[0206] In Figure 7AA, a tap input by contact 7100 is detected at a location on touchscreen 112 corresponding to emoji object 7098. In Figures 7AA and 7AB, contact 7100 moves along a path indicated by arrow 7102. As contact 7100 moves, contact 7100 "drags" emoji object 7098 such that the movement of emoji object 7098 corresponds to the movement of contact 7100. In Figure 7AC, contact 7100 lifts off of touchscreen display 112, while emoji object 7098 is suspended in space above surfaces (e.g., couch 7060 and floor 7094) in the third previously captured image. In response to the liftoff of contact 7100, emoji object 7098 descends under the influence of pseudo-gravity, as shown in Figures 7AC-7AE. In Figure 7AC, glyph object 7098 encounters physical object sofa 7060, causing the glyph object's orientation to change as glyph object 7098 rolls up the arm of sofa 7060 and continues to descend to floor 7094, as shown in Figure 7AD.

[0207] In Figure 7AF, an input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to confetti insertion control 7020, as indicated by contact 7104. In response to the input, emoji object 7098 disappears and a confetti object (e.g., confetti object 7106) is added to the third previously captured image, as shown in Figure 7AG.

[0208] 7AG-7AT show how virtual confetti objects interface with the surfaces of physical objects depicted in previously captured images and videos displayed within media object display area 7002. In FIG.

[0209] In Figures 7AG-7AJ, virtual confetti objects are continuously added (e.g., as shown in Figures 7AH, 7AI, and 7AJ) and, under the influence of pseudo-gravity, collect on detected surfaces (e.g., substantially horizontal surfaces, such as sofa 5060 and horizontal surfaces on floor 7094) in the third previously captured image.

[0210] 7AJ, an input (e.g., a tap input) is detected at a location on touch screen display 112 corresponding to the previous media object control 5006 as indicated by contact 7104. In response to the input, the display of the third previously captured image shown in FIG. 7AJ is replaced by a display of the second previously captured image, as shown in FIG. 7AK. The confetti added to the third previously captured image displayed within media object display area 7002 is added to the second previously captured image as shown in FIG. 7AK (e.g., displayed in the same location where the confetti appeared in FIG. 7AJ).

[0211] In Figures 7AK-7AL, virtual confetti objects are continually added and, under the influence of pseudo-gravity, gather on detected surfaces in the second previously captured image.

[0212] 7AJ, an input (e.g., a tap input) is detected at a location on touch screen display 112 corresponding to the previous media object control 5006 as indicated by contact 7104. In response to the input, the display of the third previously captured image shown in FIG. 7AJ is replaced by a display of the second previously captured image, as shown in FIG. 7AK. The confetti added to the third previously captured image displayed within media object display area 7002 is added to the second previously captured image as shown in FIG. 7AK (e.g., displayed in the same location where the confetti appeared in FIG. 7AJ).

[0213] In Figures 7AK-7AL, virtual confetti objects are continually added and, under the influence of pseudo-gravity, gather on detected surfaces in the second previously captured image.

[0214] In Figure 7AL, an input (e.g., a tap input) is detected at a location on touch screen display 112 corresponding to subsequent media object control 7010 as indicated by contact 7110 (e.g., multiple taps are received to advance the currently displayed media object twice). In response to the input, the display of the second previously captured image shown in Figure 7AL is replaced by a display of the previously captured video, as shown in Figure 7AM. The confetti added to the second previously captured image displayed within media object display area 7002 is added to the video as shown in Figure 7AM (e.g., displayed in the same location where the confetti was displayed in Figure 7AL).

[0215] 7AM-7AT, virtual confetti objects are continually added and, under the influence of pseudo gravity, collect at detected surfaces in the video (e.g., the edges of kite object 7112). For example, as video playback progresses through Figures 7AN, 7AO, and 7AP, virtual confetti objects are continually added and, under the influence of pseudo gravity, collect at the edges of kite object 7112 and the bottom edge of the video frame.

[0216] In Figure 7AP, an input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to playback control 7114 as indicated by contact 7116. In response to the input, playback of the video is repeated. The confetti added to the video of Figures 7AM-7AT is added to the video, at which point playback begins in Figure 7AQ. For example, when playback of the video begins, the confetti accumulated on kite object 7112 falls from the position shown in Figure 7AP, and kite object 7112 is shown in a different position within the video of Figure AQ.

[0217] 7AQ-7AT, virtual confetti objects are continually added and, under the influence of pseudo gravity, collect at detected surfaces in the video (e.g., the edges of kite object 7112). For example, as video playback progresses through Figures 7AR, 7AS, and 7AT, virtual confetti objects are continually added and, under the influence of pseudo gravity, collect at the edges of kite object 7112 and the bottom edge of the video frame.

[0218] In some embodiments, the displayed virtual confetti objects are gradually made invisible and / or ceased to be displayed (e.g., when the elapsed time since the virtual confetti objects were displayed exceeds a threshold time).

[0219] 7AU-7AX show how a virtual spotlight object 7118 interfaces with physical objects depicted in a previously captured image displayed in the media object display area 7002.

[0220] In FIG. 7AU, a second previously captured image is displayed within the media object display area 7002, and spotlight mode has been activated (e.g., in response to input received at spotlight control 7022). In spotlight mode, a spotlight virtual object 7118 is shown illuminating a portion of the image (e.g., person 7058) and an area of ​​the image beyond the spotlight virtual object 7118. In this manner, the spotlight virtual object 7118 allows for attention to be focused on a particular portion of the image. In some embodiments, the first physical object to be illuminated is selected automatically (e.g., based on a determination of the closest physical object in the foreground). The spotlight virtual object 7118 includes a simulated light beam 7122 and a simulated illumination spot 7124 that illuminates a portion of the floor in the previously captured image. In some embodiments, the light beam 7122 illuminates at least a portion of a representation of the physical object in the previously captured image. In some embodiments, the illumination spot 7124 illuminates a portion of the image that corresponds to a horizontal plane detected in the image, such as the floor 7124.

[0221] In Figure 7AV, a tap input by contact 7128 is detected at a location on touch screen 112 corresponding to spotlight object 7118. In Figures 7AV and 7AW, contact 7128 is moving along the path indicated by arrow 7102. As contact 7128 moves, spotlight object 7118 is "dragged" by contact 7128 such that the movement of spotlight object 7118 corresponds to the movement of contact 7128. In Figure 7AW, the position of spotlight object 7118 has shifted so that person 7056 is illuminated by spotlight object 7118. As spotlight object 7118 moves, the size of illumination spot 7124 has changed.

[0222] 7AX, an input (e.g., a tap input) is detected at a location on touchscreen display 112 corresponding to measurement control 7020 indicated by contact 7026. In response to the input, spotlight object 7118 disappears.

[0223] In Figure 7AY, input is detected at a location on touchscreen display 112 indicated by contacts 7132 and 7134. In Figure 7AZ, in response to detecting contacts 7132 and 7134, virtual measurement object 7136 is displayed at a location that spans the distance between contacts 7132 and 7134 (e.g., corresponding to the height of person 7058). Measurement indicator 7138 indicates that the distance between points in the physical environment captured in the image corresponding to contacts 7132 and 7134 is 1.8 m (e.g., determined using depth data stored with previously captured images).

[0224] 7AZ through 7BA, contact 7132 moves along the path indicated by arrow 7140. As contact 7132 moves, the size of virtual measurement object 7136 adjusts to span the adjusted distance between contacts 7132 and 7134, and the measurement indicated by measurement indicator 7138 is updated.

[0225] In Figure 7BB, contacts 7132 and 7134 have lifted off the touchscreen 112. Virtual measurement object 7136 and measurement indicator 7138 remain displayed. Input is detected at a location on the touchscreen display 112 indicated by contact 7142. From Figures 7BB to 7BC, contact 7142 moves along the path indicated by arrows 7144 and 7146. In response to movement of contact 7142 (e.g., beyond a threshold amount of movement), virtual measurement object 7148 and measurement indicator 7150 are displayed. From Figures 7BC to 7BD, as contact 7142 continues to move, the size of virtual measurement object 7148 is adjusted and the measurement indicated by measurement indicator 7150 is updated. The dotted portion of virtual measurement object 7148 indicates a portion of virtual measurement object 7148 passing through a physical object (e.g., person 7056).

[0226] In Figure 7BE, contact 7142 lifts off the touchscreen 112. The virtual measurement object 7148 and measurement indicator 7148 remain displayed. In Figure 7BF, an input is detected at a location on the touchscreen display 112 indicated by contact 7152. In response to the input, an end of the virtual measurement object 7148 (e.g., the end closest to the received input) moves to a position corresponding to the location of contact 7152.

[0227] 8A-8W show exemplary user interfaces for initiating a shared annotation session, according to some embodiments. The user interfaces in those figures are used to illustrate processes described below, including those in FIGS. 9A-9F, 10A-10B, 11A-11F, and 12A-12D. For ease of explanation, some of the embodiments are discussed with reference to operations performed on a device having touch-sensitive display system 112. In such embodiments, the focus selector is optionally a respective finger or stylus contact, a representative point corresponding to a finger or stylus contact (e.g., the centroid of each contact or a point associated with each contact), or the centroid of two or more contacts detected on touch-sensitive display system 112. However, similar operations, along with a focus selector, are optionally performed on a device having display 450 and separate touch-sensitive surface 451 in response to detecting contacts on touch-sensitive surface 451 while displaying the user interface shown in the figures on display 450.

[0228] 8A-8G illustrate the establishment of a shared annotation session between two devices.

[0229] 8A shows a physical environment 8000 in which a first user operates a first device 100-1 (e.g., device 100) and a second user operates a second device 100-2 (e.g., device 100). A collaborative user interface 8002 displayed by device 100-1 is shown in inset 8004 corresponding to device 100-1. Inset 8006 shows a web browser user interface currently displayed by device 100-2. A prompt 8008 displayed within collaborative user interface 8002 includes instructions to initiate a shared annotation session.

[0230] In Figure 8B, input by contact 8012 is received at a location corresponding to control 8010 displayed by device 100-1 to initiate a shared annotation session. In response to the input, a request is sent from device 100-1 to a remote device (e.g., device 100-2) to initiate a shared annotation session. After the request is sent and a response indicating acceptance of the request has not been received, notification 8014 is displayed by device 100-1, as shown in Figure 8C.

[0231] 8C , in response to receiving a request to initiate a shared annotation session, device 100-2 displays prompt 8016 including instructions to accept the request for the shared annotation session. Input by contact 8020 is received at a location corresponding to control 8018 displayed by device 100-2 to accept the request for the shared annotation session. In response to the input, an acceptance of the request is transmitted from device 100-2 to a remote device (e.g., device 100-1).

[0232] 8D , an indication of acceptance of the request to initiate a shared annotation session has been received by device 100-1. Device 100-1 displays prompt 8022 including instructions to move device 100-1 toward device 100-2. Prompt 8022 includes a representation 8026 of device 100-1 and a representation 8028 of device 100-2. Device 100-2 displays prompt 8024 including instructions to move device 100-2 toward device 100-1. Prompt 8024 includes a representation 8030 of device 100-1 and a representation 8032 of device 100-2.

[0233] 8D-8E show animations displayed in prompts 8022 and 8024. In prompt 8022, representation 8026 of device 100-1 is animated to move toward representation 8028 of device 100-2. In prompt 8024, representation 8032 of device 100-2 is animated to move toward representation 8030 of device 100-1.

[0234] 8F , the connection criteria are met (e.g., first device 100-1 and second device 100-2 are moving toward each other and / or at least a portion of physical space 8000 captured within the field of view of one or more cameras of first device 100-1 corresponds to at least a portion of physical space 8000 captured within the field of view of one or more cameras of device 100-2). Notification 8034 displayed by first device 100-1 and notification 8036 displayed by second device 100-2 each include an indication that a shared annotation session has begun. First device 100-1 displays (e.g., overlaid by notification 8034) a representation of the field of view of one or more cameras of first device 100-1. Poster 8038a in physical environment 8000 is visible within the field of view of one or more cameras of first device 100-1, as indicated by representation 8038b of poster 8038a displayed by first device 100-1. Second device 100-2 displays a representation of the field of view of one or more cameras of second device 100-2 (e.g., overlaid by notification 8036). Devices 100-1 and 100-2 display a shared field of view (e.g., at least a portion of physical space 8000 captured within the field of view of one or more cameras of first device 100-1 corresponds to at least a portion of physical space 8000 captured within the field of view of one or more cameras of device 100-2). For example, poster 8038a in physical environment 8000 is visible within the field of view of one or more cameras of second device 100-2, as shown by representation 8038c of poster 8038a displayed by second device 100-2. In Figure 8G, the respective fields of view of the cameras are displayed by first device 100-1 and second device 100-2 without notifications 8034 and 8036 (e.g., notifications 8034 and 8036 are no longer displayed).

[0235] Figures 8H-8M show annotation input received during a shared annotation session, Figures 8H-8J show annotation input provided to second device 100-2, and Figures 8K-8M show annotation input provided to first device 100-1.

[0236] In FIG. 8H , input via contact 8044 (e.g., input received on a touchscreen display of second device 100-2) is detected by second device 100-2. While the input is being detected by second device 100-2, first device 100-1 displays avatar 8048 at a location within the shared field of view that corresponds to the location within the shared field of view at which the input is received at second device 100-2. As contact 8044 moves along the path indicated by arrow 8046 as shown in FIGS. 8H-8I , annotations corresponding to the movement of contact 8044 are displayed by first device 100-1 (as annotation 8050-1) and second device 100-2 (as annotation 8050-2). In FIG. 8J , further annotation input is provided via further movement of contact 8044.

[0237] In Figure 8K, input by contact 8052 (e.g., input received on a touchscreen display of first device 100-1) is detected by first device 100-1. While the input is being detected by first device 100-1, second device 100-2 displays avatar 8054 at a location within the shared field of view that corresponds to the location within the shared field of view where the input was received at first device 100-1. As contact 8052 moves, annotations corresponding to the movement of contact 8052 are displayed by first device 100-1 (as annotation 8056-1) and second device 100-2 (as annotation 8056-2), as shown in Figures 8K-8M.

[0238] 8M-8P, the movement of first device 100-1 increases the distance between first device 100-1 and second device 100-2. In FIG. 8N, as first device 100-1 moves away from second device 100-2, the representation of the field of view of the camera of first device 100-1 displayed by first device 100-1 is adjusted (e.g., such that the portion of representation 8038b of physical poster 8038a displayed by first device 100-1 is reduced).

[0239] In some embodiments, one or more annotations (e.g., 8050-1, 8050-2, 8056-1, and / or 8056-2) have a fixed spatial relationship with respect to a portion of physical environment 8000 (e.g., such that movement of a device camera relative to physical environment 8000 changes the display position of the annotations). In Figure 8O, as first device 100-1 continues to move away from second device 100-2 such that annotation 8056-1 is no longer displayed by first device 100-1, visual indication 8058 corresponding to annotation 8056-1 is displayed by first device 100-1 (e.g., to indicate the direction of movement of first device 100-1 necessary to re-display annotation 8056-1).

[0240] 8P, movement of first device 100-1 away from second device 100-2 increases the distance between first device 100-1 and second device 100-2 to the point where the shared annotation session is broken (e.g., above a threshold distance). Device 100-1 displays prompt 8060 including instructions to move device 100-1 toward device 100-2. Device 100-2 displays prompt 8062 including instructions to move device 100-2 toward device 100-1. In some embodiments, prompt 8060 includes an animated element (e.g., as described with respect to FIGS. 8D-8E).

[0241] From Figure 8P to Figure 8Q, the movement of first device 100-1 decreases the distance between first device 100-1 and second device 100-2. In Figure 8Q, the distance between first device 100-1 and second device 100-2 has decreased sufficiently for the shared annotation session to be restored. When devices 100-1 and 100-2 finish displaying their respective prompts 8060 and 8062, the respective fields of view of their respective device cameras are redisplayed.

[0242] 8R-8W illustrate a gaming application using a shared session between first device 100-1 and second device 100-2 (eg, a shared session established as described above with respect to FIGS. 8A-8G).

[0243] 8R shows a physical environment 8068 in which a first user 8064a operates a first device 100-1 (e.g., device 100) and a second user 8066a operates a second device 100-2 (e.g., device 100). A game application user interface 8063-1 displayed by the first device 100-1 is shown in inset 8070, which corresponds to the first device 100-1. A game application user interface 8063-2 displayed by the second device 100-2 is shown in inset 8072, which corresponds to the second device 100-2. User 8064a faces user 8066a, as shown in user interface 8063-1, such that a representation 8066b of user 8066a is visible within the field of view of one or more cameras (e.g., rear-facing cameras) of device 100-1. Representation 8064b of user 8064a is similarly visible within the field of view of one or more cameras (eg, rear-facing cameras) of device 100-2, as shown in user interface 8063-2.

[0244] Game application user interfaces 8063-1 and 8063-2 display basketball hoops 8074 and 8076, respectively. By providing input within the respective game application user interfaces, users activate a virtual basketball object in the displayed representation of the respective device camera's field of view to create a basket within the respective basketball hoop. In some embodiments, the respective basketball hoops are fixed to the spatial location of the respective device so that the respective users can move their devices to create challenges to opponents. Game application user interfaces 8063-1 and 8063-2 also display game data areas 8078 and 8080, respectively. Game data displayed in the game data areas may include, for example, a record of points scored by successful baskets, as well as the distance between devices 100-1 and 100-2 (e.g., for use as a basis for assigning a score to a given shot).

[0245] In Figure 8S, input (e.g., tap input) by contact 8082 is detected by device 100-2 to activate a virtual basketball object. In response to detecting the input by contact 8082, virtual basketball object 8084 is added to the field of view of the camera of device 100-2 displayed within game application user interface 8063-2, as shown in Figure 8T. Figure 8T further shows input (e.g., tap input) by contact 8086 detected by device 100-2 to activate the virtual basketball object. In response to detecting the input by contact 8086, virtual basketball object 8088 is added to the field of view of the camera of device 100-1 displayed within game application user interface 8063-1, as shown in Figure 8V. Figure 8V shows input (e.g., tap input) by contact 8090 detected by device 100-2 to activate the virtual basketball object. In response to detecting input by contact 8090, a virtual basketball object 8092 is added to the field of view of the camera of device 100-2 displayed within game application user interface 8063-2, as shown in Figure 8W. From Figure 8V to Figure 8W, user 8064a lowers device 100-1 such that the displayed position of ring 8076 and representation 8064b of user 8064 is changed within user interface 8063-2.

[0246] 9A-9F are flow diagrams illustrating a method 900 for repositioning an annotation, according to some embodiments. Method 900 is performed on an electronic device (e.g., device 300 of FIG. 3 or portable multifunction device 100 of FIG. 1A ) having a display generating element (e.g., a display, a projector, a head-up display, etc.), one or more input devices (e.g., a touch-sensitive surface such as a touch-sensitive remote control, or a touchscreen display that also serves as a display generating element, a mouse, a joystick, a wand controller, and / or a camera that tracks the position of one or more features of a user, such as the user's hands), and one or more cameras (e.g., one or more rear-facing cameras on a side of the device opposite the display and touch-sensitive surface). In some embodiments, the display is a touchscreen display, and the touch-sensitive surface is on or integrated into the display. In some embodiments, the display is separate from the touch-sensitive surface. Some operations of method 900 are optionally combined, and / or the order of some operations is optionally changed.

[0247] The device, via a display generation element, displays (902) a first user interface region (e.g., user interface 5003) that includes a representation of the field of view of one or more cameras that is updated with changes in the field of view of the one or more cameras over time (e.g., the representation of the field of view is continuously updated (e.g., with a preset frame rate, such as 24, 48, or 60 fps) according to changes occurring in the physical environment around the cameras and according to movement of the cameras relative to the physical environment). For example, as shown in FIGS. 5A-5B, the view of the physical environment 5000 within the field of view of one or more cameras is updated according to changes in the position of the cameras of device 100 as device 100 is moved.

[0248] While displaying a first user interface region including a representation of the field of view of one or more cameras, the device receives (904) a first request to add an annotation (e.g., text or a drawing generated and / or placed by moving a contact (e.g., a finger or stylus contact) on the touch-sensitive surface) to the displayed representation of the field of view of the one or more cameras via one or more input devices (e.g., the first request is a contact input detected within the first user interface region on the touchscreen display (e.g., at a location corresponding to a control for initiating an annotation, or a location within the first user interface region (e.g., the location where the annotation is to be initiated)). For example, the request to add an annotation is a stylus 5012 input (e.g., an input to initiate an annotation as described with reference to FIG. 5C and / or a contact input at a location corresponding to a control (e.g., control 5004 for toggling between still and video mode) as described with reference to FIG. 5H).

[0249] In response to a first request to add an annotation to the displayed representation of the field of view of the one or more cameras (e.g., including in response to detecting a finger contact or touchdown or movement of a stylus on the touch-sensitive surface at a location corresponding to a portion of the physical environment captured in the representation of the camera's field of view, or a user interface object (e.g., a button that activates AR annotation mode)), the device replaces (906) the representation of the field of view of the one or more cameras in the first user interface area with a still image of the field of view of the one or more cameras captured at a time corresponding to receiving the first request to add the annotation (e.g., pausing a live feed of the field of view of the one or more cameras (e.g., displaying a still image of the current field of view) while the field of view continues to change with movement of the device) and displays in the first user interface area a still image corresponding to the paused live feed of the field of view of the one or more cameras. For example, in response to input by stylus 5012 (e.g., as described with respect to FIG. 5C), the representation of the device camera's field of view (e.g., as described with respect to FIGS. 5A-5B) is replaced by a display of a still image of one or more camera's field of view captured at a time corresponding to receipt of a first request to add an annotation (e.g., as described with respect to FIGS. 5C-5D).

[0250] While displaying the still image within the first user interface region, the device receives (908) a first annotation (e.g., drawing input) via one or more input devices on a first portion of the still image, the first portion of the still image corresponding to a first portion of the physical environment captured in the still image. For example, annotation 5018 is received on a portion of the still image (e.g., a portion including representation 5002b of physical mug 5002a) that corresponds to a portion of the physical environment captured in the still image (e.g., a portion of physical environment 5000 including physical mug 5002a), as described with respect to FIGS. 5D-5G . In some embodiments, while displaying the still image and receiving annotation input on the still image, the device continues to track the position of the camera relative to the surrounding physical environment (e.g., based on changes in the camera's field of view and input from other sensors (e.g., motion sensors, gravity sensors, gyroscopic sensors, etc.)). In some embodiments, the device determines whether the physical location or object corresponding to the annotation's location in the still image has moved out of the camera's field of view, and if so, the device also determines the spatial relationship between the portion of the physical environment currently in the camera's field of view and the physical location or object that is the target of the annotation. For example, mug 5002 is the target of annotation 5018.

[0251] While displaying the first annotation on the first portion of the still image in the first user interface region (e.g., after receiving the first annotation on the first portion of the still image), the device receives a first request via one or more input devices to redisplay representations of the fields of view of the one or more cameras in the first user interface region (910). For example, the request to redisplay representations of the fields of view of the one or more cameras in the first user interface region is input by a stylus 5012 on a control 5004 for toggling between still image mode and video mode, as described with respect to FIG. 5H.

[0252] In response to receiving a first request to re-display a representation of the field of view of the one or more cameras in the first user interface region (e.g., in response to detecting a lack of contact on the touch-sensitive surface for a threshold amount of time (e.g., a drawing session is deemed to have ended) or in response to detecting a tap on a user interface object (e.g., a button for deactivating AR annotation mode)), the device replaces the display of the still image with a representation of the field of view of the one or more cameras in the first user interface region (910) (e.g., the representation of the field of view is continuously updated at a preset frame rate (e.g., 24, 48, or 60 fps, etc.) in accordance with changes occurring in the physical environment around the camera and in accordance with movement of the camera relative to the physical environment). For example, as described with respect to FIG. 5H, in response to input with stylus 5012 on control 5004, the still image displayed in FIGS. 5C-5H is replaced with a representation of the field of view of the one or more cameras in first user interface region 5003, as described with respect to FIGS. 5I-5N. Pursuant to a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is currently outside the field of view of one or more cameras (e.g., as a result of movement of the device that occurred after the live feed of the camera views was stopped), the device displays (e.g., as part of the computing system) an indication of the current spatial relationship of the one or more cameras with respect to the first portion of the physical environment captured in the still image, concurrently with a representation of the field of view of the one or more cameras (e.g., displaying a visual indication such as a dot or other shape on the edge of the displayed camera's field of view and at the location of the edge nearest the annotated first portion of the physical environment, or displaying a simplified map of the physical environment concurrently with the representation of the cameras' field of view and marking the relative position of the first portion of the physical environment and the device on the map).For example, as described with respect to FIG. 5L , a portion of the physical environment captured in the still image (e.g., including mug 5002 relative to which annotation 5018 was received) is currently outside the field of view of one or more cameras (e.g., in FIGS. 5J-5L , physical mug 5002 a is outside the field of view of the camera displayed in user interface 5003), and an indication of the current spatial relationship of the one or more cameras relative to a first portion of the physical environment captured in the still image (e.g., indicator dots 5022) is displayed concurrently with a representation of the field of view of the one or more cameras (e.g., as shown in FIG. 5L ). Following a determination that the first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is currently within the field of view of one or more cameras, the device refrains from displaying the indication. For example, in response to an input to redisplay the camera's field of view (e.g., an input by stylus 5012 on control 5004 for toggling between still image mode and video mode, as described with respect to FIG. 5T), in response to a determination that a portion of the physical environment captured in the still image (e.g., including mug 5002 for received annotations 5018 and 5028) is currently within the field of view of one or more cameras (e.g., in FIG. 5T, physical mug 5002a is visible (as visual representation 5002b) within the field of view of the camera displayed in user interface 5003), an indication of the current spatial relationship of the one or more cameras to a first portion of the physical environment captured in the still image is not displayed (e.g., as shown in FIG. 5T). Displaying an indication of the current spatial relationship of one or more cameras to the portion of the physical environment captured in the still image pursuant to a determination that the portion of the physical environment captured in the still image is currently outside the field of view of one or more cameras provides improved visual feedback to the user (e.g., indicating that movement of the camera is necessary to view the physical environment captured in the still image).Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotated portion of the still image), makes the user-device interface more efficient, and also reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0253] In some embodiments, displaying (912) an indication of the current spatial relationship of the one or more cameras with respect to a first portion of the physical environment captured in the still image includes displaying an indicator proximate an edge of a representation of the field of view of the one or more cameras and moving the indicator along the edge according to movement of the one or more cameras with respect to the physical environment. For example, as described with respect to FIG. 5L, indicator 5022 is displayed proximate a left edge of the field of view of the camera displayed in user interface 5003, and as described with respect to FIGS. 5L-5M, the indicator is moved along the edge according to movement of the camera of device 100. In some embodiments, in the case of a rectangular representation of the field of view of the one or more cameras, the indicator is a visual indication such as a dot or other shape that moves along an edge of the rectangular representation of the field of view, where the visual indication can move along one straight edge according to a first movement of the one or more cameras and the visual indication may also hop from one straight edge to another straight edge according to a second movement of the one or more cameras. Moving the indicator along the edge of the camera view provides visual feedback to the user (e.g., indicating the direction of camera movement required to view the portion of the physical environment captured in the still image). Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotated portion of the still image), makes the user-device interface more efficient, and additionally reduces device power usage and improves battery life by allowing the user to use the device more quickly and efficiently.

[0254] In some embodiments, while displaying an indication of the current spatial relationship of the one or more cameras with respect to a first portion of the physical environment captured in the still image, the device (e.g., as part of a computing system) detects a first movement of the one or more cameras (914). In response to detecting the first movement of the one or more cameras, the device updates the representation of the field of view of the one or more cameras according to a change in the field of view of the one or more cameras caused by the first movement. In response to determining that the first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including an object for which the first annotation was received) is currently outside the field of view of the one or more cameras, the device (e.g., as part of a computing system) updates the indication of the current spatial relationship of the one or more cameras with respect to the first portion of the physical environment captured in the still image according to the first movement of the one or more cameras (e.g., moving a visual indication, such as a dot or other shape along an edge of the representation of the field of view according to the first movement of the cameras). Pursuant to a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is now within the field of view of one or more cameras, the device ceases displaying the indication. For example, as the camera of device 100 moves, indication 5022 is updated (e.g., moved upward in user interface 5003) while the portion of the physical environment captured in the still image is outside the camera's field of view (e.g., mug 5002a is outside the camera's field of view, as described with respect to FIGS. 5L-5M), and indication 5022 is no longer displayed when the portion of the physical environment captured in the still image is within the camera's field of view (e.g., as described with respect to FIGS. 5L-5N). In some embodiments, a scaled-down representation of the still image with the first annotation (e.g., scaled-down representation 5020) is displayed adjacent to a position on an edge of the representation of the field of view where the visual indicator was last displayed, and the still image view scales down and moves toward the position of the first annotation shown in the representation of the camera's field of view).In some embodiments, the indication re-appears when the first portion of the physical environment moves out of the camera's field of view with further movement of the camera relative to the physical environment (e.g., as described with respect to FIGS. 5U-5V). Terminating the display of the indication of the one or more cameras' current spatial relationship to the portion of the physical environment pursuant to determining that the first portion of the physical environment captured in the still image is currently within the one or more cameras' field of view provides visual feedback to the user (e.g., indicating that no further movement is required to view the portion of the physical environment captured in the still image). Providing the user with improved visual feedback enhances device usability (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotation in the still image), provides a more efficient user-device interface, and reduces device power usage and improves battery life by allowing the user to use the device more quickly and efficiently.

[0255] In some embodiments, in response to receiving a first request to redisplay the representation of the field of view of one or more cameras in the first user interface area (916), the device displays a first annotation on the first portion of the physical environment captured in the representation of the field of view of one or more cameras in accordance with a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is currently within the field of view of one or more cameras. For example, as described with respect to FIG. 5N , the device displays annotation 5018 on the portion of the physical environment captured in the representation of the field of view of one or more cameras (e.g., annotation 5018 is displayed at a position corresponding to visual representation 5002b of physical mug 5002a in accordance with a determination that a first portion of the physical environment captured in the still image (e.g., a portion of physical environment 5000 including physical mug 5002a) is currently within the field of view of one or more cameras). Displaying the annotation of the portion of the physical environment captured in the annotated still image provides visual feedback to the user (e.g., indicating that the portion of the physical environment captured in the still image is currently within the field of view of one or more cameras). Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotated portion of the still image), makes the user-device interface more efficient, and also reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0256] In some embodiments, the first annotation is displayed (918) as a two-dimensional object (e.g., annotation 5018) on a first depth plane within a first portion of the physical environment captured within a representation of the field of view of one or more cameras. In some embodiments, the first depth plane is detected according to detecting a physical object (e.g., physical mug 5002a) or object feature at the first depth plane in the first portion of the physical environment. Displaying the annotation on a depth plane within the portion of the physical environment within the field of view of one or more cameras provides improved visual feedback to the user (e.g., indicating that the annotation has a fixed spatial relationship to the physical environment). Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to imbue physical world objects with additional information contained in the annotation), makes the user device interface more efficient, and further reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0257] In some embodiments, a first annotation (e.g., annotation 5018) is displayed (920) at a location in space within a first portion of the physical environment captured within a representation of the field of view of one or more cameras. In some embodiments, the first annotation floats in space separate from any physical objects detected in the first portion of the physical environment. Displaying the annotation at a location in space within a portion of the physical environment captured in a still image provides improved visual feedback to the user (e.g., indicating that the annotation has a fixed spatial relationship to the physical environment). Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to imbue physical world objects with additional information contained in the annotation), makes the user device interface more efficient, and further reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0258] In some embodiments, a first annotation (e.g., annotation 5018) is displayed (922) at a location on a physical object (e.g., physical mug 5002a) detected in a first portion of the physical environment captured within a representation of the field of view of one or more cameras. In some embodiments, the first annotation is attached to a physical object (or a feature of a physical object) detected in the first portion of the physical environment. Displaying the annotation of a portion of the physical environment at a location on a physical object detected in the portion of the physical environment captured in the still image provides improved visual feedback to the user (e.g., indicating that the annotation has a fixed spatial relationship to the physical environment). Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to imbue physical world objects with additional information contained in the annotation), makes the user device interface more efficient, and further reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0259] In some embodiments, in response to receiving (924) a first request to re-display a representation of the field of view of the one or more cameras in the first user interface area, in accordance with a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is not currently within the field of view of the one or more cameras, the device displays a visual representation of the first annotation (e.g., drawing input) drawn on the first portion of the still image simultaneously with the representation of the field of view of the one or more cameras in the first user interface area (e.g., a miniature representation of the still image with the first annotation displayed adjacent to a position on an edge of the representation of the field of view that is closest to the first portion of the physical space currently represented within the field of view of the camera). For example, in response to a request to redisplay a representation of the field of view of one or more cameras in the first user interface area (e.g., input by stylus 5012 on control 5004 for toggling between still image mode and video mode, as described with reference to FIG. 5H), and following a determination that a first portion of the physical environment captured in the still image is currently within the field of view of one or more cameras (e.g., a portion of the physical environment captured in the still image (e.g., including mug 5002 for received annotation 5018) is currently within the field of view of one or more cameras), the device displays annotation 5018 drawn on the first portion of the still image simultaneously with the representation of the field of view of the one or more cameras in the first user interface area (e.g., as described with reference to FIG. 5N). In some embodiments, in response to receiving a first request to re-display a representation of the field of view of one or more cameras in the first user interface area, in accordance with a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object on which the first annotation was received) is currently outside the field of view of the one or more cameras, the computing system refrains from displaying a visual representation of the first annotation drawn on the first portion of the still image (e.g., annotation 5018 is not displayed, as shown in Figures 5L and 5M).Displaying a visual representation of the annotation of the still image simultaneously with the representation of the camera's field of view pursuant to determining that the first portion of the physical environment is not within the camera's field of view provides improved visual feedback to the user (e.g., an indication of the current spatial relationship (e.g., a dot) corresponding to the received annotation input). Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotated portion of the still image), makes the user-device interface more efficient, and additionally reduces device power usage and improves battery life by allowing the user to use the device more quickly and efficiently.

[0260] In some embodiments, the first annotation shown in the representation of the field of view has a first perspective based on the current spatial relationship of one or more cameras with respect to a first portion of the physical environment captured in the still image (926), which differs from a second perspective of the first annotation displayed on the still image (e.g., a perspective of the first annotation shown in a scaled-down representation of the still image displayed adjacent to the representation of the field of view). In some embodiments, an animated transition is displayed showing the scaled-down representation of the still image being transformed into a representation of the current field of view. Displaying the annotation at a different perspective from that shown in the still image provides improved visual feedback to the user (e.g., indicating that the annotation is anchored to a portion of the physical environment captured in the still image). FIGS. 5AE and 5AF provide an example of an annotation 5018 shown in the representation of the field of view from a different perspective. Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to quickly and accurately locate a portion of the physical environment that corresponds to the annotation portion of the still image), provides a more efficient user-device interface, and reduces device power usage and improves battery life by allowing the user to use the device more quickly and efficiently.

[0261] In some embodiments, in response to receiving a first request to re-display the representations of the fields of view of the one or more cameras in the first user interface region (928), in accordance with a determination that a first portion of the physical environment captured in the still image (e.g., a portion of the physical environment including the object for which the first annotation was received) is currently outside the field of view of the one or more cameras, a visual representation of a first annotation drawn on the first portion of the still image (e.g., drawing input) is displayed concurrently with the representations of the fields of view of the one or more cameras in the first user interface region (e.g., a scaled-down representation of the still image with the first annotation is displayed adjacent to a position on an edge of the representation of the field of view that is closest to the first portion of the physical space currently displayed in the field of view of the camera), and the visual representation of the first annotation drawn on the first portion of the still image (e.g., scaled-down representation 5020) is converted into an indication (e.g., indication 5022) of the current spatial relationship of the one or more cameras (e.g., as part of a computing system) with respect to the first portion of the physical environment captured in the still image (e.g., as described with respect to Figures 5I-5L). For example, the indication may be a visual indication such as a dot or other shape displayed on the edge of the camera's displayed field of view at a location on that edge closest to the first portion of the physical environment where the annotation was made, and before the visual indication is displayed, a miniature view of the still image with the first annotation is displayed at that location and converted into the visual indication. Converting the visual representation of the annotation (e.g., a scaled-down representation of the still image) into an indication (e.g., a dot) of the camera's current spatial relationship to the portion of the physical environment captured in the still image provides improved visual feedback to the user (e.g., indicating that the indication (e.g., dot) and the visual representation of the annotation (e.g., a reduced-scale image) are different representations of the same annotation).Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the annotated portion of the still image), makes the user-device interface more efficient, and also reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0262] In some embodiments, while displaying a first user interface region including a representation of the field of view of one or more cameras, and prior to detecting a first request to add an annotation to the displayed view of the displayed view, the device displays (930) an indication (e.g., indication 5022) of the current spatial relationship (e.g., as part of a computing system) of the one or more cameras relative to a previously received second portion of the physical environment to which the second annotation was previously added (e.g., displaying a visual indication such as a dot or other shape on an edge of the displayed camera's field of view and at a location on the edge closest to the second portion of the physical environment where the second annotation was made, or displaying a simplified map of the physical environment simultaneously with the representation of the camera's field of view, marking the relative positions of the physical environment and the second portion of the device on the map). In some embodiments, the second annotation was added to the second portion of the physical environment shown in the representation of the field of view of the one or more cameras in the same manner as the first annotation was added to the first portion of the physical environment shown in the representation of the camera's field of view. Displaying an indication of the camera's current particular relationship to the second portion of the physical environment to which the previously received annotation was added provides improved visual feedback to the user (e.g., indicating that camera movement is necessary to view the previously received annotation for the portion of the physical environment captured in the still image). For example, indication 5022 is displayed prior to detecting a request to add annotation 5028 as described with respect to FIGS. 5P-5R. Providing improved visual feedback to the user enhances device usability (e.g., by allowing the user to quickly and accurately locate the portion of the physical environment that corresponds to the previously annotated portion of the still image), makes the user-device interface more efficient, and, in addition, reduces device power usage and improves battery life by allowing the user to use the device more quickly and efficiently.

[0263] In some embodiments, after receiving a first request to redisplay a representation of the field of view of one or more cameras within the first user interface region, in accordance with a determination that both the first portion of the physical environment and the second portion of the physical environment are outside the field of view of the one or more cameras, the device simultaneously displays an indication of the current spatial relationship of the one or more cameras to the first portion of the physical environment and an indication of the current spatial relationship of the one or more cameras to the second portion of the physical environment (932). For example, if multiple annotations (e.g., annotations 5018 and 5036) have been added to different portions of the physical environment, indicators corresponding to the different annotations (e.g., indicators 5018 and 5040, FIG. 5AA) are simultaneously displayed around the edges of the representation of the field of view of the camera at their respective positions closest to the corresponding portions of the physical environment. Simultaneously displaying an indication of the current spatial relationship of the camera to the first portion of the physical environment and an indication of the current spatial relationship of the camera to the second portion of the physical environment provides the user with improved visual feedback (e.g., indicating the direction of camera movement required to view one or more of the multiple received annotations). Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate a portion of the physical environment that corresponds to a previously annotated portion of a still image), makes the user-device interface more efficient, and additionally reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0264] In some embodiments, while displaying an indication of the current spatial relationship of the one or more cameras relative to a first portion of the physical environment and an indication of the current spatial relationship of the one or more cameras relative to a second portion of the physical environment, the device detects (934) a second movement of the one or more cameras relative to the physical environment, and in response to detecting the second movement of the one or more cameras relative to the physical environment, in accordance with a determination that both the first portion of the physical environment and the second portion of the physical environment are outside the field of view of the one or more cameras, the device updates the indication of the current spatial relationship of the one or more cameras relative to the first and second portions of the physical environment, respectively, in accordance with the second movement of the one or more cameras relative to the physical environment (e.g., by moving virtual indicators in different directions and / or at different speeds along the edges of the representation of the field of view). For example, respective visual indicators corresponding to different annotations covering different portions of the physical environment are displayed at different positions on the edge of the representation of the camera's field of view, and movement of the device causes each visual indicator to move in different directions and at different speeds according to changes in the current spatial relationship of their corresponding annotations to the device (e.g., indicators 5018 and 5040 move according to movement of device 100 as described with respect to FIGS. 5AA-5AD). The visual indicators may move together or separately and / or at different speeds depending on the actual spatial relationship between different portions of the physical environment marked with different annotations. Updating the indication of the camera's current spatial relationship to the first and second portions of the physical environment according to camera movement provides improved visual feedback to the user (e.g., camera movement moves the camera closer to or further away from portions of the physical environment corresponding to the annotated portions of the still image).Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate a portion of the physical environment that corresponds to a previously annotated portion of a still image), makes the user-device interface more efficient, and additionally reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0265] In some embodiments, in accordance with a determination that the first portion and the second portion of the physical environment are within a predetermined range of the one or more cameras, an indication of the current spatial relationship of the one or more cameras to the first portion of the physical environment and an indication of the current spatial relationship of the one or more cameras to the second portion of the physical environment are displayed (936). In some embodiments, the user interface provides a method for selecting a subset of annotations from all annotations added to various portions of the physical environment, and only indications corresponding to the selected subset of annotations are displayed along with the representations of the cameras' fields of view. Displaying an indication for the first portion of the physical environment and for the second portion of the physical environment in accordance with a determination that the first portion and the second portion of the physical environment are within a predetermined range of the one or more cameras provides improved visual feedback to the user (e.g., by reducing cluttering of the user interface with indicators when the first and second portions are outside the predetermined range). Providing improved visual feedback to the user enhances usability of the device (e.g., by allowing the user to quickly and accurately locate a portion of the physical environment that corresponds to a previously annotated portion of a still image), makes the user-device interface more efficient, and additionally reduces power usage and improves battery life of the device by allowing the user to use the device more quickly and efficiently.

[0266] It should be understood that the particular order described for the operations in Figures 9A-9F is merely an example, and that the described order is not intended to indicate the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. Furthermore, it should be noted that other process details described herein with respect to other methods described herein (e.g., methods 1000, 1100, and 1200) may also be applied in methods similar to method 9000 described above with respect to Figures 9A-9F. For example, the contact, input, annotation, physics object, user interface area, field of view, movement, and / or animation described above with reference to method 9000 optionally have one or more of the contact, input, annotation, physics object, user interface area, field of view, movement, and / or animation characteristics described herein with reference to other methods described herein (e.g., methods 1000, 1100, and 1200). For the sake of brevity, those details will not be repeated here.

[0267] 10A-10B are flow diagrams illustrating a method 1000 for receiving an annotation on a portion of a physical environment captured in a still image corresponding to a pause point in a video, according to some embodiments. Method 1000 is performed on an electronic device (e.g., device 300 of FIG. 3 or portable multifunction device 100 of FIG. 1A) having a display generation element (e.g., a display, a projector, a heads-up display, etc.) and one or more input devices (e.g., a touchscreen display that also functions as a display generation element). In some embodiments, the display is a touchscreen display, and the touch-sensitive surface is on or integrated into the display. In some embodiments, the display is separate from the touch-sensitive surface. Some operations of method 1000 are optionally combined, and / or the order of some operations is optionally changed.

[0268] The device displays, via a display generation element, a user interface including a video playback area (1002). For example, device 100 displays, via touchscreen display 112, user interface 6000 including video playback area 6002, as described with respect to FIG. 6A.

[0269] While displaying a playback of a first portion of the video in the video playback area, the device receives (1004) a request to add an annotation to the video playback via one or more input devices (e.g., the request is input by a contact detected within a video playback user interface on a touchscreen display (e.g., a location corresponding to a control to begin the annotation or a location within the video playback area (e.g., the location where the annotation begins)). For example, the request to add an annotation to the video playback is input by contact 6024 detected at a location corresponding to markup control 6010.

[0270] In response to receiving the request to add the annotation, the device pauses (1006) playback of the video at a first position in the video (e.g., identifies a current frame of the video (e.g., a pause position) and stops playback of the video at the current frame).

[0271] The device displays (1008) a still image (e.g., a frame of the video) corresponding to the first pause position in the video (e.g., displaying the current frame of the video shown at the time the request was received). For example, as described with respect to FIGS. 6C-6D , in response to input by contact 6024 at a location corresponding to markup control 6010 in the video, playback of the video is paused and a still image corresponding to the pause position in the video is displayed.

[0272] While displaying the still image (e.g., in the video playback area), the device receives (1008) an annotation (e.g., touch-based drawing input) via one or more input devices on a first portion of the physical environment captured in the still image. For example, as described with respect to FIGS. 6E-6F , annotation 6030 is received on a portion of the physical environment corresponding to kite object 6028 in the displayed still image. When a "physical environment" is referred to herein, it will be understood that a non-physical environment (e.g., a computer-generated environment) may be included in the still image. For example, the annotation is received on a portion of the image (e.g., a portion of the computer-generated video) corresponding to a pause point in the video. In some embodiments, the computer-generated image includes depth data and / or objects for which annotations and / or virtual objects are suitable.

[0273] After receiving the annotation, the device displays (1010) in the video playback area (e.g., during continuous playback of the video or while receiving input (e.g., on a timeline) such as scrubbing the video back and forth) a second portion of the video that corresponds to a second location in the video that is different from the first location in the video (e.g., before or after a paused location in the video), where a first portion of the physical environment is captured in the second portion of the video, and the annotation is displayed in the second portion of the video. For example, after annotation 6030 is received as described with respect to FIGS. 6E-6F, annotation 6030 is displayed in the second portion of the video (e.g., as described with respect to FIG. 6J). In some embodiments, while the second portion of the video is displayed, the annotation is displayed in a second location in the video playback area that is different from the first location of the annotation received while the still image was displayed (e.g., the annotation is "sticky" to a location (e.g., a physical object) in the physical environment that was captured in the video clip, such that the location (e.g., the physical object) moves simultaneously with the annotation as the video progresses). Displaying annotations on a portion of the video that is different from the portion of the video to which the annotations were applied without requiring further input (e.g., to identify the surface to which the annotations should be applied) enhances the usability of the device. Performing actions without requiring further user input enhances the usability of the device and makes the user device interface more efficient (e.g., by allowing a user to add information to previously captured video via direct annotation of the video without having to re-record the video or annotate multiple portions of the video), as well as reducing the device's power usage and improving battery life by allowing a user to use the device more quickly and efficiently.

[0274] In some embodiments, video is captured (1012) by a camera during relative movement of the camera and the physical environment (e.g., during video capture, camera movement data and depth data of the physical environment are simultaneously captured and stored along with the simultaneously captured image data), and a third portion of the video is captured between the first and second portions of the video and does not include the first portion of the physical environment during the relative movement of the camera and the physical environment. In some embodiments, an annotation (e.g., annotation 6030) received on a still image (e.g., as shown in FIGS. 6E-6F) targets a first object (e.g., kite 6028) located in the first portion of the physical environment and appears in a position corresponding to the first object in the second portion of the video (e.g., the annotation is not made directly on a still image of any frame in the second portion of the video). For example, in FIG. 6J, annotation 6030 appears in the second portion of the video in a position corresponding to kite 6028. In some embodiments, the annotation is not displayed during the third portion of the video that does not include the first object (e.g., the annotation is not displayed persistently and is only shown when the current frame includes the first object). In some embodiments, the annotation is rotated and scaled to appear at the position of the first object according to the distance and field of view of the first object. For example, in FIGS. 6J-6N, annotation 6030 is rotated and scaled to appear at the position of kite 5026 in the video. In some embodiments, the first, third, and second portions of the video are consecutively captured portions of the video, or the second, third, and first portions of the video are consecutively captured portions of the video. Because there is a discontinuity in the captured subject matter within the camera's field of view (e.g., during the capture of the third portion of the video), the first portion of the physical environment captured in the first portion of the video cannot be recognized as the same first portion of the physical environment captured in the third portion of the video based solely on the image data of the video (e.g., identifying tracking points across consecutive frames via frame-to-frame comparison).The camera movement and / or depth data is used (optionally along with image data) to create a three-dimensional or quasi-three-dimensional model of the physical environment captured in the video, so that specific locations within the physical environment can be recognized in each frame of the video, regardless of their appearance in the frame or the viewing perspective. Displaying annotations at locations corresponding to the object to which they are pointed and not displaying the annotations in portions of the video that do not include the object provides improved visual feedback to the user (e.g., by providing an indication that the annotation is anchored at a location corresponding to the object). Performing operations without requiring further user input improves device usability, makes the user device interface more efficient (e.g., by allowing a user to add information to a video without re-recording the video or annotating multiple portions of the video), and additionally reduces device power usage and improves battery life by allowing users to use the device more quickly and efficiently.

[0275] In some embodiments, the device displays (1014) a timeline of the video (e.g., a scrub bar having a position indicator indicating the position of the currently displayed frame, a scrollable sequence of reduced-scale images of sample frames from successive segments of the video) (e.g., concurrently with the display of the video (e.g., while the video is playing and / or while the video is paused)), and displaying a second portion of the video is performed in response to user input that scrubs through the video timeline to a second position in the video (e.g., user input that drags the position indicator along the scrub bar or user input that scrolls the sequence of reduced-scale images of sample frames past a still marker of the currently displayed frame). For example, input by contact 6038 is received at a position corresponding to a timeline 6004 that includes a series of sample frames (e.g., sample frame 6006 as described with respect to FIG. 6A ), and the video displayed in the video playback area 6002 is played in response to the input (e.g., as described with respect to FIGS. 6L-6N ). Displaying a timeline for scrubbing through the video without requiring further user input enhances the usability of the device (e.g., by allowing a user to access annotations displayed in a second portion of the video using existing scrubbing controls without requiring a separate control or input). Reducing the number of user inputs required to perform an action enhances the usability of the device and makes the user device interface more efficient. This further reduces power usage and improves the device's battery life by allowing a user to use the device more quickly and efficiently.

[0276] In some embodiments, displaying the second portion of the video is performed as a result of rewinding (1016) the video from a first position to a second position (e.g., the second position precedes the first position on the video's timeline). For example, as described with respect to FIGS. 6L-6N , input by contact 6038 is received to rewind the video displayed in video playback area 6002. Displaying the annotations in a portion of the video different from the portion of the video to which the annotations were applied in response to the rewind input, without requiring further user input, enhances device usability (e.g., by allowing a user to access the annotations displayed in the second portion of the video using existing scrub controls without requiring a separate control or input). Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0277] In some embodiments, displaying the second portion of the video is performed as a result of fast-forwarding (1018) the video from a first position to a second position (e.g., the second portion precedes the first portion on the video timeline and plays at a faster than normal playback speed). Displaying the annotations in a portion of the video different from the portion of the video to which the annotations were applied in response to the fast-forward input, without requiring further user input, enhances device usability (e.g., by allowing a user to access annotations displayed in the second portion of the video using existing scrub controls without requiring a separate control or input). Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0278] In some embodiments, displaying the second portion of the video is performed as a result of normal playback (1020) of the video from the first position to the second position (e.g., the second portion precedes the first portion on the video's timeline and plays at a faster playback speed than the normal playback speed). In some embodiments, when the user finishes providing annotations on the still image, the user exits the annotation mode by selecting a “Done” button displayed with the still image (e.g., as described with respect to FIG. 6G). As a result, the device continues playing the video from the first position, and annotations are displayed in each subsequent frame including the first portion of the physical environment at the same physical location (e.g., physical object), even if the first portion of the physical environment is captured from a different distance and / or a different perspective compared to the still image. Displaying annotations in a portion of the video different from the portion of the video to which the annotations were applied in response to normal playback of the video without requiring further user input enhances usability of the device (e.g., by allowing the user to access annotations displayed in the second portion of the video using existing scrub controls without requiring a separate control or input). Performing actions without requiring further user input makes the user device interface more efficient and also allows the user to use the device more quickly and efficiently, thereby reducing power usage and improving the device's battery life.

[0279] In some embodiments, the device displays (1022) a list of media content objects including a video (e.g., displaying a representation of the video in a media library) via a display generation element. The device receives input to select a video from the list of media content objects, and in response to receiving the input to select the video, displays a user interface object (e.g., a “markup” button for adding annotations) with a representation of the video in a video playback area, the user interface object configured to receive a request to add an annotation to the video while the video is playing (e.g., a tap input to activate the button to add an annotation). Displaying the user interface object configured to receive a request to add an annotation to a video while the video is playing provides improved feedback (e.g., showing options for adding annotations to portions of the video while it is playing). In some embodiments, the markup button provides improved feedback (e.g., along with other playback controls) when touch input is detected on the video playback area during video playback; when the markup button is activated, the currently displayed frame is shown in a markup-enabled state, ready to receive drawing annotations directly on the image of the currently displayed frame. Providing improved feedback enhances usability of the device (e.g., by allowing a user to access annotations displayed in a second portion of the video using existing scrub controls without requiring a separate control or input). Performing actions without requiring further user input makes the user device interface more efficient and further reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.

[0280] It should be understood that the particular order described for the operations in FIGS. 10A-10B is merely an example, and that the described order is not intended to indicate the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. Furthermore, it should be noted that other process details described herein with respect to other methods described herein (e.g., methods 900, 1100, and 1200) may also be applicable in methods similar to method 1000 described above with respect to FIGS. 10A-10B. For example, the contact, input, annotation, physics object, user interface area, field of view, movement, and / or animation described above with reference to method 1000 optionally have one or more of the contact, input, annotation, physics object, user interface area, field of view, movement, and / or animation characteristics described herein with reference to other methods described herein (e.g., methods 900, 1100, and 1200). For the sake of brevity, those details will not be repeated here.

[0281] 11A-11F are flow diagrams illustrating a method 1100 of adding a virtual object to a previously captured media object. Method 1100 is performed on an electronic device (e.g., device 300 of FIG. 3 or portable multifunction device 100 of FIG. 1A ) having a display generating element (e.g., a display, a projector, a heads-up display, etc.) and one or more input devices (e.g., a touch-sensitive surface such as a touch-sensitive remote control, or a touchscreen display that also functions as a display generating element, a mouse, a joystick, a wand controller, and / or a camera that tracks the position of one or more features of a user, such as the user's hand). In some embodiments, the display is a touchscreen display, and the touch-sensitive surface is on or integrated into the display. In some embodiments, the display is separate from the touch-sensitive surface. Some operations of method 1100 are optionally combined, and / or the order of some operations is optionally changed.

[0282] The device displays (1102), via a display generation element, a first previously captured media object including one or more first images (e.g., still photographs, live photographs, or a video including a sequence of image frames), the first previously captured media object being recorded and stored along with first depth data corresponding to the first physical environment captured in each of the one or more first images (e.g., first depth data generated by one or more depth sensors of the device (e.g., emitter / detector systems such as infrared, sonar, and / or lidar, and / or image analysis systems (e.g., video segment analysis and / or stereo image / video analysis)) at a time corresponding to the time the first media object was captured by one or more cameras. For example, as described with respect to FIG. 7A , the device 100 displays the previously captured images within the media object display area 7002 via the touchscreen display 112.

[0283] While displaying a first previously captured media object (e.g., while displaying a still image, while displaying a representative image of a live photo, while displaying a frame of a video while the video is playing, or while displaying a frame of a video while the video is paused or stopped), the device receives a first user request via one or more input devices to add a first virtual object (e.g., a falling ball, confetti, text, a spotlight, an emoji, paint, a measurement graphic) to the first previously captured media object (1104). For example, the request to add a virtual object to the previously captured media object is a tap input received on the touchscreen display 112 to add the virtual object to the previously captured image, as described with respect to FIG. 7C .

[0284] In response to adding a first virtual object to a first previously captured media object in response to a first user request, the device displays the first virtual object over at least a portion of each image in the first previously captured media object (1106), with the first virtual object displayed with at least a first position or orientation (or movement path) determined using first depth data corresponding to each image in the first previously captured media object. For example, as described with respect to FIG. 7C , in response to a tap input received on the touchscreen display 112, a virtual object (e.g., a virtual ball object 7030) is added to the previously captured image. Adding a virtual object to a previously captured media object using depth data from the previously captured media object without requiring user input (e.g., indicating a planar location of the media object) enhances device usability. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0285] In some embodiments, displaying (1108) a first virtual object over at least a portion of the respective image in the first previously captured media object includes displaying a first movement of the first virtual object relative to a first physical surface captured in the first previously captured media object after the first virtual object is placed over a corresponding one of the one or more first images, wherein the first movement of the first virtual object is constrained by a first pseudo surface corresponding to the first physical surface determined based on the depth data and a pseudo gravity direction (e.g., the pseudo gravity direction is, optionally, determined based on the direction of gravity recorded in the depth data at the time the first image was captured, or is the same as the actual gravity direction for a current orientation of a device displaying the respective image). For example, as described with respect to FIG. 7C , a virtual ball object 7030 is constrained by a pseudo gravity direction by a pseudo surface corresponding to a physical surface determined based on the depth data (e.g., floor surface 7040). In some embodiments, different types of virtual objects have different simulated physical properties (e.g., shape, size, weight, elasticity, etc.) that interact with the simulated surface in different ways. In one example, if a still image captures a couch with curved armrests and a flat seating area, a virtual rubber ball is shown falling from the top of the image, landing on the curved surface of the armrests, bouncing off the curved surface of the armrests, bouncing off the land on the flat seating area, and then rolling onto the floor. In contrast, confetti flakes are shown fluttering down from the top of the image, landing on the curved surface of the armrests, sliding off the curved surface of the armrests onto the flat seating area, and remaining on the flat seating area. In another example, a 3D letter "A" is placed on the curved surface of a user's arm by a user's finger, and the 3D letter "A" falls sideways and lands on the flat surface of the seating area when the user's finger is lifted off the touchscreen.In some embodiments, a surface mesh corresponding to the physical environment captured in the still image is generated based on the depth data, and virtual objects inserted into the still image are animated during insertion and / or after initial placement to exhibit movement and final positions / or orientations that conform to basic physics, such as laws related to gravity, forces, and physical interactions between objects. By moving virtual objects onto the physical plane of a previously captured media object to indicate the media object's planar position (e.g., according to simulated gravity), movement of the virtual object can be performed automatically without requiring further input (e.g., without the need to provide a user with input indicating the virtual object's movement path), enhancing device usability. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0286] In some embodiments, displaying 1110 the first virtual object over at least a portion of the respective image in the first previously captured media object displays a change in shape of the first virtual object according to a first physical surface captured in the first previously captured media object after the first virtual object is placed over a corresponding one of the one or more first images, the change in shape of the first virtual object being constrained by a first pseudo-surface corresponding to the first physical surface determined based on the first depth data. For example, as described with respect to FIGS. 7U-7X, the shape of the virtual decal object 7084 changes as the object moves across the surface of the couch 7060 onto the floor 7004 depicted in the previously captured image. In some embodiments, if a still image captures a couch with curved armrests and a flat seating area, a virtual paintball is shown shooting in the image, with the virtual paintball spreading over the curved surface of the armrest. In contrast, if a virtual paintball is shown spraying onto an image and lands on a flat surface of the seating area, the virtual paint spreads across the flat surface of the seating area. In another example, a long virtual streamer falls onto the armrest over the curved surface of the armrest, while a long virtual streamer that fell on a flat seating area lies flat on the flat surface of the seating area. By displaying changes in the shape of a virtual object according to surfaces in a previously captured media object, the changes in the shape of the virtual object can occur automatically without requiring further input (e.g., without requiring the user to provide input directing the change in the shape of the virtual object), enhancing device usability. Performing actions without requiring further user input makes the user device interface more efficient and also reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0287] In some embodiments, while displaying the first virtual object over at least a portion of a respective image in the first previously captured media object, the device detects (1112) a second user request to switch from displaying the first previously captured media object to displaying a second previously captured media object (e.g., a horizontal swipe input on the first virtual object to indicate a previous or next item in a horizontally arranged list of media objects, a vertical swipe on the first virtual object to indicate a previous or next item in a vertically arranged list of media objects, or a vertical swipe on the first virtual object to switch to the next or previous media object). a tap on the front or back button for capturing a second previously captured media object, the second previously captured media object including one or more second images, the second previously captured media object recorded and stored along with second depth data corresponding to the second physical environment captured in each of the one or more second images (e.g., second depth data generated by one or more depth sensors of the device (e.g., emitter / detector systems such as infrared, ultrasonic, and / or lidar, and / or image analysis systems (e.g., video segment analysis and / or stereo image / video analysis)) at a time corresponding to the time the second media object was captured by one or more cameras). For example, as described with respect to FIGS. 7E-7F, while a virtual ball object is displayed over the first previously captured image, as shown in FIG. 7E, a request (e.g., input on a subsequent media object control 7010) to switch from displaying the first previously captured image to displaying the second previously captured image (as shown in FIG. 7F) is detected.In response to receiving a second user request to switch from displaying a first previously captured media object to displaying a second previously captured media object (e.g., a first previously captured media object), the device replaces the display of the first previously captured media object with a display of the second previously captured media object, sliding out the first previously captured media object and sliding in the second previously captured media object in the direction of the swipe input (e.g., a horizontal swipe or a vertical swipe input). For example, in response to an input received as described with respect to FIG. 7E, the device may switch from displaying a first previously captured media object in media object display area 7002, as shown in FIG. 7E, to displaying a second previously captured media object in media object display area 7002, as shown in FIG. 7F. 7E, the device switches to displaying a second previously captured image in area 7002. The device displays a first virtual object over at least a portion of each image in the second previously captured media object, the first virtual object being displayed with at least a second position or orientation (or path of movement) determined based on a first position or orientation (or path of movement) of the first virtual object in each image of the first previously captured media object and based on second depth data corresponding to each image in the second previously captured media object. For example, virtual ball objects 7034 and 7044 added to the first previously captured image displayed in FIG. 7E are displayed over the second previously captured image displayed in FIG. 7F.In some embodiments, if a first virtual object is a piece of virtual confetti or a virtual ball falling in a first, previously captured image of a media object and landing on a first surface of the image (e.g., a flat surface of a core seating area), when a user switches to display a second image by swiping horizontally across the first image, the second image slides horizontally and the virtual confetti or virtual ball begins to move from its position in the first image (e.g., falling downward from a position corresponding to the surface of the core seating area) and lands on a second surface (e.g., falling downward from the surface of the floor, or the surface of a cushion on the floor, etc.). In other words, the virtual object persists when switching between media objects, and the position, orientation, and movement path of the virtual object in the next image are influenced by the position, orientation, and movement path of the virtual object in the previous image. Switching from displaying a virtual object over a first previously captured media object to displaying the virtual object over a second previously captured media object at a position determined based on the position or orientation of the virtual object in the first previously captured media object without requiring user input (e.g., to indicate the position of the virtual object in the second previously captured media object) enhances the usability of the device. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves the device's battery life by allowing a user to use the device more quickly and efficiently.

[0288] In some embodiments, the first user request is a request (1114) to add multiple instances of a first type of virtual object (e.g., virtual ball objects 7034 and 7044 described with respect to FIGS. 7C-7E) to a previously captured media object over time (e.g., adding virtual confetti or virtual balls falling to an image over time), e.g., the first virtual object is one of multiple instances of the first type of virtual object added to the first previously captured media object. In response to receiving a second user request to switch to displaying a second previously captured media object, the device displays a second virtual object over at least a portion of each image in the second previously captured media object, the second virtual object being an instance of a first type of virtual object that is distinct from the first virtual object and not added to the first previously captured media object, the second virtual object being displayed with at least a third position or orientation or movement path determined using second depth data corresponding to each image in the second previously captured media object. For example, in some embodiments, the first user request is a request to add a series of virtual objects of the same type over time (e.g., in a sequential manner) to create an effect in the image, such as falling confetti, raindrops, or fireworks. While an effect is being applied to a first image or video (e.g., when multiple instances of virtual confetti, raindrops, or fireworks are added to the first image), if the user switches to a next image or video (e.g., by swiping horizontally or vertically across the first image or video), the effect is also automatically applied to the next image or video (e.g., adding new instances of virtual confetti, raindrops, or fireworks to the image) without the user explicitly invoking the effect for the next image or video (e.g., activating a control for that effect). Switching from displaying a first virtual object over a first previously captured media object to displaying a second virtual object over a second previously captured media object at a position determined based on the position or orientation of the virtual object in the first previously captured media object without requiring user input (e.g., to indicate the position of the virtual object in the second previously captured media object) enhances usability of the device.Performing actions without requiring further user input makes the user device interface more efficient and also allows the user to use the device more quickly and efficiently, thereby reducing power usage and improving the device's battery life.

[0289] In some embodiments, the first previously captured media object and the second previously captured media object (1116) are two separate still images (e.g., a first previously captured image as shown in media object display area 7002 of FIG. 7E and a second previously captured image as shown in media object display area 7002 of FIG. 7F) previously recorded and stored with different depth data corresponding to different physical environments and / or different views of the same physical environment. For example, the second still image need not have any connection to the first still image in terms of subject matter captured within the image to have the same effects (e.g., falling confetti, virtual balls, fireworks, virtual block letters, etc.) continue to be applied to the second still image. Switching from displaying a virtual object over a first previously captured media object to displaying the second virtual object over a second previously captured media object with different depth data than the first previously captured media object without requiring further user input (e.g., to indicate the location of the virtual object within the second previously captured media object) enhances the usability of the device. Performing actions without requiring further user input makes the user device interface more efficient and also allows the user to use the device more quickly and efficiently, thereby reducing power usage and improving the device's battery life.

[0290] In some embodiments, the first previously captured media object is a video including a sequence of consecutive image frames (e.g., as described with respect to FIGS. 7AM-7AT), and displaying the first virtual object over at least a portion of each image in the first previously captured media object includes displaying the first virtual object over a first portion of the first image frame while displaying the first image frame of the first previously captured media object during playback of the first previously captured media object (1118), wherein the first virtual object is displayed with a position or orientation (or movement path) determined according to a portion of the first depth data corresponding to the first image frame of the first previously captured media object; and immediately after displaying the first image frame, displaying the first virtual object over a first portion of the first image frame of the first previously captured media object (1118). and while displaying a second image frame of the media object (e.g., the second image frame immediately precedes the first image frame in the media object in normal or fast-forward playback of the media object, the second image frame immediately precedes the first image frame in the media object in reverse playback of the media object, or the second image is an initial frame of the media object and the first image is a last frame of the media object during looped playback of the media object), displaying a first virtual object over a second portion of the second image frame, wherein the first virtual object is displayed with a position or orientation (or movement path) determined according to a position or orientation (or movement path) of the first virtual object in the first image frame and according to a portion of the first depth data corresponding to the first previously captured second image frame of the media object.For example, if the first virtual object is virtual confetti or a virtual ball falling onto a surface (e.g., the surface of a moving or stationary object) in a first image frame of a video, as the video continues to play and the surface is shown in the next image frame, the position and / or orientation and / or path of the virtual confetti or virtual ball will change depending on the position and orientation of the surface in the new image frame. For example, if the surface is a stationary table surface, the virtual confetti will appear to be located in the same position on the stationary table surface, and the virtual ball will appear to roll along the stationary table surface, even though the surface is now viewed from a different perspective and occupies a different area in the second image frame compared to the first image frame. Similarly, if the surface is the top of a trapdoor that suddenly appears in the video, the virtual confetti will begin to fall gradually from its stationary position on top of the trapdoor, and the virtual ball will appear to descend with acceleration from its position on top of the trapdoor due to pseudo-gravity. In some embodiments, the first user request is a request to add a series of virtual objects of the same type over time (e.g., in a sequential manner) to create an effect on the image, such as falling confetti or fireworks. For example, the virtual confetti objects placed on the edge of the physical kite object 7112 are displayed with altered positions, orientations, and movement paths as video playback occurs in FIGS. 7AM-7AT. While the effect is being applied to the first image frame (e.g., as multiple instances of virtual confetti or fireworks are added to the first image frame), as video playback continues, the effect is also automatically applied to the next image frame (e.g., new instances of virtual confetti or fireworks are also added to the next image frame). In some embodiments, at the end of the video, the virtual object added to the final image frame comprises a virtual object that was added to multiple previous image frames and that has settled into a final position and orientation in the final image frame based on prior interactions with pseudo surfaces corresponding to the physical environment depicted in the previous image frames and the pseudo surfaces corresponding to the physical environment depicted in the last image frame.Displaying a virtual object over a second frame of video displayed immediately after displaying a first frame of video, where the position or orientation of the virtual object in the second image frame is determined using depth data from the second image frame without requiring further input (e.g., without requiring a user to provide input indicating the location of the virtual object in each frame of the video), improves device usability. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing a user to use the device more quickly and efficiently.

[0291] In some embodiments, displaying the first previously captured media object includes playing the video according to a first timeline including at least one of looping, fast-forwarding, or reversing the sequence of consecutive image frames (1120), and displaying the first virtual object over at least a portion of each image in the first previously captured media object includes displaying, during playback of the video according to the first timeline, changes in the position or orientation (or movement path) of the first virtual object according to a forward timeline associated with the actual order of the sequence of image frames displayed during playback of the video (including, for example, looping from the end of the video to the beginning, switching frames at a non-uniform rate during video playback, playing the video backward from a later frame to an earlier frame, etc.) (e.g., the previous position and orientation of the virtual object in each previously displayed image frame affects the position and orientation of the currently displayed image frame). In other words, the timeline of the movement of the virtual object in the displayed image frames is independent of the timeline in which the media object is played. Displaying changes in the position or orientation of a virtual object according to a timeline associated with the order of the sequence of image frames without requiring further input (e.g., without requiring a user to provide input indicating changes in the virtual object's position in each frame of a video) improves the usability of the device. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves the device's battery life by allowing a user to use the device more quickly and efficiently.

[0292] In some embodiments, displaying 1122 the first virtual object over at least a portion of the respective image in the first previously captured media object includes displaying a shadow of the first virtual object according to a first physical surface captured in the first previously captured media object while the first virtual object is positioned over a corresponding one of the one or more first images, wherein the shadow of the first virtual object is constrained by a first pseudo-surface corresponding to the first physical surface determined based on the first depth data. For example, if the still image captures a couch with curved armrests and a flat seating area, a virtual letter A that is positioned on the curved surface of the armrest and then falls sideways onto the flat surface of the seating area has a shading that changes according to the surface that virtual letter A is currently on and the current orientation of virtual letter A relative to the surface. In some embodiments, a three-dimensional or quasi-three-dimensional mesh is generated based on depth data associated with the image or video, the mesh surface assuming geometric characteristics of the physical environment captured in the image or video, and shading is cast on the mesh surface based on the position and orientation of the virtual object relative to a simulated light source and the mesh surface. Displaying virtual objects with shading constrained by simulated surfaces corresponding to physical surfaces based on the depth data of previously captured media objects without requiring further input (e.g., to identify surfaces within the previously captured media object) enhances device usability. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0293] In some embodiments, the first user request is a user request to place a virtual first text object at a first position within the respective image within the first previously captured media object (1124). For example, the user request is an input provided at the virtual text object 7064 to initiate an edit mode for the virtual text object 7064, as described with respect to FIGS. 7Q-7S. The device receives user input to update the virtual first text object (1126), including adding a first virtual character to the virtual first text object (e.g., editing the text entry area by typing the character at the end of the existing text entry), and in response to receiving the user input, the device displays the first virtual character at a second position within the respective image within the first previously captured media object adjacent to a preceding virtual character in the virtual first text object and according to a portion of the first depth data corresponding to the second position within the respective image. In some embodiments, the text has lighting and shadows that are generated based on a surface mesh of the environment captured in the respective image. Displaying a text object at a location within a previously captured media object without requiring further input (e.g., to identify depth data within the previously captured media object) and positioning characters of the text object according to the depth data within the media object enhances usability of the device. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.

[0294] In some embodiments, displaying the first virtual object over at least a portion of each image in the first previously captured media object (1128) includes: displaying the first virtual object above a horizontal surface (e.g., rather than below the horizontal surface) in accordance with a determination that the pseudo surface proximate the current position of the first virtual object in the respective image is a horizontal surface; and displaying the first virtual object in front of a vertical surface in accordance with a determination that the pseudo surface proximate the current position of the first virtual object in the respective image is a vertical surface. Displaying the virtual object above or in front of a surface proximate the virtual object depending on whether the surface is a horizontal or vertical surface without requiring further input (e.g., to indicate whether the surface is a horizontal or vertical surface) enhances device operability. Performing operations without requiring further user input makes the user device interface more efficient and further reduces power usage and improves device battery life by allowing users to use the device more quickly and efficiently.

[0295] In some embodiments, displaying (1130) the first virtual object over at least a portion of the respective image in the first previously captured media object includes, in accordance with a determination that the respective image includes a first pseudo surface (e.g., a foreground object) and a second pseudo surface (e.g., a background object) having different depths proximate to the current position of the first virtual object in the respective image, displaying the first virtual object at a depth between the first pseudo surface and the second pseudo surface (e.g., at least a first portion of the first virtual object is occluded by the first pseudo surface, at least a second portion of the first virtual object occludes at least a portion of the second pseudo surface, or below an object represented by the first pseudo surface). For example, as described with respect to FIG. 7G , a virtual ball object 7045 is displayed at a depth between a first pseudo surface (e.g., a back wall of a room depicted in a previously captured image) and a second pseudo surface (e.g., a table 7054). In some embodiments, a complete three-dimensional model of the physical environment cannot be established based on depth data from the image alone. Spatial information about the space between the first and second pseudo surfaces is absent. The first virtual object is positioned between the first and second pseudo surfaces regardless of the absence of spatial information within the depth range between the first and second pseudo surfaces. Displaying a virtual object with depth between the first and second pseudo surfaces without requiring additional input (e.g., to indicate the depth of the virtual object, the depth of t...

Claims

1. 1. A computer system in communication with a display generation element and one or more sensors that detect user input, comprising: receiving a request to display a first virtual effect in a view of a physical environment, the view of the physical environment being visible through the display generation element; in response to detecting the request to display the first virtual effect in the view of the physical environment, displaying one or more virtual objects superimposed on the view of the physical environment; Displaying the one or more virtual objects includes: displaying animated movements of each of the one or more virtual objects in at least a portion of the view of the physical environment, wherein the animated movements of each are constrained according to a direction of pseudo-gravity associated with the view of the physical environment; and constraining the animated movement of the first one of the one or more virtual objects according to a determination that a current position of the first virtual object corresponds to a first surface detected in the view of the physical environment during the animated movement of the first one of the one or more virtual objects; and constraining the animated movement of the respective second virtual object of the one or more virtual objects in accordance with a determination that a current position of the second virtual object corresponds to a second surface detected in the view of the physical environment, wherein the second surface is lower than the first surface in the direction of pseudo gravity; This includes: the view of the physical environment includes a view of one or more physical objects moving within the physical environment; 10. The method of claim 1, wherein displaying the respective animated movement of the one or more virtual objects includes moving at least one of the one or more virtual objects according to a movement of the one or more physical objects moving in the physical environment.

2. displaying one or more virtual objects superimposed on the view of the physical environment, 10. The method of claim 1 , further comprising continuing to add additional instances of a first type of virtual object superimposed on the view of the physical environment, the additional instances of the first type of virtual object moving in the direction of the pseudo-gravity associated with the view of the physical environment, with a first plurality of the first type of virtual objects clustered proximate to the first surface of the view of the physical environment and a second plurality of the first type of virtual objects clustered proximate to the second surface of the view of the physical environment.

3. detecting a first event that changes the view of the physical environment from a first view of a first physical environment to a second view of a second physical environment that is different from the first view of the first physical environment while displaying the respective animated movements of the one or more virtual objects in at least a portion of the view of the physical environment; In response to detecting the first event, terminating the display of the first view of the first physical environment; and displaying the animated movement of each of the one or more virtual objects in at least a portion of the second view of the second physical environment; and The method of claim 1 further comprising:

4. simultaneously displaying, with the view of the physical environment visible through the display generation element, a first control corresponding to the first virtual effect and a second control corresponding to a second virtual effect different from the first virtual effect, wherein receiving the request to display the first virtual effect in the view of the physical environment includes detecting a first user input selecting the first control; maintaining display of the second control corresponding to the second virtual effect while displaying the animated movement of each of the one or more virtual objects in at least the portion of the view of the physical environment; The method of claim 1 further comprising:

5. detecting a second user input selecting the second control while displaying the animated movement of each of the one or more virtual objects in at least the portion of the view of the physical environment and maintaining display of the second control corresponding to the second virtual effect; in response to detecting a second user input selecting the second control; terminating the display of the one or more virtual objects in at least the portion of the view of the physical environment; and displaying a virtual object for each of the second virtual effects in at least the portion of the view of the physical environment; The method of claim 4 further comprising:

6. 10. The method of claim 1, wherein displaying the respective animated movements of the one or more virtual objects comprises displaying animated movements of two or more virtual objects at different depths with respect to the view of the physical environment.

7. 1. A computer system comprising: a display generating element; one or more sensors for detecting user input; one or more processors; a memory for storing one or more programs; The one or more programs are configured to run on the one or more processors, the one or more programs comprising: receiving a request to display a first virtual effect in a view of a physical environment, the view of the physical environment being visible through the display generation element; displaying one or more virtual objects superimposed on the view of the physical environment in response to detecting the request to display the first virtual effect in the view of the physical environment; displaying animated movements of each of the one or more virtual objects in at least a portion of the view of the physical environment, wherein the animated movements of each are constrained according to a direction of pseudo-gravity associated with the view of the physical environment; and constraining the animated movement of the first one of the one or more virtual objects according to a determination that a current position of the first virtual object corresponds to a first surface detected in the view of the physical environment during the animated movement of the first one of the one or more virtual objects; and constraining the animated movement of the respective second virtual object of the one or more virtual objects in accordance with a determination that a current position of the second virtual object corresponds to a second surface detected in the view of the physical environment, wherein the second surface is lower than the first surface in the direction of pseudo gravity; including instructions for, including displaying, the view of the physical environment includes a view of one or more physical objects moving within the physical environment; 11. The computer system of claim 10, wherein displaying the respective animated movement of the one or more virtual objects includes moving at least one of the one or more virtual objects according to a movement of the one or more physical objects moving within the physical environment.

8. 8. The computer system of claim 7, wherein the one or more programs include instructions for carrying out the method of any one of claims 2 to 6.

9. 1. A computer program having instructions that, when executed by a computer system in communication with a display generation element and one or more sensors that detect user input, cause the computer system to: receiving a request to display a first virtual effect in a view of a physical environment, the view of the physical environment being visible through the display generation element; displaying one or more virtual objects superimposed on the view of the physical environment in response to detecting the request to display the first virtual effect in the view of the physical environment; Displaying the one or more virtual objects includes: displaying animated movements of each of the one or more virtual objects in at least a portion of the view of the physical environment, wherein the animated movements of each are constrained according to a direction of pseudo-gravity associated with the view of the physical environment; and in accordance with a determination that a current position of the first virtual object corresponds to a first surface detected in the view of the physical environment during the respective animated movement of a first virtual object of the one or more virtual objects, constraining the respective animated movement of the first virtual object according to the first surface detected in the view of the physical environment; and during the respective animated movement of a second virtual object of the one or more virtual objects, in accordance with a determination that a current position of the second virtual object corresponds to a second surface detected in the view of the physical environment, constraining the respective animated movement of the second virtual object in accordance with the second surface detected in the view of the physical environment, wherein the second surface is lower than the first surface in the direction of pseudo gravity. This includes: the view of the physical environment includes a view of one or more physical objects moving within the physical environment; 11. The computer program product of claim 10, wherein displaying the respective animated movement of the one or more virtual objects comprises moving at least one of the one or more virtual objects according to a movement of the one or more physical objects moving within the physical environment.

10. A computer program as described in claim 9, comprising instructions which, when executed by the computer system, cause the computer system to perform a method as described in any one of claims 2 to 6.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    CN102194050A

  • Terminal device, server device, and image processing method

    JP2004145448A

  • Image display device, method, and program

    JP2011070495A

  • Information processing device, information processing method and program

    JP2011197777A

  • Display device and control method

    JP2014071499A