System, method and graphical user interface for modeling, measuring and mapping
By combining the computer system of camera and input devices, providing intuitive virtual/augmented reality modeling and annotation methods, solving the problems of cumbersome input and inefficiency in the prior art, achieving efficient modeling and annotation, and extending battery life.
Patent Information
- Application Number
- CN202210629947.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-23
- Filing Date
- 2020-09-25
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-09-25
AI Technical Summary
When modeling and annotating physical environments and objects, existing virtual/augmented reality technologies have problems such as cumbersome input, low efficiency, limited real-time and limited interactive views, especially on battery-driven devices.
Using computer systems combined with cameras and input devices, by reducing user input, provides intuitive modeling, measurement and drawing functions, utilizing graphical user interfaces to achieve efficient interaction, including touchpads, touch-sensitive displays and graphical user interfaces, supports multiple applications, and displays the interaction of different views of the physical environment and virtual objects through augmented reality technology.
Improves efficiency and user experience of virtual/augmented reality technology, reduces power consumption, extends battery life, and provides more intuitive modeling and annotation.
Smart Images

Figure CN114967929B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of September 25, 2020, application number 202080064101.9, and invention name “System, method and graphical user interface for modeling, measurement and mapping using augmented reality”.
[0002] Related patent applications
[0003] This application claims priority to U.S. Provisional Application Serial No. 62 / 965,710 filed on January 24, 2020, U.S. Provisional Application Serial No. 62 / 907,527 filed on September 27, 2019, and U.S. Patent Application Serial No. 17 / 030,209 filed on September 23, 2020, and is a continuation of U.S. Patent Application Serial No. 17 / 030,209 filed on September 23, 2020. Technical Field
[0004] This generally relates to computer systems for virtual / augmented reality, including but not limited to electronic devices for modeling and annotating physical environments and / or objects using a virtual / augmented reality environment. Background Art
[0005] Augmented and / or virtual reality environments can be used to model and annotate physical environments and objects therein by providing different views of the physical environment and objects therein and enabling users to overlay annotations, such as measurements and drawings, on the physical environment and objects therein and visualize the interactions between the annotations and the physical environment and objects therein. However, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality are cumbersome, inefficient, and limited. In some cases, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality are limited in functionality. In some cases, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality require multiple separate inputs (e.g., a series of gestures and button presses, etc.) to achieve the desired result (e.g., by activating multiple displayed user interface elements to access different modeling, measurement, and / or drawing functions). In some cases, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality are limited to real-time implementations; in other cases, conventional methods are limited to implementations using previously captured media. In some embodiments, conventional methods for modeling and annotating physical environments and objects provide only a limited view of the physical environment / objects and the interactions between virtual objects and the physical environment / objects. Furthermore, conventional methods take longer than necessary, thereby wasting energy. This latter consideration is particularly important in battery-powered devices. Summary of the Invention
[0006] Therefore, there is a need for computer systems with improved methods and interfaces for modeling, measuring, and mapping using virtual / augmented reality environments. Such methods and interfaces optionally supplement or replace conventional methods for modeling, measuring, and mapping using virtual / augmented reality environments. Such methods and interfaces reduce the amount, extent, and / or nature of input from the user and produce a more efficient human-computer interface. For battery-powered devices, such methods and interfaces can save power and increase the time between battery charges.
[0007] The disclosed computer system is utilized to reduce or eliminate the above-mentioned defects and other problems associated with the user interface for modeling, measuring and drawing using virtual / augmented reality. In some embodiments, the computer system includes a desktop computer. In some embodiments, the computer system is portable (e.g., a laptop computer, a tablet computer or a handheld device). In some embodiments, the computer system includes a personal electronic device (e.g., a wearable electronic device, such as a watch). In some embodiments, the computer system has a touchpad (and / or communicates with the touchpad). In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display") (and / or communicates with the touch-sensitive display). In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory and one or more modules, a program or instruction set for performing multiple functions stored in the memory. In some embodiments, the user interacts with the GUI in part by stylus and / or finger contact and gestures on the touch-sensitive surface. In some embodiments, in addition to virtual / augmented reality-based modeling, measurement, and drawing functions, these functions optionally include playing games, image editing, drawing, presentations, word processing, spreadsheet creation, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, note taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] According to some embodiments, a method is performed at a computer system having a display generation component, an input device, and one or more cameras in a physical environment. The method includes capturing a representation of the physical environment via the one or more cameras, including updating the representation to include representations of corresponding portions of the physical environment in the field of view of the one or more cameras as the field of view of the one or more cameras moves. The method includes, after capturing the representation of the physical environment, displaying a user interface including an activatable user interface element for requesting display of a first orthogonal view of the physical environment. The method includes receiving, via the input device, user input corresponding to the activatable user interface element for requesting display of the first orthogonal view of the physical environment; and in response to receiving the user input, displaying the first orthogonal view of the physical environment based on the captured representation of the one or more portions of the physical environment.
[0009] According to some embodiments, a method is performed at a computer system having a display generation component, an input device, and one or more cameras in a physical environment. The method includes capturing information indicative of the physical environment via the one or more cameras, the information including information indicating a corresponding portion of the physical environment in the field of view of the one or more cameras when the field of view of the one or more cameras moves. The corresponding portion of the physical environment includes a plurality of primary features of the physical environment and one or more secondary features of the physical environment. The method includes, after capturing the information indicative of the physical environment, displaying a user interface, including simultaneously displaying: a plurality of primary features generated at a first fidelity level corresponding to the plurality of primary features of the physical environment; and one or more graphical representations of secondary features generated at a second fidelity level corresponding to the one or more secondary features of the physical environment, wherein the second fidelity level is lower than the first fidelity level.
[0010] According to some embodiments, a method is performed on a computer system having a display generation component and one or more input devices. The method includes displaying, via the display generation component, a representation of a physical environment, wherein the representation of the physical environment includes a representation of a first physical object, the first physical object occupying a first physical space in the physical environment and having first corresponding object properties; and a virtual object at a location in the representation of the physical environment corresponding to a second physical space in the physical environment that is different from the first physical space. The method includes detecting a first input corresponding to the virtual object, wherein movement of the first input corresponds to a request to move the virtual object in the representation of the physical environment relative to the representation of the first physical object. The method includes, upon detecting the first input, at least partially moving the virtual object in the representation of the physical environment based on the movement of the first input. Based on determining that the movement of the first input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to physical spaces in the physical environment that are not occupied by physical objects having the first corresponding object properties, at least partially moving the virtual object in the representation of the physical environment includes moving the virtual object by a first amount. Based on determining that movement of the first input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to a physical space in the physical environment that at least partially overlaps with a first physical space of the first physical object, at least partially moving the virtual object in the representation of the physical environment includes moving the virtual object a second amount that is less than the first amount, through at least a subset of the one or more locations corresponding to the physical space in the physical environment that at least partially overlaps with the first physical space of the first physical object.
[0011] According to some embodiments, a method is performed on a computer system having a display generation component and one or more input devices. The method includes displaying a first representation of a first previously captured media via the display generation component, wherein the first representation of the first media includes a representation of a physical environment. The method includes, while displaying the first representation of the first media, receiving input corresponding to a request to annotate a portion of the first representation corresponding to a first portion of the physical environment. The method includes, in response to receiving the input, displaying an annotation on the portion of the first representation corresponding to the first portion of the physical environment, the annotation having one or more of a position, orientation, or scale determined based on the physical environment. The method includes, after receiving the input, displaying an annotation on a portion of a second representation of a second previously captured media that is displayed, wherein the second previously captured media is different from the first previously captured media and the portion of the second representation corresponds to the first portion of the physical environment.
[0012] According to some embodiments, a method is performed at a computer system having a display generation component, an input device, and one or more cameras in a physical environment. The method includes displaying a first representation of a field of view of the one or more cameras via the display generation component, and receiving a first drawing input via the input device, the first drawing input corresponding to a request to add a first annotation to the first representation of the field of view. The method includes, in response to receiving the first drawing input: displaying a first annotation along a path of movement corresponding to the first drawing input in the first representation of the field of view of the one or more cameras; and after displaying the first annotation along the path of movement corresponding to the first drawing input, displaying an annotation constrained to correspond to an edge of a physical object in the physical environment based on determining that a corresponding portion of the first annotation corresponds to one or more locations within a threshold distance from an edge of the physical object.
[0013] According to some embodiments, a method is performed on a computer system having a display generation component and one or more input devices. The method includes displaying a representation of a first previously captured media item via the display generation component. The representation of the first previously captured media item is associated with (e.g., includes) depth information corresponding to a physical environment in which the first media item was captured. The method includes, while displaying the representation of the first previously captured media item, receiving, via the one or more input devices, one or more first inputs corresponding to a request to display, in the representation of the first previously captured media item, a first representation of a first measurement corresponding to a first corresponding portion of the physical environment captured in the first media item. The method includes, in response to receiving the one or more first inputs corresponding to a request to display the first representation of the first measurement in the representation of the first previously captured media item: displaying, via the display generation component, the first representation of the first measurement on at least a portion of the representation of the first previously captured media item corresponding to the first corresponding portion of the physical environment captured in the representation of the first media item based on the depth information associated with the first previously captured media item; and displaying, via the display generation component, a first label corresponding to the first representation of the first measurement, the first label describing the first measurement based on the depth information associated with the first previously captured media item.
[0014] According to some embodiments, a method is performed on a computer system having a display generation component and one or more input devices. The method includes displaying, via the display generation component, a representation of a first previously captured media item including a representation of a first physical environment from a first viewpoint. The method includes receiving, via the one or more input devices, an input corresponding to a request to display, from a second viewpoint, a representation of a second previously captured media item including a representation of a second physical environment. The method includes, in response to receiving the input corresponding to the request to display the representation of the second previously captured media item, displaying an animated transition from the representation of the first previously captured media item to the representation of the second previously captured media item based on determining that one or more attributes of the second previously captured media item meet a proximity criterion relative to one or more corresponding attributes of the first previously captured media item. The animated transition is based on a difference between a first viewpoint of the first previously captured media item and a second viewpoint of the second previously captured media item.
[0015] According to some embodiments, a method is performed on a computer system having a display generation component and one or more cameras. The method includes displaying a representation of a field of view of the one or more cameras via the display generation component. The representation of the field of view includes a representation of a first body in a physical environment in the field of view of the one or more cameras, and a corresponding portion of the representation of the first body in the representation of the field of view corresponds to a first anchor point on the first body. The method includes, while displaying the representation of the field of view: updating the representation of the field of view over time based on a change in the field of view. The change in the field of view includes movement of the first body that moves the first anchor point, and as the first anchor point moves along a path in the physical environment, the corresponding portion of the representation of the first body corresponding to the first anchor point changes along the path in the representation of the field of view corresponding to the movement of the first anchor point. The method includes displaying an annotation in the representation of the field of view corresponding to at least a portion of the path of the corresponding portion of the representation of the first body that corresponds to the first anchor point.
[0016] In accordance with some embodiments, a computer system (e.g., an electronic device) includes a display generating component (e.g., a display, a projector, a head-mounted display, a head-up display, etc.), one or more cameras (e.g., a camera that continuously or at fixed intervals provides a real-time preview of at least a portion of the content within the camera's field of view and optionally generates a video output comprising one or more image frames capturing the content within the camera's field of view), and one or more input devices (e.g., a touch-sensitive surface, such as a touch-sensitive remote control, or a touch screen display that also serves as a display generating component, a mouse, a joystick, a wand controller, and / or a camera that tracks one or more features of a user, such as the position of a user's hands), optionally one or more gesture sensors, optionally one or more sensors that detect intensity of contact with the touch-sensitive surface, optionally one or more tactile output generators, one or more processors, and a memory that stores one or more programs (and / or communicates with these components); the one or more programs are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the performance of operations of any of the methods described herein. In accordance with some embodiments, a computer-readable storage medium has stored therein instructions that, when executed by a computer system comprising a display generating component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting the intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators (and / or in communication with these components), causes the computer system to perform the operations of any of the methods described herein or causes the operations of any of the methods described herein to be performed. In accordance with some embodiments, a graphical user interface on a computer system comprising a display generating component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting the intensity of contact with a touch-sensitive surface, optionally one or more tactile output generators, a memory, and one or more processors for executing one or more programs stored in the memory (and / or in communication with these components) includes one or more elements displayed in any of the methods described herein, and the one or more elements are updated in response to input, as described in any of the methods described herein. In accordance with some embodiments, a computer system includes a display generating component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, optionally one or more tactile output generators, and means for performing or causing the operations of any method described herein to be performed (and / or communicating with these components).In accordance with some embodiments, an information processing device for use in a computer system that includes display generating components, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators (and / or communicates with these components) includes a device for performing the operations of any method described herein or causing the operations of any method described herein to be performed.
[0017] Thus, a computer system having (and / or communicating with) a display generation component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators has improved methods and interfaces for modeling, measuring, and mapping using virtual / augmented reality, thereby increasing the effectiveness, efficiency, and user satisfaction of such computer systems. Such methods and interfaces can supplement or replace conventional methods of modeling, measuring, and mapping using virtual / augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.
[0019] Figure 1A is a block diagram illustrating a portable multifunction device with a touch-sensitive display according to some embodiments.
[0020] Figure 1B is a block diagram illustrating example components for event processing according to some embodiments.
[0021] Figure 2A A portable multifunction device with a touch screen is shown according to some embodiments.
[0022] Figure 2B A portable multifunction device with an optical sensor and a time-of-flight sensor is shown according to some embodiments.
[0023] Figure 3A is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments.
[0024] Figures 3B to 3C is a block diagram of an exemplary computer system according to some embodiments.
[0025] Figure 4A An exemplary user interface for an application menu on a portable multifunction device according to some embodiments is shown.
[0026] Figure 4B An exemplary user interface for a multifunction device having a touch-sensitive surface separate from the display is shown according to some embodiments.
[0027] Figures 5A to 5LL An exemplary user interface for interacting with an augmented reality environment is shown according to some embodiments.
[0028] 6A to 6T An exemplary user interface for adding annotations to a media item is shown according to some embodiments.
[0029] 7A to 7B is a flowchart of a process for providing different views of a physical environment according to some embodiments.
[0030] Figures 8A to 8C is a flow diagram of a process for providing representations of a physical environment at different fidelity levels to the physical environment, according to some embodiments.
[0031] Figures 9A to 9G is a flow diagram of a process for displaying modeled spatial interactions between virtual objects / annotations and a physical environment, according to some embodiments.
[0032] Figures 10A to 10E is a flow diagram of a process for applying modeled-space interaction with virtual objects / annotations to multiple media items, according to some embodiments.
[0033] Figures 11A to 11JJ An exemplary user interface for scanning a physical environment and adding annotations to captured media items of the physical environment is shown according to some embodiments.
[0034] Figures 12A to 12RR An exemplary user interface for scanning a physical environment and adding measurements of objects in a captured media item corresponding to the physical environment is shown according to some embodiments.
[0035] Figures 13A to 13HH An exemplary user interface for transitioning between a displayed media item and a different media item selected by a user for viewing is shown in accordance with some embodiments.
[0036] Figures 14A to 14SS An exemplary user interface for viewing motion tracking information corresponding to a representation of a moving individual is shown according to some embodiments.
[0037] FIG. 15A to FIG. 15B Flowchart of a process for scanning a physical environment and adding annotations to captured media items of the physical environment according to some embodiments.
[0038] 16A to 16Eis a flowchart of a process for scanning a physical environment and adding measurements of objects in a captured media item that correspond to the physical environment, according to some embodiments.
[0039] 17A to 17D is a flowchart of a process for transitioning between a displayed media item and a different media item selected by a user for viewing, according to some embodiments.
[0040] 18A to 18B is a flowchart of a process for viewing motion tracking information corresponding to a representation of a moving individual according to some embodiments. DETAILED DESCRIPTION
[0041] As described above, augmented reality environments can be used to model and annotate physical environment spaces and objects therein by providing different views of the physical environment and objects therein and enabling users to overlay annotations such as measurements and drawings on the physical environment and objects therein, and visualizing the interactions between annotations and the physical environment and objects therein. Conventional methods for modeling and annotating using augmented reality environments are generally limited in functionality. In some cases, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality require multiple separate inputs (e.g., a series of gestures and button presses, etc.) to achieve the desired results (e.g., accessing different modeling, measurement, and / or drawing functions by activating multiple displayed user interface elements). In some cases, conventional methods for modeling and annotating physical environments and objects using augmented and / or virtual reality are limited to real-time implementations; in other cases, conventional methods are limited to implementations using previously captured media. In some embodiments, conventional methods for modeling and annotating physical environments and objects only provide limited views of the physical environment / objects and the interactions between virtual objects and the physical environment / objects. The embodiments disclosed herein provide a user with an intuitive way to model and annotate a physical environment using augmented and / or virtual reality (e.g., by enabling the user to perform different operations in the augmented and / or virtual reality environment with less input and / or by simplifying the user interface). Additionally, the embodiments disclosed herein provide improved feedback that provides the user with additional information about the physical environment and additional information and views of interactions with virtual objects, as well as information about operations performed in the augmented / virtual reality environment.
[0042] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways. For example, they make it easier to model and annotate physical environments by providing options for different views of the physical environment, presenting intuitive interactions between physical and virtual objects, and allowing annotations made in one view of the physical environment to be applied to other views of the physical environment.
[0043] under, Figure 1A to Figure 1B 、 Figures 2A to 2B as well as Figures 3A to 3B A description of an exemplary apparatus is provided. Figures 4A to 4B 、 Figures 5A to 5LL as well as 6A to 6T An exemplary user interface for interacting with and annotating an augmented reality environment and media items is shown. 7A to 7B A flow chart illustrating a method of providing different views of a physical environment is shown. Figures 8A to 8C A flow chart illustrating a method of providing representations of a physical environment at different fidelity levels thereto is shown. Figures 9A to 9G A flow chart showing a method of modeling spatial interactions between virtual objects / annotations and a physical environment is shown. Figures 10A to 10E A flow chart illustrating a method of applying modeled-space interaction with virtual objects / annotations to multiple media items. Figures 11A to 11JJ An exemplary user interface for scanning a physical environment and adding annotations to captured media items of the physical environment is shown. Figures 12A to 12RR An exemplary user interface is shown for scanning a physical environment and adding measurements of objects in a captured media item corresponding to the physical environment. Figures 13A to 13HH An exemplary user interface for transitioning between a displayed media item and a different media item selected by a user for viewing is shown. Figures 14A to 14SS An exemplary user interface for viewing motion tracking information corresponding to a representation of a moving individual is shown. FIG. 15A to FIG. 15B A flow chart illustrating a method of scanning a physical environment and adding annotations to captured media items of the physical environment is shown. 16A to 16E A flow chart illustrating a method of scanning a physical environment and adding measurements of objects in a captured media item corresponding to the physical environment. 17A to 17D A flow chart illustrating a method of transitioning between a displayed media item and a different media item selected by a user for viewing is shown. 18A to 18B A flow chart illustrating a method of viewing motion tracking information corresponding to a representation of a moving individual is shown. Figures 5A to 5LL 、 6A to 6T 、 Figures 11A to 11JJ 、 Figures 12A to 12RR 、 Figures 13A to 13HH and Figures 14A to 14SS The user interface in 7A to 7B 、 Figures 8A to 8C 、 Figures 9A to 9G 、 Figures 10A to 10E 、 FIG. 15A to FIG. 15B 、 16A to 16E 、 17A to 17D and 18A to 18B in the process.
[0044] Exemplary devices
[0045] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to one of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other cases, well-known methods, processes, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure various aspects of the embodiments.
[0046] It will also be understood that, although in some cases, the terms "first," "second," etc., are used to describe various elements herein, these elements should not be limited by these terms. These terms are simply used to distinguish one element from another. For example, a first element can be named a second element and similarly, a second element can be named a first element without departing from the scope of the various described embodiments. The first element and the second element are both contacts, but they are not the same element unless the context clearly indicates otherwise.
[0047] The terms used in the description of the various embodiments described herein are only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms "a" and "the" are intended to also include plural forms unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" used herein refer to and encompass any and all possible combinations of one or more items in the associated listed items. It will also be understood that the terms "includes," "including," "comprises," and / or "comprising" when used in this specification specify the presence of stated features, integers, steps, operations, elements, and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their groupings.
[0048] As used herein, the term "if" is optionally interpreted to mean "when" followed by "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined that" or "if [stated condition or event] is detected" are optionally interpreted to mean "upon determining" or "in response to determining" or "upon detecting [stated condition or event]" or "in response to detecting [stated condition or event]," depending on the context.
[0049] A computer system for virtual / augmented reality includes an electronic device that generates a virtual / augmented reality environment. Embodiments of electronic devices, user interfaces for such devices, and processes associated with using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as a PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, but are not limited to, the Apple Watch from Apple Inc. (Cupertino, California). iPod and Device. Other portable electronic devices, such as laptop computers or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or trackpads) are optionally used. It should also be understood that in some embodiments, the device is not a portable communication device, but rather a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or trackpads) that also includes or communicates with one or more cameras.
[0050] In the following discussion, a computer system is described that includes an electronic device having a display and a touch-sensitive surface (and / or communicating with these components). However, it should be understood that the computer system can optionally include one or more other physical user interface devices, such as a physical keyboard, a mouse, a joystick, a stylus controller, and / or a camera that tracks one or more characteristics of a user, such as the position of the user's hand.
[0051] The device typically supports a variety of applications, such as one or more of the following: a gaming application, a note-taking application, a drawing application, a presentation application, a word processing application, a spreadsheet application, a telephony application, a video conferencing application, an email application, an instant messaging application, a fitness support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0052] Various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed by the device are optionally adjusted and / or varied for different applications and / or adjusted and / or varied within the respective applications. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally supports the various applications with a user interface that is intuitive and clear to the user.
[0053] Attention is now turned to embodiments of portable devices having touch-sensitive displays. Figure 1A1 is a block diagram illustrating a portable multifunction device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display system 112 is sometimes referred to as a "touch screen" for convenience, and sometimes simply as a touch-sensitive display. The device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input or control devices 116, and external ports 124. The device 100 optionally includes one or more optical sensors 164 (e.g., as part of one or more cameras). The device 100 optionally includes one or more intensity sensors 165 for detecting the intensity of contact on the device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100). Device 100 optionally includes one or more tactile output generators 163 for generating tactile output on device 100 (e.g., generating tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0054] As used in this specification and claims, the term "tactile output" refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of a device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device that will be detected by a user using the user's sense of touch. For example, when a device or a component of the device is in contact with a surface that is touch-sensitive to a user (e.g., a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation that corresponds to a perceived change in a physical characteristic of the device or component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a "press click" or "release click" on a physical actuation button. In some cases, the user will feel a tactile sensation, such as a "press click" or "release click," even when the physical actuation button associated with the touch-sensitive surface that was physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when the smoothness of the touch-sensitive surface does not change, movement of the touch-sensitive surface may optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface. Although such interpretation of touch by the user will be limited by the user's individualized sensory perception, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of a user (e.g., "press click," "release click," "roughness"), unless otherwise stated, the tactile output generated corresponds to a physical displacement of the device or a component thereof that would generate the sensory perception of a typical (or average) user. Providing tactile feedback to the user using tactile output enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0055] It should be understood that device 100 is merely one example of a portable multifunction device and that device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. Figure 1A The various components shown in the drawings are implemented in hardware, software, firmware, or any combination thereof, including one or more signal processing circuits and / or application specific integrated circuits.
[0056] Memory 102 optionally includes high-speed random access memory and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Access to memory 102 by other components of device 100 (such as CPU 120 and peripheral device interface 118) is optionally controlled by memory controller 122.
[0057] Peripherals interface 118 may be used to couple the device's input and output peripherals to CPU 120 and memory 102. One or more processors 120 run or execute various software programs and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data.
[0058] In some embodiments, peripherals interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0059] RF (radio frequency) circuitry 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuitry 108 converts electrical signals into / from electromagnetic signals and communicates with a communication network and other communication devices via the electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, memory, and the like. RF circuitry 108 optionally communicates with networks and other devices via wireless communications, such as the Internet (also known as the World Wide Web (WWW)), an intranet, and / or a wireless network (such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN)). The wireless communication optionally uses any of a variety of communication standards, protocols and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), Evolution-Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Utilizing Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols not yet developed on the date of filing of this document.
[0060] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts the audio data into electrical signals, and transmits the electrical signals to the speaker 111. The speaker 111 converts the electrical signals into sound waves audible to humans. The audio circuit 110 also receives electrical signals converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signals into audio data and transmits the audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., Figure 2A The headset jack provides an interface between the audio circuit 110 and a removable audio input / output peripheral device, such as an output-only headset or a headset with both output (e.g., a single or dual-ear headset) and input (e.g., a microphone).
[0061] The I / O subsystem 106 couples input / output peripherals on the device 100, such as a touch-sensitive display system 112 and other input or control devices 116, to a peripherals interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, an intensity sensor controller 159, a tactile feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive / send electrical signals from / to other input or control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some alternative embodiments, the one or more input controllers 160 are optionally coupled to any (or none) of the following: a keyboard, an infrared port, a USB port, a stylus, and / or a pointer device such as a mouse. The one or more buttons (e.g., Figure 2A 208) optionally includes an up / down button for volume control of the speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., Figure 2A 206 in ).
[0062] The touch-sensitive display system 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch-sensitive display system 112 and / or sends electrical signals to the touch-sensitive display system. The touch-sensitive display system 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some visual outputs or all of the visual outputs correspond to user interface objects. As used herein, the term "indicator" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object that is configured to respond to input directed to a graphical user interface object). Examples of user-interactive graphical user interface objects include, but are not limited to, buttons, sliders, icons, selectable menu items, switches, hyperlinks, or other user interface controls.
[0063] The touch-sensitive display system 112 has a touch-sensitive surface, sensor, or sensor group that accepts input from the user based on tactile and / or haptic contact. The touch-sensitive display system 112 and the display controller 156 (together with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of that contact) on the touch-sensitive display system 112 and convert the detected contact into interaction with a user interface object (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch-sensitive display system 112. In some embodiments, the point of contact between the touch-sensitive display system 112 and the user corresponds to the user's finger or stylus.
[0064] The touch-sensitive display system 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies are used in other embodiments. The touch-sensitive display system 112 and display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch-sensitive display system 112 to detect contact and any movement or interruption thereof. In some embodiments, projected mutual capacitance sensing technology is used, such as the ATmega 2000 from Apple Inc. (Cupertino, California). iPod and Technology found in.
[0065] The touch-sensitive display system 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touch screen video resolution exceeds 400 dpi (e.g., 500 dpi, 800 dpi, or greater). The user optionally uses any suitable object or appendage, such as a stylus, finger, or the like, to contact the touch-sensitive display system 112. In some embodiments, the user interface is designed to work with finger-based contacts and gestures, which may not be as precise as stylus-based input due to the larger contact area of a finger on the touch screen. In some embodiments, the device converts rough finger-based input into precise pointer / cursor positions or commands for performing the actions desired by the user.
[0066] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad is optionally a touch-sensitive surface that is separate from the touch-sensitive display system 112 or an extension of the touch-sensitive surface formed by the touch screen.
[0067] Device 100 also includes a power system 162 for powering the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, power fault detection circuitry, a power converter or inverter, a power status indicator (e.g., a light emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0068] Device 100 optionally also includes one or more optical sensors 164 (eg, as part of one or more cameras). Figure 1A An optical sensor coupled to an optical sensor controller 158 in the I / O subsystem 106 is shown. One or more optical sensors 164 optionally include a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The one or more optical sensors 164 receive light projected from the environment through one or more lenses and convert the light into data representing an image. In conjunction with the imaging module 143 (also called a camera module), the one or more optical sensors 164 optionally capture still images and / or video. In some embodiments, the optical sensor is located on the rear portion of the device 100 opposite the touch-sensitive display system 112 on the front of the device, so that the touch screen can be used as a viewfinder for still image and / or video image acquisition. In some embodiments, another optical sensor is located on the front of the device to obtain an image of the user (e.g., for selfies, for video conferencing when the user views other video conference participants on the touch screen, etc.).
[0069] Device 100 optionally also includes one or more contact intensity sensors 165 . Figure 1A A contact force sensor is shown coupled to a force sensor controller 159 in the I / O subsystem 106. One or more contact force sensors 165 optionally include one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other force sensors (e.g., sensors for measuring the force (or pressure) of a contact on a touch-sensitive surface). One or more contact force sensors 165 receive contact force information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact force sensor is juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact force sensor is located on the rear portion of the device 100 opposite the touch-sensitive display system 112 located on the front portion of the device 100.
[0070] Device 100 optionally also includes one or more proximity sensors 166 . Figure 1A A proximity sensor 166 is shown coupled to the peripherals interface 118. Alternatively, the proximity sensor 166 is coupled to the input controller 160 in the I / O subsystem 106. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is on a phone call), the proximity sensor turns off and disables the touch-sensitive display system 112.
[0071] Device 100 optionally also includes one or more tactile output generators 163 . Figure 1AA tactile output generator is shown coupled to a tactile feedback controller 161 in the I / O subsystem 106. In some embodiments, one or more tactile output generators 163 include one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile outputs on the device). One or more tactile output generators 163 receive tactile feedback generation instructions from the tactile feedback module 133 and generate tactile outputs on the device 100 that can be felt by a user of the device 100. In some embodiments, at least one tactile output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112) and optionally generates tactile output by moving the touch-sensitive surface vertically (e.g., inward / outward toward the surface of the device 100) or laterally (e.g., back and forth in the same plane as the surface of the device 100). In some embodiments, at least one tactile output generator sensor is located on the back of the device 100 opposite the touch-sensitive display system 112 located on the front of the device 100.
[0072] The device 100 optionally also includes one or more accelerometers 167, gyroscopes 168, and / or magnetometers 169 (e.g., as part of an inertial measurement unit (IMU)) for obtaining information about the device's posture (e.g., position and orientation or location). Figure 1A Sensors 167, 168, and 169 are shown coupled to peripherals interface 118. Alternatively, sensors 167, 168, and 169 are optionally coupled to input controller 160 in I / O subsystem 106. In some embodiments, information is displayed on the touch screen display in a portrait view or a landscape view based on analysis of data received from the one or more accelerometers. Device 100 optionally includes a GPS (or GLONASS or other global navigation system) receiver for obtaining information about the location of device 100.
[0073] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a tactile feedback module (or instruction set) 133, a text input module (or instruction set) 134, a global positioning system (GPS) module (or instruction set) 135, and applications (or instruction sets) 136. In addition, in some embodiments, memory 102 stores a device / global internal state 157, as shown in Figures 1A and 3. The device / global internal state 157 includes one or more of the following: an active application state, which indicates which applications (if any) are currently active; a display state, which indicates what applications, views, or other information occupy various areas of the touch-sensitive display system 112; a sensor state, including information obtained from various sensors and other input or control devices 116 of the device; and position and / or location information about the device's posture (e.g., position and / or orientation).
[0074] The operating system 126 (e.g., iOS, Android, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.
[0075] The communication module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are suitable for coupling directly to other devices or indirectly through a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external ports are compatible with some of the Apple Inc. (Cupertino, California) iPod and In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or compatible with the 30-pin connector used in some Apple Inc. (Cupertino, California) iPod and In some embodiments, the external port is a USB Type-C connector that is the same as, similar to, and / or compatible with the Lightning connector used in some electronic devices from Apple Inc. (Cupertino, California).
[0076] The contact / motion module 130 optionally detects contact with the touch-sensitive display system 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection (e.g., by a finger or stylus), such as determining whether contact has occurred (e.g., detecting a finger press event), determining the strength of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there has been movement of the contact and tracking movement across the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has ceased (e.g., detecting a finger lift event or contact break). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of a contact point optionally includes determining a rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to a single point of contact (e.g., a single finger contact or stylus contact) or multiple points of simultaneous contact (e.g., "multi-touch" / multi-finger contact). In some embodiments, the contact / motion module 130 and display controller 156 detect contact on the touchpad.
[0077] The contact / motion module 130 optionally detects gesture input by the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of the detected contacts). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a single-finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at an icon location). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and then detecting a finger lift (lift-off) event. Similarly, taps, swipes, drags, and other gestures of the stylus are optionally detected by detecting a specific contact pattern of the stylus.
[0078] In some embodiments, detecting a finger tap gesture depends on detecting the length of time between a finger press event and a finger lift event, but is independent of detecting the intensity of the finger contact between the finger press event and the finger lift event. In some embodiments, a tap gesture is detected based on determining that the length of time between the finger press event and the finger lift event is less than a predetermined value (e.g., less than 0.1, 0.2, 0.3, 0.4, or 0.5 seconds), regardless of whether the intensity of the finger contact during the tap reaches a given intensity threshold (greater than a nominal contact detection intensity threshold), such as a light press or deep press intensity threshold. Thus, a finger tap gesture can meet a specific input criterion that does not require the characteristic intensity of the contact to meet a given intensity threshold to meet the specific input criterion. For clarity, the finger contact in a tap gesture is generally required to meet a nominal contact detection intensity threshold to detect a finger press event, below which the contact is not detected. A similar analysis applies to detecting a tap gesture by a stylus or other contact. In cases where the device is capable of detecting contact by a finger or stylus hovering over the touch-sensitive surface, the nominal contact detection intensity threshold optionally does not correspond to physical contact between the finger or stylus and the touch-sensitive surface.
[0079] The same concepts apply in a similar manner to other types of gestures. For example, a swipe gesture, a pinch gesture, an expand gesture, and / or a long press gesture may be optionally detected based on satisfying criteria that are independent of the intensity of the contacts included in the gesture or that do not require one or more contacts performing the gesture to reach an intensity threshold in order to be recognized. For example, a swipe gesture is detected based on the amount of movement of one or more contacts; a zoom gesture is detected based on the movement of two or more contacts toward each other; a zoom gesture is detected based on the movement of two or more contacts away from each other; and a long press gesture is detected based on the duration of a contact on the touch-sensitive surface having less than a threshold amount of movement. Thus, the statement that a particular gesture recognition criterion does not require the intensity of the contacts to meet the corresponding intensity threshold in order to satisfy the particular gesture recognition criterion means that the particular gesture recognition criterion can be satisfied when the contacts in the gesture do not reach the corresponding intensity threshold, and can also be satisfied when one or more contacts in the gesture reaches or exceeds the corresponding intensity threshold. In some embodiments, a tap gesture is detected based on determining that a finger down event and a finger up event are detected within a predefined time period, regardless of whether the contact is above or below a corresponding intensity threshold during the predefined time period, and a swipe gesture is detected based on determining that the contact moves by more than a predefined amount, even if the contact is above the corresponding intensity threshold at the end of the contact movement. Even in specific implementations where detection of gestures is affected by the intensity of the contact performing the gesture (e.g., the device detects a long press more quickly when the intensity of the contact is above an intensity threshold, or the device delays detection of a tap input when the intensity of the contact is higher), detection of these gestures does not require the contact to reach a particular intensity threshold (e.g., even if the amount of time required to recognize the gesture varies), as long as the criteria for recognizing the gesture can be met without the contact reaching the particular intensity threshold.
[0080] In some cases, contact intensity thresholds, duration thresholds, and movement thresholds are combined in various combinations to create heuristic algorithms that distinguish between two or more different gestures directed to the same input element or area, enabling multiple different interactions with the same input element to provide a richer set of user interactions and responses. A statement that a particular set of gesture recognition criteria does not require the intensity of one or more contacts to meet the corresponding intensity threshold in order to meet the particular gesture recognition criteria does not preclude the simultaneous evaluation of other intensity-related gesture recognition criteria to identify other gestures whose criteria are met when the gesture includes a contact with an intensity above the corresponding intensity threshold. For example, in some cases, a first gesture recognition criterion for a first gesture (which does not require the intensity of the contact to meet the corresponding intensity threshold to meet the first gesture recognition criterion) competes with a second gesture recognition criterion for a second gesture (which depends on the contact meeting the corresponding intensity threshold). In such a competition, if the second gesture recognition criterion for the second gesture is met first, the gesture is optionally not recognized as meeting the first gesture recognition criterion for the first gesture. For example, if the contact reaches the corresponding intensity threshold before the contact moves a predefined amount, a deep press gesture is detected instead of a swipe gesture. Conversely, if the contact moves a predefined amount before the contact reaches the corresponding intensity threshold, a swipe gesture is detected instead of a deep press gesture. Even in such cases, the first gesture recognition criterion for the first gesture still does not require the intensity of the contact to meet the corresponding intensity threshold in order to satisfy the first gesture recognition criterion because, if the contact remains below the corresponding intensity threshold until the gesture ends (e.g., a swipe gesture with an intensity that does not increase above the corresponding intensity threshold), the gesture will be recognized as a swipe gesture by the first gesture recognition criterion. Thus, a particular gesture recognition criterion that does not require the intensity of the contact to meet the corresponding intensity threshold in order to satisfy the particular gesture recognition criterion may (A) in some cases ignore the intensity of the contact relative to the intensity threshold (e.g., for a tap gesture) and / or (B) in some cases fail to satisfy the particular gesture recognition criterion (e.g., for a long press gesture) if a set of competing intensity-related gesture recognition criteria (e.g., for a deep press gesture) recognizes the input as corresponding to an intensity-related gesture before the particular gesture recognition criterion recognizes the gesture corresponding to the input, and in this sense still depends on the intensity of the contact relative to the intensity threshold (e.g., for a long press gesture competing with a deep press gesture for recognition).
[0081] In conjunction with accelerometer 167, gyroscope 168, and / or magnetometer 169, posture module 131 optionally detects posture information about the device, such as the posture (e.g., roll, pitch, yaw, and / or position) of the device in a particular reference frame. Posture module 131 includes software components for performing various operations related to detecting device position and detecting changes in device posture.
[0082] The graphics module 132 includes various known software components for rendering and displaying graphics on the touch-sensitive display system 112 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual attributes) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0083] In some embodiments, the graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes specifying the graphics to be displayed from an application program or the like, along with coordinate data and other graphic attribute data, if necessary, and then generates screen image data for output to the display controller 156.
[0084] The tactile feedback module 133 includes various software components for generating instructions (e.g., instructions used by the tactile feedback controller 161) to produce tactile output at one or more locations on the device 100 using one or more tactile output generators 163 in response to user interaction with the device 100.
[0085] Text input module 134, which is optionally a component of graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).
[0086] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for location-based dialing; to the camera 143 as picture / video metadata; and to applications that provide location-based services such as the weather widget, the local yellow pages widget, and the map / navigation widget).
[0087] The virtual / augmented reality module 145 provides virtual and / or augmented reality logic to the applications 136 that implement augmented reality features, and in some embodiments, virtual reality features. The virtual / augmented reality module 145 facilitates the overlay of virtual content, such as virtual user interface objects, on a representation of at least a portion of the field of view of one or more cameras. For example, with the assistance of the virtual / augmented reality module 145, the representation of at least a portion of the field of view of one or more cameras can include corresponding physical objects, and the virtual user interface objects can be displayed in a displayed augmented reality environment at locations determined based on the corresponding physical objects in the field of view of the one or more cameras, or in a virtual reality environment determined based on a posture of at least a portion of the computer system (e.g., a posture of a display device used to display a user interface to a user of the computer system).
[0088] Application 136 optionally includes the following modules (or instruction sets), or a subset or superset thereof:
[0089] Contacts module 137 (sometimes called address book or contact list);
[0090] Telephone module 138;
[0091] Video conferencing module 139;
[0092] Email client module 140;
[0093] Instant messaging (IM) module 141;
[0094] Fitness support module 142;
[0095] A camera module 143 for still and / or video images;
[0096] Image management module 144;
[0097] Browser module 147;
[0098] Calendar module 148;
[0099] Widget module 149, which optionally includes one or more of the following: weather widget 149-1, stock widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets obtained by the user, and user-created widgets 149-6;
[0100] A widget creator module 150 for forming user-created widgets 149-6;
[0101] Search module 151;
[0102] Video and music player module 152, optionally consisting of a video player module and a music player module;
[0103] Memo module 153;
[0104] Map module 154;
[0105] Online Video Module 155
[0106] Modeling and annotation module 195; and / or
[0107] • Time of Flight (“ToF”) sensor module 196 .
[0108] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice replication.
[0109] In combination with the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, and the text input module 134, the contacts module 137 includes executable instructions for managing an address book or contact list (e.g., stored in the application internal state 192 of the contacts module 137 in memory 102 or memory 370), including: adding names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers and / or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing 139, email 140, or instant messaging 141; and so on.
[0110] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, phone module 138 includes executable instructions for entering a character sequence corresponding to a phone number, accessing one or more phone numbers in address book 137, modifying an entered phone number, dialing the corresponding phone number, conducting a conversation, and disconnecting or hanging up when the conversation is complete. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0111] In combination with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, one or more optical sensors 164, optical sensor controller 158, contact module 130, graphics module 132, text input module 134, contact list 137, and phone module 138, video conferencing module 139 includes executable instructions for initiating, conducting, and terminating a video conference between a user and one or more other participants in accordance with user instructions.
[0112] In conjunction with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. In conjunction with image management module 144, email client module 140 makes it very easy to create and send emails with still images or video images captured by camera module 143.
[0113] In conjunction with the RF circuitry 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, and the text input module 134, the instant messaging module 141 includes executable instructions for entering a character sequence corresponding to an instant message, modifying previously entered characters, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for phone-based instant messaging or using XMPP, SIMPLE, Apple Push Notification Service (APNs), or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, APNs, or IMPS).
[0114] In combination with the RF circuitry 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, the GPS module 135, the map module 154, and the video and music player module 152, the fitness support module 142 includes executable instructions for creating a workout (e.g., with time, distance, and / or calorie burn goals); communicating with fitness sensors (in sports equipment and smart watches); receiving fitness sensor data; calibrating sensors for monitoring a workout; selecting and playing music for a workout; and displaying, storing, and transmitting fitness data.
[0115] In conjunction with touch-sensitive display system 112, display controller 156, one or more optical sensors 164, optical sensor controller 158, contact module 130, graphics module 132, and image management module 144, camera module 143 includes executable instructions for capturing still images or videos (including video streams) and storing them in memory 102, modifying characteristics of still images or videos, and / or deleting still images or videos from memory 102.
[0116] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and camera module 143, image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0117] In combination with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132 and the text input module 134, the browser module 147 includes executable instructions for browsing the Internet (including searching for, linking to, receiving and displaying web pages or portions thereof, as well as attachments and other files linked to web pages) in accordance with user instructions.
[0118] In combination with the RF circuitry 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, the email client module 140, and the browser module 147, the calendar module 148 includes executable instructions for creating, displaying, modifying, and storing calendars and data associated with the calendars (e.g., calendar entries, to-do items, etc.) in accordance with user instructions.
[0119] In conjunction with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, and the browser module 147, the desktop widget module 149 is a mini-application that is optionally downloaded and used by the user (e.g., the weather desktop widget 149-1, the stock desktop widget 149-2, the calculator desktop widget 149-3, the alarm desktop widget 149-4, and the dictionary desktop widget 149-5) or a mini-application created by the user (e.g., the user-created desktop widget 149-6). In some embodiments, the desktop widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the desktop widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., the Yahoo! desktop widget).
[0120] In combination with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, and the browser module 147, the desktop widget creator module 150 includes executable instructions for creating a desktop widget (e.g., transferring a user-specified portion of a web page into a desktop widget).
[0121] In combination with the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, and the text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.
[0122] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats (such as MP3 or AAC files), as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch-sensitive display system 112 or on an external display wirelessly connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (trademark of Apple Inc.).
[0123] In conjunction with the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, and the text input module 134, the memo module 153 includes executable instructions for creating and managing memos, to-do lists, etc. according to user instructions.
[0124] In combination with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 includes executable instructions for receiving, displaying, modifying, and storing maps and data associated with the maps (e.g., driving directions; data about stores and other points of interest at or near a particular location; and other location-based data) in accordance with user instructions.
[0125] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, online video module 155 includes executable instructions that allow a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on touch screen 112 or on an external display connected wirelessly or via external port 124), send an email with a link to a particular online video, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 141 is used instead of email client module 140 to send a link to a particular online video.
[0126] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, camera module 143, image management module 152, video and music player module 152, and virtual / augmented reality module 145, modeling and annotation module 195 includes executable instructions that allow a user to model a physical environment and / or physical objects therein and annotate (e.g., measure, draw, and / or add virtual objects and manipulate virtual objects therein) representations (e.g., real-time or previously captured) of the physical environment and / or physical objects therein in an augmented and / or virtual reality environment, as described in more detail herein.
[0127] In conjunction with the camera module 143, the ToF sensor module 196 includes executable instructions for capturing depth information of the physical environment. In some embodiments, the ToF sensor module 196 operates in conjunction with the camera module 143 to provide depth information of the physical environment.
[0128] Each module and application identified above corresponds to a set of executable instructions for performing one or more functions and methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (i.e., instruction sets) do not have to be implemented as independent software programs, processes, or modules, so various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 102 optionally stores a subset of the above-mentioned modules and data structures. In addition, memory 102 optionally stores other modules and data structures not described above.
[0129] In some embodiments, device 100 is a device in which operation of a predefined set of functions on the device is performed exclusively through a touch screen and / or a touchpad. By using a touch screen and / or a touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on device 100 is optionally reduced.
[0130] A predefined set of functions that are exclusively performed through the touch screen and / or trackpad optionally includes navigation between user interfaces. In some embodiments, the trackpad, when touched by the user, navigates the device 100 from any user interface displayed on the device 100 to a main menu, home menu, or root menu. In such embodiments, a touch-sensitive surface is used to implement a "menu button." In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touch-sensitive surface.
[0131] Figure 1B is a block diagram illustrating exemplary components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3A ) includes an event classifier 170 (e.g., in the operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 136, 137 to 155, 380 to 390).
[0132] The event classifier 170 receives event information and determines the application 136-1 and the application view 191 of the application 136-1 to which the event information is to be delivered. The event classifier 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, the application 136-1 includes an application internal state 192 that indicates one or more current application views displayed on the touch-sensitive display system 112 when the application is active or executing. In some embodiments, the device / global internal state 157 is used by the event classifier 170 to determine which application(s) is currently active, and the application internal state 192 is used by the event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0133] In some embodiments, the application internal state 192 includes additional information, such as one or more of the following: resumption information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by the application 136-1, a state queue for enabling the user to return to a previous state or view of the application 136-1, and a redo / undo queue of previous actions taken by the user.
[0134] Event monitor 171 receives event information from peripherals interface 118. The event information includes information about sub-events (e.g., a user touch on touch-sensitive display system 112 as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 167, and / or microphone 113 (through audio circuit 110). The information received by peripherals interface 118 from I / O subsystem 106 includes information from touch-sensitive display system 112 or a touch-sensitive surface.
[0135] In some embodiments, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other embodiments, peripheral device interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or receiving an input for more than a predetermined duration).
[0136] In some embodiments, the event classifier 170 also includes a hit view determination module 172 and / or an active event identifier determination module 173.
[0137] When the touch-sensitive display system 112 displays more than one view, the hit view determination module 172 provides software procedures for determining where within one or more views a sub-event has occurred. A view consists of controls and other elements that a user can see on the display.
[0138] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of the respective application) in which a touch is detected optionally correspond to programmatic levels within the application's programmatic or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as a hit view, and the set of events recognized as correct input is optionally determined based at least in part on the hit view of the initial touch that started the touch-based gesture.
[0139] Hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy where the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (i.e., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once a hit view is identified by the hit view determination module, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.
[0140] Active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with one particular view, views higher in the hierarchy will still remain actively participating views.
[0141] Event dispatcher module 174 dispatches event information to event recognizers (e.g., event recognizer 180). In embodiments that include active event recognizer determination module 173, event dispatcher module 174 delivers the event information to the event recognizer determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores the event information in an event queue, which is retrieved by corresponding event receiver module 182.
[0142] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or part of another module stored in memory 102, such as contact / motion module 130.
[0143] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events that occur within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes multiple event recognizers 180. In other embodiments, one or more of event recognizers 180 is part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handler 190 optionally utilizes or calls data updater 176, object updater 177, or GUI updater 178 to update the application's internal state 192. Alternatively, one or more of the application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0144] A corresponding event identifier 180 receives event information (e.g., event data 179) from event classifier 170 and identifies an event from the event information. Event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event identifier 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0145] The event receiver 182 receives event information from the event classifier 170. The event information includes information about sub-events such as touches or touch movements. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current posture (e.g., position and orientation) of the device.
[0146] Event comparator 184 compares event information with predefined event or sub-event definitions and determines an event or sub-event based on the comparison, or determines or updates the state of an event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 includes the definition of an event (e.g., a predefined sequence of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, the sub-events in event 187 include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, the definition of event 1 (187-1) is a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on a displayed object, a first lift (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on a displayed object, and a second lift (touch end) of a predetermined duration. In another example, the definition of event 2 (187-2) is a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on a displayed object, movement of the touch on the touch-sensitive display system 112, and lifting of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0147] In some embodiments, event definition 187 includes definitions of events for corresponding user interface objects. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display system 112, when a touch is detected on touch-sensitive display system 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects an event handler that is associated with the sub-event and the object that triggered the hit test.
[0148] In some embodiments, the definition of the corresponding event 187 also includes delay actions that delay the delivery of the event information until it has been determined that the sub-event sequence does or does not correspond to the event type of the event identifier.
[0149] When a corresponding event recognizer 180 determines that a sequence of sub-events does not match any event in event definitions 186, the corresponding event recognizer 180 enters the event impossible, event failed, or event ended state, after which subsequent sub-events of the touch-based gesture are ignored. In this case, other event recognizers (if any) that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.
[0150] In some embodiments, corresponding event recognizers 180 include metadata 183 with configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery for actively participating event recognizers. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact or can interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0151] In some embodiments, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates an event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating an event handler 190 is different from sending (and deferred sending) sub-events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with the flag obtains the flag and performs a predefined process.
[0152] In some embodiments, event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or with the actively participating view receives the event information and performs a predetermined process.
[0153] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video or music player module 152. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the location of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on a touch-sensitive display.
[0154] In some embodiments, event handler 190 includes or has access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0155] It should be understood that the above discussion of event handling for user touches on a touch-sensitive display also applies to other forms of user input utilizing input devices to operate the multifunction device 100, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses, optionally in conjunction with single or multiple keyboard presses or holddowns; contact movement on a trackpad, such as taps, drags, scrolls, etc.; stylus input; input based on real-time analysis of video images obtained by one or more cameras; movement of the device; spoken commands; detected eye movement; biometric input; and / or any combination thereof, are optionally used as input corresponding to sub-events defining the event to be recognized.
[0156] Figure 2A A touch screen (e.g., touch-sensitive display system 112, Figure 1A) of a portable multifunction device 100 (e.g., a view of the front of the device 100). The touch screen optionally displays one or more graphics within the user interface (UI) 200. In these embodiments, and in other embodiments described below, a user can select one or more of the graphics by, for example, making a gesture on the graphics using one or more fingers 202 (not drawn to scale in the figures) or one or more styluses 203 (not drawn to scale in the figures). In some embodiments, selection of the one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or a scrolling of a finger that has made contact with the device 100 (from right to left, from left to right, up and / or down). In some specific implementations or in some cases, inadvertent contact with a graphic does not select the graphic. For example, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.
[0157] The device 100 optionally also includes one or more physical buttons, such as a "home" or menu button 204. As previously described, the menu button 204 is optionally used to navigate to any application 136 in a set of applications that are optionally executed on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen display.
[0158] In some embodiments, the device 100 includes a touch screen display, a menu button 204 (sometimes referred to as a home button 204), a push button 206 for powering the device on / off and for locking the device, a volume adjustment button 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and holding the button in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing the button before the predefined time interval has passed; and / or to unlock the device or initiate an unlocking process. In some embodiments, the device 100 also accepts voice input for activating or deactivating certain functions through the microphone 113. The device 100 also optionally includes one or more contact force sensors 165 for detecting contact force on the touch-sensitive display system 112, and / or one or more tactile output generators 163 for generating tactile output for the user of the device 100.
[0159] Figure 2BPortable multifunction device 100 is shown (e.g., a view of the back of device 100) optionally including optical sensors 164-1 and 164-2 and a time-of-flight ("ToF") sensor 220. When optical sensors (e.g., cameras) 164-1 and 164-2 simultaneously capture representations of a physical environment (e.g., images or video), the portable multifunction device can determine depth information based on the difference between the information simultaneously captured by the optical sensors (e.g., the difference between the captured images). The depth information provided by the difference (e.g., images) determined using optical sensors 164-1 and 164-2 may lack accuracy, but typically provides high resolution. To improve the accuracy of the depth information provided by the difference between the images, a time-of-flight sensor 220 is optionally used in conjunction with optical sensors 164-1 and 164-2. ToF sensor 220 transmits a waveform (e.g., light from a light-emitting diode (LED) or laser) and measures the time it takes for a reflection of the waveform (e.g., light) to return to ToF sensor 220. Depth information is determined based on the measured time it takes for the light to return to ToF sensor 220. ToF sensors typically provide high accuracy (e.g., 1 cm or better with respect to the measured distance or depth), but may lack high resolution. Therefore, combining depth information from the ToF sensor with depth information provided by disparity determined using an optical sensor (e.g., a camera) (e.g., an image) provides a depth map that is both accurate and has high resolution.
[0160] Figure 3A is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments. The device 300 does not have to be portable. In some embodiments, the device 300 is a laptop, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a gaming system, or a control device (e.g., a home controller or an industrial controller). The device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, a memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuits (sometimes referred to as a chipset) that interconnect system components and control communications between system components. The device 300 includes an input / output (I / O) interface 330 having a display 340, which is optionally a touch screen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 for generating tactile output on the device 300 (e.g., similar to the above referenced devices). Figure 1A The tactile output generator 163), sensor 359 (e.g., similar to the above reference Figure 1AOptical, acceleration, proximity, touch and / or contact intensity sensors of the type described above, and optionally the above referenced Figure 2B The time-of-flight sensor 220 is described. The memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 370 optionally includes one or more storage devices located remotely from the CPU 310. In some embodiments, the memory 370 stores data related to the portable multifunction device 100 ( Figure 1A ), or a subset thereof. In addition, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while portable multifunction device 100( Figure 1A )'s memory 102 optionally does not store these modules.
[0161] Figure 3A Each element in the above-mentioned identified element is optionally stored in one or more memory devices in the previously mentioned memory device.Each module in the above-mentioned identified module corresponds to the instruction set for performing the above-mentioned functions.The above-mentioned identified module or program (that is, instruction set) need not be implemented as an independent software program, process or module, so the various subsets of these modules are optionally combined or otherwise rearranged in various embodiments.In some embodiments, memory 370 optionally stores the subset of above-mentioned modules and data structure.In addition, memory 370 optionally stores additional modules and data structure not described above.
[0162] Figures 3B-3C is a block diagram of an exemplary computer system 301 according to some embodiments.
[0163] In some embodiments, computer system 301 includes and / or is in communication with the following components:
[0164] Input devices (302 and / or 307, e.g., a touch-sensitive surface such as a touch-sensitive remote control, or a touch screen display that also serves as a display generating component, a mouse, a joystick, a stylus controller, and / or a camera that tracks one or more characteristics of a user, such as the position of a user's hand);
[0165] Virtual / augmented reality logic 303 (e.g., virtual / augmented reality module 145);
[0166] Display generation components (304 and / or 308, e.g., displays, projectors, head-mounted displays, heads-up displays, etc.) for displaying virtual user interface elements to a user;
[0167] A camera (e.g., 305 and / or 311 ) for capturing images of the device's field of view, e.g., for determining placement of virtual user interface elements, determining the device's pose, and / or displaying an image of a portion of the physical environment in which the camera is located; and
[0168] • Posture sensors (eg, 306 and / or 311 ) to determine the posture of the device relative to the physical environment and / or changes in the posture of the device.
[0169] In some computer systems, the camera (e.g., 305 and / or 311) includes a time-of-flight sensor (e.g., time-of-flight sensor 220, Figure 2B ), used as reference above Figure 2B The depth information is captured.
[0170] On some computer systems (e.g. Figure 3B In 301-a), the input device 302, the virtual / augmented reality logic component 303, the display generation component 304, the camera 305; and the posture sensor 306 are all integrated into the computer system (e.g., Figure 1A To the picture Figure 1B 3 or device 300, such as a smart phone or tablet computer).
[0171] In some computer systems (e.g., 301-b), in addition to the integrated input device 302, virtual / augmented reality logic component 303, display generation component 304, camera 305; and gesture sensor 306, the computer system also communicates with additional devices independent of the computer system, such as an independent input device 307, such as a touch-sensitive surface, stylus, remote control, etc. and / or an independent display generation component 308, such as a virtual reality headset or augmented reality glasses that overlay virtual objects on the physical environment.
[0172] On some computer systems (e.g. Figure 3CIn 301-c), input device 307, display generation component 309, camera 311, and / or gesture sensor 312 are separate from and in communication with the computer system. In some embodiments, other combinations of components in computer system 301 and in communication with the computer system are used. For example, in some embodiments, display generation component 309, camera 311, and gesture sensor 312 are incorporated into a headset that is integrated with or in communication with the computer system.
[0173] In some embodiments, the following reference Figures 5A to 5LL and 6A to 6T All operations described are performed on a single computing device with virtual / augmented reality logic 303 (e.g., Figure 3B However, it should be understood that multiple different computing devices are often linked together to execute the following referenced computer system 301-a). Figures 5A to 5LL and 6A to 6T In any of these embodiments, the following reference is made to the operations described (e.g., a computing device having virtual / augmented reality logic 303 communicating with a separate computing device having display 450 and / or a separate computing device having touch-sensitive surface 451). Figures 5A to 5LL and 6A to 6T The computing devices depicted are one or more computing devices that include virtual / augmented reality logic 303. Additionally, it should be understood that in various embodiments, virtual / augmented reality logic 303 may be divided among multiple different modules or computing devices; however, for the purposes of this description, virtual / augmented reality logic 303 will primarily be referred to as residing in a single computing device to avoid unnecessarily obscuring other aspects of the embodiments.
[0174] In some embodiments, the virtual / augmented reality logic 303 includes one or more modules that receive and interpret input (e.g., one or more event handlers 190, including those described above with reference to FIG). Figure 1B In some embodiments, the graphical user interface may be updated based on the interpreted inputs (e.g., one or more object updaters 177 and one or more GUI updaters 178, described in more detail herein), and in response to the interpreted inputs, generates instructions for updating the graphical user interface based on the interpreted inputs, which are then used to update the graphical user interface on the display. Figure 1A and contact motion module 130 in FIG. 3 ), identifying (e.g., by Figure 1B event identifier 180 in ) and / or assigned (e.g., via Figure 1BThe interpreted input of the input received by the event classifier 170 in the touch-sensitive surface 451 is used to update the graphical user interface on the display. In some embodiments, the interpreted input is generated by a module on the computing device (e.g., the computing device receives raw contact input data to identify gestures from the raw contact input data). In some embodiments, some or all of the interpreted input is received as interpreted input by the computing device (e.g., the computing device including the touch-sensitive surface 451 processes the raw contact input data to identify gestures from the raw contact input data and sends information indicative of the gesture to the computing device including the virtual / augmented reality logic component 303).
[0175] In some embodiments, both the display and the touch-sensitive surface are connected to a computer system (e.g., Figure 3B For example, the computer system may be a desktop computer or laptop computer with an integrated display (e.g., 340 in FIG. 3 ) and a touchpad (e.g., 355 in FIG. 3 ). For another example, the computing device may be a computer with a touch screen (e.g., Figure 2A 112) of a portable multifunction device 100 (e.g., a smart phone, a PDA, a tablet computer, etc.).
[0176] In some embodiments, the touch-sensitive surface is integrated with the computer system, while the display is not integrated with the computer system including the virtual / augmented reality logic 303. For example, the computer system can be a device 300 (e.g., a desktop or laptop computer, etc.) with an integrated touchpad (e.g., 355 in FIG. 3 ), where the integrated touchpad is connected (via a wired or wireless connection) to a separate display (e.g., a computer monitor, a television, etc.). For another example, the computer system can be a device 300 (e.g., a desktop or laptop computer, etc.) with an integrated touchpad (e.g., 355 in FIG. 3 ). Figure 2A 112) of a portable multifunction device 100 (e.g., a smart phone, PDA, tablet computer, etc.), wherein the touch screen is connected (via a wired or wireless connection) to a separate display (e.g., a computer monitor, a television, etc.).
[0177] In some embodiments, a display is integrated with a computer system, while a touch-sensitive surface is not integrated with the computer system including the virtual / augmented reality logic 303. For example, the computer system can be a device 300 (e.g., a desktop computer, a laptop computer, a television with an integrated set-top box) with an integrated display (e.g., 340 in FIG. 3 ), where the integrated display is coupled (via a wired or wireless connection) to a separate touch-sensitive surface (e.g., a remote touchpad, a portable multifunction device, etc.). As another example, the computer system can be a portable multifunction device 100 (e.g., a smartphone, a PDA, a tablet computer, etc.) with a touch screen (e.g., 112 in FIG. 2 ), where the touch screen is coupled (via a wired or wireless connection) to a separate touch-sensitive surface (e.g., a remote touchpad, a portable multifunction device with another touch screen acting as a remote touchpad, etc.).
[0178] In some embodiments, neither the display nor the touch-sensitive surface is connected to a computer system (e.g., Figure 3C For example, the computer system can be a standalone computing device 300 (e.g., a set-top box, a game console, etc.) connected (via a wired or wireless connection) to a standalone touch-sensitive surface (e.g., a remote trackpad, a portable multifunction device, etc.) and a standalone display (e.g., a computer monitor, a television, etc.).
[0179] In some embodiments, the computer system has an integrated audio system (e.g., audio circuitry 110 and speakers 111 in portable multifunction device 100). In some embodiments, the computing device communicates with an audio system that is independent of the computing device. In some embodiments, the audio system (e.g., an audio system integrated into a television unit) is integrated with a separate display. In some embodiments, the audio system (e.g., a stereo system) is a separate system from the computer system and display.
[0180] Attention is now turned to an embodiment of a user interface (“UI”) that is optionally implemented on portable multifunction device 100 .
[0181] Figure 4A An exemplary user interface for an application menu on portable multifunction device 100 is shown according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof:
[0182] Signal strength indicators for wireless communications, such as cellular and Wi-Fi signals;
[0183] ·time;
[0184] Bluetooth indicator;
[0185] Battery status indicator;
[0186] A tray 408 with common application icons, such as:
[0187] o An icon 416 labeled “Phone” of the phone module 138 , which optionally includes an indicator 414 of the number of missed calls or voice messages;
[0188] o An icon 418 of the email client module 140 labeled “Mail,” which optionally includes an indicator 410 of the number of unread emails;
[0189] o An icon 420 labeled "Browser" of the browser module 147; and
[0190] o Icon 422 labeled “Music” of the video and music player module 152;
[0191] as well as
[0192] Icons for other applications, such as:
[0193] o Icon 424 labeled “Messages” of IM module 141;
[0194] o Icon 426 labeled “Calendar” of calendar module 148;
[0195] o Icon 428 labeled “Photos” of the image management module 144;
[0196] o An icon 430 labeled “Camera” of the camera module 143;
[0197] o Icon 432 labeled “Online Video” of the online video module 155;
[0198] o Icon 434 labeled “Stock Market” of the Stock Market widget 149-2;
[0199] o Icon 436 labeled “Map” of the map module 154;
[0200] o Icon 438 labeled “Weather” of the weather widget 149-1;
[0201] o Icon 440 labeled “Clock” of the alarm clock widget 149-4;
[0202] o An icon 442 labeled “Fitness Support” of the fitness support module 142 ;
[0203] o An icon 444 labeled “Memo” of the memo module 153 ; and
[0204] o An icon 446 of a settings application or module labeled “Settings” that provides access to settings for the device 100 and its various applications 136 .
[0205] It should be noted that Figure 4A The icon labels shown in the are merely exemplary. For example, other labels are optionally used for various application icons. In some embodiments, the label of a respective application icon includes the name of the application corresponding to the respective application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.
[0206] Figure 4B A touch-sensitive surface 451 (e.g., Figure 3A a device (e.g., a tablet or touchpad 355) Figure 3A Although many of the examples that follow will be given with reference to input on the touch screen display 112 (where the touch-sensitive surface and the display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, such as Figure 4B In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a main axis (e.g., Figure 4B 453) corresponding to the main axis (for example, Figure 4B According to these embodiments, the device detects contact with the touch-sensitive surface 451 at a location that corresponds to a corresponding location on the display (e.g., Figure 4B 460 and 462 in (e.g., in Figure 4B , 460 corresponds to 468 and 462 corresponds to 470). Thus, on a touch-sensitive surface (e.g., Figure 4B 451) and a display of a multi-function device (e.g., Figure 4B When the user interface 450 in FIG. 4 is separated, the user input detected by the device on the touch-sensitive surface (e.g., contacts 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0207] Additionally, although the following examples are primarily given with reference to finger inputs (e.g., finger contacts, single-finger tap gestures, finger swipe gestures, etc.), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). For another example, a tap gesture is optionally replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., instead of detecting the contact, followed by ceasing to detect the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or mice and finger contacts are optionally used simultaneously.
[0208] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 3A Touchpad 355 or Figure 4B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 451 in the display, the particular user interface element is adjusted according to the detected input. In the case that the touch-sensitive surface 451 in the display is included, the touch-sensitive surface 451 can be used to directly interact with the user interface elements on the touch-screen display. Figure 1A touch-sensitive display system 112 or Figure 4AIn some implementations (such as touch screens in a touchscreen display), a contact detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by a contact) is detected on the touchscreen display at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted based on the detected input. In some implementations, focus moves from one area of the user interface to another area of the user interface without corresponding movement of a cursor or movement of a contact on the touchscreen display (e.g., by using a tab key or arrow keys to move focus from one button to another); in these implementations, the focus selector moves based on the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user interface element (or contact on the touchscreen display) that is controlled by the user to convey the user's desired interaction with the user interface (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a trackpad or touchscreen), the position of a focus selector (e.g., a cursor, contact, or selection box) over a corresponding button will indicate that the user desires to activate the corresponding button (rather than other user interface elements shown on the device display). In some embodiments, a focus indicator (e.g., a cursor or selection indicator) is displayed via a display device to indicate the current portion of the user interface that will be affected by input received from one or more input devices.
[0209] User interface and associated processes
[0210] Attention is now directed to a computer system (e.g., an electronic device, such as portable multifunction device 100 ( )) that includes a display generating component (e.g., a display, a projector, a head-mounted display, a heads-up display, etc.), one or more cameras (e.g., a video camera that continuously provides a real-time preview of at least a portion of content within the camera's field of view and optionally generates a video output comprising one or more image frames capturing the content within the camera's field of view), and one or more input devices (e.g., a touch-sensitive surface, such as a touch-sensitive remote control, or a touch screen display that also serves as the display generating component, a mouse, a joystick, a wand controller, and / or a camera that tracks one or more features of a user, such as the position of the user's hands), optionally one or more gesture sensors, optionally one or more sensors that detect intensity of contact with the touch-sensitive surface, and optionally one or more tactile output generators (and / or communicates with these components). Figure 1A ) or device 300( Figure 3A ) or computer system 301( Figures 3B to 3C )) is an implementation plan of the user interface ("UI") and related processes implemented on.
[0211] Figures 5A to 5LL and 6A to 6T Figures 1 and 2 illustrate exemplary user interfaces for interacting with and annotating an augmented reality environment and media items according to some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including 7A to 7B 、 Figures 8A to 8C 、 Figures 9A to 9G 、 Figures 10A to 10E 、 FIG. 15A to FIG. 15B 、 16A to 16E 、 17A to 17D and 18A to 18B For ease of explanation, some embodiments in the implementation scheme will be discussed with reference to operations performed on a device with touch-sensitive display system 112. In such embodiments, the focus selector is optionally: a corresponding finger or stylus contact, a representative point corresponding to the finger or stylus contact (e.g., the center of gravity of the corresponding contact or a point associated with the corresponding contact), or the center of gravity of two or more contacts detected on touch-sensitive display system 112. However, in response to detecting a contact on touch-sensitive surface 451 when the user interface shown in the figure is displayed on display 450 together with the focus selector, similar operations are optionally performed on a device with display 450 and a separate touch-sensitive surface 451.
[0212] Figures 5A to 5LL A user is shown accessing a camera 305 (eg, Figure 3B These cameras (optionally in conjunction with a time-of-flight sensor, such as time-of-flight sensor 220, Figure 2B ) obtains depth data of the room, which is used to create a three-dimensional representation of the scanned room. The scanned room can be simplified by removing some non-essential aspects of the scanned room, or the user can add virtual objects to the scanned room. The three-dimensional depth data is also used to enhance interaction with the scanned environment (e.g., preventing virtual objects from overlapping with real-world objects in the room). In addition, to enhance the realism of the virtual objects, the virtual objects can cause deformation of the real-world objects (e.g., the virtual bowling ball 559 deforms the pillows 509 / 558, as shown in FIG. Figures 5II to 5KK as shown below).
[0213] Figure 5AUser 501 is shown scanning room 502 via camera 305 of computer system 301-b. To illustrate that computer system 301-b is scanning the room, a shaded area 503 is projected onto room 502. Room 502 includes a plurality of structural features (e.g., walls and windows) and non-structural features. Room 502 includes four walls 504-1, 504-2, 504-3, and 504-4. Wall 504-2 includes window 505, which shows a view of an area outside room 502. In addition, room 502 includes floor 506 and ceiling 507. Room 502 also includes a plurality of items resting on floor 506 of room 502. These items include a floor lamp 508, pillows 509, rug 510, and wooden table 511. Wooden table 511 is also shown in room 502, causing indentations 512-1 and 512-2 on rug 510. The room also includes a cup 513, a smart home control device 514, and a magazine 515, all of which rest on top of a wooden table 511. Additionally, natural light entering through window 505 causes shadows 516-1 and 516-2 to be cast on floor 506 and carpet 510, respectively.
[0214] Figure 5A Also shown is the display generation component 308 of the computer system 301-b that the user 501 is currently looking at while scanning the room 502. The display 308 shows a user interface 517 that shows a real-time representation 518 of what the camera 305 is currently capturing (sometimes referred to herein as a real-time view representation). The user interface 517 also includes instructions 519 and / or directional markings 520 for guiding the user as to which parts of the room still need to be scanned. In some embodiments, the user interface 517 also includes a "floor plan" visualization 521 to indicate to the user which parts of the room 502 have been scanned. This "floor plan" visualization 521 is shown as an isometric view, but an orthogonal view or an oblique top-down view may be shown instead. In some embodiments, more than one view may be displayed.
[0215] Figure 5BUser 501 is shown still performing a scan of room 502, but now with computer system 301-b in a different orientation (e.g., user 501 is following scan instructions 519 and directional marker 520, and is moving the device up and to the right). To represent this change in position, the shaded area 503 is now oriented according to how much the device has moved. Because the device has moved, the real-time representation 518 is also updated to show the new portion of room 502 that is currently being scanned. Additionally, the "floor plan" visualization 521 is now updated to show the new portion of room 502 that has been scanned. The "floor plan" visualization 521 also aggregates all of the portions that have been scanned so far. Finally, a new directional marker 522 is displayed that shows the user which portion of the room needs to be scanned next.
[0216] In response to room 502 being scanned, a simplified representation of the room is shown. A simplified representation is a representation of room 502 or other physical environment in which some details are removed from features and nonessential, non-structural features (e.g., a cup) are not shown. Multiple levels of simplification may be possible, Figure 5C-1 、 Figure 5C-2 、 Figure 5C-3 and Figure 5C-4 Examples showing some of these simplifications are given below.
[0217] Figure 5C-1 Computer system 301-b is shown displaying user interface 523 including a simplified representation of room 524-1. The portion of room 502 displayed in the simplified representation of room 524-1 corresponds to the user's orientation of the devices in room 502. The orientation of user 501 in room 502 is displayed in a small user-side orientation depiction. If user 501 changes the orientation of their computer system 301-b, the simplified representation of room 524-1 will also change. User interface 523 includes three controls in control area 525, each of which adjusts the view of room 502. A comparison is as follows:
[0218] A "first person view" control 525-1, when selected, orients the displayed representation of the room in a first person view to mimic what a user would see at their orientation in the room. The placement of the device in the room controls what is shown.
[0219] A "top-down view" control 525-2, when selected, orients the displayed representation of the room in a top-down view. In other words, the user interface will display a bird's-eye view (e.g., a top-down orthographic view).
[0220] An "Isometric View" control 525-3, when selected, orients the displayed representation of the room in an isometric view.
[0221] A "side view" control 525-4, which, when selected, displays a flat orthographic side view of the environment. Although this mode switches to an orthographic side view, it may also be another control for changing the view to another orthographic view (e.g., another side view or a bottom view).
[0222] although Figure 5C-1 、 Figure 5C-2 、 Figure 5C-3 and Figure 5C-4 A first-person view is depicted, but it should be understood that simplification may occur in any other view displayed by the device (eg, orthographic and / or isometric views).
[0223] exist Figure 5C-1 In the simplified representation of room 524-1, Figures 5A to 5B 502 in the simplified representation of room 524-1. In this simplified representation of room 524, pillows 509, rug 510, cups 513, magazines 514 have all been removed. However, some larger non-structural features still exist, such as floor lamp 508 and wooden table 511. These remaining larger non-structural features are now displayed without their texture. Specifically, the shade color of floor lamp 508 is removed, wooden table 511 no longer displays its wood grain, and window 505 no longer displays a view of the area outside the room. In addition, detected building / home automation objects and / or smart objects are displayed as icons in the simplified representation of room 524-1. In some cases, the icons completely replace the detected objects (e.g., the "Home Control" icon 526-1 replaces the home control device 514). However, it is also possible to display icons of detected building / home automation objects and / or smart objects and the corresponding objects at the same time (e.g., floor lamp 508 and corresponding smart light icon 526-2 in the simplified representation of room 524-1). Figure 5C-1 In some embodiments, while one object and its corresponding automated or intelligent object are displayed simultaneously, for another object (e.g., also considering camera 305 of computer system 301-b), only the corresponding automated or intelligent object is displayed. In some embodiments, predefined criteria are used to determine whether to replace an object with its corresponding automated or intelligent object or to display both simultaneously. In some embodiments, the predefined criteria depend in part on the selected or determined level of simplification.
[0224] Figure 5C-2 Another simplified representation of room 524-2 is shown with several items removed. The difference between the simplified representation of room 524-2 and the simplified representation of room 524-1 is that floor lamp 508 is no longer displayed. However, icons corresponding to detected building / home automation objects and / or smart objects (e.g., smart light icon 526-2) are still displayed.
[0225] Figure 5C-3 Another simplified representation of room 524-3 is shown, with all items removed. Instead, bounding boxes are placed in the simplified representation of room 524-3 to illustrate large, non-structural features. Here, a bounding box for wooden table 527-1 and a bounding box for floor lamp 527-2 are shown. These bounding boxes illustrate the size of these non-structural features. Icons corresponding to detected building / home automation objects and / or smart objects (e.g., "Home Control" icon 526-1 and smart light icon 526-2) are still displayed.
[0226] Figure 5C-4 Another simplified representation of room 524-4 is shown, in which larger non-structural features are replaced with computer-aided design ("CAD") representations. Here, a CAD representation of a wooden table 528-1 and a CAD representation of a floor lamp 528-2 are shown. These CAD representations show computerized renderings of some of the non-structural features. In some embodiments, the CAD representations are displayed only when the computer system 301-b identifies a non-structural object as an item corresponding to the CAD representation. Icons corresponding to detected building / home automation objects and / or smart objects (e.g., "Home Control" icon 526-1 and smart light icon 526-2) are still displayed. Figure 5C-4 Also shown is a CAD chair 529, which is an example of placeholder furniture. In some embodiments, placeholder furniture is placed in one or more rooms (e.g., one or more other empty rooms) to virtually "showcase" the one or more rooms.
[0227] Figure 5D A real-time representation 518 of the content displayed in the display room 502 is shown based on the user's location. If the user were to move the computer system 301-b, the content displayed would change according to such movement (e.g., as Figures 5KK to 5LL ). The real-time representation 518 is not a simplified view and shows all textures captured by the camera 305. Figure 5D Also shown is a user input 530 on a "top-down view" control 525-2. In this example, icons corresponding to detected building / home automation objects and / or smart objects (e.g., a "Home control" icon 526-1 and a smart light icon 526-2) are displayed in the augmented reality representation of the live view. Figure 5E A response to user input 530 on the "Top Down View" control 525-2 is shown. Figure 5E A top-down view of a simplified representation of room 531 is shown. Figures 5C-1 to 5C-4 (and Figure 5D518 shown in the augmented reality representation of the room 531, the simplified representation of the room 531 is fully displayed relative to the orientation of the user 501 in the room 502. In this orthogonal top-down view of the simplified representation of the room 531, representations of the window 505, the floor lamp 508, and the wooden table 511 are still displayed, but are displayed in a "top-down view". In addition, icons corresponding to detected building / home automation objects and / or smart objects are displayed (e.g., the same icons as shown in the augmented reality representation of the real-time view) (e.g., the "Home Control" icon 526-1 and the smart light icon 526-2). Although the top-down view is shown without textures (e.g., the representations of the objects in the room are shown without textures), it should be understood that in other embodiments, the top-down view includes representations of objects with textures.
[0228] Figure 5F Shown with Figure 5E The simplified representation of room 531 shown in FIG. 5 is the same top-down view. However, Figure 5F User input 532 on the "Isometric View" control 525-3 is shown. Figure 5G An isometric view of a simplified representation of room 533 is shown. Figures 5C-1 to 5C-4 Unlike the real-time representation 518, the isometric view of the simplified representation of room 533 is fully displayed relative to (e.g., independent of) the orientation of the user 501 in room 502. In this orthogonal isometric view of the simplified representation of room 533, representations of window 505, floor lamp 508, and wooden table 511 are still displayed, but are displayed in an isometric view. In addition, icons corresponding to detected building / home automation objects and / or smart objects are displayed (e.g., the same icons as shown in the augmented reality representation of the real-time view) (e.g., "Home Controls" icon 526-1 and Smart Light icon 526-2). Although the isometric view is displayed without textures, it should be understood that in other embodiments, the isometric view includes representations of objects with textures.
[0229] Figure 5H A simplified representation of room 533 is shown with Figure 5G The same isometric view shown in . However, Figure 5H User input 534 on the "First Person View" control 525-1 is shown. In response, Figure 5I A real-time representation 518 of content displayed in the display room 502 based on the location of the user 501 is shown.
[0230] Figure 5J User input 535 is shown above the smart light icon 526 - 2 in the real-time representation 518 . Figure 5KThe resulting user interface in response to user input 535 (e.g., a long press) on the smart light icon 526-2 is shown. In some embodiments, tapping the smart light icon 526-2 turns the smart light on or off, while a different input gesture, such as a long press, results in displaying a light control user interface 536 that includes a color control 537 for adjusting the color of the light output by the floor lamp 508, or a brightness control 538 for controlling the brightness of the light output by the floor lamp 508, or both (e.g., Figure 5K ). Color control 537 includes a plurality of available colors from which the user can select the light emitted by floor lamp 508. In addition, the light control user interface optionally includes an exit user interface element 539, which is shown in the upper left corner of light control user interface 536 in this example.
[0231] Figure 5L A drag user input 540-1 (e.g., a drag gesture) is shown starting from the brightness control 538 in the light control user interface 536. In response to the drag input, the brightness of the light increases. Figure 5L In FIG. 5 , this is represented by shadows 516 - 3 and 516 - 4 caused by the lampshade interfering with the light emitted by the floor lamp 508 and the wooden table 511 interfering with the light emitted by the floor lamp 508 , respectively.
[0232] Figure 5M Shown drag user input 540-1 continues to the second position 540-2 on the brightness control. In response to the drag user input at the second position 540-2, the brightness of the light in floor lamp 508 increases. In addition, light bulb symbol 541 is also updated to show that it is emitting brighter light. Figure 5N An input 561 on the exit user interface element 539 is shown, which closes the light control user interface 536, as shown in FIG. Figure 5O shown.
[0233] Figures 5O to 5R The interaction of a virtual object interacting with a representation of a real-world object is shown. As described above, the camera 305 on the computer system 301-b, optionally in conjunction with a time-of-flight sensor, is capable of recording depth information, and because of this, the computer system 301-b can cause the virtual object to resist moving into the real-world object. Figures 5O to 5R The discussion of shows that virtual objects resist (e.g., input moves at a different rate than the rate at which the virtual object moves into the real-world object) entering the space of the real-world object. Figures 5O to 5Z Throughout this discussion, the virtual stool 542 should be understood as an example of a virtual object, and the real-world wooden table 511 / 544 as an example of a real-world object.
[0234] Figure 5OAlso shown is an example of a virtual object added to the real-time representation 518 of the room 502 , in this case a virtual stool 542 . Figure 5O Also shown is a start-drag gesture 543-1 to move a virtual object (e.g., virtual stool 542) within the real-time representation 518. For purposes of explanation, the virtual stool 542 is shown without texture, but it should be understood that in some embodiments, one or more instances of virtual furniture are displayed in the real-time representation or other representation with texture.
[0235] Figure 5P A representation of input 543 - 2 continuing to wooden table 544 , which corresponds to real-world wooden table 511 , is shown. Figure 5P A virtual stool 542 is shown beginning to enter the representation of a wooden table 544. As the virtual stool 542 begins to enter the representation of the wooden table 544, the input 543-2 no longer moves at the same rate as the virtual stool 542 (e.g., the virtual stool moves at a slower rate than the input). The virtual stool 542 now overlaps the representation of the wooden table 544. When a virtual object (e.g., a virtual stool 542) overlaps a representation of a real-world object (e.g., a table 544), a portion of the virtual object (e.g., the virtual stool 542) disappears, or alternatively is displayed in a semi-transparent, de-emphasized state, or, in yet another alternative, an outline of the portion of the virtual object that overlaps the representation of the real-world object is displayed.
[0236] Figure 5Q Input 543-3 is shown continuing, but the virtual stool 542 again does not move at the same rate as the input 543-3 was moving. The virtual stool 542 does not pass a certain threshold into the representation of the wooden table 544. In other words, at first the virtual object (e.g., virtual stool 542) will resist moving into the representation of the real-world object (e.g., wooden table 544) (e.g., allowing some overlap), but after a certain amount of overlap is satisfied, the movement of the input no longer causes the virtual object (e.g., virtual stool 542) to move.
[0237] Figure 5R The input is shown to no longer be received, and in response to no longer receiving the input, the virtual object (eg, virtual stool 542) will appear at a location away from the real-world object (eg, table) that no longer causes an overlap.
[0238] Figure 5S Another drag input 545-1 is shown, different from the previous one in Figures 5O to 5R Unlike the previous gestures, the following series of gestures shows that if the user drags far enough on the virtual object (e.g., virtual stool 542), the virtual object (e.g., virtual stool 542) will quickly move through the representation of the real-world object (e.g., wooden table 544). Figure 5S to Figure 5V However, as with the previous drag input, the virtual stool 542 moves at the location where the input moves, unless the input causes an overlap with a representation of a real-world object (e.g., the representation of the wooden table 544).
[0239] Figure 5T Input 545 - 2 is shown continuing into the representation of wooden table 544 . Figure 5T A virtual object is shown (e.g., virtual stool 542_begins to enter the representation of a real-world object (e.g., wooden table 544)). As virtual stool 542 begins to enter the representation of wooden table 544, input 545-2 no longer moves at the same rate as virtual stool 542 (e.g., the virtual stool moves at a slower rate than the input). Virtual stool 542 now overlaps the representation of wooden table 544. As described above, when virtual stool 542 overlaps table 544, a portion of virtual stool 542 disappears, or alternatively is displayed in a semi-transparent, de-emphasized state, or, in another alternative, an outline of the portion of the virtual object that overlaps the representation of the real-world object is displayed.
[0240] Figure 5U Input 545-3 is shown continuing, but virtual stool 542 again does not move at the same rate as input 545-3 was moving. Virtual stool 542 does not pass a certain threshold into the representation of wooden table 544. In other words, initially virtual stool 542 will resist moving into the representation of wooden table 544 (e.g., allowing some overlap), but after a certain amount of overlap is met, the movement of the input no longer causes virtual stool 542 to move.
[0241] Figure 5V Input 545-4 is shown to satisfy a threshold distance (and also not interfere with any other real-world objects). In response to satisfying the threshold distance passing through the representation of the real-world object (e.g., wooden table 544) (and not interfering with any other real-world objects), a virtual object (e.g., virtual stool 542) is caused to rapidly move through the representation of the real-world object (e.g., wooden table 544). As the virtual object (e.g., virtual stool 542) rapidly moves through the representation of the real-world object (e.g., wooden table 544), the virtual object (e.g., virtual stool 542) aligns itself with input 545-4.
[0242] Figure 5W The virtual object (e.g., virtual stool 542) is shown now in a rapidly moving position (e.g., a position where lift-off occurs after satisfying a distance threshold after passing a representation of a real-world object (e.g., wooden table 544)).
[0243] Figure 5X Another drag input 546-1 is shown, different from the previous one in Figures 5S to 5WUnlike the previous gestures, the following series of gestures shows that if the user drags on the virtual object (e.g., virtual stool 542) quickly enough (e.g., high acceleration and / or high speed), the virtual object (e.g., virtual stool 542) will move quickly through the representation of the real-world object (e.g., wooden table 544). Figure 5S to Figure 5V However, as with the previous drag input, the virtual object (e.g., virtual stool 542) moves at the location where the input moves, unless the input results in an overlap with a real-world object or item (e.g., the representation of a wooden table 544).
[0244] Figure 5Y Input 546 - 2 is shown continuing into the representation of wooden table 544 . Figure 5Y Virtual stool 542 is shown beginning to enter the representation of wooden table 544. As virtual stool 542 begins to enter the representation of wooden table 544, input 546-2 no longer moves at the same rate as virtual stool 542 (e.g., the virtual stool moves at a slower rate than the input). Virtual stool 542 now overlaps the representation of wooden table 544. When virtual stool 542 overlaps the table, a portion of virtual stool 542 disappears (or appears in a semi-transparent, de-emphasized state), but the outline of virtual stool 542 remains.
[0245] Figure 5Z The result is shown when input 546-3 satisfies a threshold acceleration, velocity, or combination of acceleration and velocity (and also does not interfere with any other real-world objects). In response to satisfying the threshold acceleration, velocity, or combination (and not interfering with any real-world objects), computer system 301-b causes virtual stool 542 to rapidly move across the representation of wooden table 544. As virtual stool 542 rapidly moves across the representation of wooden table 544, computer system 301-b aligns virtual stool 542 with input 546-3.
[0246] Figures 5AA to 5CC The user interaction of adding a virtual table 547 to the real-time representation 518 and adjusting the size of the virtual table 547 is shown. Figure 5CC A representation of the virtual table 547 automatically resizing to abut the wooden table 544 is shown. Figures 5AA to 5LL Throughout this discussion, virtual table 547, virtual stool 542, and virtual bowling ball 559 should be understood to be examples of virtual objects, and wooden table 544, rug 558, and pillows 562 are examples of real-world objects.
[0247] Specifically, Figure 5AAA drag input is shown at a first position 548-1 above a virtual table 547 inserted into the real-time representation 518. The drag input 548 is moving in a direction toward the representation of the wooden table 544. Additionally, the drag input 548-1 appears at an edge 549 of the virtual table 547, and the direction of movement of the input 548-1 is away from a position inside the virtual table 547.
[0248] Figure 5BB The drag input is shown continuing to the second location 548-2, and the virtual table 547 resizes to a position corresponding to the drag input at the second location 548-2. The resizing of the virtual table 547 corresponds to the direction of the drag inputs 548-1 and 548-2 (e.g., if the drag gesture is to the right, the furniture item will resize to the right). Although the virtual table 547 is Figure 5BB 548 - 1 and 548 - 2 (e.g., in a direction toward a location within virtual table 547 ) will cause the size of virtual table 547 to decrease.
[0249] Figure 5CC Drag input 548-2 is shown no longer being received, but the table is shown expanding (e.g., quickly moving) to abut the representation of wooden table 544. This expansion occurs automatically when the user's input (e.g., drag input 548-2) meets a threshold proximity to the edge of another (virtual or real-world) object, without requiring any additional input from the user to cause the virtual object to expand precisely the correct amount to abut the representation of the other (e.g., virtual or real-world) object.
[0250] Figures 5DD to 5HH The ability to switch between orthogonal views (e.g., a top-down view and a side view) is shown. In addition, these figures also show how virtual objects are retained when different views are selected.
[0251] Figure 5DD Controls are shown for selecting a "first person view" control 525-1, a "top down view" control 525-2, or a "side view" control 525-4. The "side view" control 525-4 displays a flat side view (e.g., a side orthographic view) of the environment when selected. Figure 5DD Input 551 on the "Side View" control 525-4 is shown. Although Figure 5DD The side view shown in the example is without texture (e.g., representations of objects in the side view are displayed without texture), but in some embodiments, the side view is displayed with texture (e.g., representations of one or more objects in the side view are displayed with texture).
[0252] Figure 5EEThe user interface generated in response to receiving input 551 on the "Side View" control 525-4 is shown. Figure 5EE A side view simplified representation 552 is shown with virtual furniture (e.g., virtual stool 542 and virtual table 547) and real-world furniture. In addition, icons (e.g., the same icons as shown in the augmented reality representation of the live view) corresponding to detected building / home automation objects and / or smart objects are displayed (e.g., "Home Controls" icon 526-1 and smart light icon 526-2).
[0253] Figure 5FF Shown with Figure 5EE The same is shown in simplified side view 552 . Figure 5FF Also shown is input 553 on the "Top Down View" control 525-2.
[0254] Figure 5GG The user interface generated in response to receiving input 553 on the "Top Down View" control 525-2 is shown. Figure 5GG A simplified top-down view representation 554 is shown with virtual furniture (e.g., virtual stool 542 and virtual table 547) and real-world furniture. In addition, icons corresponding to detected building / home automation objects and / or smart objects are displayed (e.g., the same icons as shown in the augmented reality representation of the real-time view) (e.g., "Home Control" icon 526-1 and smart light icon 526-2). Although the top-down view is shown without texture (e.g., the representation of the objects in the top-down view is shown without texture), it should be understood that the top-down view can also be shown with texture (e.g., the representation of the objects in the top-down view is shown with texture).
[0255] Figure 5HH Shown with Figure 5GG The same top-down view simplified representation 554. Figure 5FF Also shown is input 555 on the "First Person View" control 525-1.
[0256] Figures 5II to 5KK An augmented reality representation (e.g., a real-time representation 518 with one or more added virtual objects) of a room 502 is shown. These figures illustrate how virtual objects change the visual appearance of real-world objects. Figure 5II All textures of features in room 502 are shown, and non-essential features are also shown (eg, a representation of pillow 562 corresponding to pillow 508). Figure 5IIAlso shown are virtual objects (e.g., virtual stool 542 and virtual table 547) in real-time representation 518 of room 502. These virtual objects interact with physical objects and affect their appearance. In this example, virtual table 547 creates an indentation 557-1 (e.g., compression of the carpet) on the representation of carpet 558 (corresponding to carpet 509). Virtual stool 542 also creates an indentation 557-2 (e.g., compression of the carpet) on the representation of carpet 558. Figure 5II Also shown is a virtual bowling ball 559 being inserted into the non-simplified representation of room 502 above the representation of pillow 562 via input 563 .
[0257] Figure 5JJ The user is shown releasing (e.g., no longer receiving input 563) virtual bowling ball 559 in real-time representation 518 of room 502 above representation of pillow 562. In response to releasing virtual bowling ball 559, virtual bowling ball 559 begins to fall to representation of pillow 562, following the physics of room 502.
[0258] Figure 5KK A virtual bowling ball 559 is shown landing on a representation of a pillow 562, and in response to the landing, the representation of the pillow 562 shows that it is deformed. The deformation of the representation of the pillow 562 is shown by a compression line 560 of the representation of the pillow 562.
[0259] Figure 5LL The user is shown moving to a new location within the room while the computer system is displaying a real-time representation 518 of the room 502. Consequently, the real-time representation 518 of the room 502 has been updated to reflect what the user 501 would perceive from the new location. Consequently, the computer system 301-b displays a new real-time view representation 564. As the position, yaw, pitch, and roll of the device change, the real-time representation of the room will update in real time to correspond to the device's current position, yaw, pitch, and roll. Thus, the display of the real-time representation 518 continuously adjusts to small changes in the device's position and orientation as the user holds the device. In other words, the real-time representation will appear as if the user is looking through a camera's viewfinder, but virtual objects may be added or included in the real-time representation.
[0260] 6A to 6T An exemplary user interface is shown that allows a user to insert a virtual object into a first representation of the real world. If the device detects that the first representation of the real world corresponds to (e.g., a portion of the other representation matches the first representation) another representation, the virtual object will be placed in the corresponding other representation. Although the following photos are photos of vehicles, it should be understood that the photos shown can be Figures 5A to 5LL Photos of the environment shown in .
[0261] Figure 6ADepicted is a "Media" user interface 601-1 that includes four media thumbnail items: media thumbnail item 1 602-1, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4. These media thumbnail items can represent photos, photos with live content (e.g., LIVE PHOTO, a registered trademark of Apple Inc. of Cupertino, California), videos, or Graphics Interchange Format ("GIF"). In this example, the media thumbnail items are representations of a car 624.
[0262] Figure 6B An input 603 is shown on media thumbnail item 3 602 - 3 depicted in the “Media” user interface 601 - 1 . Figure 6C A response to input 603 above media thumbnail item 3 602 - 3 is shown. Figure 6C Another "Media" user interface 601-2 is shown showing an expanded media item 604 corresponding to the media thumbnail item 3 602-3 depicted in the "Media" user interface 601-1. Figure 6C Also shown is a media thumbnail eraser 605 containing media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4. The thumbnail displayed in the center of the media thumbnail eraser 605 indicates which thumbnail item 602 corresponds to the displayed expanded media item 604. The media thumbnail items shown in the media thumbnail eraser 605 can be scrolled through and / or clicked to change between the media items displayed as the expanded media items 604.
[0263] Figure 6D An annotation 606 is shown indicating that “EV Turbo” was added to the expanded media item 604 via input 607-1. Without receiving a lift-off of input 607, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4 depicted in the media thumbnail eraser 605 are updated to display the annotation 606 indicating that “EV Turbo” was added to the expanded media item 604. Figure 6D It is also shown that input 607 - 1 is a drag input in a substantially rightward direction.
[0264] Figure 6EInput 607-1 is shown continuing to a second input position 607-2. In response to the change in position of the drag input (607-1 to 607-2), annotation 606 indicating "EV Turbo" moves to a second position within expanded media item 604. Additionally, without receiving a lift-off of input 607, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4 depicted in media thumbnail eraser 605 are updated to display annotation 606 indicating "EV Turbo" at a second position corresponding to the second position within expanded media item 604.
[0265] Figure 6F exist Figure 6M An exemplary user interface for switching between different media items is shown. In some embodiments, when switching between media items, a transitional representation of a three-dimensional model of the physical environment illustrating the change in orientation is displayed. This type of interaction allows the user to see how the media items relate to each other in three-dimensional space.
[0266] Figure 6F A drag input 608 is shown starting at position 608-1 above media thumbnail item 3 602-3 depicted in media thumbnail eraser 605. The drag input moves to the right and pulls the media thumbnail item to the right. As the media thumbnail item moves to the right, portions of the media thumbnail item cease to be displayed, and new portions of the media thumbnail item that were not displayed come into view.
[0267] Figure 6G Show that drag input 608 continues to position 608-2.When position changes, extended media item 604 begins to fade out, and the representation of the bottom three-dimensional model of the physical environment of extended media item 609 begins to fade in.In some embodiments, the representation of the three-dimensional model of the physical environment of extended media item 609 is based on the unstructured three-dimensional model of physical environment, and this unstructured three-dimensional model can use the combination of two-dimensional cell type (for example, triangle and / or quadrilateral) to approximate any geometric shape.In some embodiments, use three-dimensional cell type (for example, tetrahedron, hexahedron, pyramid and / or wedge).When drag input occurs, eraser 605 is scrolled through, and media thumbnail item 602-1 begins to appear, and media thumbnail item 602-4 begins to disappear.
[0268] Figure 6HDrag input 608 is shown continuing to position 608-3. As the position changes, expanded media item 604 fades out completely, and the representation of the three-dimensional model of the physical environment of expanded media item 609 fades in completely. As drag input 608 occurs, eraser 605 is scrolled through and additional portions of media thumbnail item 602-1 are displayed, while less media thumbnail item 602-4 (or a lesser portion thereof) is displayed.
[0269] Figure 6I Shown drag input 608 continues to position 608-4. When the position changes, the intermediate representation of the three-dimensional model of physical environment 610 is displayed. This intermediate representation of the three-dimensional model of physical environment 610 shows that the viewing angle of the media project changes as another media project begins to be displayed. In some embodiments, the intermediate representation of the three-dimensional model of physical environment 610 is an unstructured three-dimensional model that can use a combination of two-dimensional unit types (e.g., triangles and / or quadrilaterals) to approximate any geometric shape. In some embodiments, three-dimensional unit types (e.g., tetrahedrons, hexahedrons, pyramids and / or wedges) are used. When drag input 608 is carried out, eraser 605 is scrolled through, showing more media thumbnail items 602-1, and showing fewer media thumbnail items 602-4.
[0270] Figure 6J Shown drag input 608 continues to position 608-5. When the position changes, the intermediate representation of the three-dimensional model 610 of the physical environment is no longer displayed, and another representation of the three-dimensional model of the physical environment of another extended media item 612 is displayed. In this example, the other extended media item corresponds to media thumbnail item 2 602-2. In some embodiments, another representation of the three-dimensional model of the physical environment of another extended media item 612 is an unstructured three-dimensional model that can approximate any geometric shape using a combination of two-dimensional unit types (e.g., triangles and / or quadrilaterals). In some embodiments, three-dimensional unit types (e.g., tetrahedrons, hexahedrons, pyramids and / or wedges) are used. When drag input 608 is carried out, eraser 605 is scrolled through, showing more media thumbnail items 602-1, and showing fewer media thumbnail items 602-4.
[0271] Figure 6K Drag input 608 is shown continuing to position 608-6. As the position changes, another representation of the three-dimensional model of the physical environment of another expanded media item 612 begins to fade out, and another expanded media item 613 corresponding to media thumbnail item 2 602-2 begins to fade in. As drag input 608 progresses, eraser 605 is scrolled across, displaying more of media thumbnail item 602-1 and fewer of media thumbnail item 602-4.
[0272] Figure 6L Drag input 608 is shown continuing to position 608-7. As the position changes, other expanded media items 613 fade in completely, and the representation of the three-dimensional model of the physical environment of expanded media item 612 fades out completely. Figure 6M The display drag input 608 is no longer displayed, and the other expanded media items 613 are completely faded in. When the drag input 608 ends, the scrubber 605 stops scrolling, the media thumbnail item 602-1 is fully displayed, and the media thumbnail item 602-4 no longer disappears from the scrubber at all.
[0273] Figure 6N An input 614 is shown on a back button 615 within another "Media" user interface 601-2. Figure 6O Shown is that in response to receiving input 614 via the back button 615 , “Media” user interface 601 - 3 is displayed.
[0274] Figures 6P to 6T Another virtual object is shown added to the expanded media item 616 corresponding to media thumbnail item 4 602-4. These figures illustrate that annotations can be added to any one of the media items, and these annotations are displayed in the other associated media items.
[0275] Figure 6P Input 617 is shown over media thumbnail item 4 602-4. Figure Q shows that in response to receiving input 617 over media thumbnail item 4 602-4, the "Media" user interface 601-4 is displayed. Figure 6Q A response to input 617 above media thumbnail item 4 602-4 is shown. Figure 6Q The "Media" user interface 601-4 is shown, which includes the same interface as the "Media" user interface 601-3 ( Figure 6P ) corresponds to the expanded media item 618 of the media thumbnail item 3 602-4 depicted in FIG. Figure 6Q Also shown is a media thumbnail eraser 619 containing media thumbnail item 3 602-3 and media thumbnail item 4 602-4. The thumbnails displayed in the center of the media thumbnail eraser 619 indicate that the corresponding expanded media items are being displayed. The media thumbnail items shown in the media thumbnail eraser 619 can be scrolled through and / or clicked to change between the media items.
[0276] Figure 6RA wing element (e.g., spoiler) annotation 620 is shown added to a car 624 via input 621. Without receiving a liftoff of input 621, item 3 602-3 and media thumbnail item 4 602-4 depicted in media thumbnail eraser 619 are updated to display the wing element annotation 620 added to the expanded media item 618.
[0277] Figure 6S A wing element (eg, spoiler) annotation 620 is shown added to a car 624, and input 622 on a back button 623 within the "Media" user interface 601-3 is also shown. Figure 6T The “Media” user interface 601-5 is shown displayed in response to receiving input 622 via the return button 623. Within the “Media user interface 601-5,” media thumbnail item 1 602-1, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4 all include a wing element (e.g., spoiler) annotation 620.
[0278] 7A to 7B is a flow chart illustrating a method 700 for providing different views of a physical environment according to some embodiments. The method 700 is performed on a computer system (e.g., portable multifunction device 100 ( Figure 1A ), device 300( Figure 3A ) or computer system 301( Figure 3B )) is executed at the computer system, the computer system having a display generation component (for example, a touch screen 112 ( Figure 1A )、Display 340( Figure 3A ) or display generation component 304 ( Figure 3B )), an input device (e.g., one or more input devices) (e.g., a touch screen 112 ( Figure 1A )、Touchpad 355( Figure 3A ), input device 302 ( Figure 3B ) or physical buttons separate from the display), and one or more cameras in the physical environment (e.g., optical sensor 164 ( Figure 1A ) or camera 305( Figure 3B ))(The one or more cameras optionally include one or more depth sensors such as a time-of-flight sensor 220( Figure 2B ) or communicate with it) (702). Some operations in method 700 are optionally combined, and / or the order of some operations is optionally changed.
[0279] As described below, method 700 describes a user interface and interactions that occur after capturing a representation of a physical environment via one or more cameras. The user interface displayed after capturing the representation of the physical environment includes an activatable user interface element for displaying the captured physical environment in an orthogonal view. The activatable user interface element provides a simple control for manipulating the view of the representation of the physical environment and does not require the user to make multiple inputs to achieve the orthogonal view. Reducing the number of inputs required to view the representation of the physical environment in the orthogonal view enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0280] The method includes capturing (704) a representation of a physical environment via one or more cameras and, optionally, via one or more depth sensors, including updating the representation to include representations of corresponding portions of the physical environment that are within (e.g., enter) the field of view of the one or more cameras as the field of view of the one or more cameras moves; in some embodiments, the representation includes depth data corresponding to a simulated three-dimensional model of the physical environment. In some embodiments, capturing is performed in response to activation of a captured affordance.
[0281] The method also includes, after capturing the representation of the physical environment, displaying (706) a user interface that includes an activatable user interface element for requesting display of a first orthogonal view (e.g., a front orthogonal view, also referred to as a front elevation view) of the physical environment; and in some embodiments, the front orthogonal view of the physical environment is a two-dimensional representation of the physical environment in which the physical environment is projected onto a plane that is located in front of (e.g., and parallel to) its front plane (e.g., a frontal plane of a person standing and directly viewing the physical environment, such as Figure 5E 、 Figure 5EE and Figure 5FF In some embodiments, the front orthographic view is neither an isometric view nor a perspective view of the physical environment. In some embodiments, the front orthographic view of the physical environment is different from one or more (e.g., any) views of the physical environment that any of the one or more cameras had during capture.
[0282] The method also includes receiving (708) via an input device an activatable user interface element (e.g., Figure 5C-1 and in response to receiving the user input, displaying (710) a first orthogonal view of the physical environment based on the captured representation of one or more portions of the physical environment.
[0283] In some embodiments, the first orthogonal view of the physical environment based on the captured representation of one or more portions of the physical environment is a simplified orthogonal view, wherein the simplified orthogonal view simplifies the appearance of the representation of the one or more portions of the physical environment (712). In some embodiments, when physical items within the captured representation of the one or more portions of the physical environment are below a certain size threshold, the simplified orthogonal view removes the physical items from the representation of the one or more portions of the physical environment (e.g., a wall with pictures hanging on it disappears when viewed in the simplified orthogonal view). In some embodiments, when physical items are identified as items (e.g., appliances and furniture (e.g., wooden table 511 and floor lamp 508)), the computer system replaces the physical items in the physical environment with simplified representations of the physical items (e.g., replacing a physical refrigerator with a simplified refrigerator (e.g., a smoothed refrigerator with only minimal features to identify it as a refrigerator). For example, see Figure 5B Wooden table 511 and Figure 5GG 5. The simplified wooden table 511 in FIG. Automatically displaying a simplified orthogonal view of the appearance of a representation of one or more portions of a simplified physical environment provides the user with the ability to quickly identify the representation of the physical item they are interacting with. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the amount of input required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0284] In some embodiments, one or more walls, one or more floors, and / or one or more ceilings in the physical environment are identified (714) (e.g., in conjunction with or after capturing a representation of the physical environment), as well as edges of features of the physical environment. A first orthogonal view of the physical environment includes representations of the identified one or more walls, floors, ceilings, and features, represented by projection lines displayed perpendicular to the identified one or more walls, floors, ceilings, and features (e.g., Figures 5EE to 5HH). In some embodiments, walls, floors, or ceilings are identified based at least in part on determining that a physical object in the physical environment exceeds a predefined size threshold (e.g., to avoid confusing smaller flat surfaces such as a table with walls, floors, and ceilings). One or more walls, one or more floors, and / or one or more ceilings in the physical environment are automatically identified, and one or more walls, one or more floors, and / or one or more ceilings are automatically aligned in the physical environment when a first orthogonal view is selected, providing the user with the ability to not have to manually identify the walls, floors, or ceilings. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0285] In some embodiments, the first orthogonal view is based on the first perspective (716). And, in some embodiments, after displaying the first orthogonal view of the physical environment based on the captured representation of one or more portions of the physical environment, a second user input corresponding to a second activatable user interface element for requesting display of a second orthogonal view of the physical environment (e.g., input 553 on control 525-2) is received via the input device. In response to receiving the second user input (e.g., input 553 on control 525-2), the first orthogonal view of the physical environment is displayed based on the captured representation of one or more portions of the physical environment (e.g., Figure 5GG In some embodiments, the first orthogonal view is a front orthogonal view (e.g., a top-down view in FIG. 4 ), wherein the second orthogonal view is based on a second perspective (of the physical environment) that is different from the first perspective (of the physical environment). Figure 5HH ), the second orthogonal view is a top orthogonal view (e.g., Figure 5GG ), side orthogonal views (e.g., Figure 5HH ) or an isometric orthographic view (e.g. Figure 5H). A user interface displayed after capturing the representation of the physical environment includes a second activatable user interface element for displaying the captured physical environment in another orthogonal view. The second activatable user interface element provides a simple control for manipulating the view of the representation of the physical environment and does not require the user to make multiple inputs to achieve the orthogonal view. Reducing the number of inputs required to view the representation of the physical environment in the orthogonal view enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0286] In some embodiments, the captured representation of the field of view includes one or more edges (e.g., edges of the representation of the physical object), each edge forming a corresponding (e.g., non-zero and, in some embodiments, oblique) angle with an edge of the captured representation of the field of view (e.g., due to the viewing angle) (718). For example, because the user is viewing the physical environment from an angle, the lines in the user's field of view are not parallel; however, the orthogonal projection shows the projection of the representation of the physical environment such that the lines presented to the user at an angle are parallel in the orthogonal projection. Furthermore, the one or more edges that each form a corresponding angle with an edge of the captured representation of the field of view correspond to the one or more edges displayed parallel to the edges of the first orthogonal view. Displaying the edges of the captured representation of the field of view parallel to the edges of the first orthogonal view allows the user to understand the geometric properties of the captured representation (e.g., by displaying the representation without the viewing angle), which provides the user with a desired orthogonal view without having to provide multiple inputs to change the view of the captured representation. Reducing the number of inputs required to view a representation of the physical environment in a desired orthogonal view enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0287] In some embodiments, the captured representation of the field of view includes at least one set (e.g., two or more) of edges that form an oblique angle (e.g., a non-zero angle that is not a right angle or a multiple of a right angle) (e.g., before the user views the physical environment from an oblique angle, so the lines are not right angles, but the orthogonal projection displays the representation of the physical environment from an angle parallel to the lines). In some embodiments, the at least one set of edges that form an oblique angle in the captured representation of the field of view corresponds to at least one set of vertical edges in the orthogonal view. Displaying the set of edges that form an oblique angle in the captured representation of the field of view as vertical edges in the orthogonal view enables the user to view the desired orthogonal view without having to provide multiple inputs to change the view of the captured representation. Reducing the number of inputs required to view the representation of the physical environment in the desired orthogonal view enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0288] It should be understood that 7A to 7B The specific order in which the operations are described in the description is merely exemplary and is not intended to indicate that the order is the only order in which the operations may be performed. A person of ordinary skill in the art will appreciate a variety of ways to reorder the operations described herein. In addition, it should be noted that the details of other processes described herein with respect to other methods described herein (e.g., methods 800, 900, 1000, 1500, 1600, 1700, and 1800) also apply in a similar manner to the details of the processes described above with respect to 7A to 7B Method 700 described herein. For example, the physical environment, features and objects, virtual objects, inputs, user interfaces, and views of the physical environment described above with reference to method 700 optionally have one or more of the characteristics of the physical environment, features and objects, virtual objects, inputs, user interfaces, and views of the physical environment described with reference to other methods described herein (e.g., methods 800, 900, 1000, 1500, 1600, 1700, and 1800). For the sake of brevity, these details are not repeated here.
[0289] Figures 8A to 8C 8 is a flowchart illustrating a method 800 for providing a representation of a physical environment at different fidelity levels of the physical environment according to some embodiments. The method 800 is performed on a computer system (e.g., portable multifunction device 100 ( Figure 1A ), device 300( Figure 3A ) or computer system 301( Figure 3B )) is executed at the computer system, the computer system having a display generation component (for example, a touch screen 112 ( Figure 1A )、Display 340( Figure 3A) or display generation component 304 ( Figure 3B )), input devices (e.g., touch screen 112 ( Figure 1A )、Touchpad 355( Figure 3A ), input device 302 ( Figure 3B ) or physical buttons separate from the display), and one or more cameras in the physical environment (e.g., optical sensor 164 ( Figure 1A ) or camera 305( Figure 3B )), optionally in combination with one or more depth sensors (e.g., a time-of-flight sensor 220 ( Figure 2B ))(802).
[0290] As described below, method 800 automatically distinguishes primary and secondary features of a physical environment, where the primary and secondary features are identified via information provided by a camera. After distinguishing the primary and secondary features, a user interface is displayed that includes the primary features (e.g., structural, immovable features such as walls, floors, ceilings, etc.) and secondary features (e.g., discrete fixtures and / or movable features such as furniture, appliances, and other physical objects). In the representation of the physical environment, the primary features are displayed at a first fidelity, and the secondary features are displayed at a second fidelity. Distinguishing between primary and secondary features provides the user with the ability to identify (e.g., categorize) items in the physical environment without having to identify (e.g., classify) them (e.g., a device will identify a chair without the user needing to specify that the item is a chair). Performing an action (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the action, which enhances device operability and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the device's battery life by enabling the user to use the device more quickly and efficiently.
[0291] The method includes capturing (804) via one or more cameras information indicative of a physical environment, including information indicative of a respective portion of the physical environment that is within (e.g., enters) a field of view of the one or more cameras as the field of view of the one or more cameras moves, wherein the respective portion of the physical environment includes a plurality of primary features of the physical environment (e.g., Figure 5A 504-1, 504-2, 504-3, and 504-4) and one or more secondary features of the physical environment (e.g., Figure 5A In some embodiments, the information indicative of the physical environment includes information about the use of one or more cameras (e.g., optical sensor 164 ( Figure 1A ) or camera 305( Figure 3B)) and one or more depth sensors (e.g., time-of-flight sensor 220 ( Figure 2B In some embodiments, the information indicative of the physical environment is used to generate one or more representations of primary (e.g., structural) features and one or more secondary (e.g., non-structural) features (e.g., discrete fixtures and / or movable features, such as furniture, appliances, and other physical objects) of the physical environment.
[0292] The method includes displaying a user interface after capturing (806) information indicative of a physical environment (e.g., in response to capturing the information indicative of the physical environment or in response to a request to display a representation of the physical environment based on the information indicative of the physical environment). The method includes simultaneously displaying: graphical representations of a plurality of primary features of the physical environment generated at a first level of fidelity for the corresponding plurality of primary features (808); and one or more graphical representations of secondary features of the physical environment generated at a second level of fidelity for the corresponding one or more secondary features, wherein the second level of fidelity is lower than the first level of fidelity in the user interface (810).
[0293] In some embodiments, the plurality of primary features of the physical environment includes one or more walls and / or one or more floors (e.g., Figure 5A In some embodiments, when walls and / or floors are classified as primary features and represented at a first level of fidelity, decorative items such as picture frames hung on the walls or textiles placed on the floor (e.g., Figures 5II to 5KK 558) is categorized as a secondary feature and is represented at a second level of fidelity. In some embodiments, the plurality of primary features in the physical environment include one or more ceilings. Identifying the walls and floors as primary features of the physical environment provides the user with an environment that can quickly indicate which objects can be operated and which items cannot be operated, which provides the user with the ability to not have to identify the walls and floors because this is done automatically by the computer system. Performing an operation when a set of conditions have been met (e.g., automatically) without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve the desired results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0294] In some embodiments, primary features of the physical environment include one or more doors and / or one or more windows (e.g., Figure 5A505) (814). Automatically identifying doors and windows as primary features of the physical environment provides the user with the convenience of not having to specify what each feature in the physical environment is and how it interacts with other features in the physical environment. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0295] In some embodiments, one or more secondary features of the physical environment include one or more pieces of furniture (e.g., Figures 5II to 5KK (representation of wooden table 544 in ) (816). In some embodiments, a physical object (such as a piece of furniture) is classified as a secondary feature based on determining that the object is within a predefined threshold size (e.g., volume). In some embodiments, a physical object that meets or exceeds the predefined threshold size is classified as a primary feature. Furniture can include any item such as a table, a lamp, a desk, a sofa, a chair, and a lamp. Automatically identifying multiple pieces of furniture as primary features of the physical environment provides the user with the advantage of not having to specify what each feature in the physical environment is and how it interacts with other features in the physical environment. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve expected results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0296] In some embodiments, the one or more graphical representations of the one or more secondary features generated at the second fidelity level for the corresponding one or more secondary features of the physical environment include one or more icons (818) representing the one or more secondary features (e.g., a chair icon representing a chair in the physical environment, optionally displayed in the user interface at a position relative to the graphical representations of the plurality of primary features that corresponds to the position of the chair relative to the plurality of primary features in the physical environment) (e.g., Figure 5C-1 and Figures 5J to 5N In some embodiments, the icon representing the secondary feature is selectable, and in some embodiments, in response to selection of the icon, a user interface is displayed that includes one or more user interface elements for interacting with the secondary feature (e.g., Figures 5K to 5N536 in the light control user interface). The corresponding user interface element allows control of an aspect of the secondary feature (e.g., the brightness or color of the smart light). In some embodiments, information about the secondary feature is displayed (e.g., a description of the secondary feature, a link to a website for the identified (known) furniture, etc.). Automatically displaying icons representing one or more secondary features (e.g., smart lights, smart speakers, etc.) provides the user with an indication that the secondary features have been identified in the physical environment. In other words, the user does not have to navigate to a different user interface to control each secondary feature (e.g., smart smart lights, smart speakers, etc.), but can control the building automation device with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0297] In some embodiments, the one or more graphical representations of the one or more secondary features include a corresponding three-dimensional geometric shape (820) that outlines a corresponding area of the physical environment occupied by the one or more secondary features of the physical environment in the user interface. The corresponding three-dimensional geometric shape (e.g., sometimes referred to as a bounding box, see e.g., Figure 5C-3 In some embodiments, the corresponding three-dimensional geometric shape is displayed as a wireframe, as partially transparent, as a dashed or dotted outline, and / or in any other manner suitable for indicating that the corresponding three-dimensional geometric shape is merely an outline representation of the corresponding secondary feature of the physical environment. In some embodiments, the corresponding three-dimensional geometric shape is based on or corresponds to a three-dimensional model of the physical environment (e.g., generated based on depth data included in the information indicating the physical environment). Displaying a bounding box for representing one or more secondary features (e.g., a table lamp, chair, and other furniture of predetermined size) enables the user to quickly understand the size of the secondary features in the physical environment, which allows the user to manipulate the physical environment with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating the device / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0298] In some embodiments, wherein the one or more graphical representations of the one or more secondary features include predefined placeholder furniture (e.g., Figure 5C-4CAD chair 529 in the image) (822). In some embodiments, based on determining that the room does not contain furniture, the room is automatically populated with furniture. In such embodiments, it is determined what type of room is being captured (e.g., kitchen, bedroom, living room, office, cab, or commercial space) and corresponding placeholder furniture is located / placed in the determined type of room.
[0299] In some embodiments, the one or more graphical representations of the one or more secondary features include a computer-aided design (CAD) representation (e.g., Figure 5C-4 CAD representation of wooden table 528-1 and CAD representation of floor lamp 528-2 in the physical environment) (824). In some embodiments, the one or more CAD representations of one or more secondary features provide a predefined model for the one or more secondary features, which predefined model includes additional shape and / or structural information beyond that captured in the information indicating the physical environment. Displaying predefined placeholder furniture in the representation of the physical environment provides an efficient way for the user to see how the physical space can be utilized. The user no longer needs to navigate menus to add furniture (e.g., secondary non-structural features), but can instead automatically populate the physical environment with placeholder furniture. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0300] In some embodiments, one or more graphical representations of one or more secondary features are partially transparent (826). In some embodiments, the graphical representations of the secondary features are displayed as partially transparent, while the graphical representations of the primary features are not displayed as partially transparent. In some embodiments, the secondary features are partially transparent in certain views (e.g., simplified views) but are opaque in fully textured views. It can sometimes be difficult to understand the size of the physical environment, and providing the user with partially transparent secondary features allows the user to view the constraints of the representation of the physical environment. This reduces the need for the user to move the secondary features around to understand the physical environment. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0301] In some embodiments, one or more secondary features include one or more building automation devices (e.g., also known as home automation devices or smart home devices, particularly when installed in a user's home), such as smart lights (e.g., Figure 5C-1 Smart light icon 526-2 in ), smart TV, smart refrigerator, smart thermostat, smart speaker (e.g., Figure 5C-1 In some embodiments, the computer system displays a "home control" icon 526-1 in the user interface, the ... In some embodiments, when the state of a corresponding building automation device changes (e.g., a smart light turns on), the icon representing the corresponding building automation device is updated to reflect the change in state (e.g., the icon of the smart light changes from a graphic of a turned-off light bulb to a graphic of a turned-on light bulb). Displaying icons representing one or more building automation devices (e.g., smart lights, smart speakers, etc.) provides a user with a single user interface for interacting with all building automation devices in the physical environment. In other words, the user does not have to navigate to a different user interface to control each building automation device (e.g., a smart smart light, a smart speaker, etc.), but can control the building automation devices with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0302] In some embodiments, in response to receiving input at a corresponding graphical indication corresponding to a corresponding building automation device, at least one control for controlling at least one aspect of the corresponding building automation device (e.g., changing the temperature on a smart thermostat, or changing the brightness and / or color of a smart light (e.g., Figures 5J to 5N 536 in the light control user interface). In some embodiments, in response to a detected input corresponding to a corresponding icon representing a corresponding building automation device, at least one control for controlling at least one aspect of the corresponding building automation device is displayed. The icon changes appearance in response to being selected (e.g., is displayed with a selection indication). Displaying controls for representing one or more building automation devices (e.g., smart lights, smart speakers, etc.) enables a user to quickly change multiple aspects of the building automation device without having to open multiple user interfaces in a settings menu. In other words, the user does not have to navigate to a different user interface to control each building automation object (e.g., smart smart lights, smart speakers, etc.), but can control the building automation devices with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0303] In some embodiments, the user interface is a first user interface that includes a first view of the physical environment (e.g., an isometric view) and a first user interface element (e.g., Figure 5DD In some embodiments, in response to detecting activation of the first user interface element, a user interface is displayed as a first user interface that includes a first view of the physical environment (e.g., an isometric view) (832). The user interface includes a second user interface element (e.g., Figure 5DD 525-2), wherein in response to activation of a second user interface element, a second view of the physical environment that is different from the first view (e.g., a bird's-eye view, a three-dimensional top view, or a two-dimensional (orthogonal) blueprint or floor plan view) is displayed; (in some embodiments, in response to detecting the second user interface element (e.g., Figure 5DDIn some embodiments, in response to detecting activation of the third user interface element, a user interface is displayed as a third view of the physical environment that is different from the first view (e.g., a bird's-eye view, a three-dimensional top view, or a two-dimensional (orthogonal) blueprint or floor plan view). In some embodiments, in response to detecting activation of the third user interface element, a user interface is displayed as a third view of the physical environment that is different from the first view and the second view (e.g., a simplified wireframe view with at least some texture and detail removed from the physical environment). In some embodiments, the user interface includes a fourth user interface element, in response to activation of the fourth user interface element, a fourth view of the physical environment that is different from the first view, the second view, and the third view (e.g., a side view, a three-dimensional side view, or a two-dimensional (orthogonal) side view) and the fourth user interface element are displayed. A user interface having multiple user interface elements, each of which corresponds to changing a view of the physical environment, provides a user with simple controls for manipulating views of the physical environment. With such controls, the user does not need to manually make multiple inputs to change the view of the physical environment. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0304] It should be understood that Figures 8A to 8C The specific order in which the operations are described in the description is merely exemplary and is not intended to indicate that the order is the only order in which the operations may be performed. A person of ordinary skill in the art will appreciate a variety of ways to reorder the operations described herein. In addition, it should be noted that the details of other processes described herein with respect to other methods described herein (e.g., methods 700, 900, 1000, 1500, 1600, 1700, and 1800) also apply in a similar manner to the details of the processes described above with respect to Figures 8A to 8C For example, the physical environment, features and objects, virtual objects, inputs, user interfaces, and views of the physical environment described above with reference to method 800 optionally have one or more of the characteristics of the physical environment, features and objects, virtual objects, inputs, user interfaces, and views of the physical environment described with reference to other methods described herein (e.g., methods 700, 900, 1000, 1500, 1600, 1700, and 1800). For the sake of brevity, these details are not repeated here.
[0305] Figures 9A to 9G is a flow chart illustrating a method 900 for displaying modeled spatial interactions between virtual objects / annotations and a physical environment according to some embodiments. The method 900 is performed on a computer system (e.g., a portable multifunction device 100 ( Figure 1A ), device 300( Figure 3A ) or computer system 301( Figure 3B )) is executed at the computer system, the computer system having a display generation component (for example, a touch screen 112 ( Figure 1A )、Display 340( Figure 3A ) or display generation component 304 ( Figure 3B )) and one or more input devices (e.g., touch screen 112 ( Figure 1A )、Touchpad 355( Figure 3A ), input device 302 ( Figure 3B ) or a physical button separate from the display) (902). Some operations in method 900 are optionally combined, and / or the order of some operations is optionally changed.
[0306] As described below, method 900 describes adding a virtual object to a representation of a physical environment and indicating to the user that the virtual object is interacting with (e.g., partially overlapping) a physical object in the physical environment. One indication of such interaction is that when a virtual object partially overlaps a physical object (e.g., a real-world object) in the physical environment, the virtual object is shown to be moving at a slower rate (e.g., being dragged by the user). Such interaction indicates to the user that the virtual object is interacting with the physical object occupying the physical space in the physical environment. Providing such feedback to the user helps the user orient the virtual object so that it does not overlap with the real-world object, as overlapping of the virtual object with the real-world object is unlikely to occur in the physical environment. Without such a feature, the user would have to make multiple inputs to avoid overlapping of the virtual object with the physical object. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0307] The method includes displaying (904) via a display generation component a representation of a physical environment (e.g., a two-dimensional representation, such as a real-time view from one or more cameras, or a previously captured still image or a frame of a previously captured video). The representation of the physical environment includes a representation of a first physical object occupying a first physical space in the physical environment (e.g., Figure 5OIn some embodiments, the representation of the physical environment is a real-time view of the field of view of one or more cameras of the computer system (e.g., Figure 5O In some embodiments, the representation of the physical environment is a previously captured still image (e.g., a previously captured photograph or a frame of a previously captured video) (906). The method also includes a virtual object (e.g., a virtual stool 542) at a location in the representation of the physical environment that corresponds to a second physical space in the physical environment that is different from the first physical space (908). In some embodiments, the representation of the physical environment includes or is associated with the use of one or more cameras (e.g., optical sensors 164 ( Figure 1A ) or camera 305( Figure 3B )) and one or more depth sensors (e.g., time-of-flight sensor 220 ( Figure 2B )) is associated with the depth information of the physical environment captured.
[0308] The method includes detecting (910) a first input corresponding to a virtual object, wherein movement of the first input corresponds to a request to move the virtual object in a representation of a physical environment relative to a representation of a first physical object. The method includes, upon detecting (912) the first input, at least partially moving the virtual object in the representation of the physical environment based on the movement of the first input. Based on determining that the movement of the first input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to physical space in the physical environment that is not occupied by a physical object having first corresponding object properties, at least partially moving the virtual object in the representation of the physical environment includes moving the virtual object by a first amount (e.g., Figure 5O Based on determining that the movement of the first input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to a physical space in the physical environment that at least partially overlaps with a first physical space of the first physical object, at least partially moving the virtual object in the representation of the physical environment includes moving the virtual object a second amount, less than the first amount, through at least a subset of the one or more locations corresponding to the physical space in the physical environment that at least partially overlaps with the first physical space of the first physical object (e.g., before no longer overlapping with the first physical space). Figures 5P to 5Q543-1 at a location where the virtual stool 542 in the representation of the physical environment overlaps. In some embodiments, an overlap between a virtual object and one or more representations of physical objects in the representation of the physical environment (e.g., visual overlap in a two-dimensional representation) does not necessarily mean that the virtual object moves through a location corresponding to an overlapping physical space (e.g., spatial overlap in a three-dimensional space). For example, a virtual object moves through a corresponding space behind or below a physical object (e.g., at a different depth than a portion of the physical object), creating (e.g., visual) overlap between the representations of the virtual object and the physical object in the (e.g., two-dimensional) representations of the physical environment and the virtual object (e.g., due to the virtual object being significantly obscured by at least a portion of the physical object representation), even if the virtual object (e.g., in a virtual sense) does not occupy any of the same physical space occupied by any physical object. In some embodiments, the magnitude of the virtual object's movement (through a location in the representation of the physical environment corresponding to a physical space that at least partially overlaps with a first physical space of a first physical object) decreases as the degree of overlap between the corresponding physical space and the first physical space increases.
[0309] In some embodiments, the representation of the physical environment corresponds to a first (e.g., perspective) view of the physical environment, and the method includes: based on determining that the virtual object is located at a corresponding location in the representation of the physical environment, causing one or more portions of the virtual object to overlap with one or more representations of corresponding physical objects in the physical environment and corresponding to a physical space in the physical environment that is occluded by the one or more corresponding physical objects from the first perspective of the physical environment, thereby changing (918) the appearance (e.g., the virtual stool 542 is occluded from the first perspective) of the virtual object. Figures 5P to 5QIn some embodiments, a portion of the virtual object that is obscured is de-emphasized, while another portion of the virtual object is not obscured and is therefore not de-emphasized (e.g., displayed in a de-emphasized state). The virtual object is displayed in a de-emphasized state so as to de-emphasize (e.g., such as by being displayed as at least partially transparent, forgoing display, and / or displaying (e.g., only an outline)) one or more portions of the virtual object that overlap with one or more representations of the corresponding physical object (e.g., to indicate that a portion of the virtual object is partially obscured (e.g., occluded) from view by the representation of the first physical object). In some embodiments, the obscured portion of the virtual object is de-emphasized, while another portion of the virtual object is not obscured and is therefore not de-emphasized (e.g., displayed as a texture). The visual change (e.g., de-emphasis) of the virtual object shows the user that the virtual object is obscured by the physical object, which provides the user with depth information about the virtual object in the physical environment. By providing the user with enhanced depth information about the virtual object, the user can easily place the virtual object in the physical environment with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0310] In some embodiments, the representation of the physical environment corresponds to a first (e.g., perspective) view of the physical environment. The embodiment includes, in response to detecting a first input corresponding to a virtual object, (920) displaying an outline around the virtual object. While continuing to display the outline around the virtual object, based on determining the corresponding position of the virtual object in the representation of the physical environment such that one or more portions of the virtual object overlap with one or more representations of corresponding physical objects in the physical environment, and corresponding to a physical space in the physical environment that is occluded by the one or more corresponding physical objects from the first perspective of the physical environment, forgoing display of the one or more portions of the virtual object that overlap with the one or more representations of the corresponding physical objects (e.g., while maintaining display of the outline around the virtual object, e.g., regardless of whether the outline portion of the virtual object is displayed). In some embodiments, forgoing display of the one or more portions of the virtual object that overlap with the one or more representations of the corresponding physical objects includes displaying the one or more overlapping portions of the virtual object without texture (e.g., including visual characteristics other than the shape of the virtual object, such as material properties, pattern, design, finish. In some embodiments, light and / or shadows are not considered texture and remain displayed). In some embodiments, non-overlapping portions of the virtual object are displayed with texture and outline (e.g., based on determining that those portions of the virtual object are not occluded). By giving up displaying a portion of the virtual object but still displaying the outline around the virtual object, the user is shown that the virtual object is partially obscured by the physical object, thereby providing the user with depth information of the virtual object in the physical environment. By providing the user with enhanced depth information about the virtual object, the user can easily place the virtual object in the physical environment with minimal input. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating the device / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0311] In some embodiments, detecting the first input is stopped (922); in response to stopping detecting the first input, moving the virtual object (e.g., as if Figure 5R) to a position in the representation of the physical environment that corresponds to a physical space near a first physical object in the physical environment, which physical space does not overlap with the first physical space of the first physical object (e.g., or with the physical space of any corresponding physical object in the physical environment). (In some embodiments, the physical space near the first physical object is a physical space that does not overlap with the physical space of the first physical object (or any physical object) and is subject to various constraints such as the path of the first input and / or the shapes of the first physical object and the virtual object. In some embodiments, the physical space near the first physical object is a non-overlapping physical space that has a minimum distance (e.g., in the physical environment) from the physical space corresponding to the corresponding position of the virtual object (e.g., such that the least amount of movement would be required to move the object directly (e.g., along the shortest path) from the physical space corresponding to the virtual object to the nearest physical space). In some embodiments, the physical space near the first physical object is determined based on a vector between the initial position of the input and the ending position of the input, and corresponds to a point along the vector that is as close as possible to the ending position of the input without causing the virtual object to occupy space overlapping with one or more (e.g., any) physical objects. In some embodiments, the speed and / or acceleration of the virtual object to quickly move (e.g., return) to the position corresponding to the unoccupied physical space near the first physical object depends on the speed and / or acceleration corresponding to the virtual object. The degree of overlap between the physical space of the object and the physical space occupied by the physical object (e.g., how many of the one or more locations correspond to the overlapping physical space through which the virtual object moves). For example, the virtual object starts at a position corresponding to a larger overlap between the physical space occupied by the physical object and the physical space corresponding to the virtual object (e.g., "occupied" by it in a virtual sense), and moves quickly with a higher speed and / or acceleration to a position corresponding to unoccupied virtual space near the first physical object; in another example, the virtual object starts at a position corresponding to a smaller overlap between the physical space occupied by the physical object and the physical space corresponding to the virtual object, and moves quickly with a lower speed and / or acceleration to a position corresponding to unoccupied virtual space near the first physical object. After detecting a gesture that causes the virtual object to partially overlap with the physical object, the computer device automatically moves the virtual object to the nearest physical space in the physical environment that does not cause overlap, which provides the user with context about how to place virtual items in the physical environment and is consistent with physical aspects of the physical environment (e.g., physical objects cannot overlap with each other).Performing an operation without further user input when a set of conditions have been met (e.g., automatically) reduces the amount of input required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0312] In some embodiments, after the virtual object moves a second amount, which is less than the first amount, through at least a subset of one or more locations corresponding to physical space in the physical environment that at least partially overlap with the first physical space of the first physical object, the virtual object is moved (924) through the first physical space of the first physical object (e.g., to a location in the representation of the physical environment that corresponds to physical space not occupied by a physical object in the physical environment, as determined by determining that the movement of the first input satisfies the distance threshold). Figure 5S to Figure 5V ). In some embodiments, the distance threshold is satisfied when the movement of the first input corresponds to a request to move a virtual object through one or more locations corresponding to a physical space that at least partially overlaps with a first physical space of a first physical object to a location corresponding to a physical space that does not overlap with the physical space of any physical object (e.g., and optionally is at least a threshold distance away). For example, if a user attempts to drag a virtual chair through a space corresponding to a space occupied by a physical table, then based on a determination that the user is attempting to move the chair through the table to place the chair on the other side of the table, the chair is displayed as moving quickly (e.g., moving) "through" the table after previously being displayed as resisting movement "into" the table. After detecting that the gesture that causes the virtual object to partially overlap with the physical object satisfies the distance threshold, the virtual object is moved through the first physical space of the first physical object, providing the user with the ability to move the object through the physical object with one input when the device determines that this is in fact the user's intention. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0313] In some embodiments, based on determining that movement of the first input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to physical spaces in the physical environment that at least partially overlap with the first physical space of the first physical object. Based on determining that the first input satisfies a speed threshold (e.g., and / or in some embodiments, such as Figures 5X to 5ZIn some embodiments, the virtual object is moved by a second amount less than the first amount based on a determination that the input does not satisfy a speed threshold (e.g., and / or an acceleration threshold and / or a distance threshold). After detecting that the input partially overlaps the virtual object with the physical object, a determination is made as to whether the input satisfies the speed threshold, and if so, the virtual object is moved through the first physical space of the first physical object. This interaction provides the user with the ability to make a single input to move the object through the physical object when the device determines that this is the user's intention. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0314] In some embodiments, based on determining that movement of the first input corresponds to a request to move a virtual object through one or more locations in a representation of the physical environment, the one or more locations corresponding to a physical space in the physical environment that at least partially overlaps with a first physical space of the first physical object. Based on determining that the first input does not satisfy a speed threshold (e.g., and / or in some embodiments, an acceleration threshold and / or a distance threshold), and / or determining that the first input corresponds to a request to move the virtual object to a corresponding location corresponding to a physical space in the physical environment that does not overlap with the first physical space of the first physical object, moving the virtual object through one or more locations corresponding to a physical space in the physical environment that at least partially overlaps with the first physical space of the first physical object to the corresponding location is abandoned (928). After detecting that the input causes the virtual object to partially overlap with the physical object, determining whether the input satisfies the speed threshold, and if the input does not satisfy the speed threshold, not moving the virtual object through the first physical space of the first physical object. This interaction provides the user with the ability to not accidentally move an object through a physical object when the user does not intend to do so. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0315] In some embodiments, the initial position of the first input is within the display area of the virtual object (e.g., and at least a predefined threshold distance from an edge of the virtual object) (930). In some embodiments, movement of the first input corresponds to moving the virtual object in the representation of the physical environment based at least in part on determining that the initial position of the first input is within the display area of the virtual object (e.g., and not at or near an edge of the virtual object, as in Figure 5O and Figure 5X The position of the virtual object is changed in response to input within the display area of the virtual object, providing a user with an intuitive and unique area to provide input that causes the change in the position of the object. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0316] In some embodiments, a second input corresponding to the virtual object is detected. In response to detecting the second input, the virtual object is adjusted in the representation of the physical environment (e.g., Figure 5AA 547) size request (e.g., Figure 5AA 549 in (eg, relative to a representation of a first physical object (eg, Figure 5AA ), resizing the virtual object in the representation of the physical environment based on movement of the second input (932), wherein based on determining that the movement of the second input corresponds to a request to resize the virtual object so that at least a portion (e.g., an edge) of the virtual object is within a predefined distance threshold of an edge (e.g., or multiple edges) of the first physical object, (e.g., automatically) resizing the virtual object in the representation of the physical environment based on the movement of the second input (932) includes resizing the virtual object to quickly move to the edge (e.g., or multiple edges) of the first physical object (e.g., as Figures 5AA to 5CC, which resizes in response to a drag input and snaps to a representation of a wooden table 544). In some embodiments, if resizing is requested in multiple directions (e.g., length and width), the virtual object is resized in each requested direction and can optionally be "snapped" to abut one or more virtual objects in each direction. In some embodiments, the object is moved to a predefined position relative to the object as long as it is within a threshold distance to which the object is snapped. When the input corresponds to a request to resize a virtual object, and the request to resize the virtual object ends within a predefined distance threshold of the edge of a physical object, the device automatically resizes the virtual object to abut the physical object. This provides the user with the ability to resize objects to align edges without requiring the user to make fine adjustments to achieve the desired virtual object size. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0317] In some embodiments, determining (934) that the second input corresponds to a request to resize the virtual object includes determining that the initial position of the second input corresponds to an edge of the virtual object (e.g., input 548-1 at Figure 5AA In some embodiments, movement of an input (e.g., drag) initiated from a corner of a virtual object resizes the virtual object in multiple directions (e.g., any combination of length, width, and / or height). Inputs occurring at the edges of a virtual object that result in resizing of the virtual object provide the user with intuitive control over resizing and reduce the amount of input required to resize the virtual object. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0318] In some embodiments, based on determining that the second input includes movement in a first direction, adjusting the size of the virtual object includes adjusting (936) the size of the virtual object in the first direction (e.g., without adjusting the size of the virtual object in one or more other directions, or in other words, without maintaining an aspect ratio between the size of the virtual object in the first direction (e.g., a first dimension, such as length) and the size of the virtual object in other directions (e.g., along other dimensions, such as width or height)). Based on determining that the drag gesture includes movement in a second direction, adjusting the size of the virtual object includes adjusting the size of the virtual object in the second direction (e.g., without adjusting the size of the virtual object in one or more other directions, or in other words, without maintaining an aspect ratio between the size of the virtual object in the second direction (e.g., a second dimension, such as width) and the size of the virtual object in other directions (e.g., along other dimensions, such as length or height)). Adjusting the size of the virtual object in the direction the input is moving provides the user with intuitive controls for adjusting the size of different parts of the virtual object (e.g., dimensions such as width or height). Intuitive controls result in fewer erroneous inputs. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0319] In some embodiments, the representation of the physical environment is displayed (938) from changing the representation of the first physical object and the virtual object (e.g., Figures 5II to 5KK The method further comprises: providing light from a light source (e.g., light from a physical environment or a virtual light) that enhances the visual appearance of virtual objects (e.g., virtual table 547 and virtual stool 542) in the physical environment casting shadows on a physical object (e.g., a representation of carpet 558). Based on determining that the virtual object is located in the representation of the physical environment corresponding to a physical space between the light source in the physical environment and a first physical object (e.g., a first physical space occupied by the virtual object) (e.g., the virtual object at least partially “blocks” a path of light that would otherwise be cast on the physical object by the light source), displaying a shadow region (e.g., a simulated shadow) over at least a portion of the representation of the first physical object (e.g., over a portion of the representation of the first physical object that is “obscured” from light by the virtual object, as if the virtual object casts a shadow over the first physical object).
[0320] Based on a determination (e.g., a first physical space occupied by a first physical object) between a light source and a physical space corresponding to the position of a virtual object in a representation of a physical environment (e.g., the physical object at least partially blocks the path of light that would otherwise be "cast" onto the virtual object by the light source), a shadow region (e.g., a simulated shadow) is displayed on at least a portion of the virtual object (e.g., over the portion of the virtual object that is "blocked" by the first physical object, as if the first physical object casts a shadow over the virtual object). In some embodiments, the light source can be a virtual light source (e.g., a simulated light, including, for example, a simulated colored light displayed in response to a color change of a smart light). In some embodiments, the light source is in the physical environment (e.g., sunlight, lighting from a physical light bulb). In some embodiments, where the light source is in the physical environment, the computer system determines the position of the light source and displays the shadow based on the determined position of the light source. Automatically displaying shadow regions on both physical and virtual objects provides the user with a representation of the realistic physical environment. When the virtual object appears realistic, the user does not need to enter a separate application and edit the representation of the physical environment to enhance the "realism" of the virtual object. Performing an operation without further user input when a set of conditions have been met (e.g., automatically) reduces the amount of input required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0321] In some embodiments, light from a light source that changes the visual appearance of the representation of the first physical object and the virtual object (e.g., light from the physical environment or from a virtual light) is displayed (940) in the representation of the physical environment (e.g., as Figure 5L ; Figures 5II to 5KK As shown). Based on determining a position of the virtual object in the representation of the physical environment that corresponds to a physical space in the physical environment, the physical space is in the path of light from the light source, thereby increasing the brightness of the area of the virtual object (e.g., showing at least some light projected from the light source onto the virtual object). Automatically increasing the brightness of the area of the virtual object in response to the light from the light source provides the user with the ability to make the virtual object appear real without having to enter a separate media editing application. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve intended results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0322] In some embodiments, the representation of the physical environment includes a representation of a second physical object that occupies a third physical space in the physical environment and has a second corresponding object property (e.g., is a "soft" or springy object, as indicated by the virtual table 547 and the virtual stool 542, causing the object to deform). Upon detecting (942) a corresponding input corresponding to a request to move the virtual object in the representation of the physical environment relative to the representation of the second physical object (e.g., the corresponding input is a portion of the first input, or is different from the first input), the virtual object in the representation of the physical environment is at least partially moved based on the movement of the corresponding input. Based on a determination that the movement of the corresponding input corresponds to a request to move the virtual object through one or more locations in the representation of the physical environment, the one or more locations corresponding to a physical space in the physical environment that at least partially overlaps with the third physical space of the second physical object (e.g., and optionally based on a determination that the virtual object has the first corresponding object property). The virtual object is at least partially moved through at least a subset of one or more locations of physical space in the physical environment corresponding to a third physical space that at least partially overlaps with a second physical object; and in some embodiments, for a given amount of movement of the corresponding input, the amount of physical space that the virtual object moves through that overlaps with the physical object having the second corresponding object attribute is greater than the corresponding amount of physical space that the virtual object would have moved through that overlaps with the physical object having the first corresponding object attribute. For example, in response to the same degree of overlap requested by the movement input, the virtual object can be moved to appear more "embedded" in the soft physical object than in the rigid physical object.
[0323] In some embodiments, one or more changes in the visual appearance (e.g., simulated deformation) of at least a portion of a representation of a second physical object that at least partially overlaps with the virtual object are displayed. In some embodiments, the change in visual appearance (e.g., the extent of the simulated deformation) is based at least in part on a second corresponding object property of the second virtual object, and optionally also on simulated physical properties of the virtual object, such as stiffness, weight, shape, and velocity and / or acceleration of motion. In some embodiments, the deformation is maintained while the virtual object remains in a position corresponding to a physical space that at least partially overlaps with the physical space of the second physical object. In some embodiments, after the virtual object is moved so that the virtual object no longer "occupies" one or more physical spaces that overlap with the physical space occupied by the second physical object, the one or more changes in the visual appearance of at least a portion of the representation of the second physical object cease to be displayed (e.g., and optionally, a different set of changes are displayed (e.g., one or more changes are reversed) so that the second physical object appears to have returned to its original appearance before being simulated deformed by the virtual object). For example, a simulated deformation of a physical sofa cushion is displayed when a (e.g., rigid, heavy) virtual object is placed on the sofa cushion, and the deformation gradually decreases (e.g., reverses) as the sofa cushion returns to its shape after the virtual object is removed. In some embodiments, the object properties of the physical object (e.g., physical characteristics, such as material hardness, stiffness, elasticity, etc.) are determined by the computer system, and based on the determined object properties of the physical object, different simulated interactions between the physical object and the virtual object will be displayed. Automatic deformation of the physical object in response to the virtual object provides the user with a more realistic representation of the physical environment. When the virtual object looks very real, the user does not need to enter a separate application and edit the representation of the physical environment to enhance the "reality" of the virtual object. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user obtain the expected results and reducing user errors when operating the device / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0324] In some embodiments, one or more changes in the visual appearance of at least a portion of a representation of a second physical object that at least partially overlaps with the virtual object are displayed (944) based on one or more object properties (e.g., physical characteristics) of the second physical object in the physical environment (e.g., based on depth data associated with the second physical object). Automatically detecting the object properties of the physical object provides the user with a representation of the real physical environment. When the virtual object appears real, the user does not need to enter a separate application and edit the representation of the physical environment to enhance the "realism" of the virtual object. Performing an operation (e.g., automatically) when a set of conditions have been met without further user input reduces the number of inputs required to perform the operation, which enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve the desired results and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0325] It should be understood that Figures 9A to 9G The specific order in which the operations are described in the description is merely exemplary and is not intended to indicate that the order is the only order in which the operations may be performed. A person of ordinary skill in the art will appreciate a variety of ways to reorder the operations described herein. In addition, it should be noted that the details of other processes described herein with respect to other methods described herein (e.g., methods 700, 800, 1000, 1500, 1600, 1700, and 1800) also apply in a similar manner to the details of the processes described above with respect to Figures 9A to 9G Method 900 is described. For example, the physical environment, features and objects, virtual objects, object properties, inputs, user interfaces, and views of the physical environment described above with reference to method 900 optionally have one or more of the characteristics of the physical environment, features and objects, virtual objects, object properties, inputs, user interfaces, and views of the physical environment described with reference to other methods described herein (e.g., methods 700, 800, 1000, 1500, 1600, 1700, and 1800). For the sake of brevity, these details are not repeated here.
[0326] Figures 10A to 10E 1 is a flow chart illustrating a method 1000 for applying modeled space interaction with virtual objects / annotations to multiple media items according to some embodiments. The method 1000 is performed on a computer having a display generation component and one or more input devices (and optionally, one or more cameras (e.g., optical sensors 164 ( Figure 1A ) or camera 305( Figure 3B )) and one or more depth sensing devices (e.g., time of flight sensor 220 ( Figure 2B ))) is executed at a computer system (1002).
[0327] As described below, method 1000 describes annotating a representation of a physical environment (e.g., marking a photo or video), wherein the annotation position, orientation, or scale is determined within the physical environment. Using the annotation position, orientation, or scale in the representation, subsequent representations that include the same physical environment can be updated to include the same annotation. The annotation will be placed in the same position relative to the physical environment. Having such features avoids the need for the user to repeatedly annotate multiple representations of the same environment. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating the device / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0328] The method includes displaying (1004) via a display generation component a first representation of first previously captured media (eg, including one or more images (eg, Figure 6C
[0064] The method further includes providing an expanded media item 604 in the first representation of the first media and, in some embodiments, storing the expanded media item 604 with the depth data, wherein the first representation of the first media includes a representation of the physical environment. While displaying the first representation of the first media, receiving (1006) a request corresponding to annotating a portion of the first representation corresponding to the first portion of the physical environment (e.g., by adding a virtual object or modifying an existing displayed virtual object, such as Figures 6D to 6E As shown, where an annotation 606 is added to the input of a request to expand a media item 604).
[0329] In response to receiving the input, an annotation is displayed on a portion of the first representation corresponding to the first portion of the physical environment, the annotation having one or more of a position, orientation, or scale determined based on the physical environment (e.g., its physical characteristics and / or physical objects) (e.g., using depth data corresponding to the first media) (1008).
[0330] After (e.g., in response to) receiving the input, displaying an annotation on a portion of a displayed second representation of second previously captured media, wherein the second previously captured media is different from the first previously captured media and the portion of the second representation corresponds to the first portion of the physical environment (e.g., the annotation is displayed on the portion of the second representation of the second previously captured media having one or more of a position, orientation, or scale determined based on the physical environment, as shown in the second representation of the second previously captured media) (1010) (see, e.g., Figure 5O, which shows media thumbnail item 1 602-1, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4, each including annotation 606). In some embodiments, the annotation is displayed on a portion of a second representation of the second previously captured media using depth data corresponding to the second media. In some embodiments, where the viewpoint of the first portion of the physical environment represented in the second representation of the second previously captured media is different from the viewpoint of the first representation of the first media relative to the first portion of the physical environment, the position, orientation, and / or scale of the annotation (e.g., virtual object) differs between the first representation and the second representation based on their respective viewpoints.
[0331] In some embodiments, after (e.g., in response to) receiving (1012) input corresponding to a request to annotate the portion of the first representation and before displaying the annotation on the portion of the displayed second representation of the second media, the method includes displaying a first animated transition from display of the first representation of the first media (e.g., in response to input corresponding to selection of the second media) to displaying a first representation of the three-dimensional model of the physical environment represented in the first representation of the first media (e.g., Figures 6F to 6H In some embodiments, the first animated transition is displayed in response to input corresponding to selection of the second media. In some embodiments, the transition representation of the three-dimensional model of the physical environment is simplified relative to the first and second representations of the media. In some embodiments, the transition representation of the three-dimensional model of the physical environment is a wireframe representation of the three-dimensional model of the physical environment generated based on detected physical features such as edges and surfaces.
[0332] In some embodiments, the method further includes displaying a second animated transition from displaying the first representation of the three-dimensional model to displaying a second representation of the three-dimensional model of the physical environment represented in a second representation of the second media (e.g., Figures 6I to 6Jshowing such an animated transition) (e.g., generated by a computer system from depth information indicative of a physical environment associated with (e.g., stored together with) the second media), and representing one or more (e.g., any) annotations displayed at least partially in the second representation of the second media. In some embodiments, the animated transition from the first representation of the 3D model to the second representation of the 3D model includes performing one or more transformations of the 3D model, including, for example, rotation, translation, and / or rescaling of the 3D model.
[0333] In some embodiments, the method also includes displaying a third animated transition (e.g., 6K to 6L shows such an animated transition) from display of the second representation of the three-dimensional model to display of the second representation of the second media (e.g., including stopping the display, such as by fading out the second representation of the three-dimensional model, and optionally by (e.g., simultaneously) fading in the second representation of the second media). Displaying the animated transition when switching between different representations of the physical environment (e.g., representations of the three-dimensional model of the physical environment) provides the user with contextual information about the different positions, orientations, and / or magnifications of the captured representations. Providing improved feedback enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0334] In some embodiments, an annotation (e.g., Figures 6F to 6M Annotations are shown being displayed 606. Displaying animated transitions with annotations when switching between different representations of the physical environment (e.g., representations of a three-dimensional model of the physical environment) provides the user with contextual information about the different positions, orientations, and magnifications at which the representations are captured, as well as where the annotations will be placed. Providing improved feedback enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0335] In some embodiments, the display of the annotation is updated (1016) in response to the first change in perspective (e.g., in response to an input corresponding to a selection of the second media) based on determining that the first animated transition from the display of the first representation of the first media (e.g., in response to an input corresponding to a selection of the second media) to the display of the first representation of the three-dimensional model of the physical environment represented in the first representation of the first media includes a first change in perspective. Figures 6F to 6Hshowing such an animated transition) (e.g., displaying an annotation having one or more of a position, orientation, or scale determined based on the physical environment, as represented during the first change in perspective during the first animated transition).
[0336] In some embodiments, based on determining that the second animated transition from displaying the first representation of the three-dimensional model to displaying the second representation of the three-dimensional model of the physical environment represented in the second representation of the second medium includes a second change in perspective, the method includes updating the display of the annotation in response to the second change in perspective (e.g., Figures 6I to 6J showing such an animated transition) (e.g., displaying an annotation having one or more of a position, orientation, or scale determined based on the physical environment as represented during the first change in perspective during the second animated transition).
[0337] In some embodiments, based on determining that the third animated transition from displaying the second representation of the three-dimensional model to displaying the second representation of the second media includes a third change in perspective, updating the display of the annotation in response to the third change in perspective (e.g., Figures 6K to 6L showing such animated transitions) (e.g., displaying an annotation with one or more of a position, orientation, or scale that is determined based on the physical environment, as represented during a third change of perspective during the first animated transition). Displaying multiple animations including changes in perspective of the annotation provides the user with the ability to see how annotations made in one representation will appear in the other representations. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0338] In some embodiments, after (e.g., in response to) receiving input corresponding to a request to annotate a portion of the first representation, and before displaying the annotation on the portion of the displayed second representation of the second media: receiving (1018) input corresponding to a selection of the second media; and in response to receiving input corresponding to the selection of the second media, displaying the corresponding representation of the second media (e.g., Figure 6FIn some embodiments, displaying the annotation on the portion of the second representation of the second media is performed after (e.g., in response to) receiving an input corresponding to a selection of the second media. In some embodiments, the corresponding representation of the second media is an image or representation (e.g., an initial frame) of a video that at least partially corresponds to at least a first portion of the physical environment. In some embodiments, the corresponding representation of the second media is a second representation of the second media, and in some such embodiments, in response to receiving an input corresponding to a selection of the second media, the annotation is displayed on the portion of the second representation (e.g., which serves as the corresponding representation). In some embodiments, the corresponding representation of the second media is a different frame of video than the second representation of the second media, and in some such embodiments, the second representation of the second media is displayed after at least a portion of the video is played. In some such embodiments where the corresponding representation of the second media does not correspond to the first portion of the physical environment, the annotation is not displayed in the corresponding representation of the second media. In some such embodiments, the annotation is not displayed until playback of the video reaches an initial frame (e.g., the second representation) of the second media, the initial frame corresponding to at least the first portion of the physical environment, and in response to receiving an input corresponding to a selection of the second media, the annotation is displayed on the portion of the second representation of the second media in conjunction with displaying the second media. In some embodiments, a computer system displays a first representation of a first media in a user interface that also includes a media (e.g., image or video) selector that includes one or more corresponding media representations, such as thumbnails (e.g., displayed in a scrollable list or array), and input corresponding to selecting a second media is input corresponding to a representation of the second media displayed in the media selector. In some embodiments, after receiving input corresponding to a request to annotate a portion of the first representation corresponding to the first portion of the physical environment, the corresponding representation of the media corresponding to at least the first portion of the physical environment is updated to reflect the addition of the annotation to the first representation.
[0339] Including a media selector (e.g., media thumbnail eraser 605) in the user interface provides the user with a quick control for switching between media items and does not require the user to navigate multiple user interfaces to interact with each media item. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0340] In some embodiments, after displaying the annotation on the portion of the first representation corresponding to the first portion of the physical environment, an input corresponding to a request to view a real-time representation of the physical environment (e.g., a real-time feed from a camera) is received (1020). In response to receiving the request to view a representation of the current state of the physical environment (e.g., a representation of the field of view of one or more cameras that changes as the physical environment changes within the field of view of the one or more cameras or as the field of view of the one or more cameras changes around the physical environment): a representation of the current state of the physical environment is displayed. Based on determining that the representation of the current state of the physical environment corresponds to at least the first portion of the physical environment, an annotation is displayed on a portion of the representation of the current state of the physical environment corresponding to the first portion of the physical environment, wherein the annotation is displayed in one or more of a position, orientation, or scale determined based on the physical environment as represented in the representation of the current state of the physical environment (e.g., the annotation appears different in the real-time representation than in the first or second representation based on a difference between the viewpoint of the real-time representation and the viewpoint of the first or second representation). Displaying multiple representations provides the user with the ability to view how annotations made in one representation will appear in the other representations. Viewing the representations simultaneously avoids the user having to switch between the representations. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0341] In some embodiments, the first representation and the second representation are displayed simultaneously (e.g., Figure 6O showing simultaneous display of media items, and Figures 6D to 6EThe media item is shown moving simultaneously with the annotation); the input corresponding to the request to annotate a portion of the first representation includes movement of the input; and in response to receiving the input, simultaneously: modifying (1022) (e.g., moving, resizing, expanding, etc.) the first representation of the annotation in the portion of the first representation corresponding to the first portion of the physical environment based at least in part on the movement of the input. Also simultaneously modifying (e.g., moving, resizing, expanding, etc.) the second representation of the annotation in the portion of the second representation corresponding to the first portion of the physical environment based at least in part on the movement of the input. In some embodiments, where the view of the first portion of the physical environment represented in the second representation of the second media is different from the viewpoint of the first representation of the first media relative to the first portion of the physical environment, the position, orientation, and / or scale of the annotation (e.g., virtual object) differs between the first representation and the second representation based on the respective viewpoints of the first and second representations. Displaying multiple representations provides the user with the ability to view how annotations made in one representation will appear in the other representations. Viewing the representations simultaneously avoids the user having to switch between the representations. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0342] In some embodiments, a portion of the annotation is not displayed on the second representation from the second media (1024) (e.g., because the annotation is obscured, occluded, or exceeds the frame of the second representation, such as Figure 6O 2-2).
[0343] In some embodiments, while displaying the second representation from the second previously captured media, a second input corresponding to a request to annotate (e.g., by adding a virtual object or modifying an existing displayed virtual object) a portion of the second representation corresponding to a second portion of the physical environment is received (1026). In response to receiving the second input, a second annotation is displayed on the portion of the second representation corresponding to the second portion of the physical environment, the second annotation having one or more of a position, orientation, or scale (e.g., a position, orientation, or scale) determined based on the physical environment (e.g., its physical properties and / or physical objects) (e.g., using depth data corresponding to the first image). Figure 6R A wing element (ie, a spoiler) is shown added to a car 624 via input 621 (annotation 620 ).
[0344] After (e.g., in response to) receiving the second input, a second annotation is displayed on (e.g., added to) a portion of the first representation of the first media corresponding to the second portion of the physical environment (e.g., Figure 6T Wing element annotation 620 is displayed in media thumbnail item 1 602-1, media thumbnail item 2 602-2, media thumbnail item 3 602-3, and media thumbnail item 4 602-4. In some embodiments, the second annotation is displayed on the first representation in response to receiving a second input (e.g., when the first representation is displayed simultaneously with the second representation). In some embodiments, the second annotation is displayed on the first representation in response to receiving an intermediate input corresponding to a request to (re)display the first representation of the first media. When annotating in a representation of a physical environment (e.g., tagging a photo or video), the annotation position, orientation, or scale is determined within the physical environment. Using the annotation position, orientation, or scale, subsequent representations that include the same physical environment can be updated to include the annotation. The annotation will be placed in the same position relative to the physical environment. Having such features avoids the need for the user to repeatedly annotate multiple representations of the same environment. Reducing the number of inputs required to perform an operation enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), thereby further reducing power usage and extending the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0345] In some embodiments, at least a portion of the second annotation is not displayed on the first representation of the first media (e.g., because it is obscured or outside the frame of the first representation) (see, e.g., Figure 6T, where the wing element annotation 620 is partially hidden from view in media thumbnail item 1 602-1) (1028). Making a portion of the annotation invisible provides the user with information as to whether the annotation is obscured by the physical environment. Providing improved feedback enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0346] It should be understood that Figures 10A to 10E The specific order in which the operations are described in the description is merely exemplary and is not intended to indicate that the order is the only order in which the operations may be performed. A person of ordinary skill in the art will recognize many ways to reorder the operations described herein. In addition, it should be noted that the details of other processes described herein with respect to other methods described herein (e.g., methods 700, 800, 900, 1500, 1600, 1700, and 1800) also apply in a similar manner to the details of the processes described above with respect to Figures 10A to 10E Method 1000 is described. For example, the physical environment, features and objects, virtual objects and annotations, inputs, user interfaces, and views of the physical environment described above with reference to method 1000 optionally have one or more of the characteristics of the physical environment, features and objects, virtual objects and annotations, inputs, user interfaces, and views of the physical environment described with reference to other methods described herein (e.g., methods 700, 800, 900, 1500, 1600, 1700, and 1800). For the sake of brevity, these details are not repeated here.
[0347] Figures 11A to 11JJ 、 Figures 12A to 12RR 、 Figures 13A to 13HH and Figures 14A to 14SS Example user interfaces for interacting with and annotating augmented reality environments and media items are shown in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described herein, including 7A to 7B 、 Figures 8A to 8C 、 Figures 9A to 9G 、 Figures 10A to 10E 、 FIG. 15A to FIG. 15B 、 16A to 16E 、 17A to 17D and 18A to 18B For ease of explanation, reference will be made to the process of Figure 1AIn some embodiments, the focus selector is optionally: a corresponding finger or stylus contact, a representative point corresponding to the finger or stylus contact (e.g., a center of gravity of the corresponding contact or a point associated with the corresponding contact), or a center of gravity of two or more contacts detected on the touch-sensitive display system. However, in response to detecting a contact on the touch-sensitive surface when the user interface shown in the figures is displayed on the display together with the focus selector, optionally on a device having a display (e.g., Figure 4B display 450) and a separate touch-sensitive surface (e.g., Figure 4B Similar operations are performed on a device having a touch-sensitive surface 451).
[0348] Figures 11A to 11JJ shows a method of transmitting data to a computer system (e.g., Figure 1A device 100 or Figure 3B One or more cameras (e.g., Figure 1A Optical sensor 164 or Figure 3B The camera 305 of the user's computer 306 scans the physical environment and captures various media items of the physical environment (e.g., representations of the physical environment, such as images (e.g., photos)). Depending on the capture mode selected, the captured media items can be real-time view representations of the physical environment or still view representations of the physical environment. In some embodiments, the display generation component (e.g., Figure 1A touch-sensitive display 112 or Figure 3B In some embodiments, a camera (optionally in combination with one or more time-of-flight sensors, such as Figure 2B The time-of-flight sensor 220 of the camera captures depth data of the physical environment, and the depth data is used to create a media item having a representation of the scanned physical environment (e.g., based on a three-dimensional model of the physical environment generated using the depth data). The depth data is also used to enhance user interaction with the scanned physical environment, for example, by displaying dimensional information (e.g., measurements) of physical objects, constraining user-provided annotations to predefined locations in the representation of the physical environment, displaying different views of the physical environment, etc. During scanning and capture, some features of the physical environment may be removed to provide a simplified representation of the physical environment to be displayed on a display (e.g., a simplified live view representation or a simplified still view representation).
[0349] Figure 11AA physical environment 1100 is shown. Physical environment 1100 includes a number of structural features, including two walls 1102-1 and 1102-2 and a floor 1104. In addition, physical environment 1100 includes a number of non-structural objects, including a table 1106 and a cup 1108 placed on top of table 1106. The system displays a user interface 1110-1 that includes a real-time view representation of physical environment 1100. User interface 1110-1 also includes a control bar 1112 having a number of controls for interacting with the real-time view representation of physical environment 1100 and for switching to different views of physical environment 1100.
[0350] To illustrate the position and orientation of the system's camera in the physical environment 1100 during scanning, Figure 11A Also included is a legend 1114 showing a top-down schematic diagram of the camera relative to the physical environment 1100. The top-down schematic diagram indicates a camera position 1116-1, a camera field of view 1118-1, and a schematic representation 1120 of the table 1106.
[0351] Figure 11B The device 100 and its camera are shown to have moved to different positions during the scan. To illustrate this change in camera position, legend 1114 shows an updated top-down schematic diagram, including updated camera position 1116-2. As a result of the camera moving to camera position 1116-2, the system displays a different real-time view representation of the physical environment 1100 from the perspective of the camera at camera position 1116-2 in updated user interface 1110-2.
[0352] Figure 11C A user is shown capturing a still view representation of the physical environment 1100 by placing contact 1122 on a record control 1124 in the user interface 1110 - 2 .
[0353] Figure 11D In response to user input from contact 112 on record control 1124, the still view representation of physical environment 1100 is captured, and the system displays updated user interface 1110-3, displaying the still view representation and updated control bar 1112. In control bar 1112, record control 1124 (e.g., Figure 11C ) is replaced by a return control 1126 which, when activated, causes the display (e.g., redisplay) of the real-time view representation of the physical environment 1110.
[0354] Figure 11E Shows that even when the camera moves to a different camera position 1116-3 (e.g., a real-time view representation updated as the camera moves ( Figures 11A to 11B ) different), the display of the static view representation of the physical environment 1110 (such as Figures 11C to 11D) is retained in user interface 1110-3.
[0355] Figure 11F A zoomed-in view of a user interface 1110-3 is shown that includes a stationary view representation of a physical environment 1110. A control bar 1112 includes a plurality of controls for interacting with the stationary view representation of the physical environment 1110, including a measurement control 1128, an annotation control 1130, a "slide to fade" control 1132, a "first person view" control 1134, a "top down view" control 1136, and a "side view" control 1138. Figure 11F In the example shown, the "first person view" control 1134 is selected (e.g., highlighted in the control bar 1112). Thus, the user interface 1110-3 includes a first person view (e.g., a front perspective view) of the physical environment 1100 captured from the perspective of the camera at the camera position 1116-3 ( Figure 11E ).
[0356] Figures 11G to 11L Adding annotations to a still view representation of the physical environment 1100 is shown. Figure 11G Selection of annotation control 1130 (eg, by a user) via contact 1140 is shown. Figure 11H The annotation control 1130 is shown highlighted in response to selection of the annotation control 1130. The user interface 1110-3 then enters an annotation session (eg, annotation mode) that allows the user to add annotations to the stationary view representation of the physical environment 1100.
[0357] Figure 11H The contact 1142 that initiates the annotation of the user interface 1110-3 is shown starting at a location in the user interface 1110-3 that corresponds to a position in the physical environment 1100 along the table 1148 (e.g., a representation of the table 1106, Figure 11A )'s physical location of edge 1146 (e.g., a representation of an edge). Figure 11H In some embodiments, in response to contact 1142, the system displays bounding box 1144 over a portion of the stationary view representation of physical environment 1100 (e.g., edge 1146 including table 1148). In some embodiments, bounding box 1144 is not displayed and is included in the Figure 11H In some embodiments, one or more attributes (e.g., size, position, orientation, etc.) of bounding box 1144 are determined based on the position and / or movement of contact 1142 and corresponding (e.g., nearest) features in the static view representation of physical environment 1100. For example, in Figure 11H, edge 1146 is the feature closest to the location of contact 1142. Thus, bounding box 1144 includes edge 1146. In some embodiments, when two or more features (e.g., two edges of a table) are equidistant from the location of contact 1142 (or within a predefined threshold distance from the location of contact 1142), two or more bounding boxes can be used, each containing a corresponding feature. In some embodiments, depth data recorded from the physical environment (e.g., by one or more cameras and / or one or more depth sensors) is used to identify features in a still view (or real-time view) representation of the physical environment 1100.
[0358] Figures 11I to 11J Movement of contact 1142 across user interface 1110-3 is shown along a path corresponding to edge 1146. The path of contact 1142 is entirely within bounding box 1144. As contact 1142 moves, annotations 1150 are displayed along the path of contact 1142. In some embodiments, annotations 1150 are displayed in a predefined (or, in some embodiments, user-selected) color and thickness to distinguish annotations 1150 from features (e.g., edge 1146) included in the static view representation of physical environment 1110-3.
[0359] Figures 11K to 11L The process of converting annotation 1150 to be constrained to correspond to edge 1146 is shown. Figure 11K , the user completes adding annotation 1150 by lifting contact 1142 off the display. Since annotation 1150 is completely contained within bounding box 1144, after lifting, annotation 1150 is converted into a different annotation 1150' constrained to edge 1146, as shown in FIG. Figure 11L 144 is shown. In addition, in embodiments where bounding box 1144 is displayed when adding annotation 1150, upon lifting of contact 1142, bounding box 1144 ceases to be displayed in user interface 1110-3. A label 1152 indicating the length measurement of annotation 1150' (e.g., the physical length of edge 1146 to which annotation 1150' corresponds) is displayed in a portion of user interface 1110-3 adjacent to annotation 1150'. In some embodiments, a user can add an annotation corresponding to a two-dimensional physical area or a three-dimensional physical space, optionally with corresponding labels indicating measurements of other physical properties of the annotation, such as area (e.g., for an annotation corresponding to a two-dimensional area) or volume (e.g., for an annotation corresponding to a three-dimensional space).
[0360] Figures 11M to 11P 100. FIG. 100 illustrates a process in which a user adds a second annotation 1154 to the still view representation of the physical environment 1100. Figure 11M, a bounding box 1144 is again shown to indicate the area within a threshold distance of an edge 1146 (e.g., where annotations will be converted, as described herein with reference to Figures 11K to 11L described). Figures 11N to 11O , as contact 1142-2 moves across the display, annotation 1154 is displayed along the path of contact 1142-2. Portions of the path of contact 1142-2 extend beyond bounding box 1144. Thus, after lift-off of contact 1142-2, Figure 11P , annotation 1154 remains in its original location and is not transformed to be constrained to edge 1146. In some embodiments, a label indicating the measurement of a physical property (e.g., length) of the physical space to which annotation 1154 corresponds is not displayed (e.g., because annotation 1154 is not constrained to any feature in the static view representation of physical environment 1100).
[0361] Figures 11Q to 11T A display (e.g., redisplay) of a real-time view representation of the physical environment 1100 is shown. Figure 11Q , the user uses contact 1156 on user interface 1110-3 to select return control 1126. In response, as Figure 11R As shown, the system stops displaying user interface 1110 - 3 , which includes a static view representation of physical environment 1100 , and instead displays user interface 1110 - 4 , which includes a real-time view representation of physical environment 1100 . Figure 11R Also shown is a legend 1114 indicating the camera position 1116-4 corresponding to the viewpoint for capturing the real-time view representation of the physical environment 1100 in the user interface 1110-4. In some embodiments, after switching to the real-time view representation, previously added annotations (e.g., annotations 1150' and 1154) and markers (e.g., label 1152) remain displayed relative to features (e.g., edges) in the stationary view representation of the physical environment 1100 (e.g., as the camera moves relative to the physical environment 1100, the annotations move in the displayed representation of the physical environment 1100 such that the annotations continue to be displayed over corresponding features in the physical environment 1100). Figures 11R to 11S As shown, when the real-time view representation of the physical environment 1100 is displayed, changes in the camera's field of view (eg, the soccer ball 1158 rolling into the camera's field of view 1118 - 4 ) are reflected (eg, displayed in real-time) in the user interface 1110 - 4 . Figure 11T The camera is shown having moved to a different location 1116-5 within the physical environment 1100. Thus, the user interface 1110-4 displays a different real-time view representation of the physical environment (eg, a different portion of the physical environment 1100) from the camera's perspective from location 1116-5.
[0362] Figures 11U to 11XAdding annotations to a real-time view representation of the physical environment 1100 is shown. Figure 11U 1144 , in some embodiments, the border 1162 is not displayed and is included in the image. Figure 11U 1162 is merely intended to indicate the invisible threshold, and in other embodiments, boundary 1162 is displayed (e.g., optionally upon detecting contact 1160). Figure 11V The user is shown having added an annotation 1164 around the rim of the cup 1159 (e.g., by moving contact 1160 along a path around the rim). The annotation 1164 is completely within the boundary 1162. Thus, Figures 11W to 11X It is shown that after lifting off contact 1160, annotation 1164 is converted to annotation 1164' that is constrained to the rim of cup 1159. Label 1166 is displayed next to annotation 1164' to indicate the circumference of the rim of cup 1159 to which annotation 1164' corresponds.
[0363] Figures 11Y to 11Z Switching between different types of representation views (e.g., a real-time view representation or a static view representation) of the physical environment 1100 is shown. Figure 11Y , when the “first person view” control 1134 is selected (e.g., highlighted) and a first person view (e.g., a front perspective view) of the physical environment 1100 captured from the perspective of the camera at camera position 1116-4 is displayed, the user uses contact 1166 to select the “top down view” control 1136. In response to the selection of the “top down view” control 1136 using contact 1166, the system displays an updated user interface 1110-5 showing a top down view representation of the physical environment 1100. In user interface 1110-5, the “top down” control 1136 is highlighted and the previously added annotations and labels are displayed at corresponding locations in the top down view representation of the physical environment 1100 that correspond to the locations that they were added to. Figure 11Y146 in the first-person view is also displayed as constrained to edge 1146 in the top-down view. In another example, annotation 1154 that extends unconstrained along edge 1146 in the first-person view is also displayed along edge 1146 in the top-down view, but is not limited to that edge. In some embodiments, the top-down view is a simulated representation of the physical environment using depth data collected by the camera. In some embodiments, one or more actual features of the physical environment, such as object surface textures or surface patterns, are omitted in the top-down view.
[0364] Figures 11AA to 11FF 1 shows the process of a user adding annotations to a top-down view representation of a physical environment 1100. Figure 11AA , while annotation control 1130 remains selected, the user initiates adding an annotation to the top-down view representation of physical environment 1100 using contact 1170. Bounding box 1172 indicates an area within a threshold distance of edge 1174. Figures 11BB to 11CC The movement of contact 1170 along edge 1174 is shown, along with the display of annotation 1176 along a path corresponding to the movement of contact 1170. Because annotation 1176 is completely contained within bounding box 1172, the Figure 11DD After lifting off contact 1170, annotation 1176 is converted to annotation 1176' constrained to correspond to edge 1174, as shown. Figure 11EE Additionally, a label 1178 is displayed next to the annotation 1176' to indicate the measurement of the physical region (eg, edge 1174) to which the annotation 1176' corresponds, which in this case is a length.
[0365] Figures 11FF to 11GG A switch back to displaying the first person view representation of the physical environment 1100 is shown. Figure 11FF In , the user selects the "first person view" control 1134 using contact 1180. In response, Figure 11GG In , the system is represented by a first-person view of the physical environment 1100 (e.g., Figure 11GG The first-person view shows all previously added annotations and labels in their respective locations, including the annotations 1150', 1154 and 1164' ( Figure 11R and Figure 11Z ), also includes annotation 1176'( Figure 11EE ).
[0366] Figures 11HH to 11JJThe use of a "slide to fade" control 1132 to transition a representation of the physical environment 1100 between a photorealistic view of the camera's field of view and a model view (e.g., a canvas view) of the camera's field of view is shown. When a user drags a slider thumb along the "slide to fade" control 1132 with input 1180 (e.g., a touch input, such as a drag input), one or more features of the representation of the physical environment 1100 fade out of view. The extent to which the features fade is proportional to the extent of movement of the slider thumb along the "slide to fade" control 1132, as controlled by input 1180 (e.g., a touch input, such as a drag input). Figure 11II shows a partial transition to the model view according to movement of the slider thumb from the left end to the middle of the "slide to fade" control 1132, and Figure 11JJ 136).
[0367] 12A to 12O Measuring one or more properties of one or more objects (eg, an edge of a table) in a representation of a physical environment (eg, a still view representation or a real-time view representation) is shown.
[0368] Figure 12A Shown is a physical environment 1200. The physical environment 1200 includes a number of structural features, including two walls 1202-1 and 1202-2 and a floor 1204. Additionally, the physical environment 1200 includes a number of non-structural objects, including a table 1206 and a cup 1208 placed on top of the table 1206.
[0369] The device 100 in the physical environment 1200 displays a real-time view representation of the physical environment 1200 in a user interface 1210-1 on the touch-sensitive display 112 of the device 100. The device 100 displays a real-time view representation of the physical environment 1200 via one or more cameras of the device 100 (e.g., Figure 1A Optical sensor 164 or Figure 3B 305) captures a real-time view representation of the physical environment 1200 (or on a computer system such as Figure 3BIn some embodiments of the computer system 301 of FIG. 1 , the user interface 1210 - 1 also includes a control bar 1212 having a plurality of controls for interacting with the real-time view representation of the physical environment 1200 .
[0370] Figure 12A Also included is a legend 1214 showing a top-down schematic diagram of physical environment 1200. The top-down schematic diagram in legend 1214 indicates the position and field of view of one or more cameras of device 100 relative to physical environment 1200 via a schematic diagram 1220 of camera position 1216-1 and camera field of view 1218, respectively, relative to table 1206.
[0371] Figure 12B Contact 1222 is shown selecting record control 1224 in control bar 1212. In response to selection of record control 1224, device 100 captures a still view representation (e.g., an image) of physical environment 1200 and displays updated user interface 1210-2, as shown. Figure 12C As shown, a still view representation of a captured physical environment 1200 is included.
[0372] Figure 12C Shows the response to user selection Figure 12B 12. The user interface 1210-2 includes a zoomed-in view of the user interface 1210-2 generated by the recording control 1224 in FIG. Figure 12A ). The control bar 1212 in the user interface 1210-2 includes a measurement control 1228 for activating (e.g., entering) a measurement mode to add measurements to objects in the captured still view representation of the physical environment 1200. Figure 12C Selection of measurement control 1228 via contact 1222-2 is shown.
[0373] Figure 12D Shows the response to Figure 12C Upon selection of the measurement control 1228 by contact 1222-2 in the measurement mode, device 100 displays an updated user interface 1210-3. In user interface 1210-3, measurement control 1228 is highlighted to indicate that measurement mode has been activated. In addition, user interface 1210-3 includes a plurality of controls for performing a measurement function to measure an object, including a guideline 1229 located in the center of user interface 1210-3, an add measurement point control 1233 for...
Claims
1. A method for displaying a user interface, comprising: At a computer system having a display generating component, an input device, and one or more cameras in a physical environment: capturing, via the one or more cameras, information indicative of the physical environment, the information indicative of the physical environment comprising information indicative of a corresponding portion of the physical environment within the field of view of the one or more cameras as the field of view of the one or more cameras moves, wherein the corresponding portion of the physical environment comprises a plurality of structural features of the physical environment and one or more non-structural features of the physical environment; as well as After capturing the information indicative of the physical environment, displaying a user interface based on the identified features in the captured information, wherein: The identified features in the captured information include one or more identified structural features and one or more identified non-structural features; and Displaying the user interface based on the identified features includes simultaneously displaying: a graphical representation of the plurality of structural features generated at a first level of fidelity based on the identification of the plurality of structural features; and One or more graphical representations of the one or more non-structural features are generated at a second fidelity level based on the recognition of the one or more non-structural features, wherein the second fidelity level is lower than the first fidelity level.
2. The method of claim 1 , wherein the plurality of structural features of the physical environment comprises one or more walls, one or more floors, one or more doors, and / or one or more windows; and wherein, The one or more non-structural features of the physical environment include one or more pieces of furniture. 3 . The method of claim 1 , wherein the one or more graphical representations of the one or more non-structural features generated at the second level of fidelity include one or more icons representing the one or more non-structural features.
4. A method according to any one of claims 1 to 2, wherein the one or more graphical representations of the one or more non-structural features include corresponding three-dimensional geometric shapes that outline corresponding areas of the physical environment occupied by the one or more non-structural features of the physical environment in the user interface.
5. The method of any one of claims 1 to 2, wherein the one or more graphical representations of the one or more non-structural features include predefined placeholder furniture.
6. The method of any one of claims 1 to 2, wherein the one or more graphical representations of the one or more non-structural features comprise computer-aided design (CAD) representations of the one or more non-structural features.
7. The method of any one of claims 1 to 2, wherein the one or more graphical representations of the one or more non-structural features are partially transparent.
8. A method according to any one of claims 1 to 2, wherein the one or more non-structural features include one or more building automation devices, and the one or more graphical representations of the one or more non-structural features include graphical indications that the graphical representations correspond to the one or more building automation devices.
9. The method according to claim 8, comprising: In response to receiving input at a respective graphical indication corresponding to a respective building automation device, at least one control for controlling at least one aspect of the respective building automation device is displayed.
10. The method according to any one of claims 1 to 2, wherein: the user interface being a first user interface, the first user interface comprising a first view of the physical environment and a first user interface element, wherein the first view of the physical environment is displayed in response to activation of the first user interface element; the user interface comprising a second user interface element, wherein in response to activation of the second user interface element, a second view of the physical environment different from the first view is displayed; as well as The user interface includes a third user interface element, wherein in response to activation of the third user interface element, a third view of the physical environment that is different from the first view and the second view is displayed.
11. The method according to any one of claims 1 to 2, wherein: the user interface comprising a representation of the field of view of the one or more cameras as the field of view of the one or more cameras moves; the graphical representations of the plurality of structural features being displayed at locations in the representation of the field of view of the one or more cameras that correspond to respective locations of the plurality of structural features in the physical environment; as well as The graphical representations of the one or more non-structural features are displayed at locations in the representation of the field of view of the one or more cameras that correspond to respective locations of the one or more non-structural features in the physical environment.
12. The method according to any one of claims 1 to 2, wherein the user interface comprises: A three-dimensional representation of at least a first portion of the physical environment, wherein the three-dimensional representation of at least the first portion of the physical environment includes three-dimensional representations of the plurality of structural features of the physical environment.
13. The method of any one of claims 1 to 2, wherein the plurality of structural features comprises structural features in the physical environment, and the plurality of non-structural features comprises non-structural features in the physical environment.
14. A computer system comprising: Display generated components; Input devices; One or more cameras in the physical environment; one or more processors; as well as A memory storing one or more programs, wherein the one or more programs are configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 13.
15. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by a computer system comprising a display generating component, an input device, and one or more cameras in a physical environment, cause the computer system to perform the method according to any one of claims 1 to 13.
16. A computer program product comprising a computer program which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and apparatus for presenting location-based content
US20110279445A1
Constructing augmented reality environment with pre-computed lighting
US20140125668A1
Coordinating multiple virtual environments
US20170021273A1
Augmented reality e-commerce for home improvement
US20170132841A1
Method and device for obtaining real time status and controlling of transmitting devices
US20180204385A1