Systems, methods, and graphical user interfaces for scanning and modeling environment
By designing an improved computer system and user interface, the problem of cumbersome, inefficient and high energy consumption in scanning and modeling of enhanced and/or virtual reality environments in the prior art is solved, achieving a more efficient scanning process and a better user experience.
Patent Information
- Application Number
- CN202510124526.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-08
- Filing Date
- 2023-05-09
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems such as cumbersome input, inefficient, limited functionality, lack of feedback and high energy consumption when scanning and modeling the environment using enhanced and/or virtual reality environments.
An improved computer system and user interface is designed to provide a more efficient human-machine interface by reducing user input and provide real-time feedback and progress indications during the scanning process to improve scanning efficiency and user experience.
This enables more efficient scanning and modeling of the environment in enhanced and/or virtual reality environments, reducing energy consumption, especially in battery-driven devices, and extending device usage time.
Smart Images

Figure CN120045105A_ABST
Abstract
Description
[0001] This application is an application with application number 202380048408.3, application date May 9, 2023, and invention name “ Used for System, method and graphical user interface for scanning and modeling an environment " is a divisional application for a patent application.
[0002] Related patent applications
[0003] This application is a continuation of U.S. patent application No. 18 / 144,746 filed on May 8, 2023, which claims priority to U.S. provisional patent application No. 63 / 340,444 filed on May 10, 2022. Technical Field
[0004] The present invention generally relates to computer systems for augmented and / or virtual reality, including but not limited to electronic devices for scanning and modeling an environment (such as a physical environment and / or objects therein) using an augmented and / or virtual reality environment. Background Art
[0005] In recent years, the development of computer systems for augmented and / or virtual reality has increased significantly. Augmented reality environments can be used to annotate and model physical environments and objects therein. Before generating a model of a physical environment, the user needs to use a depth and / or image sensing device to scan the physical environment. Conventional methods for scanning and modeling using augmented and / or virtual reality are cumbersome, inefficient and limited. In some cases, conventional methods for scanning and modeling using augmented reality are functionally limited due to not providing enough feedback and requiring the user to specify what type of features are being scanned. In some cases, conventional methods for scanning using augmented reality do not provide enough guidance to help users scan the environment successfully and efficiently. In some cases, when scanning is in progress, conventional methods for scanning and modeling using augmented reality do not provide enough feedback to the user about the progress, quality and results of the scan. In addition, conventional methods take longer than the required time, thereby wasting energy. This latter consideration is particularly important in battery-driven devices. Summary of the invention
[0006] Therefore, there is a need for a computer system with improved methods and interfaces for scanning and modeling an environment using an augmented and / or virtual reality environment. Such methods and interfaces optionally supplement or replace conventional methods for scanning and modeling an environment using an augmented and / or virtual reality environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from a user and produce a more efficient human-computer interface. For battery-powered devices, such methods and interfaces can save power and increase the time between battery charges.
[0007] The computer system disclosed in the present invention reduces or eliminates the above-mentioned defects and other problems associated with the user interface for enhancement and / or virtual reality. In some embodiments, the computer system includes a desktop computer. In some embodiments, the computer system is portable (e.g., a laptop, a tablet computer, or a handheld device). In some embodiments, the computer system includes a personal electronic device (e.g., a wearable electronic device, such as a watch). In some embodiments, the computer system has a touch pad (and / or communicates with the touch pad). In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display") (and / or communicates with a touch-sensitive display). In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or instruction set stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI in part by touching a stylus and / or a finger and by gestures on a touch-sensitive surface. In some embodiments, in addition to the measurement function based on augmented reality, these functions optionally include playing games, image editing, drawing, presentation, word processing, spreadsheet making, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, note taking and / or digital video playback. Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] According to some embodiments, a method is performed at a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes displaying a first user interface via the display generation component, wherein the first user interface simultaneously includes: a representation of a field of view of one or more cameras, the representation of the field of view including a first view of a physical environment corresponding to a first viewpoint of a user in the physical environment; and a preview of a three-dimensional model of the physical environment. The preview includes a partially completed three-dimensional model of the physical environment, the partially completed three-dimensional model being displayed in a first orientation corresponding to the first viewpoint of the user. The method includes, while displaying the first user interface, detecting a first movement of the one or more cameras in the physical environment, the first movement changing the current viewpoint of the user in the physical environment from the first viewpoint to a second viewpoint. The method also includes, in response to detecting the first movement of the one or more cameras: updating the preview of the three-dimensional model in the first user interface according to the first movement of the one or more cameras, including: adding additional information to the partially completed three-dimensional model; and rotating the partially completed three-dimensional model from the first orientation corresponding to the first viewpoint of the user to the second orientation corresponding to the second viewpoint of the user. The method includes, while displaying the first user interface, detecting a first input to the preview of the three-dimensional model in the first user interface, where the representation of the field of view includes a second view of the physical environment corresponding to the second viewpoint of the user, and the preview of the three-dimensional model includes the partially completed model having the second orientation. The method includes, in response to detecting the first input to the preview of the three-dimensional model in the first user interface: updating the preview of the three-dimensional model in the first user interface based on the first input, including: based on determining that the first input satisfies a first criterion, rotating the partially completed three-dimensional model from the second orientation corresponding to the second viewpoint of the user to a third orientation that does not correspond to the second viewpoint of the user.
[0009] According to some embodiments, a method is performed at a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes displaying a first user interface via the display generation component. The first user interface includes a representation of a field of view of one or more cameras, and the representation of the field of view includes a corresponding view of a physical environment corresponding to a current viewpoint of a user in the physical environment. The method includes, when displaying the first user interface, based on determining that a first object has been detected in the field of view of the one or more cameras: at a first time, displaying a first representation of the first object at a location corresponding to the position of the first object in the physical environment in the representation of the field of view. One or more spatial attributes of the first representation of the first object have values corresponding to one or more spatial dimensions of the first object in the physical environment. The method includes replacing the display of the first representation of the first object in the representation of the field of view with a display of a second representation of the first object at a second time later than the first time. The second representation of the first object does not spatially indicate the one or more spatial dimensions of the first object in the physical environment.
[0010] According to some embodiments, a method is performed at a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes during scanning of a physical environment to obtain depth information of at least a portion of the physical environment: displaying a first user interface via the display generation component. The first user interface includes a representation of a field of view of one or more cameras, and the representation of the field of view includes a corresponding view of the physical environment corresponding to a current viewpoint of a user in the physical environment. The method includes detecting movement of the one or more cameras in the physical environment while displaying the first user interface, the detection including: detecting a first movement that changes the current viewpoint of the user from a first viewpoint in the physical environment to a second viewpoint in the physical environment. The method also includes, in response to detecting the movement of the one or more cameras in the physical environment, the movement including the first movement changing the current viewpoint of the user from the first viewpoint in the physical environment to the second viewpoint in the physical environment, displaying in the first user interface a first visual indication of the representation covering the field of view of the one or more cameras based on determining that there is a corresponding portion of the physical environment that has not yet been scanned and is located between the first portion of the physical environment that has been scanned and the second portion of the physical environment that has been scanned, wherein the first visual indication indicates a position of the corresponding portion of the physical environment in the field of view of the one or more cameras, and the corresponding portion of the physical environment is not visible in the representation of the field of view of the one or more cameras.
[0011] According to some embodiments, a method is performed at a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes displaying a first user interface via the display generation component during scanning of a physical environment to obtain depth information of at least a portion of the physical environment, wherein the first user interface includes a representation of a field of view of one or more cameras. The method includes displaying a plurality of graphical objects covering the representation of the field of view of the one or more cameras, the display including: displaying at least a first graphical object at a first position and a second graphical object at a second position, the first graphical object representing one or more estimated spatial attributes of a first physical feature that has been detected in a corresponding portion of the physical environment in the field of view of the one or more cameras, and the second graphical object representing one or more estimated spatial attributes of a second physical feature that has been detected in the corresponding portion of the physical environment in the field of view of the one or more cameras. The method includes, when displaying the plurality of graphical objects covering the representation of the field of view of the one or more cameras, changing one or more visual attributes of the first graphical object according to a change in the corresponding prediction accuracy of the estimated spatial attributes of the first physical feature, and changing one or more visual attributes of the second graphical object according to a change in the corresponding prediction accuracy of the estimated spatial attributes of the second physical feature.
[0012] According to some embodiments, a computer system includes a display generation component (also referred to as a display device, e.g., a display, a projector, a head-mounted display, a head-up display, etc.), one or more cameras (e.g., a camera that continuously or repeatedly at fixed intervals provides a real-time preview of at least a portion of the content within the camera's field of view and optionally generates a video output comprising one or more image frame streams capturing the content within the camera's field of view), and one or more input devices (e.g., a touch-sensitive surface, such as a touch-sensitive remote control, or a touch screen display that also serves as a display generation component, a mouse, a joystick, a wand controller, and / or one or more cameras that track one or more features of a user such as the positioning of the user's hands), optionally one or more posture sensors, optionally one or more depth sensors, optionally one or more sensors that detect the intensity of contact with the touch-sensitive surface, optionally one or more tactile output generators, one or more processors, and a memory that stores one or more programs (and / or communicates with these components); the one or more programs are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the performance of operations of any of the methods described herein. According to some embodiments, a computer-readable storage medium has stored therein instructions that, when executed by a computer system comprising a display generation component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting the intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators (and / or communicating with these components), cause the computer system to perform the operations of any of the methods described herein or cause the operations of any of the methods described herein to be performed. According to some embodiments, a graphical user interface on a computer system comprising a display generation component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting the intensity of contact with a touch-sensitive surface, optionally one or more tactile output generators, a memory, and one or more processors for executing one or more programs stored in the memory (and / or communicating with these components) includes one or more elements displayed in any of the methods described herein, and the one or more elements are updated in response to input, as described in any of the methods described herein. According to some embodiments, a computer system includes (and / or communicates with) a display generating component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, optionally one or more tactile output generators, and components for performing or causing the operations of any of the methods described herein to be performed.According to some embodiments, an information processing device for use in a computer system that includes a display generating component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators (and / or communicates with these components) includes components for performing the operations of any of the methods described herein or causing the operations of any of the methods described herein to be performed.
[0013] Thus, a computer system having a display generation component, one or more cameras, one or more input devices, optionally one or more gesture sensors, optionally one or more sensors for detecting intensity of contact with a touch-sensitive surface, and optionally one or more tactile output generators (and / or in communication therewith) has improved methods and interfaces for annotating, measuring, and modeling an environment (such as a physical environment and / or objects therein) using an augmented and / or virtual reality environment, thereby increasing the effectiveness, efficiency, and user satisfaction of such computer systems. Such methods and interfaces may supplement or replace conventional methods for annotating, measuring, and modeling an environment (such as a physical environment and / or objects therein) using an augmented and / or virtual reality environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.
[0015] Figure 1A is a block diagram illustrating a portable multifunction device with a touch-sensitive display according to some embodiments.
[0016] Figure 1B is a block diagram illustrating example components for event handling according to some embodiments.
[0017] Figure 2A A portable multifunction device with a touch screen according to some embodiments is shown.
[0018] Figure 2B A portable multifunction device with an optical sensor and a time-of-flight sensor is shown according to some embodiments.
[0019] Figure 3A is a block diagram of an example multifunction device with a display and a touch-sensitive surface according to some embodiments.
[0020] FIG. 3B to FIG. 3C is a block diagram of an example computer system according to some embodiments.
[0021] Figure 4AAn example user interface for presenting a menu of applications on a portable multifunction device according to some embodiments is shown.
[0022] Figure 4B An example user interface is shown for a multifunction device having a touch-sensitive surface that is separate from the display according to some embodiments.
[0023] FIG. 5A to FIG. 5AD An example user interface for scanning and modeling an environment and interacting with a generated schematic thereof is shown according to some embodiments.
[0024] 6A to 6F is a flow chart of a method of displaying a preview of a three-dimensional model of an environment during scanning and modeling of the environment, according to some embodiments.
[0025] FIG. 7A to FIG. 7D is a flowchart of a method for displaying representations of objects identified in an environment during scanning and modeling of the environment, according to some embodiments.
[0026] FIG. 8A to FIG. 8D is a flow chart of a method of providing guidance indicating locations of missing portions of a presumed completed portion of an environment during scanning and modeling of the environment, according to some embodiments.
[0027] 9A to 9E is a flow chart of a method of displaying a scan progress indication during scanning and modeling of an environment according to some embodiments. DETAILED DESCRIPTION
[0028] As described above, an augmented reality environment can be used to facilitate scanning and modeling of a physical environment and objects therein by providing different views of the physical environment and objects therein and guiding a user to move through the physical environment to capture the data necessary to generate a model of the physical environment. Conventional methods for scanning and modeling using augmented and / or virtual reality environments are generally limited in functionality. In some cases, conventional methods for scanning and modeling a physical environment using augmented reality do not provide a preview of a three-dimensional model generated based on the scan until the scan is fully completed. In some cases, conventional methods for scanning and modeling a physical environment using augmented reality display a three-dimensional representation of the physical environment during the scan of the physical environment, but do not allow the user to manipulate or view the three-dimensional representation from different angles during the scan of the physical environment. In some cases, conventional methods for scanning and modeling a physical environment do not simultaneously scan and model the structural and non-structural elements of the physical environment during the same scan, and do not display annotations based on the identification of structural elements and non-structural elements in the augmented reality environment and the preview of the three-dimensional model of the physical environment. The embodiments disclosed herein provide a user with an intuitive way to scan and model an environment using augmented and / or virtual reality (e.g., by providing more intelligent and complex functionality, by enabling the user to perform different operations in the augmented reality environment with less input, and / or by simplifying the user interface). In addition, the embodiments disclosed herein provide improved feedback that provides the user with additional information about the physical objects being scanned or modeled and about the operations performed in the virtual / augmented reality environment.
[0029] The systems, methods, and GUIs described herein improve user interface interactions with augmented and / or virtual reality environments in a variety of ways. For example, they make it easier to scan and model a physical environment by providing automatic detection of features in a physical space and annotating different types of detected features, improved guidance, ..., by providing improved feedback to the user about the progress of the modeling process while the environment is being modeled.
[0030] under, Figure 1A to Figure 1B , FIG. 2A to FIG. 2B as well as FIG. 3A to FIG. 3C A description of an example device is provided. FIG. 4A to FIG. 4B and FIG. 5A to FIG. 5AD Example user interfaces for interacting with, annotating, scanning, and modeling environments, such as augmented reality environments, are shown. 6A to 6F is a flow chart of a method of displaying a preview of a three-dimensional model of an environment during scanning and modeling of the environment, according to some embodiments. FIG. 7A to FIG. 7D is a flowchart of a method for displaying representations of objects identified in an environment during scanning and modeling of the environment, according to some embodiments. FIG. 8A to FIG. 8Dis a flow chart of a method of providing guidance indicating locations of missing portions of a presumed completed portion of an environment during scanning and modeling of the environment, according to some embodiments. 9A to 9E is a flow chart of a method of displaying a scan progress indication during scanning and modeling of an environment according to some embodiments. FIG. 5A to FIG. 5AD The user interface in 6A to 6F , FIG. 7A to FIG. 7D , FIG. 8A to FIG. 8D as well as 9A to 9E in the process.
[0031] Example Device
[0032] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. Many specific details are shown in the following detailed description in order to provide a full understanding of the various described embodiments. However, it will be apparent to one of ordinary skill in the art that the various described embodiments can be practiced without these specific details. In other cases, well-known methods, processes, components, circuits, and networks are not described in detail, so as not to unnecessarily obscure various aspects of the embodiments.
[0033] It will also be understood that, although in some cases, the terms "first", "second", etc. are used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, the first contact can be named as the second contact, and similarly, the second contact can be named as the first contact without departing from the scope of the various described embodiments. The first contact and the second contact are both contacts, but they are not the same contact unless the context clearly indicates otherwise.
[0034] The terms used in the description of various described embodiments herein are only for the purpose of describing specific embodiments, and are not intended to be limiting. As used in the description of various described embodiments and in the appended claims, the singular forms "one" and "the" are intended to also include plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more items in the associated listed items. It will also be understood that the terms "including" and / or "comprising" when used in this specification specify the presence of stated features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or their grouping.
[0035] As used herein, the term "if" is optionally interpreted to mean "when" followed by "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined that" or "if [a stated condition or event] is detected" are optionally interpreted to mean "upon determining that" or "in response to determining that" or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.
[0036] Computer systems for augmented and / or virtual reality include electronic devices that generate augmented and / or virtual reality environments. Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as a PDA and / or music player functions. Example embodiments of portable multifunction devices include, but are not limited to, the Apple iPod from Apple Inc., Cupertino, California. iPod and Device. Other portable electronic devices, such as laptop computers or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or trackpads) are optionally used. It should also be understood that in some embodiments, the device is not a portable communication device, but rather a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or trackpads) that also includes or communicates with one or more cameras.
[0037] In the following discussion, a computer system is described that includes an electronic device having a display and a touch-sensitive surface (and / or communicating with these components). However, it should be understood that the computer system may optionally include one or more other physical user interface devices, such as a physical keyboard, a mouse, a joystick, a stylus controller, and / or a camera that tracks one or more features of a user, such as the position of the user's hands.
[0038] The device typically supports a variety of applications, such as one or more of the following: a gaming application, a note-taking application, a drawing application, a presentation application, a word processing application, a spreadsheet application, a telephony application, a video conferencing application, an email application, an instant messaging application, a fitness support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0039] Various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed by the device are optionally adjusted and / or varied for different applications, and / or adjusted and / or varied within the respective applications. In this way, the common physical architecture of the device (such as a touch-sensitive surface) optionally supports various applications with a user interface that is intuitive and clear to the user.
[0040] Attention is now turned to embodiments of portable devices having touch-sensitive displays. Figure 1A 1 is a block diagram illustrating a portable multifunction device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display system 112 is sometimes called a "touch screen" for convenience, and is sometimes simply referred to as a touch-sensitive display. The device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input or control devices 116, and an external port 124. The device 100 optionally includes one or more optical sensors 164 (e.g., as part of one or more cameras). The device 100 optionally includes one or more intensity sensors 165 for detecting the intensity of contact on the device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100). Device 100 optionally includes one or more tactile output generators 163 for generating tactile output on device 100 (e.g., generating tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touch pad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0041] As used in this specification and claims, the term "tactile output" refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, which will be detected by a user using the user's sense of touch. For example, in the case where a device or a component of the device is in contact with a user's touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or a touchpad) is optionally interpreted by the user as a "press click" or "release click" to a physical actuation button. In some cases, the user will feel a tactile sensation, such as a "press click" or "release click", even when the physical actuation button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when the smoothness of the touch-sensitive surface does not change, the movement of the touch-sensitive surface may be optionally interpreted or sensed by the user as the "roughness" of the touch-sensitive surface. Although such interpretation of touch by the user will be limited by the user's individualized sensory perception, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of a user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception of a typical (or ordinary) user. Providing tactile feedback to the user using tactile output enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0042] It should be understood that device 100 is merely one example of a portable multifunction device, and that device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. Figure 1A The various components shown are implemented in hardware, software, firmware, or any combination thereof, including one or more signal processing circuits and / or application specific integrated circuits.
[0043] The memory 102 optionally includes high-speed random access memory, and optionally also includes non-volatile memory, such as one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices. Access to the memory 102 by other components of the device 100 (such as the CPU 120 and the peripheral device interface 118) is optionally controlled by a memory controller 122.
[0044] Peripheral interface 118 may be used to couple the device's input and output peripherals to CPU 120 and memory 102. One or more processors 120 run or execute various software programs and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data.
[0045] In some embodiments, peripherals interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0046] RF (radio frequency) circuit 108 receives and sends RF signals, also referred to as electromagnetic signals. RF circuit 108 converts electrical signals into / from electromagnetic signals and communicates with a communication network and other communication devices via electromagnetic signals. RF circuit 108 optionally includes well-known circuits for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, memory, and the like. RF circuit 108 optionally communicates with networks and other devices via wireless communications, such as the Internet (also referred to as the World Wide Web (WWW)), an intranet, and / or a wireless network (such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN)). The wireless communication optionally uses any of a variety of communication standards, protocols and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), Evolution-Data-only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11d), IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11d, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Utilizing Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols that have not been developed on the date of filing of this document.
[0047] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between a user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts the audio data into electrical signals, and transmits the electrical signals to the speaker 111. The speaker 111 converts the electrical signals into sound waves audible to humans. The audio circuit 110 also receives electrical signals converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signals into audio data and transmits the audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., Figure 2A The headset jack provides an interface between the audio circuit 110 and a removable audio input / output peripheral device, such as an output-only headset or a headset having both output (e.g., a single or dual-ear headset) and input (e.g., a microphone).
[0048] The I / O subsystem 106 couples input / output peripherals on the device 100, such as a touch-sensitive display system 112 and other input or control devices 116, to a peripheral device interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, an intensity sensor controller 159, a tactile feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive / send electrical signals from / to other input or control devices 116. Other input or control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some alternative embodiments, the one or more input controllers 160 are optionally coupled to any (or none) of the following: a keyboard, an infrared port, a USB port, a stylus, and / or a pointer device such as a mouse. One or more buttons (e.g., Figure 2A 208) optionally includes an up / down button for volume control of the speaker 111 and / or the microphone 113. The one or more buttons optionally include a push button (e.g., Figure 2A 206).
[0049] The touch-sensitive display system 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch-sensitive display system 112 and / or sends electrical signals to the touch-sensitive display system. The touch-sensitive display system 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some visual outputs or all of the visual outputs correspond to user interface objects. As used herein, the term "indicator" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object configured to respond to input directed to a graphical user interface object). Examples of user-interactive graphical user interface objects include, but are not limited to, buttons, sliders, icons, selectable menu items, switches, hyperlinks, or other user interface controls.
[0050] The touch-sensitive display system 112 has a touch-sensitive surface, sensor, or sensor group that accepts input from a user based on tactile and / or haptic contact. The touch-sensitive display system 112 and display controller 156 (together with any associated modules and / or instruction sets in memory 102) detect contact (and any movement or interruption of that contact) on the touch-sensitive display system 112 and convert the detected contact into interaction with a user interface object (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch-sensitive display system 112. In some embodiments, the point of contact between the touch-sensitive display system 112 and the user corresponds to the user's finger or stylus.
[0051] The touch-sensitive display system 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies are used in other embodiments. The touch-sensitive display system 112 and display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch-sensitive display system 112 to detect contact and any movement or interruption thereof. In some embodiments, projected mutual capacitance sensing technology is used, such as the . iPod and Technology found in.
[0052] The touch-sensitive display system 112 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen video resolution exceeds 400 dpi (e.g., 500 dpi, 800 dpi or greater). The user optionally uses any suitable object or attachment such as a stylus, finger, etc. to contact the touch-sensitive display system 112. In some embodiments, the user interface is designed to work with finger-based contacts and gestures, which may not be as accurate as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts rough finger-based input into precise pointer / cursor positioning or commands for performing the actions desired by the user.
[0053] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad is optionally a touch-sensitive surface that is separate from the touch-sensitive display system 112, or an extension of the touch-sensitive surface formed by the touch screen.
[0054] The device 100 also includes a power system 162 for powering the various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., a light emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0055] Device 100 optionally also includes one or more optical sensors 164 (eg, as part of one or more cameras). Figure 1A An optical sensor coupled to an optical sensor controller 158 in the I / O subsystem 106 is shown. One or more optical sensors 164 optionally include a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. One or more optical sensors 164 receive light projected through one or more lenses from the environment and convert the light into data representing an image. In conjunction with an imaging module 143 (also called a camera module), one or more optical sensors 164 optionally capture static images and / or videos. In some embodiments, an optical sensor is located on the rear portion of the device 100 opposite to the touch-sensitive display system 112 on the front of the device, so that the touch screen can be used as a viewfinder for static image and / or video image acquisition. In some embodiments, another optical sensor is located on the front of the device to obtain an image of the user (e.g., for self-portraits, for video conferencing when the user is watching other video conference participants on the touch screen, etc.).
[0056] Device 100 optionally also includes one or more contact intensity sensors 165 . Figure 1A A contact force sensor coupled to a force sensor controller 159 in the I / O subsystem 106 is shown. One or more contact force sensors 165 optionally include one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other force sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). One or more contact force sensors 165 receive contact force information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact force sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact force sensor is located on the rear of the device 100 opposite to the touch-sensitive display system 112 located on the front of the device 100.
[0057] Device 100 optionally also includes one or more proximity sensors 166 . Figure 1A A proximity sensor 166 is shown coupled to the peripherals interface 118. Alternatively, the proximity sensor 166 is coupled to the input controller 160 in the I / O subsystem 106. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is on a phone call), the proximity sensor turns off and disables the touch-sensitive display system 112.
[0058] Device 100 optionally also includes one or more tactile output generators 163. Figure 1AA tactile output generator coupled to a tactile feedback controller 161 in the I / O subsystem 106 is shown. In some embodiments, one or more tactile output generators 163 include one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile outputs on the device). One or more tactile output generators 163 receive tactile feedback generation instructions from the tactile feedback module 133 and generate tactile outputs on the device 100 that can be felt by a user of the device 100. In some embodiments, at least one tactile output generator is arranged in juxtaposition or proximity to a touch-sensitive surface (e.g., a touch-sensitive display system 112), and optionally generates tactile outputs by moving the touch-sensitive surface vertically (e.g., inward / outward of the surface of the device 100) or laterally (e.g., backward and forward in the same plane as the surface of the device 100). In some embodiments, at least one tactile output generator sensor is located on a rear portion of the device 100 opposite the touch-sensitive display system 112 located on the front of the device 100.
[0059] The device 100 optionally also includes one or more accelerometers 167, gyroscopes 168, and / or magnetometers 169 (e.g., as part of an inertial measurement unit (IMU)) for obtaining information about the device's posture (e.g., position and orientation or attitude). Figure 1A Sensors 167, 168, and 169 are shown coupled to peripheral device interface 118. Alternatively, sensors 167, 168, and 169 are optionally coupled to input controller 160 in I / O subsystem 106. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on analysis of data received from the one or more accelerometers. Device 100 optionally includes a GPS (or GLONASS or other global navigation system) receiver for obtaining information about the location of device 100.
[0060] In some embodiments, the software components stored in the memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a tactile feedback module (or instruction set) 133, a text input module (or instruction set) 134, a global positioning system (GPS) module (or instruction set) 135, and an application (or instruction set) 136. In addition, in some embodiments, the memory 102 stores a device / global internal state 157, as shown in Figures 1A and 3. The device / global internal state 157 includes one or more of the following: an active application state, which indicates which applications (if any) are currently active; a display state, which indicates what applications, views, or other information occupy various areas of the touch-sensitive display system 112; a sensor state, which includes information obtained from various sensors and other input or control devices 116 of the device; and position and / or location information about the posture (e.g., position and / or posture) of the device.
[0061] The operating system 126 (e.g., iOS, Android, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between various hardware and software components.
[0062] The communication module 128 facilitates communication with other devices through one or more external ports 124, and also includes various software components for processing data received by the RF circuit 108 and / or the external port 124. The external port 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) is suitable for coupling directly to other devices, or indirectly through a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a device that is compatible with some of Apple Inc. (Cupertino, California). iPod and In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or compatible with the 30-pin connector used in Apple Inc. (Cupertino, California). iPod and In some embodiments, the external port is a USB Type-C connector that is the same as, similar to, and / or compatible with the Lightning connector used in Apple Inc. (Cupertino, California) electronic devices.
[0063] The contact / motion module 130 optionally detects contact with the touch-sensitive display system 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection (e.g., by a finger or stylus), such as determining whether contact has occurred (e.g., detecting a finger press event), determining the strength of the contact (e.g., the force or pressure of the contact, or a substitute for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact disconnection). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of a contact point optionally includes determining a rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to a single point of contact (e.g., a single finger contact or stylus contact) or multiple points of simultaneous contact (e.g., "multi-touch" / multi-finger contact). In some embodiments, contact / motion module 130 and display controller 156 detect contact on a touch pad.
[0064] The contact / motion module 130 optionally detects gesture input by the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of the detected contacts). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a single-finger tap gesture includes detecting a finger press event, and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at an icon location). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and then detecting a finger lift (lift-off) event. Similarly, taps, swipes, drags, and other gestures of the stylus are optionally detected by detecting a specific contact pattern of the stylus.
[0065] In some embodiments, detecting a finger tap gesture depends on detecting the length of time between a finger press event and a finger lift event, but is independent of the strength of the finger contact between the finger press event and the finger lift event. In some embodiments, based on determining that the length of time between the finger press event and the finger lift event is less than a predetermined value (e.g., less than 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.4 seconds, or 0.5 seconds), a tap gesture is detected regardless of whether the strength of the finger contact during the tap reaches a given strength threshold (greater than a nominal contact detection strength threshold), such as a light press or deep press strength threshold. Thus, a finger tap gesture can meet a specific input criterion that does not require the characteristic strength of the contact to meet a given strength threshold to meet the specific input criterion. For clarity, the finger contact in a tap gesture is typically required to meet a nominal contact detection strength threshold to detect a finger press event, below which the contact is not detected. A similar analysis applies to detecting a tap gesture by a stylus or other contact. In the case where the device is capable of detecting contact with a finger or stylus hovering over the touch-sensitive surface, the nominal contact detection intensity threshold optionally does not correspond to physical contact between the finger or stylus and the touch-sensitive surface.
[0066] The same concepts apply in a similar manner to other types of gestures. For example, a swipe gesture, a pinch gesture, an expand gesture, and / or a long press gesture may be optionally detected based on satisfying criteria that are independent of the strength of the contacts included in the gesture or that do not require one or more contacts performing the gesture to reach a strength threshold in order to be recognized. For example, a swipe gesture is detected based on the amount of movement of one or more contacts; a pinch gesture is detected based on the movement of two or more contacts toward each other; an expand gesture is detected based on the movement of two or more contacts away from each other; and a long press gesture is detected based on the duration of a contact on a touch-sensitive surface having less than a threshold amount of movement. Thus, the statement that a particular gesture recognition criterion does not require the strength of the contacts to satisfy a corresponding strength threshold in order to satisfy the particular gesture recognition criterion means that the particular gesture recognition criterion can be satisfied when the contacts in the gesture do not reach the corresponding strength threshold, and can also be satisfied when one or more contacts in the gesture reach or exceed the corresponding strength threshold. In some embodiments, a tap gesture is detected based on determining that a finger down event and a finger up event are detected within a predefined time period, regardless of whether the contact is above or below a corresponding strength threshold during the predefined time period, and a swipe gesture is detected based on determining that the contact moves greater than a predefined amount, even if the contact is above the corresponding strength threshold at the end of the contact movement. Even in specific implementations where detection of gestures is affected by the strength of the contact performing the gesture (e.g., the device detects a long press more quickly when the strength of the contact is above the strength threshold, or the device delays detection of a tap input when the strength of the contact is higher), detection of these gestures does not require the contact to reach a particular strength threshold (e.g., even if the amount of time required to recognize the gesture varies), as long as the criteria for recognizing the gesture can be met without the contact reaching the particular strength threshold.
[0067] In some cases, the contact intensity threshold, duration threshold, and movement threshold are combined in various combinations to create a heuristic algorithm to distinguish between two or more different gestures for the same input element or region, so that multiple different interactions with the same input element can provide a richer set of user interactions and responses. The statement that a particular set of gesture recognition criteria does not require the intensity of one or more contacts to meet the corresponding intensity threshold in order to meet the particular gesture recognition criteria does not preclude the simultaneous evaluation of other intensity-related gesture recognition criteria to identify other gestures with criteria that are met when the gesture includes a contact with an intensity above the corresponding intensity threshold. For example, in some cases, a first gesture recognition criterion for a first gesture (which does not require the intensity of the contact to meet the corresponding intensity threshold to meet the first gesture recognition criterion) competes with a second gesture recognition criterion for a second gesture (which depends on the contact reaching the corresponding intensity threshold). In such a competition, if the second gesture recognition criterion for the second gesture is met first, the gesture is optionally not recognized as satisfying the first gesture recognition criterion for the first gesture. For example, if the contact reaches the corresponding intensity threshold before the contact moves a predefined amount of movement, a deep press gesture is detected instead of a swipe gesture. Conversely, if the contact moves the predefined amount of movement before the contact reaches the corresponding intensity threshold, a swipe gesture is detected instead of a deep press gesture. Even in such cases, the first gesture recognition criteria for the first gesture still does not require the intensity of the contact to satisfy the corresponding intensity threshold to satisfy the first gesture recognition criteria because if the contact remains below the corresponding intensity threshold until the gesture ends (e.g., a swipe gesture with an intensity of the contact that does not increase above the corresponding intensity threshold), the gesture will be recognized as a swipe gesture by the first gesture recognition criteria. Thus, a particular gesture recognition criterion that does not require the intensity of the contact to satisfy the corresponding intensity threshold to satisfy the particular gesture recognition criterion will (A) in some cases ignore the intensity of the contact relative to the intensity threshold (e.g., for a tap gesture) and / or (B) in some cases fail to satisfy the particular gesture recognition criterion (e.g., for a long press gesture) if a set of competing intensity-related gesture recognition criteria (e.g., for a deep press gesture) recognizes the input as corresponding to an intensity-related gesture before the particular gesture recognition criterion recognizes the gesture corresponding to the input, and in this sense still depends on the intensity of the contact relative to the intensity threshold (e.g., for a long press gesture competing with the deep press gesture for recognition).
[0068] In conjunction with accelerometer 167, gyroscope 168, and / or magnetometer 169, posture module 131 optionally detects posture information about the device, such as the posture of the device in a particular reference frame (e.g., roll, pitch, yaw, and / or orientation). Posture module 131 includes software components for performing various operations related to detecting device orientation and detecting changes in device posture.
[0069] Graphics module 132 includes various known software components for rendering and displaying graphics on touch-sensitive display system 112 or other displays, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual attributes) of displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, without limitation, text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0070] In some embodiments, the graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes for specifying the graphics to be displayed from an application or the like, along with coordinate data and other graphic attribute data if necessary, and then generates screen image data to be output to the display controller 156.
[0071] The tactile feedback module 133 includes various software components for generating instructions (e.g., instructions used by the tactile feedback controller 161) to produce tactile output at one or more locations on the device 100 using one or more tactile output generators 163 in response to user interaction with the device 100.
[0072] Text input module 134, which is optionally a component of graphics module 132, provides a soft keyboard for entering text in various applications (eg, contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).
[0073] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for location-based dialing; to the camera 143 as picture / video metadata; and to applications that provide location-based services such as a weather widget, local yellow pages widget, and a map / navigation widget).
[0074] The virtual / augmented reality module 145 provides virtual and / or augmented reality logic components to the application 136 that implements augmented reality features, and in some embodiments, virtual reality features. The virtual / augmented reality module 145 facilitates the overlay of virtual content, such as virtual user interface objects, on a representation of at least a portion of a field of view of one or more cameras. For example, with the assistance of the virtual / augmented reality module 145, the representation of at least a portion of a field of view of one or more cameras may include corresponding physical objects, and the virtual user interface objects may be displayed in a displayed augmented reality environment at a location determined based on the corresponding physical objects in the field of view of the one or more cameras, or in a virtual reality environment determined based on a posture of at least a portion of a computer system (e.g., a posture of a display device used to display a user interface to a user of the computer system).
[0075] Application 136 optionally includes the following modules (or instruction sets) or a subset or superset thereof:
[0076] ● Contacts module 137 (sometimes called address book or contact list);
[0077] ● Telephone module 138;
[0078] ● Video conferencing module 139;
[0079] ● Email client module 140;
[0080] ●Instant messaging (IM) module 141;
[0081] ●Fitness support module 142;
[0082] ● Camera module 143 for still images and / or video images;
[0083] ● Image management module 144;
[0084] ●Browser module 147;
[0085] ● Calendar module 148;
[0086] A widget module 149, which optionally includes one or more of the following: a weather widget 149-1, a stock market widget 149-2, a calculator widget 149-3, an alarm widget 149-4, a dictionary widget 149-5, and other widgets acquired by the user, and a user-created widget 149-6;
[0087] A widget creator module 150 for forming a user-created widget 149-6;
[0088] ●Search module 151;
[0089] ● A video and music player module 152, optionally consisting of a video player module and a music player module;
[0090] ●Note module 153;
[0091] ● Map module 154; and / or
[0092] ●Online video module 155;
[0093] ● Annotation and Modeling Module 195; and / or
[0094] ●Time of Flight (“ToF”) sensor module 196 .
[0095] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice replication.
[0096] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, the contact module 137 includes executable instructions for managing an address book or contact list (e.g., stored in the application internal state 192 of the contact module 137 in memory 102 or memory 370), including: adding names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers and / or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing 139, email 140, or IM 141; etc.
[0097] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, phone module 138 includes executable instructions for entering a character sequence corresponding to a phone number, accessing one or more phone numbers in address book 137, modifying an entered phone number, dialing a corresponding phone number, conducting a conversation, and disconnecting or hanging up when the conversation is complete. As described above, wireless communications optionally use any of a variety of communication standards, protocols, and technologies.
[0098] In combination with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, one or more optical sensors 164, optical sensor controller 158, contact module 130, graphics module 132, text input module 134, contact list 137, and phone module 138, video conferencing module 139 includes executable instructions for initiating, conducting, and terminating a video conference between a user and one or more other participants in accordance with user instructions.
[0099] In conjunction with RF circuit 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. In conjunction with image management module 144, email client module 140 makes it very easy to create and send emails with still images or video images captured by camera module 143.
[0100] In conjunction with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, and the text input module 134, the instant messaging module 141 includes executable instructions for performing the following operations: inputting a character sequence corresponding to an instant message, modifying previously input characters, transmitting the corresponding instant message (e.g., using a short message service (SMS) or multimedia message service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, Apple Push Notification Service (APNs) or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or enhanced messaging services (EMS). As used herein, "instant messaging" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, APNs or IMPS).
[0101] In combination with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and video and music player module 152, fitness support module 142 includes executable instructions for creating workouts (e.g., with time, distance, and / or calorie burn goals); communicating with fitness sensors (in sports equipment and smart watches); receiving fitness sensor data; calibrating sensors for monitoring fitness; selecting and playing music for a workout; and displaying, storing, and transmitting fitness data.
[0102] In conjunction with touch-sensitive display system 112, display controller 156, one or more optical sensors 164, optical sensor controller 158, contact module 130, graphics module 132, and image management module 144, camera module 143 includes executable instructions for capturing still images or videos (including video streams) and storing them in memory 102, modifying characteristics of still images or videos, and / or deleting still images or videos from memory 102.
[0103] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and camera module 143, image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, displaying (e.g., in a digital slide show or album), and storing still images and / or video images.
[0104] In combination with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132 and the text input module 134, the browser module 147 includes executable instructions for browsing the Internet in accordance with user instructions (including searching, linking to, receiving and displaying web pages or portions thereof, as well as attachments and other files linked to web pages).
[0105] In combination with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, the email client module 140 and the browser module 147, the calendar module 148 includes executable instructions for creating, displaying, modifying and storing calendars and data associated with the calendar (e.g., calendar entries, to-do items, etc.) in accordance with user instructions.
[0106] In conjunction with RF circuit 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is a mini-application that is optionally downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm widget 149-4, and dictionary widget 149-5) or a mini-application created by a user (e.g., user-created widget 149-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).
[0107] In conjunction with the RF circuit 108, the touch-sensitive display system 112, the display controller 156, the contact module 130, the graphics module 132, the text input module 134, and the browser module 147, the widget creator module 150 includes executable instructions for creating a widget (e.g., transferring a user-specified portion of a web page into a widget).
[0108] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132 and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.
[0109] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats (such as, MP3 or AAC files), as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch-sensitive display system 112, or on an external display connected wirelessly or via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (trademark of Apple Inc.).
[0110] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, the note module 153 includes executable instructions for creating and managing notes, to-do lists, etc. according to user instructions.
[0111] In combination with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 includes executable instructions for receiving, displaying, modifying, and storing maps and data associated with the maps (e.g., driving directions; data about stores and other points of interest at or near a particular location; and other location-based data) in accordance with user instructions.
[0112] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, email client module 140, and browser module 147, online video module 155 includes executable instructions that allow a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on touch screen 112, or on an external display connected wirelessly or via external port 124), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, instant messaging module 141 is used instead of email client module 140 to send a link to a specific online video.
[0113] In conjunction with touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, camera module 143, image management module 152, video and music player module 152, and virtual / augmented reality module 145, annotation and modeling module 195 includes executable instructions that allow a user to model a physical environment and / or physical objects therein and annotate (e.g., measure, draw, and / or add virtual objects and manipulate virtual objects therein) representations (e.g., in real time or previously captured) of the physical environment and / or physical objects therein in an augmented and / or virtual reality environment, as described in more detail herein.
[0114] In conjunction with the camera module 143, the ToF sensor module 196 includes executable instructions for capturing depth information of the physical environment. In some embodiments, the ToF sensor module 196 operates in conjunction with the camera module 143 to provide depth information of the physical environment.
[0115] Each module and application identified above corresponds to a set of executable instructions for performing one or more functions and methods described in the present application (e.g., computer-implemented methods and other information processing methods described herein). These modules (i.e., instruction sets) do not have to be implemented with independent software programs, processes or modules, so various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 102 optionally stores a subset of the modules and data structures identified above. In addition, memory 102 optionally stores other modules and data structures not described above.
[0116] In some embodiments, the device 100 is a device on which operation of a predefined set of functions is performed exclusively through a touch screen and / or a touch pad. By using a touch screen and / or a touch pad as the primary input control device for operating the device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on the device 100 is optionally reduced.
[0117] A predefined set of functions that are performed exclusively through the touch screen and / or trackpad optionally includes navigation between user interfaces. In some embodiments, the trackpad, when touched by the user, navigates the device 100 to a main menu, a home menu, or a root menu from any user interface displayed on the device 100. In such embodiments, a touch-sensitive surface is used to implement a "menu button." In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touch-sensitive surface.
[0118] Figure 1B is a block diagram illustrating example components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3A ) includes an event classifier 170 (e.g., in the operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 136, 137 to 155, 380 to 390).
[0119] The event classifier 170 receives event information and determines the application 136-1 and the application view 191 of the application 136-1 to which the event information is to be delivered. The event classifier 170 includes an event monitor 171 and an event distributor module 174. In some embodiments, the application 136-1 includes an application internal state 192 that indicates one or more current application views displayed on the touch-sensitive display system 112 when the application is active or executing. In some embodiments, the device / global internal state 157 is used by the event classifier 170 to determine which application(s) is currently active, and the application internal state 192 is used by the event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0120] In some embodiments, the application internal state 192 includes additional information, such as one or more of the following: resumption information to be used when application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by application 136-1, a state queue for enabling a user to return to a previous state or view of application 136-1, and a repeat / undo queue of previous actions taken by the user.
[0121] Event monitor 171 receives event information from peripherals interface 118. Event information includes information about sub-events (e.g., a user touch on touch-sensitive display system 112 as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 167, and / or microphone 113 (through audio circuit 110). The information received by peripherals interface 118 from I / O subsystem 106 includes information from touch-sensitive display system 112 or a touch-sensitive surface.
[0122] In some embodiments, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other embodiments, peripheral device interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or receiving an input for more than a predetermined duration).
[0123] In some embodiments, the event classifier 170 also includes a hit view determination module 172 and / or an active event identifier determination module 173.
[0124] When the touch-sensitive display system 112 displays more than one view, the hit view determination module 172 provides software procedures for determining where within one or more views a sub-event has occurred. A view consists of controls and other elements that a user can see on the display.
[0125] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of the respective application) in which a touch is detected optionally correspond to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest level view in which a touch is detected is optionally referred to as a hit view, and the set of events that are identified as correct input is optionally determined at least in part based on the hit view of the initial touch that started the touch-based gesture.
[0126] Hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy that should handle the sub-events. In most cases, the hit view is the lowest level view in which the initiating sub-event (i.e., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once a hit view is identified by the hit view determination module, the hit view typically receives all sub-events related to the same touch or input source for which it is identified as the hit view.
[0127] Active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with one particular view, higher views in the hierarchy will still remain as actively participating views.
[0128] Event distributor module 174 distributes event information to event recognizers (e.g., event recognizers 180). In embodiments including active event recognizer determination module 173, event distributor module 174 delivers the event information to the event recognizers determined by active event recognizer determination module 173. In some embodiments, event distributor module 174 stores the event information in an event queue, which is retrieved by corresponding event receiver module 182.
[0129] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or is part of another module stored in memory 102, such as contact / motion module 130.
[0130] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the user interface of the application. Each application view 191 of application 136-1 includes one or more event identifiers 180. Typically, the corresponding application view 191 includes multiple event identifiers 180. In other embodiments, one or more event identifiers in event identifiers 180 are part of an independent module, which is a higher-level object such as a user interface toolkit or application 136-1 from which methods and other properties are inherited. In some embodiments, the corresponding event handler 190 includes one or more of the following: data updater 176, object updater 177, GUI updater 178 and / or event data 179 received from event classifier 170. Event handler 190 optionally utilizes or calls data updater 176, object updater 177 or GUI updater 178 to update application internal state 192. Alternatively, one or more of the application views 191 include one or more corresponding event handlers 190. In addition, in some embodiments, one or more of the data updater 176, the object updater 177, and the GUI updater 178 are included in the corresponding application view 191.
[0131] A corresponding event identifier 180 receives event information (e.g., event data 179) from event classifier 170 and identifies an event from the event information. Event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, event identifier 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0132] Event receiver 182 receives event information from event classifier 170. Event information includes information about sub-events such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves the movement of the touch, the event information optionally also includes the speed and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another orientation (e.g., from a longitudinal orientation to a transverse orientation, or vice versa), and the event information includes corresponding information about the current posture (e.g., positioning and orientation) of the device.
[0133] Event comparator 184 compares event information with predefined event or sub-event definitions, and determines an event or sub-event based on the comparison, or determines or updates the state of an event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 includes the definition of an event (e.g., a predefined sub-event sequence), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in event 187 include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, the definition of event 1 (187-1) is a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on a displayed object, a first lift (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on a displayed object, and a second lift (touch end) of a predetermined duration. In another example, the definition of event 2 (187-2) is a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on a displayed object, movement of the touch on the touch-sensitive display system 112, and lifting of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0134] In some embodiments, event definition 187 includes definitions of events for corresponding user interface objects. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display system 112, when a touch is detected on touch-sensitive display system 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects an event handler associated with a sub-event and the object that triggered the hit test.
[0135] In some embodiments, the definition of the corresponding event 187 also includes delay actions that delay the delivery of the event information until after it has been determined that the sub-event sequence does or does not correspond to the event type of the event identifier.
[0136] When a corresponding event recognizer 180 determines that a sequence of sub-events does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events of the touch-based gesture are ignored. In this case, other event recognizers (if any) that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.
[0137] In some embodiments, the corresponding event recognizers 180 include metadata 183 with configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact or can interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0138] In some embodiments, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates an event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferred sending) the sub-events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a tag associated with the identified event, and the event handler 190 associated with the tag obtains the tag and executes a predefined process.
[0139] In some embodiments, the event delivery instructions 188 include a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with a sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or with an actively participating view receives the event information and executes a predetermined process.
[0140] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video or music player module 152. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the positioning of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on a touch-sensitive display.
[0141] In some embodiments, event handler 190 includes or has access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0142] It should be understood that the above discussion of event processing of user touches on a touch-sensitive display also applies to other forms of user input that utilize input devices to operate the multifunction device 100, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses, optionally in conjunction with single or multiple keyboard presses or holddowns; contact movement on a touchpad, such as tapping, dragging, scrolling, etc.; stylus input; input based on real-time analysis of video images obtained by one or more cameras; movement of the device; verbal commands; detected eye movement; biometric input; and / or any combination thereof are optionally used as input corresponding to sub-events that define the event to be distinguished.
[0143] Figure 2A A touch screen (e.g., touch-sensitive display system 112, Figure 1A) of a portable multifunction device 100 (e.g., a view of the front of the device 100). The touch screen optionally displays one or more graphics within the user interface (UI) 200. In these embodiments and in other embodiments described below, the user can select one or more of the graphics by, for example, making gestures on the graphics using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics will occur when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or rolling of a finger that has made contact with the device 100 (from right to left, from left to right, up and / or down). In some specific implementations or in some cases, inadvertent contact with a graphic will not select the graphic. For example, when the gesture corresponding to the selection is a tap, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application.
[0144] The device 100 optionally also includes one or more physical buttons, such as a "home desktop" or menu button 204. As previously described, the menu button 204 is optionally used to navigate to any application 136 in a set of applications optionally executed on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen display.
[0145] In some embodiments, the device 100 includes a touch screen display, a menu button 204 (sometimes referred to as a main desktop button 204), a push button 206 for powering the device on / off and for locking the device, a volume adjustment button 208, a user identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and keeping the button pressed for a predefined time interval; lock the device by pressing the button and releasing the button before the predefined time interval passes; and / or unlock the device or initiate an unlocking process. In some embodiments, the device 100 also accepts voice input for activating or deactivating certain functions through a microphone 113. The device 100 also optionally includes one or more contact strength sensors 165 for detecting contact strength on the touch-sensitive display system 112, and / or one or more tactile output generators 163 for generating tactile output for a user of the device 100.
[0146] Figure 2BPortable multifunction device 100 is shown optionally including optical sensors 164-1 and 164-2 and a time-of-flight ("ToF") sensor 220 (e.g., a view of the rear of device 100). When optical sensors (e.g., cameras) 164-1 and 164-2 simultaneously capture representations of a physical environment (e.g., images or videos), the portable multifunction device can determine depth information based on differences between the information simultaneously captured by the optical sensors (e.g., differences between the captured images). The depth information provided by the differences (e.g., images) determined using optical sensors 164-1 and 164-2 may lack accuracy, but typically provides high resolution. To improve the accuracy of the depth information provided by the differences between the images, time-of-flight sensor 220 is optionally used in conjunction with optical sensors 164-1 and 164-2. ToF sensor 220 transmits a waveform (e.g., light from a light emitting diode (LED) or laser) and measures the time it takes for a reflection of the waveform (e.g., light) to return to ToF sensor 220. Depth information is determined based on the measured time it takes for the light to return to ToF sensor 220. ToF sensors generally provide high accuracy (e.g., 1 cm or better accuracy with respect to measuring distance or depth), but may lack high resolution (e.g., the resolution of ToF sensor 220 is optionally one-quarter or less than one-quarter the resolution of optical sensor 164, or one-sixteenth or less than one-sixteenth the resolution of optical sensor 164). Thus, combining the depth information from the ToF sensor with the depth information provided by the disparity determined using an optical sensor (e.g., a camera) (e.g., an image) provides a depth map that is both accurate and has high resolution.
[0147] Figure 3A300 is a block diagram of an example multifunction device with a display and a touch-sensitive surface according to some embodiments. Device 300 does not have to be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a game system, or a control device (e.g., a home controller or an industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, a memory 370, and one or more communication buses 320 for interconnecting these components. Communication bus 320 optionally includes circuits (sometimes referred to as chipsets) that interconnect system components and control communications between system components. Device 300 includes an input / output (I / O) interface 330 with a display 340, which is optionally a touch screen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 for generating tactile output on device 300 (e.g., similar to the above reference). Figure 1A The tactile output generator 163), the sensor 359 (e.g., similar to the tactile output generator 163 described above), Figure 1A Optical, acceleration, proximity, touch-sensitive and / or contact intensity sensors of the type described, and optionally the above referenced Figure 2B The memory 370 includes a high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 370 optionally includes one or more storage devices located away from the CPU 310. In some embodiments, the memory 370 stores information related to the portable multifunction device 100 ( Figure 1A ) or a subset thereof. In addition, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while portable multifunction device 100 ( Figure 1A )'s memory 102 optionally does not store these modules.
[0148] Figure 3AEach element in the above-mentioned identified element is optionally stored in one or more memory devices in the previously mentioned memory device. Each module in the above-mentioned identified module corresponds to the instruction set for performing the above-mentioned functions. The above-mentioned identified module or program (that is, instruction set) need not be implemented as an independent software program, process or module, so the various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores the above-mentioned identified module and the subgroup of data structure. In addition, memory 370 optionally stores additional modules and data structures not described above.
[0149] FIG. 3B to FIG. 3C is a block diagram of an example computer system 301 according to some embodiments.
[0150] In some embodiments, computer system 301 includes and / or is in communication with the following components:
[0151] Input devices (302 and / or 307, e.g., a touch-sensitive surface such as a touch-sensitive remote control, or a touch screen display that also serves as a display generating component, a mouse, a joystick, a stylus controller, and / or a camera that tracks one or more features of a user such as the position of a user's hand;
[0152] Virtual / augmented reality logic 303 (e.g., virtual / augmented reality module 145);
[0153] A display generation component (304 and / or 308, e.g., a display, a projector, a head-mounted display, a heads-up display, etc.) for displaying virtual user interface elements to a user;
[0154] ● A camera (e.g., 305 and / or 311) for capturing images of the device's field of view, e.g., for determining placement of virtual user interface elements, determining the device's pose, and / or displaying an image of a portion of the physical environment in which the camera is located; and
[0155] • Posture sensors (eg, 306 and / or 311) for determining the posture of the device relative to the physical environment and / or changes in the posture of the device.
[0156] In some embodiments, computer system 301 (e.g., cameras 305 and / or 311) includes a time-of-flight sensor (e.g., time-of-flight sensor 220, Figure 2B ) ) , for reference as above Figure 2B The capturing depth information.
[0157] On some computer systems (e.g. Figure 3BIn 301-a), input device 302, virtual / augmented reality logic component 303, display generation component 304, camera 305, and gesture sensor 306 are all integrated into a computer system (e.g., Figure 1A to Figure 1B In a portable multifunction device 100 in FIG. 3 or device 300 in FIG. 3 , such as a smart phone or tablet computer).
[0158] In some computer systems (e.g., 301-b), in addition to the integrated input device 302, virtual / augmented reality logic component 303, display generation component 304, camera 305; and gesture sensor 306, the computer system also communicates with additional devices independent of the computer system, such as an independent input device 307, such as a touch-sensitive surface, a stylus, a remote control, etc. and / or an independent display generation component 308, such as a virtual reality headset or augmented reality glasses that overlay virtual objects on the physical environment.
[0159] On some computer systems (e.g. Figure 3C In 301-c), input device 307, display generation component 309, camera 311, and / or gesture sensor 312 are separate from and in communication with the computer system. In some embodiments, other combinations of components in and in communication with the computer system 301 are used. For example, in some embodiments, display generation component 309, camera 311, and gesture sensor 312 are combined in a headset that is integrated with or in communication with the computer system.
[0160] In some embodiments, the following reference FIG. 5A to FIG. 5AD All operations described are performed on a single computing device having a virtual / augmented reality logic component 303 (e.g., Figure 3B However, it should be understood that multiple different computing devices are often linked together to perform the following references. FIG. 5A to FIG. 5AD The operations described herein (e.g., a computing device having virtual / augmented reality logic 303 communicating with an independent computing device having display 450 and / or an independent computing device having touch-sensitive surface 451). In any of these embodiments, the following reference is made to FIG. 5A to FIG. 5AD The computing device described is one or more computing devices that contain virtual / augmented reality logic 303. Additionally, it should be understood that in various embodiments, virtual / augmented reality logic 303 may be divided between multiple different modules or computing devices; however, for the purposes of the description herein, virtual / augmented reality logic 303 will primarily be referred to as residing in a single computing device to avoid unnecessarily obscuring other aspects of the embodiments.
[0161] In some embodiments, the virtual / augmented reality logic component 303 includes one or more modules (e.g., one or more event handlers 190, including those described above with reference to Figure 1B In some embodiments, the one or more object updaters 177 and the one or more GUI updaters 178 described in more detail above receive the interpreted inputs and, in response to the interpreted inputs, generate instructions for updating the graphical user interface based on the interpreted inputs, which are then used to update the graphical user interface on the display. Figure 1A and contact motion module 130 in FIG. 3 ), identifying (e.g., by Figure 1B event identifier 180 in ) and / or distribution (e.g., via Figure 1B The interpreted input of the input received by the event classifier 170 in the touch-sensitive surface 451 is used to update the graphical user interface on the display. In some embodiments, the interpreted input is generated by a module on the computing device (e.g., the computing device receives the raw contact input data to identify the gesture from the raw contact input data). In some embodiments, some or all of the interpreted input is received by the computing device as interpreted input (e.g., the computing device including the touch-sensitive surface 451 processes the raw contact input data to identify the gesture from the raw contact input data and sends information indicating the gesture to the computing device including the virtual / augmented reality logic component 303).
[0162] In some embodiments, both the display and the touch-sensitive surface are connected to a computer system (e.g., Figure 3B 301-a). For example, the computer system may be a desktop computer or laptop computer with an integrated display (e.g., 340 in FIG. 3) and a touchpad (e.g., 355 in FIG. 3). For another example, the computing device may be a computer with a touch screen (e.g., Figure 2A 112) of a portable multifunction device 100 (e.g., a smart phone, a PDA, a tablet computer, etc.).
[0163] In some embodiments, the touch-sensitive surface is integrated with the computer system, while the display is not integrated with the computer system including the virtual / augmented reality logic component 303. For example, the computer system can be a device 300 (e.g., a desktop computer or laptop computer, etc.) with an integrated touchpad (e.g., 355 in FIG. 3 ), wherein the integrated touchpad is connected (via a wired or wireless connection) to a separate display (e.g., a computer monitor, a television, etc.). For another example, the computer system can be a computer with a touch screen (e.g., Figure 2AA portable multifunction device 100 (e.g., a smart phone, a PDA, a tablet computer, etc.) is provided with a touch screen 112) in which the touch screen is connected (via a wired or wireless connection) to an independent display (e.g., a computer monitor, a television, etc.).
[0164] In some embodiments, the display is integrated with the computer system, while the touch-sensitive surface is not integrated with the computer system including the virtual / augmented reality logic component 303. For example, the computer system can be a device 300 (e.g., a desktop computer, a laptop computer, a television with an integrated set-top box) with an integrated display (e.g., 340 in FIG. 3), wherein the integrated display is connected (via a wired or wireless connection) to a separate touch-sensitive surface (e.g., a remote trackpad, a portable multifunction device, etc.). For another example, the computer system can be a device 300 with a touch screen (e.g., Figure 2A A portable multifunction device 100 (e.g., a smart phone, a PDA, a tablet computer, etc.) having a touch screen 112 therein, wherein the touch screen is connected (via a wired or wireless connection) to an independent touch-sensitive surface (e.g., a remote touchpad, a portable multifunction device in which another touch screen is used as a remote touchpad, etc.).
[0165] In some embodiments, neither the display nor the touch-sensitive surface is connected to a computer system (e.g., Figure 3C For example, the computer system may be an independent computing device 300 (e.g., a set-top box, a game console, etc.) connected (via a wired or wireless connection) to an independent touch-sensitive surface (e.g., a remote trackpad, a portable multifunction device, etc.) and an independent display (e.g., a computer monitor, a television, etc.).
[0166] In some embodiments, the computer system has an integrated audio system (e.g., audio circuit 110 and speaker 111 in portable multifunction device 100). In some embodiments, the computing device communicates with an audio system that is independent of the computing device. In some embodiments, the audio system (e.g., an audio system integrated into a television unit) is integrated with a separate display. In some embodiments, the audio system (e.g., a stereo system) is a separate system from the computer system and display.
[0167] Attention is now turned to an embodiment of a user interface (“UI”) that is optionally implemented on portable multifunction device 100 .
[0168] Figure 4A An example user interface of an application menu on portable multifunction device 100 is shown according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof:
[0169] ● one or more signal strength indicators of one or more wireless communications such as cellular signals and Wi-Fi signals;
[0170] ● Time;
[0171] Bluetooth indicator;
[0172] Battery status indicator;
[0173] A tray 408 with icons for commonly used applications, such as:
[0174] o An icon 416 of the phone module 138 labeled "Phone," which optionally includes an indicator 414 of the number of missed calls or voicemails;
[0175] o an icon 418 of the email client module 140 labeled “Mail”, which optionally includes an indicator 410 of the number of unread emails;
[0176] o An icon 420 labeled "Browser" of the browser module 147; and
[0177] o An icon 422 labeled “Music” of the video and music player module 152;
[0178] as well as
[0179] ● Icons for other apps, such as:
[0180] o Icon 424 labeled "Message" of IM module 141;
[0181] o An icon 426 labeled “Calendar” of the calendar module 148;
[0182] o Icon 428 labeled “Photos” of the image management module 144;
[0183] o An icon 430 labeled “Camera” of the camera module 143;
[0184] ○ Icon 432 labeled “Online Video” of the online video module 155;
[0185] ○ Icon 434 labeled “Stock Market” of the Stock Market Widget 149 - 2 ;
[0186] o An icon 436 labeled “Map” of the map module 154;
[0187] ○ Icon 438 labeled “Weather” of the weather widget 149-1;
[0188] ○ Icon 440 labeled as “Clock” of the alarm clock widget 149 - 4 ;
[0189] o An icon 442 labeled “Fitness Support” of the fitness support module 142;
[0190] o An icon 444 labeled "Notes" of the notes module 153; and
[0191] o An icon 446 of a settings application or module labeled “Settings” that provides access to settings for the device 100 and its various applications 136;
[0192] ○ Icon 448 of the online store of the application;
[0193] ○An icon 450 for a calculator application;
[0194] ○ Icon 452 of the recording application;
[0195] o An icon 454 for a utility application; and
[0196] o Icon 504 of the Paint Designer application.
[0197] It should be noted that Figure 4A The icon labels shown in are merely examples. For example, other labels are optionally used for various application icons. In some embodiments, the label of the corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.
[0198] Figure 4B A touch-sensitive surface 451 (eg, Figure 3A A device (e.g., a tablet or touch pad 355) Figure 3A Although many of the examples that follow will be given with reference to input on a touch screen display 112 (where a touch-sensitive surface and a display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, such as Figure 4B In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a main axis (e.g., Figure 4B 453) corresponding to the main axis (for example, Figure 4B According to these embodiments, the device detects contact with touch-sensitive surface 451 at a location corresponding to a corresponding location on the display (e.g., Figure 4B 460 and 462) (e.g., in Figure 4B460 corresponds to 468 and 462 corresponds to 470). Thus, on a touch-sensitive surface (e.g., Figure 4B 451) and a display of a multi-function device (e.g., Figure 4B 450 in ) is separated, the user input detected by the device on the touch-sensitive surface (e.g., contacts 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0199] In addition, although the following examples are given primarily with reference to finger inputs (e.g., finger contacts, single-finger tap gestures, finger swipe gestures, etc.), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input, movement of the device or movement of one or more cameras of the device relative to the surrounding physical environment, and / or movement of the user relative to the device being tracked using one or more cameras). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact), or by a gesture involving the user moving his or her hand in a particular direction. For another example, a tap gesture is optionally replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., instead of detecting contact, followed by ceasing to detect contact) or by a corresponding gesture representing a tap gesture. Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple input devices of a particular type are optionally used simultaneously, or multiple input devices of different types are optionally used simultaneously.
[0200] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface that a user is interacting with. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 3A Touchpad 355 or Figure 4B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 451 in the display, the particular user interface element is adjusted according to the detected input. In the case that the touch-sensitive surface 451 in the display includes a touch-screen display (e.g., Figure 1A A touch-sensitive display system 112 or Figure 4AIn some implementations of a touch screen in a touch screen, a contact detected on the touch screen acts as a "focus selector" such that when an input (e.g., a press input by a contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted according to the detected input. In some implementations, the focus moves from one area of the user interface to another area of the user interface without corresponding movement of a cursor or movement of a contact on the touch screen display (e.g., by using a tab key or arrow keys to move the focus from one button to another); in these implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is typically a user interface element (or contact on the touch screen display) that is controlled by the user to convey the user's desired interaction with the user interface (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a trackpad or touch screen), the position of a focus selector (e.g., a cursor, contact, or selection box) over a corresponding button will indicate that the user desires to activate the corresponding button (rather than other user interface elements shown on the device display). In some embodiments, a focus indicator (e.g., a cursor or selection indicator) is displayed via a display device to indicate the current portion of the user interface that will be affected by input received from one or more input devices.
[0201] User interface and associated processes
[0202] Attention is now directed to a computer system (e.g., portable multifunction device 100( )) that may be configured to include display generating components (e.g., display devices such as displays, projectors, head mounted displays, heads up displays, etc.), one or more cameras (e.g., cameras that continuously provide a real-time preview of at least a portion of content within the camera's field of view and optionally generate a video output comprising one or more streams of image frames capturing the content within the camera's field of view), and one or more input devices (e.g., a touch-sensitive surface such as a touch-sensitive remote control, or a touch screen display that also functions as a display generating component, a mouse, a joystick, a wand controller, and / or a camera that tracks one or more features of a user such as the position of the user's hands), optionally one or more gesture sensors, optionally one or more sensors that detect intensity of contact with the touch-sensitive surface, and optionally one or more tactile output generators (and / or in communication with these components). Figure 1A )、Device 300( Figure 3A ) or computer system 301( Figure 3B )) implemented on the user interface ("UI") and related processes.
[0203] FIG. 5A to FIG. 5ADAn example user interface for scanning and modeling an environment (such as a physical environment) according to some embodiments is shown. The user interface in these figures is used to illustrate the process described below, including 6A to 6F , FIG. 7A to FIG. 7D , FIG. 8A to FIG. 8D and 9A to 9E For ease of explanation, some embodiments in the implementation scheme will be discussed with reference to operations performed on a device with touch-sensitive display system 112. In such embodiments, the focus selector is optionally: a corresponding finger or stylus contact, a representative point corresponding to the finger or stylus contact (e.g., the center of gravity of the corresponding contact or a point associated with the corresponding contact), or the center of gravity of two or more contacts detected on touch-sensitive display system 112. However, in response to detecting a contact on touch-sensitive surface 451 when the user interface shown in the figure is displayed on display 450 together with the focus selector, similar operations are optionally performed on a device with display 450 and a separate touch-sensitive surface 451.
[0204] FIG. 5A to FIG. 5AD An example user interface for scanning and modeling a physical environment using augmented reality is shown according to some embodiments.
[0205] Figure 5A An example home screen user interface (e.g., home screen 502) is shown that includes multiple application icons corresponding to different applications, including at least application icon 420 for a browser application and application icon 504 for a paint design application. As disclosed herein, the browser application and the paint design application are illustrative examples of applications published by different application vendors that utilize application programming interfaces (APIs) or developer toolkits that provide some or all of the scanning and modeling functionality described herein. In addition to the scanning and modeling functionality and user interfaces described herein, different applications provided by different application vendors may have different functionality and / or user interfaces. Different applications provided by different application vendors may provide additional user interfaces for interacting with various representations of the physical environment (e.g., two-dimensional maps, three-dimensional models, and / or images and depth data) that have been obtained using the scanning and modeling user interfaces described herein.
[0206] like Figure 5AAs shown, a corresponding input satisfying the selection criteria is detected on the application icon of the corresponding application in the home screen user interface (e.g., a tap input 506 is detected on application icon 420; a tap input 508 is detected on application icon 504; an air gesture for the application icon in a virtual or augmented reality environment; or another selection input activating the corresponding application). In response to detecting the corresponding input satisfying the selection criteria, the user interface of the corresponding application is displayed. For example, in response to tap input 506 on application icon 420, the user interface of the browser application is displayed (e.g., Figure 5B In response to a tap input 508 on the application icon 504, a user interface of the paint design application is displayed (e.g., as shown in FIG. Figure 5C shown).
[0207] In some embodiments, a user may interact with the user interface of a corresponding application to cause a change in the user interface of the corresponding application. Figure 5B As shown, in response to a user's interaction with the browser application's user interface, the browser application's user interface 510 displays a web page (e.g., having a URL of "www: / / example.com") corresponding to a seller of audio / video equipment (e.g., an online store called "Example Store") that provides functionality for selecting types and quantities of different audio / video equipment (e.g., speakers, subwoofers, cameras, and / or displays) for purchase. Figure 5C As shown, in response to user interaction with the user interface of the paint design application, the user interface 514 of the paint design application displays the interior surfaces selected by the user (e.g., a highlight wall, a wall with a window, a wall behind a TV, and / or a wall behind a sofa) and the corresponding paint / wallpaper selections.
[0208] Figure 5B and Figure 5C An example of how the scanning and modeling user interface described herein can be utilized by an application programming user interface or developer toolkit is shown. In some embodiments, such as Figure 5B As shown, the web page shown in the user interface 510 of the browser application has an embedded user interface object 512, which, when selected, causes the scanning and modeling user interface described herein to be displayed. Figure 5CAs shown, the user interface 514 of the paint design application also includes a user interface object 512, which, when selected, causes the scanning and modeling user interface described herein to be displayed. In some embodiments, the appearance of the user interface object 512 does not have to be the same in the user interfaces of different applications, as long as it is configured to trigger the same application programming user interface and / or developer toolkit for the same scanning and modeling function (e.g., "start scanning"). In some embodiments, different applications may utilize different application programming interfaces or developer toolkits to trigger different scanning and modeling user interfaces that share some or all of the features described herein.
[0209] In some embodiments, such as Figure 5B and Figure 5C As shown, in response to detecting a corresponding input (eg, Figure 5B A tap input 516 on a user interface object 512 in Figure 5C ), such as a tap input 518 on a user interface object 512 in Figure 5D As shown, device 100 displays an initial state of the scanning and modeling user interface described herein.
[0210] exist Figure 5DIn the example, device 100 is located in a physical environment (e.g., room 520 or another three-dimensional environment that includes structural elements (e.g., walls, ceiling, floor, windows, and / or doors) and non-structural elements (e.g., pieces of furniture, appliances, physical objects, pets, and / or people). A camera of device 100 (e.g., optical sensor 164, TOF sensor 220, and / or other imaging and / or depth sensor) faces a first portion of room 520, and the camera's field of view includes a first portion of room 520 corresponding to the camera's current viewpoint (e.g., the current viewpoint is determined based on the camera's current position relative to the physical environment and the current pan / tilt / yaw angles). As the camera moves in the physical environment, the camera's viewpoint and field of view change accordingly, and user interface 522 will show different portions of the physical environment corresponding to the updated viewpoint and updated field of view. In this example, the initial state of user interface 522 includes camera view 524 and user interface object 526 overlaid on camera view 524. In some embodiments, the user interface object 526 is optionally animated to indicate the movement performed by the camera relative to the physical environment 520. In some embodiments, the user interface object 526 is animated in a corresponding manner to prompt the user to begin moving the camera in the physical environment in a corresponding manner (e.g., performing a back-and-forth sideways motion or a figure-of-eight motion) to help the device 100 identify one or more cardinal directions (e.g., horizontal and / or vertical directions) and / or one or more planes (e.g., horizontal and / or vertical planes) in the physical environment. In some embodiments, the initial state of the user interface 522 also includes a prompt (e.g., a banner 528 or another type of alert or guide) that provides the user with text instructions (e.g., "find a wall to scan" or another instruction) and / or graphical guidance (e.g., an animated illustration of how to move the device or another type of illustrative guide) on how to start the scanning process.
[0211] In this example, room 520 includes several structural elements, including four walls (e.g., walls 530, 532, 534, and 536), a ceiling (e.g., ceiling 538), a floor (e.g., floor 540), a window (e.g., window 542), and an entryway (e.g., entryway 544). Room 520 also includes several non-structural elements, including various furniture (e.g., stool 546, cabinet 548, TV cabinet 550, sofa 552, and side table 554), physical objects (e.g., floor lamp 556 and table lamp 558), and other physical objects (e.g., TV 560 and box 562). For illustrative purposes, Figure 5CIncluded is a top view 564 of room 520 illustrating the relative positioning of structural and non-structural elements of room 520 and the corresponding positioning (as indicated by the rounded tip of object 566) and facing direction (e.g., as represented by the curved sides of object 566) of a camera of device 100.
[0212] like Figure 5D As shown, the camera view 524 included in the user interface 522 includes a representation of a first portion of the physical environment, the representation including a representation 530' of a wall 530, a representation 532' of a wall 532, a representation 538' of a ceiling 538, a representation 540' of a floor 540, a representation 548' of a cabinet 548, and a representation 542' of a window 542. The representation of the first portion of the physical environment corresponds to the user's current viewpoint, as indicated by the positioning and facing direction of the object 566 in the top view 564 of the room 520. Although the user interface 522 in this example includes a camera view of the physical environment as a representation of the field of view of one or more cameras, in some embodiments, the representation of the field of view included in the user interface 522 is optionally a see-through view of the physical environment as seen through a transparent or semi-transparent display generating component that displays the user interface 522. In some embodiments, the touch screen display of the device 100 in this example is optionally replaced by another type of display generating component that displays the user interface 522, such as a head-mounted display, a projector, or a heads-up display. In some embodiments, the touch input described in these examples is replaced by an air gesture or other type of user input. For ease of illustration, representations of objects, structural elements, non-structural elements that appear in a representation of the field of view (e.g., a camera view 524 or a perspective view as seen through a transparent or semi-transparent display generating component) are referred to using the same reference numerals as their counterparts in the physical environment, rather than a primed version of the reference numerals.
[0213] Figures 5E to 5W Changes in the user interface 522 during scanning and modeling of the room 520 are shown according to some embodiments. Figures 5E to 5W Device 100 is shown displaying an augmented reality view of a room 520 that includes a representation of the field of view of one or more cameras (e.g., camera view 524 or a view of the environment through a transparent or semi-transparent display generating component) and a preview of a three-dimensional model of the room 520 generated based on a scan of the room 520 (e.g., preview 568 or another preview that includes a partially completed three-dimensional model of the physical environment). In some embodiments, the preview of the three-dimensional model of the room 520 is overlaid on the representation of the field of view of the one or more cameras in the user interface 522, for example, as shown in FIG. Figures 5E to 5WIn some embodiments, a preview of the three-dimensional model of room 520 is optionally displayed in a region of user interface 522 that is separate from the representation of the field of view. In some embodiments, the augmented reality view of room 520 also includes various prompts, alerts, annotations, and / or visual guides (e.g., text and / or graphical objects that prompt and guide the user to change viewpoint, move slowly, move faster, go back to rescan a missed point, and / or perform another action to facilitate scanning) overlaid on the representation of the field of view and / or displayed separately.
[0214] Figure 5E 560 . A change in user interface 522 is shown as a scan of a first portion of a physical environment begins. As the camera of device 100 begins capturing images and depth data of a first portion of the physical environment, user interface object 526 is transformed into a preview 568 of a three-dimensional model generated based on the captured images and depth data. At the beginning, the data is limited, and the progress of the scan and model generation is shown by an expanded graphical indication (e.g., indication 570 or another graphical indication) within preview 568. In some embodiments, preview 568 has a three-dimensional shape that is typical of the physical environment (e.g., a cube shape for a room or a cuboid for a house). In some embodiments, the three-dimensional shape is modified (e.g., expanded and / or adjusted) as the shape of the physical environment is explored and ascertained based on the captured images and / or depth data during the scan. In some embodiments, the device 100 performs edge detection and surface detection (e.g., plane detection and / or curved surface detection) in the first portion of the physical environment based on the captured image and / or depth data; and as edges and surfaces are detected and characterized in the first portion of the physical environment, the device 100 displays corresponding graphical representations of the detected edges and / or surfaces in the user interface 522. Figure 5EAs shown, graphical object 571 (e.g., line and / or linear graphical object) is displayed at a position corresponding to a detected edge between wall 530 and floor 540; graphical object 572 (e.g., line and / or linear graphical object) is displayed at a position corresponding to a detected edge between wall 530 and ceiling 538; graphical object 574 (e.g., line and / or linear graphical object) is displayed at a position corresponding to a detected edge between wall 530 and wall 532; and graphical object 576 (e.g., line and / or linear graphical object) is displayed at a position corresponding to a detected edge between wall 532 and floor 540. In some embodiments, as additional portions of the detected edges are detected and / or ascertained based on the progress of scanning and model generation, the corresponding graphical representations of the detected edges (e.g., graphical objects 571, 572, 574, and 576) are extended in length and / or thickness. In some embodiments, as the precise position of the detected edges is adjusted based on the progress of scanning and model generation, the positioning of the corresponding graphical representations is adjusted (e.g., shifted and / or jittered).
[0215] exist Fig. 5F In some embodiments, as scanning and model generation continues over time and / or based on additional captured images and / or depth data, more details of the spatial properties of the first portion of the physical environment are ascertained. The progress of the scanning and model generation (e.g., changes in the predicted accuracy of the estimated spatial properties of the detected edges and / or surfaces) is shown by changes in the visual properties (e.g., length, shape, thickness, feathering, brightness, translucency, opacity, and / or sharpness) of the graphical objects displayed at the locations of the detected edges and surfaces. Figure 5E and Fig. 5F As shown, as more edges between the wall 532 and the floor 540 are detected, the graphic object 576 extends in length along the edge between the wall 532 and the floor 540. In some embodiments, as shown in FIG. Figure 5E and Fig. 5FAs shown, as the predicted accuracy of the estimated spatial properties (e.g., position, orientation, shape, size and / or spatial extent) of the edge between the wall 530 and the ceiling 538 improves (e.g., due to additional data processing and / or additional captured imagery and / or depth data), the visual characteristics (e.g., length, shape, thickness, amount of feathering, brightness, translucency, opacity and / or sharpness) of the graphical object 572 change accordingly (e.g., extending in length, becoming more detailed or clearer in shape, decreasing in thickness, reducing feathering at the border, increasing opacity, increasing brightness, reducing translucency and / or increasing sharpness). Similarly, as the predicted accuracy of estimated spatial properties (e.g., location, orientation, shape, size, and / or spatial extent) of other edges (e.g., the edge between wall 530 and floor 540, the edge between wall 530 and wall 532) improves (e.g., due to additional data processing and / or additional captured images and / or depth data), the visual characteristics (e.g., length, shape, thickness, feathering, brightness, translucency, opacity, and / or sharpness) of their corresponding graphical objects (e.g., graphical object 571 and graphical object 574) change accordingly (e.g., extending in length, becoming more detailed or clearer in shape, decreasing in thickness, reducing feathering at borders, increasing opacity, increasing brightness, reducing translucency, and / or increasing sharpness). In some embodiments, as more edges and / or surfaces are detected in the first portion of the physical environment, additional graphical objects (e.g., graphical object 578 and graphical object 580) are displayed at corresponding locations of the detected edges and / or surfaces. In some embodiments, an overlay (e.g., a color overlay and / or a texture overlay), other types of graphical objects (e.g., a point cloud, a wireframe, and / or a texture), and / or a visual effect (e.g., a blur, a saturation change, an opacity change, and / or a brightness change) is displayed at the location of the detected surface. In some embodiments, as scanning and model generation progresses and more surfaces are detected and characterized, the area covered by the overlay, other types of graphical objects, and / or visual effects is expanded. For example, in some embodiments, as scanning and model generation progresses, the overlay, point cloud, wireframe, texture, and / or visual effect spans across Figure 5E and Fig. 5F The detected surfaces corresponding to walls 530 and 532 in the image are gradually expanded. In some embodiments, as the predicted accuracy of the estimated spatial properties of the detected surfaces changes (e.g., improves or decreases), the visual properties (e.g., intensity, saturation, brightness, density, opacity, fill material type, and / or sharpness) of the overlay, point cloud, wireframe, texture, and / or visual effect applied to the location of the detected surface also change (e.g., improve, decrease, or otherwise change) accordingly.
[0216] like Fig. 5FAs shown, in addition to detecting edges and surfaces of structural elements (e.g., walls, ceilings, floors, windows, entryways, and / or doors), device 100 also simultaneously detects non-structural elements (e.g., furniture, fixtures, physical objects, and / or other types of non-structural elements) during scanning. In this example, the edges and / or surfaces of cabinet 548 have been detected, but the cabinet has not yet been identified, and device 100 displays graphical object 580 at the location of the detected cabinet 548 (e.g., including displaying segments 580-1, 580-2, 580-3, and 580-4 at the locations of the detected edges) to convey the spatial characteristics that have been estimated for the detected edges and / or surfaces of cabinet 548. In some embodiments, at a given moment (e.g., at Fig. 5F At a moment captured in the first figure, at a moment captured in another figure, or at another moment during the scan, the degree of advancement and prediction accuracy of spatial properties of detected edges, surfaces, and / or objects in different sub-portions of the first portion of the physical environment may be different. For example, the prediction accuracy of the spatial properties of the edge between wall 530 and floor 540 is greater than the prediction accuracy of the spatial properties of the edge between wall 532 and ceiling 538, and greater than the prediction accuracy of the spatial properties of the detected edge of cabinet 548. Thus, the visual properties (e.g., length, shape, thickness, feathering, brightness, translucency, opacity, and / or sharpness) of graphical object 571 are made different from those corresponding visual properties of graphical objects 578 and 580 to reflect the differences in the corresponding prediction accuracy of the spatial properties of the detected edges and / or surfaces of their corresponding physical features (e.g., the edge between wall 530 and floor 540, the edge between wall 532 and ceiling 538, and cabinet 548, respectively). In some embodiments, different portions of a graphical object displayed for different portions of a detected physical feature (e.g., an edge, surface, and / or object) optionally have different values for one or more visual attributes at a given moment, where the values of the one or more visual attributes are determined based on the corresponding predicted accuracy of the spatial attributes of the different portions of the detected physical feature. For example, at a given moment, different portions of a graphical object 580 for different portions of a detected edge and / or surface of a cabinet 548 have different values for one or more visual attributes (e.g., thickness, sharpness, feathering, and / or brightness) depending on the corresponding predicted accuracy of the spatial attributes of different portions of the detected edge and / or surface of the cabinet.
[0217] exist Fig. 5F, a preview 568 of the three-dimensional model of the room 520 is updated to show portions of the walls 530, walls 532, and floor 540 that have been detected based on the scanned image and / or depth data. The spatial relationship between the detected walls 530, walls 532, and floor 540 is shown in the preview 568 by the spatial relationship between their corresponding representations 530", 532", and 540". In some embodiments, a graphical object (e.g., an overlay 570 or another graphical object) is displayed in the preview 568 to indicate the real-time progress of the scanning and model generation (e.g., the overlay 570 extends across the surfaces of the representations 530", 532", and 540" as the spatial properties of their corresponding physical features are estimated with increasingly better accuracy). In some embodiments, as Fig. 5F As shown, preview 568 includes a partially completed three-dimensional model of room 520, and the partially completed three-dimensional model of room 520 is oriented relative to the camera's viewpoint according to the orientation of room 520 relative to the camera's viewpoint. In other words, the portion of the physical environment that is in the camera's field of view (e.g., a camera view of the physical environment, an augmented reality view of the physical environment, and / or a see-through view of the physical environment) that faces the user's viewpoint corresponds to the portion of the partially completed three-dimensional model that faces the user's viewpoint. In some embodiments, as the camera moves (e.g., translates and / or rotates in three dimensions) in the physical environment, the orientation of the camera view 524 and the partially completed three-dimensional model in preview 568 are updated accordingly to reflect the movement of the user's viewpoint.
[0218] exist Figure 5G , as scanning and model generation continues over time, more edges and / or surfaces are detected in the first portion of the physical environment, and corresponding graphical objects are displayed at the locations of the detected edges and / or surfaces to represent their spatial properties (e.g., graphical object 582 is displayed at the location of window 542). In addition, as scanning and model generation progresses over time, spatial characteristics of additional portions of the detected edges, surfaces, and / or objects are estimated, and the visual properties of their corresponding graphical objects are updated based on the changes in the spatial characteristics and their estimated accuracy (e.g., as additional portions of the edges of cabinet 548 are detected and characterized, segments 580-2 and 580-3 are extended in length, and the visual properties of segments 580-1 and 580-4 are updated based on the changes in the estimated accuracy of the corresponding edges of cabinet 548). Figure 5G , an overlay 570 extends across representations 530", 532", and 540" to indicate the progress of the scanning and model generation.
[0219] exist Figure 5HIn the example of FIG. 5 , as scanning and model generation continues over time, more edges and / or surfaces are detected in the first portion of the physical environment. Corresponding graphical objects are displayed at the locations of the detected edges and / or surfaces to represent their spatial properties (e.g., a new surface is detected corresponding to the front surface of cabinet 548, and / or a new surface is detected corresponding to the left surface of cabinet 548), and / or existing graphical objects are expanded and / or extended along newly detected portions of previously detected edges and / or surfaces (e.g., graphical object 582 is extended along a newly detected edge of window 542).
[0220] exist Figure 5H In the example of FIG. 5 , as scanning and model generation continues, detection and characterization of one or more edges and / or surfaces of one or more structural elements and non-structural elements of room 520 are completed. Figure 5H As shown, in response to detecting that detection and characterization of the edge between the wall 530 and the floor 540 has been completed (e.g., based on determining that the prediction accuracy of one or more spatial properties of the edge is above a completion threshold, and / or based on determining that the entire extent of the edge has been detected), the final state of the graphic object 571 is displayed. In some embodiments, the final state of the graphic object displayed in response to detecting that detection and characterization of the corresponding edge or surface of the graphic object in the physical environment is completed has a set of predetermined values for one or more visual attributes of the graphic object (e.g., shape, thickness, feathering, brightness, translucency, opacity, and / or sharpness). Figure 5H As shown, before the detection and characterization of the edge between the wall 530 and the floor 540 is completed, the graphic object 571 is optionally a line broken at an appropriate position, has a higher brightness, has a higher feathering degree along its border, and / or is semi-transparent; and in response to the detection and characterization of the edge between the wall 530 and the floor 540, the final state of the graphic object 571 is displayed, which is optionally a solid line without fragments, has a lower brightness, has no feathering or has a reduced feathering degree along its border, and / or is opaque. Figure 5H As shown, before the detection and characterization of the edge of the cabinet 548 is completed, the graphic object 580 is optionally a plurality of broken or dashed lines, with a plurality of brightness levels along different edges and / or different portions of the same edge, a plurality of feathering degrees along the boundaries of different edges and / or different portions of the same edge, and / or different translucency levels along different edges and / or different portions of the same edge; and in response to detecting that the detection and characterization of the edge of the cabinet 548 is completed, the final state of the graphic object 580 is displayed, which is optionally a set of solid lines (e.g., a two-dimensional bounding box, a three-dimensional bounding box, or other type of outline), with uniform and lower brightness, no feathering or reduced feathering degree along all edges, and / or is uniformly opaque. Figure 5H In the embodiment of the present invention, before detection and characterization of a surface (e.g., the surface of a wall 530 or a cabinet 548) is completed, an overlay or other type of graphical object (e.g., a wireframe, point cloud, and / or texture) displayed at the location of the surface optionally has uneven brightness, includes broken patches, is evolving and flickering, and / or is more transparent; and in response to detecting that detection and characterization of the surface are complete, a final state of the graphical object is displayed, the graphical object optionally having uniform brightness, having a continuous shape, having a stable appearance without flickering, and / or being more opaque.
[0221] In some embodiments, completion of detection and characterization of a surface is visually indicated by an animation (e.g., a sudden increase in brightness followed by a decrease in brightness of an overlay and / or graphical object displayed at the location of the surface) and / or a rapid change in a set of visual properties of an overlay and / or graphical object displayed at the location of the detected surface. In some embodiments, completion of detection and characterization of an edge is visually indicated by a change from a line having varying visual properties (e.g., varying brightness, varying thickness, varying length, varying amount of feathering, and / or varying level of sharpness) (e.g., changes and variations in the predicted accuracy of estimated spatial properties of a detected edge or different portions of the detected edge) to a stable, uniform, and solid line having a preset brightness, thickness, and / or sharpness, and / or no feathering. Figure 5H In the example shown, completion of detecting and characterizing the edge between the wall 530 and the floor 540 is indicated by displaying an animation and / or visual effect 584 that is distinct from a change in the appearance of the graphical object 571 that is displayed as a function of the progress of the scan at or near the wall 530, the floor 540, and the edges therebetween and / or as a function of a change in the predicted accuracy of the estimated spatial properties of the detected edges. In some embodiments, the rate at which the graphical object 571 extends along the detected edge is based on the predicted accuracy of the estimated spatial properties of the detected edge, e.g., the graphical object 571 initially extends along the edge at a slower rate and, as the scan progresses, the graphical object 571 extends at a faster rate and the predicted accuracy of the estimated spatial properties of the detected edge improves over time. Figure 5HIn the example shown, completion of detecting and characterizing the surfaces of cabinet 548 is indicated by displaying an animation and / or visual effect 586 (e.g., animations and / or visual effects 586-1, 586-2, and 586-3 shown on different surfaces of cabinet 548) that is distinct from an expansion and change in the appearance of a covering on cabinet 548 that is displayed based on the progress of a scan of cabinet 548 and / or based on changes in the predicted accuracy of estimated spatial properties of detected surfaces and edges of cabinet 548.
[0222] In some embodiments, a scanning progress indication (e.g., an overlay and / or visual effect) displayed at the location of the detected surface has enhanced visual properties (e.g., higher brightness, higher opacity, and / or higher color saturation) that decrease over time as the predicted accuracy of the estimated spatial properties of the detected surface increases. In some embodiments, completion of scanning and modeling the detected surface is visually indicated by an animated change showing an accelerated enhancement of the visual properties followed by a reduction in enhancement (e.g., an increase in brightness followed by a decrease in brightness, an increase in opacity followed by a decrease in opacity, and / or an increase in color saturation followed by a decrease in color saturation).
[0223] In some embodiments, a scanning progress indication (e.g., a linear graphical object and / or a bounding box) displayed at the location of a detected edge has enhanced visual properties (e.g., higher brightness, higher opacity, and / or higher color saturation) that decrease over time as the predicted accuracy of the estimated spatial properties of the detected edge increases. In some embodiments, completion of scanning and modeling of the detected edge is visually indicated by an animated change showing an accelerated enhancement of the visual properties followed by a reduction in enhancement (e.g., an increase in brightness followed by a decrease in brightness, an increase in opacity followed by a decrease in opacity, and / or an increase in color saturation followed by a decrease in color saturation).
[0224] In some embodiments, when a corner in the physical environment is detected during scanning (e.g., the corner of wall 530, wall 532, and ceiling 538), the prediction accuracy of the spatial properties of the three edges that meet at the corner is improved; and thus, the prediction accuracy applied to the graphical objects displayed at the location of the detected edge (e.g., Figure 5H 572, 574, and 578 in the example) to indicate the predicted accuracy of the detected edge (e.g., animated flickering and / or texture shifting), the amount of feathering, and / or other visual effects (e.g., such as Figure 5HIn some embodiments, when three detected edges intersect at the same corner (e.g., a threshold area around a point or a location), if the prediction accuracy of the three detected edges meets a preset threshold accuracy, the completion of the three detected edges is confirmed; and if the detected edges do not intersect at the same corner, the prediction accuracy of the detected edges will be reduced.
[0225] In this example, the edge between the wall 530 and the floor 540 is partially behind the cabinet 548, and optionally, the graphical object 571 extends along the predicted position of the edge behind the cabinet 548 based on the imaging of the wall 530 and the floor 540 and the captured depth data. In some embodiments, the portion of the graphical object 571 assumed to be behind the cabinet 548 is optionally displayed with reduced visual prominence compared to other portions of the graphical object 571 displayed along the unobstructed portion of the edge. In some embodiments, the reduced visual prominence (e.g., reduced brightness, reduced opacity, increased feathering, and / or reduced sharpness) of the portion of the graphical object 571 displayed along a portion of the edge behind the cabinet 548 corresponds to a reduced prediction accuracy of the spatial properties of that portion of the edge behind the cabinet 548.
[0226] like Figure 5G to Figure 5H As shown, as the scan of the first portion of the room 520 continues, the graphical objects 580 displayed along the detected edges and / or surfaces of the cabinet 548 gradually form a three-dimensional bounding box around the view of the cabinet 548 in the user interface 522. The spatial characteristics (e.g., size, length, height, thickness, spatial extent, dimensions, and / or shape) of the graphical objects 580 correspond to the spatial characteristics (e.g., size, length, height, thickness, spatial extent, dimensions, and / or shape) of the cabinet 548. When the edges and surfaces of the cabinet 548 are fully detected, an object 548″ representing the cabinet 548 in the three-dimensional model of the room 520 is displayed in the preview 568 as part of the partially completed three-dimensional model of the room 520. The spatial relationship between the object 548″ in the partially completed model of the room 520 in the preview 568 and the representation 530″ for the wall 530, the representation 532″ for the wall 530, and the representation 540″ for the floor 540 corresponds to the spatial relationship between the cabinet 548 and the wall 530, the wall 532, and the floor 540 in the room 520. The shape and size of the object 548″ relative to the partially completed three-dimensional model in the preview 568 corresponds to the shape and size of the cabinet 548 relative to the room 520. In some embodiments, the object 548″ is a simplified three-dimensional object relative to the cabinet 548 (for example, the detailed surface texture and decorative patterns on the surface of the cabinet 548 are not represented in the object 548″). Figure 5H568, the surfaces and edges of wall 530, the surfaces and edges of wall 532, and the surfaces and edges of floor 540 are also represented by their corresponding representations 530", 532", and 540" in preview 568. The orientation of the partially completed model of room 520 in preview 568 and the camera view 524 of the first portion of room 520 in user interface 522 correspond to the same viewpoint of the user (e.g., the viewpoint represented by the positioning and facing direction of object 566 in top view 564 of room 520).
[0227] In some embodiments, after detecting and modeling edges and surfaces in a first portion of the physical environment has been completed (e.g., including at least three edges and surfaces of wall 530, edges and surfaces of window 542, and edges and surfaces of cabinet 548), device 100 optionally displays a prompt directing the user to continue moving one or more cameras to scan new portions of the environment. Fig.5I As shown, after scanning the first portion of room 520, the user turns the camera to face a second portion of room 520 adjacent to the first portion of room 520. The user's current viewpoint is represented by Fig.5I 566 in the top view 564 of the room 520. Corresponding to the updated viewpoint of the user, the camera view 524 of the physical environment included in the user interface 522 is updated to include a second portion of the room 520, including the wall 532 and the furniture and physical objects in front of the wall 532 (e.g., the stool 546, the TV cabinet 550, the TV 560, and the floor lamp 556). Fig.5I As shown, window 542 has been moved out of the current field of view of the camera, and cabinet 548 has been shifted to the left of the field of view of the camera. After the movement of the one or more cameras and the change of the user's viewpoint, more images and depth information corresponding to the second portion of the physical environment are captured by the one or more cameras, and a model for the second portion of the physical environment is being generated based on the newly captured images and / or depth data. Fig.5IAs shown, graphical object 576 for the edge between wall 532 and floor 540 is extended along the edge between wall 532 and floor 540 based on the newly captured image and depth data from the second portion of the physical environment. The earlier displayed portion of the graphical object 576 (e.g., the left portion) is, optionally, displayed with less visual enhancement (e.g., lower brightness, lower color saturation and / or lower opacity) but with higher definition (e.g., more stable, more solid, higher sharpness, less flickering and / or less feathering) to indicate a higher predicted accuracy of the spatial characteristics of the left portion of the edge between the wall 532 and the floor 540; while the later displayed portion of the graphical object 576 (e.g., the right portion) is, optionally, displayed with more visual enhancement (e.g., higher brightness, higher color saturation and / or higher opacity) but with lower definition (e.g., more scattered, more fragmented, less sharpness, more flickering and / or more feathering) to indicate a lower predicted accuracy of the spatial characteristics of the right portion of the edge between the wall 532 and the floor 540. In addition, a graphical object 590 is displayed at the location of the stool 546 to indicate the outline of the stool 546, a graphical object 592 is displayed at the location of the television 560 to indicate the edge and surface of the television 560, and a graphical object 594 is displayed at the location of the television cabinet 550 to indicate the outline of the television cabinet 550. In some embodiments, the graphical objects 590, 592, 594 are displayed with different values or value sets of one or more visual attributes (e.g., brightness, thickness, texture, feathering, blur, sharpness, density, and / or opacity) according to the respective predicted accuracy of the estimated spatial attributes of the edges and surfaces of the stool 546, the television 560, and the television cabinet 550. As the scan continues to progress, the appearance of graphical objects 590, 592, 594 is continuously updated (e.g., expanded and / or updated in terms of the values of one or more visual properties) based on the detection of new portions of the edges and surfaces of stool 546, television 560, and television cabinet 550 and based on updates to the corresponding predicted accuracy of the estimated spatial properties of the edges and surfaces of stool 546, television 560, and television cabinet 550.
[0228] exist Fig.5I In the example, according to the change of the user's current viewpoint (for example, from Figure 5H 520 is rotated to the left by the first angular amount around the vertical axis in response to the camera's field of view rotating to the right by the first angular amount. Fig.5IAs shown, object 548″ representing cabinet 548, representation 530″ for wall 530, and representation 542″ for window 542 (e.g., a hollowed-out area, a transparent area, or another type of representation) are rotated to the left of preview 568, while cabinet 548 and wall 530 are shifted to the left of camera view 524 in user interface 522. Fig.5I , representation 530″ of wall 530, representation 532″ of wall 532, and representation 540″ of floor 540 in the partially completed three-dimensional model of room 520 displayed in preview 568 are expanded as more image and depth data of wall 530, wall 532, and floor 540 are captured by one or more cameras and processed by device 100.
[0229] exist Fig.5I , after identifying (e.g., based on scan data and / or spatial characteristics of the cabinet) a cabinet 548 (e.g., as an object of a known type, as having a corresponding label or name, as belonging to a known group, and / or otherwise identifiable with a label, icon, or another similar representation), a previously displayed graphical object 580 at the location of the cabinet 548 is gradually replaced by another representation 596 (e.g., a label, icon, avatar, text object, and / or graphical object) of the cabinet 548 that does not spatially indicate one or more spatial characteristics (e.g., size, length, height, thickness, dimensions, and / or shape) of the cabinet 548. For example, as Fig.5IAs shown, the graphic object 580 spatially indicates one or more spatial characteristics of the cabinet 548 (e.g., the size, length, height, thickness, dimensions and / or shape of the graphic object 580 corresponds to the size, length, height, thickness, dimensions and / or shape of the cabinet 548, and / or the graphic object 580 is a bounding box or outline of the cabinet 548). When another representation 596 is displayed at the location of the cabinet 548, the graphic object 580 gradually fades out from the location of the cabinet 548. In some embodiments, the spatial characteristics of the representation 596 are independent of the spatial characteristics of the cabinet 548 (e.g., the size, length, height, thickness, dimensions and / or shape of the representation 596 of the cabinet 548 does not correspond to the size, length, height, thickness, dimensions and / or shape of the cabinet 548). In some embodiments, the representation 596 is smaller than the graphic object 580 (e.g., occupies less area and / or has a smaller spatial extent). In some embodiments, representation 596 indicates the type of object that has been identified (e.g., representation 596 includes the name of cabinet 548, the model number of cabinet 548, the type of furniture that cabinet 548 has, the brand name of cabinet 548, and / or the owner or manufacturer of cabinet 548). In some embodiments, representation 596 is an icon or image that indicates the object type of cabinet 548. In some embodiments, after representation 596 is displayed, graphical object 580 is no longer displayed (e.g., as shown in FIG. 1 ). Figure 5J ). In some embodiments, after displaying representation 596, graphical object 580 is displayed in a translucent and / or dimmed state or another state with reduced visual prominence. In some embodiments, the spatial relationship between graphical object 580 and cabinet 548 is fixed after scanning and modeling of cabinet 548 is completed, regardless of the orientation of cabinet 524 relative to the user's current viewpoint (e.g., when the viewpoint changes, graphical object 580 and cabinet 548 move and rotate in the same manner in camera view 524). In some embodiments, the spatial relationship between representation 596 and cabinet 548 is not fixed and may change according to the user's current viewpoint (e.g., when the viewpoint changes, representation 596 and cabinet 548 may translate together (e.g., representation 596 is attached to the detected front surface of cabinet 548), but representation 596 will rotate to face the current viewpoint, regardless of the facing direction of cabinet 548 relative to the viewpoint).
[0230] exist Figure 5J , as the scan of the second portion of the physical environment continues, graphical object 580 ceases to be displayed at the location of cabinet 548 in camera view 524, and representation 596 remains displayed at the location of cabinet 548 (e.g., representation 596 is attached to the front surface of cabinet 548 and is rotated to face the user's viewpoint). Figure 5JAs more edges and / or surfaces are scanned and modeled, graphical objects corresponding to the newly detected edges and / or surfaces are displayed at the corresponding locations of these newly detected edges and / or surfaces in the camera view 524 (e.g., graphical object 598 is displayed at the location of floor lamp 556). Figure 5J In , as additional portions of known edges and / or surfaces are scanned and modeled, graphical objects corresponding to these known edges and / or surfaces are expanded in camera view 524 (e.g., graphical object 592 corresponding to television 560 and graphical object 594 corresponding to television cabinet 550 are expanded). Figure 5J In, as the predicted accuracy of the spatial properties of the detected edges and / or surfaces continues to change and / or improve, one or more display properties of the graphical objects corresponding to the detected edges and / or surfaces are updated based on the changes in the predicted accuracy of the spatial properties of their corresponding edges and surfaces (e.g., the display properties of graphical object 590 corresponding to stool 546, the display properties of graphical object 594 corresponding to television cabinet 550, and the display properties of graphical object 576 for the edge between wall 532 and floor 540 are updated based on the changes in the predicted accuracy of the spatial properties of their corresponding structural and / or non-structural elements). Figure 5J , as detection and modeling of edges and / or surfaces are completed, a final state of a graphical object representing the edges and / or surfaces is displayed (e.g., a final state of graphical object 592 for television 560 is displayed), and, optionally, an animated change in the appearance of the graphical object is displayed to indicate completion of scanning and modeling of the edges and / or surfaces (e.g., visual effect 598 is displayed for completion of scanning the edge between wall 532 and floor 540, and visual effect 600 is displayed for completion of scanning the surface of television 560).
[0231] exist Figure 5J In some embodiments, device 100 determines that an unscanned portion of room 520 exists between the first portion of the physical environment that has been modeled and the second portion of the physical environment that has been modeled based on scanning a first portion of the physical environment and optionally scanning a second portion of the physical environment. In some embodiments, device 100 determines that an unscanned portion of the physical environment exists between the two scanned portions of the physical environment based on determining that the models of the two scanned portions of the physical environment cannot be satisfactorily joined together. In this example, when the first portion of room 520 is being scanned (e.g., as shown in FIG. 1 ), device 100 determines that an unscanned portion of the physical environment exists between the two scanned portions of the physical environment based on determining that the models of the two scanned portions of the physical environment cannot be satisfactorily joined together. FIG. 5F to FIG. 5H548 is in a position that blocks a portion of the wall 530 from being captured by the camera; and when the viewpoint changes and the second portion of the room is being scanned, the cabinet 548 still blocks the view of the missing portion of the wall 530, and when the second portion of the room 520 is in the field of view of the camera, the missing portion of the wall 530 is almost completely moved out of the field of view of the camera. It should be clarified that the missing portion of the wall 530 that has not yet been scanned refers to the portion of the wall 530 that includes the entryway 544 (which is visually blocked by the cabinet 548 from certain viewing angles), rather than the portion of the wall 530 that is directly behind the rear surface of the cabinet 548, which is not visible from any viewing angle. The device 100 determines, for example, based on the above information that the user may have presumed that scanning and modeling of the first wall 530 of the physical environment has been completed and the user has continued to scan the second portion of the physical environment. Based on the above determination, the device 100 displays a prompt (e.g., a banner 602 and / or another alert or notification) for the user to scan the missing points in the presumed completed portion of the physical environment. In some embodiments, the prompt is updated to provide more detailed and up-to-date guidance on how the user can move to scan the missing portion of the presumed completed portion of the physical environment (e.g., an updated banner that reads "move forward," "move left," "turn so the camera faces left," and / or other appropriate instructions). In some embodiments, in addition to the prompt, device 100 displays one or more visual guides to help the user find the location of the missing portion of the already scanned portion of the physical environment. For example, Figure 5J As shown, a visual indication (e.g., arrow 604 and / or another type of visual indication or graphical object) is displayed to indicate the position of the missing portion of the wall 530 hidden behind the cabinet 548 (e.g., arrow 604 points to the position of the missing portion of the wall 530 behind the cabinet 548 from the current viewing angle). In some embodiments, the visual indication is an animated object (e.g., an animated arrow and / or an animated icon), and the animation (e.g., the direction of movement of the animated object and / or the movement pattern of the animated object) indicates the position of the missing portion of the wall 530 behind the cabinet 548, as viewed from the current viewing angle. In some embodiments, the device 100 displays the visual indication at a position on one side of the camera view closest to the missing portion of the wall 530. In some embodiments, as the field of view of the camera changes, the visual indication is optionally updated depending on the relative spatial positioning of the missing portion of the wall 530 and the currently displayed portion of the physical environment. In some embodiments, the visual indication is displayed at a visual depth corresponding to the missing portion of the presumably completed portion of the physical environment (e.g., arrow 604 is displayed at a depth corresponding to the depth of the missing portion of wall 530 that is hidden behind cabinet 548 from the current viewing angle).
[0232] In some embodiments, the device 100 further displays a visual indication (e.g., a dot 606 or another type of visual indication) at a location in the camera view that corresponds to a location from which the missing portion of the wall 530 can be captured by the camera. For example, a dot 606 of the camera view 524 at a location overlaid on the floor 540 is displayed to indicate that if the user were to stand near the stool 546 and point the camera in the direction indicated by the arrow 604, the image and depth data of the missing portion of the wall 530 would be captured. In some embodiments, the visual indication is an animated object (e.g., a bouncing ball, another type of animated object, or a visual effect). In some embodiments, the visual indication is displayed at a visual depth corresponding to a location where the missing portion of the presumed completed portion of the physical environment can be scanned (e.g., a dot 606 is displayed at a depth corresponding to a depth at which the missing portion of the wall 530 behind the cabinet 548 can be scanned).
[0233] In some embodiments, a visual indication indicating the location of the missing portion of the wall 530 is displayed in the preview 568 of the three-dimensional model of the room 520. Figure 5J As shown, arrow 608 is displayed in the partially completed model of room 520 at a location next to representation 548" for cabinet 548 and points to a portion of representation 530" for wall 530 that has not yet been scanned and modeled (e.g., the unscanned portion of wall 530 is shown as a flat portion regardless of what structural and / or non-structural elements exist in the unscanned portion of wall 530 and the space in front of it). In some embodiments, the appearance of visual indication 608 corresponds to the appearance of visual indication 604. In some embodiments, the appearance of visual indication 608 is different from that of visual indication 604, wherein the respective appearances of visual indication 608 and visual indication 604 are optionally customized to their respective surroundings to enhance visibility of the visual indication.
[0234] In some embodiments, a visual indication is displayed in the preview 568 of the three-dimensional model of the room 520 indicating where the user can place the camera to capture the missing portion of the wall 530. Figure 5JAs shown, point 610 is displayed in the partially completed model of room 520 at a location on representation 540″ of floor 540 that is adjacent to representation 548″ of cabinet 548. In some embodiments, the appearance of visual indication 610 corresponds to the appearance of visual indication 606. In some embodiments, the appearance of visual indication 610 is different from that of visual indication 606, wherein the respective appearances of visual indication 610 and visual indication 606 are optionally customized to their respective surroundings to enhance visibility of the visual indication. In some embodiments, visual indication 608 and / or visual indication 610 are animated. In some embodiments, visual indication 608 and / or visual indication 610 are stationary relative to preview 568 of the three-dimensional model of room 520.
[0235] exist Figure 5J , as scanning and modeling of the second portion of the physical environment continues, representations of the newly detected objects are added to the partially completed three-dimensional model of room 520 in preview 568. For example, object 560" is added for television 560 at a location in the partially completed three-dimensional model that corresponds to the location of television 560 in the physical environment. In some embodiments, representations of non-structural elements (such as a piece of furniture, physical object, and / or appliance) are not added to the partially completed model in preview 568 until detection and characterization of the non-structural elements are complete (e.g., representations for stool 546, television cabinet 550, and floor lamp 598 are not added to the model yet).
[0236] Figure 5K to Figure 5P Interaction with a partially completed three-dimensional model of room 520 in preview 568 is shown while a scan of a second portion of the physical environment is ongoing and progressing. For example, during the scan of the second portion of the physical environment, more objects are identified and their corresponding spatial representations (e.g., bounding boxes or other graphical objects that spatially indicate the spatial dimensions of the objects) are replaced by their corresponding non-spatial representations (e.g., icons, labels, and / or other graphical objects that do not spatially indicate the spatial dimensions of the objects). In addition, the spatial characteristics and / or predicted accuracy of the spatial characteristics of one or more edges and / or surfaces have changed, and the spatial characteristics and visual attributes of their spatial representations have been updated accordingly. In some embodiments, as detection and modeling of edges and / or surfaces are completed, corresponding visual effects are displayed to indicate that detection and modeling of these edges and / or surfaces are completed. In addition, as detection and modeling of objects are completed, their corresponding representations (e.g., three-dimensional representations or two-dimensional representations) are added to the partially completed three-dimensional model in preview 568.
[0237] For example, in Figure 5K, when the scanning and modeling of stool 546 is complete, graphical object 590 of stool 546 is updated to its final state that spatially represents the spatial characteristics of stool 546 (e.g., graphical object 590 is displayed as a bounding box or another shape representing the spatial extent of stool 546), and a corresponding three-dimensional representation 546" of stool 546 is added to the partially completed model of room 520 at a position to the left of representation 560" for television 560. In addition, in Figure 5K In the present invention, when the scanning and modeling of the TV cabinet 550 are completed, the graphic object 594 of the TV cabinet 550 is updated to its final state in which it spatially represents the spatial characteristics of the TV cabinet 550 (for example, the graphic object 594 is displayed as a bounding box or another shape representing the spatial extent of the TV cabinet 550), and a corresponding three-dimensional representation 550" of the TV cabinet 550 is added to the partially completed model of the room 520 at a position below the representation 560" for the TV 560 (for example, a cube representing the shape and spatial extent of the TV cabinet 550).
[0238] exist Figure 5K , after the television 560 is identified (e.g., based on scan data and / or spatial characteristics of the television) (e.g., as an object of a known type, as having a corresponding label or name, as belonging to a known group, and / or otherwise identifiable with a label, icon, or another similar representation), the previously displayed graphical object 592 at the location of the television 560 is gradually replaced by another representation 612 of the television 560 (e.g., a label, icon, avatar, text object, and / or graphical object) that does not spatially indicate one or more spatial characteristics (e.g., size, length, height, thickness, dimensions, and / or shape) of the television 560. For example, as Figure 5JAs shown, the graphical object 592 spatially indicates one or more spatial characteristics of the television 560 (e.g., the size, length, height, thickness, dimensions, and / or shape of the graphical object 592 corresponds to the size, length, height, thickness, dimensions, and / or shape of the television 560, and / or the graphical object 592 is a bounding box or outline of the television 560). When another representation 612 is displayed at the location of the television 560, the graphical object 592 gradually fades out from the location of the television 560. In some embodiments, the spatial characteristics of the representation 612 are independent of the spatial characteristics of the television 560 (e.g., the size, length, height, thickness, dimensions, and / or shape of the graphical object 612 does not correspond to the size, length, height, thickness, dimensions, and / or shape of the television 560). In some embodiments, the representation 612 is smaller than the graphical object 592 and smaller than the television 560 (e.g., occupies less area and / or has a smaller spatial extent). In some embodiments, representation 612 indicates the type of object that has been identified (e.g., representation 612 includes the name of television 560, the model of television 560, the type of appliances that television 560 has, the brand name of television 560, and / or the owner or manufacturer of television 560). In some embodiments, representation 612 is an icon or image that indicates the object type of television 560. In some embodiments, after representation 612 is displayed, graphical object 592 is no longer displayed (e.g., as shown in FIG. 1 ). Figure 5L ). In some embodiments, after displaying representation 612, graphical object 592 is displayed in a semi-transparent and / or dimmed state or another state with reduced visual prominence. In some embodiments, the spatial relationship between graphical object 592 and television 560 is fixed after scanning and modeling of television 560 is completed, regardless of the orientation of television 560 relative to the user's current viewpoint (e.g., graphical object 592 and television 560 move and rotate in the same manner in camera view 524 when the viewpoint changes). In some embodiments, the spatial relationship between representation 612 and television 560 is not fixed and may change depending on the user's current viewpoint (e.g., when the viewpoint changes, representation 612 and television 560 may translate together (e.g., representation 612 is attached to the detected front surface of television 560), but representation 612 will rotate to face the current viewpoint, regardless of the facing direction of television 560).
[0239] In some embodiments, such as Figure 5KAs shown, for different types of objects that have been identified, their non-spatial representations are different. In some embodiments, for the same type of objects that have been identified, their non-spatial representations are optionally the same, regardless of how different their spatial representations may be. For example, the non-spatial representation of a large chair and the non-spatial representation of a small chair are optionally the same (e.g., both are labels with stylized chair icons or text labels "Chair"), even if their spatial representations are different (e.g., one is a larger bounding box and the other is a small bounding box, or one is a large cylinder for a large round chair and one is a small cube for a small office chair). In some embodiments, the non-spatial representations of smart home devices (e.g., smart speakers, smart home devices, and / or smart lights) optionally have a similar appearance, but have different visual attributes other than spatial attributes (e.g., visual attributes such as color and / or text or graphical content) to represent different types of smart home devices.
[0240] exist Figure 5K , the non-spatial representation 596 of the cabinet 548 and the non-spatial representation 612 of the television 560 are respectively displayed at the locations of their corresponding objects, but both are rotated to face the user's current viewpoint. In some embodiments, as the user's viewpoint moves, the positioning and perspective of the cabinet 548 and the television 560 will change in the camera view 524 according to the movement of the viewpoint (e.g., the non-spatial representation 596 of the cabinet 548 will translate with the front surface of the cabinet 548 while rotating to continue facing the viewpoint, and the non-spatial representation 612 of the television 560 will translate with the front surface of the television 560 while rotating to continue facing the viewpoint (e.g., optionally rotating by an amount different from that performed by the non-spatial representation 596 and / or turning in a direction different from that performed by the non-spatial representation)).
[0241] exist Figure 5L , as scanning and modeling of the second portion of the physical environment continues, scanning and modeling of the floor lamp 556 is completed, and the final state of the graphical object 598 is displayed to indicate the spatial characteristics of the floor lamp 556. In addition, a representation 556" of the floor lamp 556 is added to the partially completed model of the room 520 in the preview 568, positioned toward the right of the representation 550" of the television cabinet 550. Figure 5LIn response to detecting that stool 546 has been identified, a non-spatial representation 614 (e.g., a label, icon, avatar, text object, and / or graphical object) of stool 546 indicating an identity (e.g., object type, model, name, owner, manufacturer, and / or text description) of stool 546 is displayed at a viewpoint-facing location of stool 546, wherein non-spatial representation 614 of stool 546 does not spatially indicate the spatial dimensions of stool 546. In some embodiments, after displaying non-spatial representation 614 of stool 546, spatial representation 590 of stool 546 ceases to be displayed or is reduced in visual prominence (e.g., displayed with lower brightness and / or color saturation, and / or greater translucency). In some embodiments, after scanning and modeling of an object is completed, but the object is not identified for a period of time, the spatial representation of the object remains displayed without being replaced by a non-spatial representation (e.g., because the TV cabinet 550 has not yet been identified by the device 100, the spatial representation 594 of the TV cabinet 550 remains displayed and is not replaced by the corresponding non-spatial representation). In some embodiments, after scanning and modeling of an object is completed, but the object is not identified for a period of time, the spatial representation of the object fades out after this period of time, even if there is no non-spatial representation to replace it.
[0242] exist Figure 5M , the spatial representation 590 of the stool 546 is replaced by the non-spatial representation 614 of the stool 546 and stops being displayed in the camera view 524. Figure 5M , a non-spatial representation 616 of the floor lamp 556 is displayed at the viewpoint-facing location of the floor lamp 556 in the camera view 524. The non-spatial representation 616 identifies the floor lamp 556 (e.g., identifies the name, object type, owner, group, manufacturer, and / or model of the floor lamp 556). In some embodiments, when the non-spatial representation 616 of the floor lamp 556 is displayed at the location of the floor lamp 556 in the camera view 524, the spatial representation 598 of the floor lamp 556 is reduced in visual prominence or stops being displayed. Figure 5N In the second portion of the physical environment, scanning and modeling are completed, and non-spatial representations of the identified objects in the first portion of the physical environment and the second portion of the physical environment are displayed in the camera view 524 at the locations of their corresponding objects, all of which are facing the viewpoint. Figure 5N , the non-spatial representation 598 of the floor lamp 556 is no longer displayed in the camera view 524. In some embodiments, one or more additional edges, surfaces, and / or objects in the second portion of the physical environment may still be in the Figure 5K to Figure 5P 564 ) during the process of inspection and modeling (e.g., while a partially completed model in preview 568 is being manipulated by a user, as described below).
[0243] exist Figure 5K , while scanning and modeling of the second portion of the physical environment is ongoing, and while updating the camera view 24 and preview 568 with the graphical objects, non-spatial representations, spatial representations, and / or three-dimensional representations of detected edges, surfaces, and / or objects, the device 100 detects the start of input to the preview 568. In some embodiments, such as Figure 5K As shown, detecting the start of input includes detecting contact 616 at a location on touch screen 220 that corresponds to a portion of the partially completed three-dimensional model in preview 568. Figure 5K , device 100 further detects movement of contact 616 across touch screen 220 in a first direction (eg, a slide input or drag input to the right on the partially completed model in preview 568 ).
[0244] exist Figure 5L In some embodiments, in response to detecting an input that includes movement in a first direction (e.g., in response to detecting a sliding input or a dragging input on the partially completed model in the preview 568 in the first direction), the device 100 moves the partially completed model in the preview 568 in a first manner according to the first input (e.g., rotating the partially completed model and / or translating the partially completed model in the first direction). In this example, in response to a rightward sliding on the partially completed model, the device 100 rotates the partially completed model around a vertical axis (e.g., an axis in the direction of gravity, and / or an axis pointing in a downward direction of the preview 568 and / or the user interface 522). In some embodiments, the amount and / or rate of rotation of the partially completed model is based on the distance and / or rate of the sliding input detected on the partially completed model. In some embodiments, during rotation of the partially completed model in preview 568, objects and / or surfaces within the partially completed model may become visually obscured by other objects and / or surfaces in the partially completed model due to the rotation (e.g., representation 550" of TV cabinet 550 obscures representation 546" of stool 546, and representation 556" of floor lamp 556 obscures representation 550" of TV cabinet 550). In some embodiments, visual instructions for guiding the user to rescan missing points in the presumably completed portions of the physical environment (e.g., object 608 and object 610) may become visually obscured by other objects and / or surfaces in the partially completed model due to the rotation (e.g., arrow 608 becomes obscured by representation 548" of cabinet 548) and / or may visually obscure other objects and / or surfaces in the partially completed model due to the rotation.
[0245] exist Figure 5L, after the partially completed model of room 520 in preview 568 is rotated in accordance with a drag input of contact 616, and before the drag input is terminated (e.g., before contact 616 is lifted off, or before other types of termination are detected depending on the input type), the partially completed model of room 520 in preview 568 is shown having an orientation that is different from the orientation of the physical environment relative to the user's viewpoint. Figure 5M , when termination of the drag input is detected, device 100 restores the orientation of the partially completed model in preview 568 so that the orientation of the partially completed model again matches the orientation of the physical environment relative to the current viewpoint. In some embodiments, if the user's viewpoint moves relative to the physical environment while the partially completed model is being rotated and / or moved in preview 568 based on user input for the partially completed model, device 100 updates camera view 524 so that the view of the physical environment in user interface 522 continues to correspond to the current viewpoint, wherein the orientation of the partially completed model after the partially completed model is rotated and / or moved by user input is not based on the current viewpoint as long as termination of the input has not been detected. Figure 5M , upon detecting termination of the drag input (e.g., lifting off contact 616, or another type of termination depending on the input type), device 100 displays the partially completed model in an orientation corresponding to the current viewpoint (e.g., the same orientation as the physical environment in camera view 524).
[0246] exist Figure 5N In , another user input is detected at the location of the partially completed model in preview 568 (e.g., an expand gesture performed by moving two contacts 618-1 and 618-2 away from each other after touching down on the partially completed model in preview 568, or another zoom input of a different input type). In response to detecting the user input, device 100 rescales the partially completed model in preview 568 according to the user input (e.g., increasing the scale of the partially completed model according to the movement of the contacts in the expand gesture, and / or decreasing the scale of the partially completed model according to the movement of the contacts in the pinch gesture). In some embodiments, the direction and magnitude of the rescaling of the partially completed model is based on the direction and magnitude of the relative movement of the user input (e.g., separating contacts causes the model to be enlarged, moving contacts together causes the model to be contracted, and / or the centers of the contacts moving in corresponding directions causes the model to be translated while the model is rescaled). In Fig.5OIn response to detecting a user input corresponding to a request to rescale the partially completed model in preview 568, the partially completed model of room 520 is enlarged. In some embodiments, the changed scale of the partially completed model in preview 568 is maintained, for example, by obscuring a larger portion of camera view 524 than before the input was detected, until termination of the user input is detected (e.g., lifting off contacts 618-1 and 618-2, or another type of termination depending on the type of input). Figure 5P In the embodiment of the present invention, after detecting termination of the user input, the device 100 displays the partially completed model at the original scale used before detecting the user input.
[0247] Figure 5Q to Figure 5R 548 and the area in front of it (e.g., visually obscured by cabinet 548 and / or behind cabinet 548 along the line of sight from the user's viewpoint when scanning the first and second portions of room 520) are shown rescanned according to some embodiments based on the guidance provided by objects 604 and 606. Figure 5Q As shown, if the user follows the guidance provided by arrow 604 and point 606 in camera view 524 (and / or follows the guidance provided by arrow 608 and point 610 in preview 568) under the prompt of banner 602, the user moves the camera toward the position indicated by point 606 and / or point 610. In some embodiments, as Figure 5Q As shown, the banner 602 is optionally updated to show updated instructions to guide the user to move the camera to a desired location and / or face a desired direction to scan the missing portion of the physical environment. Figure 5Q , the user's updated viewpoint is indicated by the position and facing direction of object 566 in top view 564 of room 520. Figure 5Q As shown, as the camera is moving toward the location in the physical environment marked by point 606 in camera view 524, camera view 524 is updated to show a closer view of cabinet 548. In some embodiments, if the location of arrow 604 in the physical environment would be visually obscured by cabinet 548 from the user's current viewpoint, arrow 604 is shown as being visually obscured by cabinet 548 (e.g., the tip of arrow 604 is not drawn, or is shown as semi-transparent). Figure 5Q , the non-spatial representations of the identified objects (eg, representation 596 for cabinet 548 and representation 614 for stool 546) are shown at the locations of their corresponding objects and are each rotated to face the current viewpoint.
[0248] exist Figure 5R, the user has moved to the location indicated by point 606 and / or point 610 and points the camera to the location indicated by arrow 604 and / or arrow 610 (e.g., the user's current location and facing direction are indicated by object 566 in top view 564 of room 520), and camera view 524 is updated to show the missing portion of wall 530 and the area in front of it. After a period of scanning and modeling, image and / or depth data of the missing portion of wall 530 and the area in front of it are captured by the camera and processed by device 100, and edges, surfaces, and / or objects in that portion of the physical environment are detected and modeled and optionally identified. In this example, a structural element (e.g., entryway 544 and / or another structural element) is detected and modeled, and a graphical object 620 is displayed at the location of the structural element in camera view 524 to spatially represent the spatial characteristics of the structural element (e.g., graphical object 620 is an outline and / or covering indicating the shape, size, and / or outline of entryway 544). In some embodiments, the spatial representation of the structural element may be optionally replaced by a non-spatial representation (e.g., an icon, a label, or another type of non-spatial representation) that does not spatially represent the spatial characteristics of the identified structural element and indicates the identity of the structural element (e.g., the type of structural element, the name of the structural element, and / or the style of the structural element). In some embodiments, the missing portion of the physical environment is scanned and modeled similarly to the above description of Figure 5F to Figure 5P Scanning and modeling of new, unscanned parts of the physical environment described.
[0249] exist Figure 5Q and Figure 5R , as the user's viewpoint changes, the size of the non-spatial representations of the identified objects (e.g., representation 596 for cabinet 548 and representation 614 for stool 546) remain constant, even though their corresponding objects may appear closer or farther away in camera view 524 due to the movement of the viewpoint. Figure 5Q and Figure 5R , as the user's viewpoint changes, the non-spatial representations of identified objects (e.g., representation 596 for cabinet 548 and representation 614 for stool 546) are rotated to continue facing that viewpoint, even though they may translate along with their corresponding objects in camera view 524 due to the movement of the viewpoint.
[0250] exist Figure 5Q and Figure 5R , as the user's viewpoint changes, the camera view is updated to show the physical environment from different perspectives and positions, and the partially completed model of room 520 in preview 568 is rotated to correspond to the current viewpoint. Figure 5Q544 , the entryway 544 is represented by a hollowed-out area or transparent area 544 ″ in the representation 530 ″ of the wall 530 having a size corresponding to the size of the entryway 544 in the physical environment. In some embodiments, portions of the camera view 524 that are located behind the representation 542 for the window 542 ″ and the representation 544 for the entryway 544 ″ are visible in the partially completed three-dimensional model in the preview 568 through the representation 542 for the window 542 ″ and the representation 544 for the entryway 544 ″.
[0251] exist Figure 5S to Figure 5W , after scanning a first portion of the physical environment and a second portion of the physical environment (including the initially omitted portion of the physical environment), the user continues to scan a third portion of the physical environment by translating one or more cameras in the room 520 and changing the facing direction of one or more cameras (e.g., as shown by the positioning and facing direction of object 566 in top view 564 of room 520).
[0252] exist Figure 5S 544, the camera is turned to face the corner between wall 532, wall 534, and floor 540 after rescanning the missing portion of the physical environment in front of entryway 544. In response to detecting movement of the one or more cameras and a corresponding movement of the user's viewpoint, device 100 updates camera view 524 to include a third portion of the physical environment corresponding to the user's current viewpoint, the third portion of the physical environment including floor lamp 556, wall 534, and sofa 552. In addition to updating camera view 524, device 100 also rotates the partially completed three-dimensional model in preview 568 to a new orientation corresponding to the user's current viewpoint.
[0253] like FIG. 5S to FIG. 5T As shown, while capturing and processing image and / or depth data of a third portion of the physical environment, graphical objects corresponding to edges, surfaces, and / or objects in the third portion of the physical environment are displayed. For example, in response to detecting one or more edges and / or surfaces of the sofa 552, a graphical object 622 is displayed at a location of the sofa 552 that overlays the camera view 524. The graphical object 622 is a spatial representation that spatially indicates the spatial characteristics of the sofa 552 in the camera view 524. During the scan, the graphical object 622 is expanded as additional edges and / or surfaces or additional portions of the detected edges and / or surfaces are detected and characterized; and the values of one or more visual attributes of the graphical object 622 are updated in real time based on changes in the predicted accuracy of the spatial characteristics of the corresponding edges and / or surfaces represented by the graphical object 622. Figure 5S, a final state of graphical object 622 is displayed in response to determining that detection and spatial representation of sofa 552 are complete, wherein the final state of graphical object 622 is a solid three-dimensional outline and / or bounding box of sofa 552. In some embodiments, completion of scanning and spatial representation of sofa 552 is indicated by an animated change (e.g., a sudden increase in brightness followed by a decrease in brightness of a covering on an edge and / or surface of sofa 552, and / or a cessation of an applied visual effect (e.g., feathering and / or flickering) on an edge and / or surface of sofa 552).
[0254] exist Figure 5S and Figure 5T , when the partially completed three-dimensional model of room 520 in preview 568 is rotated to an orientation that corresponds to the user's current viewpoint and to the currently displayed portion of the physical environment, representation 530″ of wall 530 is rotated to a position that will visually obscure representations of other objects and / or surfaces (e.g., representation 532″ of wall 532, representation 534″ of wall 534, representation 550″ of television cabinet 550, representation 560″ of television 560, and / or other representations of structural elements and / or non-structural elements) that exceed a threshold portion of the interior of the three-dimensional model. In some embodiments, representation 530″ of wall 530 is made more translucent or removed entirely so that all or part of representations of other portions of the partially completed three-dimensional model become visible in preview 568 that would otherwise be visually obscured by representation 530″ of wall 530. Figure 5S As shown, representation 532" of wall 532 and representation 534" of wall 534 are visible, while representation 530" of wall 530 is removed, or rendered transparent, partially transparent, or semi-transparent. In some embodiments, representation 544" of entryway 544 and outlines of representation 542" of window 542 remain displayed as transparent, partially transparent, or hollowed-out areas in the partially completed three-dimensional model of room 520 (e.g., objects inside the partially completed model are visible through the transparent, partially transparent, or hollowed-out areas), even though representation 530" of wall 530 has been removed (e.g., optionally retaining the outline) or has been rendered transparent, partially transparent, or semi-transparent in preview 568. Figure 5T In response to detecting that scanning and modeling of the sofa 552 are complete, the device 100 displays a representation 552'' of the sofa 552 in the partially completed model in the preview 568, wherein the position of the representation 552'' of the sofa 552 in the partially completed model of the room 520 corresponds to the position of the sofa 552 in the room 520.
[0255] exist Figure 5U, after scanning the third portion of the physical environment, the camera is turned to face the side table 554 in the room 520. In response to detecting the movement of the one or more cameras and the corresponding movement of the user's viewpoint, the device 100 updates the camera view 524 to include a fourth portion of the physical environment corresponding to the user's current viewpoint, the fourth portion of the physical environment including the wall 534, the sofa 552, the side table 554, and the lamp 558. In addition to updating the camera view 524, the device 100 also rotates the partially completed three-dimensional model in the preview 568 to a new orientation corresponding to the user's current viewpoint.
[0256] like Figure 5U As shown, while capturing and processing image and / or depth data of a fourth portion of the physical environment, graphical objects corresponding to edges, surfaces, and / or objects in the fourth portion of the physical environment are displayed. For example, in response to detecting one or more edges and / or surfaces of the side table 554, the graphical object 624 is displayed at a location of the side table 554 that overlays the camera view 524. The graphical object 624 is a spatial representation that spatially indicates the spatial characteristics of the side table 554 in the camera view 524. During the scan, the graphical object 624 is expanded as additional edges and / or surfaces or additional portions of the detected edges and / or surfaces are detected and characterized; and the values of one or more visual attributes of the graphical object 624 are updated in real time based on changes in the predicted accuracy of the spatial characteristics of the corresponding edges and / or surfaces represented by the graphical object 624. Figure 5U , a final state of the graphical object 624 is displayed in response to determining that the detection and spatial representation of the side table 554 is complete, wherein the final state of the graphical object 624 is a solid three-dimensional outline and / or bounding box of the side table 554. In some embodiments, the completion of the scanning and spatial representation of the side table 554 is indicated by an animated change (e.g., a sudden increase in brightness followed by a decrease in brightness of a covering on the edge and / or surface of the side table 554, and / or a cessation of an applied visual effect (e.g., feathering and / or flickering) on the edge and / or surface of the side table 554). Similarly, in Figure 5U In response to detecting one or more edges and / or surfaces of the desk lamp 558, a graphical object 626 is displayed at a location of the desk lamp 558 that overlays the camera view 524. The graphical object 626 is a spatial representation that spatially indicates the spatial characteristics of the desk lamp 558 in the camera view 524. During the scan, the graphical object 626 is expanded as additional edges and / or surfaces or additional portions of the detected edges and / or surfaces are detected and characterized; and the values of one or more visual attributes of the graphical object 626 are updated in real time based on changes in the predicted accuracy of the spatial characteristics of the corresponding edges and / or surfaces represented by the graphical object 626. Figure 5U, a final state of the graphical object 626 is displayed in response to determining that detection and spatial representation of the desk lamp 558 is complete, wherein the final state of the graphical object is a solid three-dimensional outline and / or bounding box of the desk lamp 558. In some embodiments, completion of the scanning and spatial representation of the desk lamp 558 is indicated by an animated change (e.g., a sudden increase in brightness followed by a decrease in brightness of a covering on the edge and / or surface of the side table 554, and / or a cessation of an applied visual effect (e.g., feathering and / or flickering) on the edge and / or surface of the desk lamp 554).
[0257] exist Figure 5U , after a spatial representation of sofa 552 (e.g., graphical object 622 or another graphical object) is displayed at the location of sofa 552 in camera view 524, device 100 identifies sofa 552, e.g., determines the object type, model, style, owner, and / or category of sofa 552. In response to identifying sofa 552, device 100 replaces the spatial representation of sofa 552 (e.g., graphical object 624 or another spatial representation that spatially indicates the spatial dimensions of sofa 552) with a non-spatial representation of sofa 552 (e.g., object 628 or another object that does not spatially indicate the spatial dimensions of sofa 552).
[0258] exist Figure 5U , the graphic object 632 displayed at the position of the edge between the wall 534 and the floor 540 includes the portion behind the sofa 552 and the side table 554; and based on the lower prediction accuracy of the spatial characteristics of the portion of the edge behind the sofa 552 and the side table 554, the portion of the graphic object 632 corresponding to the portion of the edge behind the sofa 552 and the side table 554 is displayed with reduced visibility (for example, with higher translucency, reduced brightness, reduced sharpness, more feathering and / or with a larger blur radius) compared to the portion of the edge that is not obscured by the sofa 552 and the side table 554.
[0259] exist Figure 5U, when the partially completed three-dimensional model of room 520 in preview 568 is rotated to an orientation that corresponds to the user's current viewpoint and to the currently displayed portion of the physical environment, representation 530" of wall 530 remains in a position that will visually occlude representations of other objects and / or surfaces that are interior to the three-dimensional model (e.g., representation 532" of wall 532, representation 534" of wall 534, representation 546" of stool 546, representation 550" of television cabinet 550, representation 560" of television 560, representation 556" of floor lamp 556, representation 552" of sofa 552, representation 554" of side table 554, representation 558" of table lamp 558, and / or other representations of structural elements and / or non-structural elements). In some embodiments, the representation 530″ of the wall 530 is made more translucent or removed entirely so that all or part of the representation of other parts of the partially completed three-dimensional model becomes visible in the preview 568, which representation would otherwise be visually obscured by the representation 530″ of the wall 530. Figure 5U As shown, representation 532" of wall 532 and representation 534" of wall 534 are visible, while representation 530" of wall 530 is removed (optionally retaining the outline), or made transparent, partially transparent, or semi-transparent. In some embodiments, representation 544" of entryway 544 and the outline of representation 542" of window 542 remain displayed as transparent, partially transparent, or hollowed-out areas in the partially completed three-dimensional model of room 520 (e.g., objects inside the partially completed model are visible through the transparent, partially transparent, or hollowed-out areas), even though representation 530" of wall 530 has been removed or has been made transparent, partially transparent, or semi-transparent in preview 568. Figure 5U In response to detecting that scanning and modeling of the side table 554 and the desk lamp 558 are completed, the device 100 displays a representation 554'' of the side table 554 and a representation 558'' of the desk lamp 558 in the partially completed model in the preview 568, wherein the positions of the representation 554'' of the side table 554 and the representation 558'' of the desk lamp 558 in the partially completed model of the room 520 correspond to the positions of the side table 554 and the desk lamp 558 in the room 520, respectively.
[0260] exist Figure 5V564 , after scanning the fourth portion of the physical environment, the camera is moved and rotated to face the last unscanned wall of the room 520, namely, wall 536. The current position and facing direction of the camera is indicated by the position and facing direction of the object 566 in the top view 564 of the room 520. In response to detecting the movement of one or more cameras and the corresponding movement of the user's viewpoint, the device 100 updates the camera view 524 to include a fifth portion of the physical environment corresponding to the user's current viewpoint, the fourth portion of the physical environment including the wall 536 and the box 562. In addition to updating the camera view 524, the device 100 also rotates the partially completed three-dimensional model in the preview 568 to a new orientation corresponding to the user's current viewpoint.
[0261] like Figure 5V As shown, while capturing and processing image and / or depth data of a fifth portion of the physical environment, graphical objects corresponding to edges, surfaces, and / or objects in the fifth portion of the physical environment are displayed. For example, in response to detecting one or more edges and / or surfaces of box 562, graphical object 630 is displayed at a location of box 562 that overlays camera view 524. Graphic object 630 is a spatial representation that spatially indicates the spatial characteristics of box 562 in camera view 524. During scanning, graphical object 630 is expanded as additional edges and / or surfaces or additional portions of detected edges and / or surfaces are detected and characterized; and the values of one or more visual attributes of graphical object 630 are updated in real time based on changes in the predicted accuracy of the spatial characteristics of the corresponding edges and / or surfaces represented by graphical object 630. Figure 5V , a final state of the graphical object 630 is displayed in response to determining that the detection and spatial representation of the box 562 is complete, wherein the final state of the graphical object 630 includes a solid three-dimensional outline and / or bounding box of the box 562. In some embodiments, the completion of the scanning and spatial representation of the box 562 is indicated by an animated change (e.g., a sudden increase in brightness followed by a decrease in brightness of an overlay on the edge and / or surface of the box 562, and / or a cessation of an applied visual effect (e.g., feathering and / or flickering) on the edge and / or surface of the box).
[0262] exist Figure 5VIn some embodiments, after the spatial representation of box 562 (e.g., graphical object 630 or another graphical object) is displayed at the location of box 562 in camera view 524, device 100 is unable to identify box 562 (e.g., determine the object type, model, style, owner, and / or category of box 562). Therefore, as long as box 562 remains in the field of view, the spatial representation of box 562 remains displayed in camera view 524. In some embodiments, graphical object 630 stops being displayed or becomes less visible (e.g., becomes more translucent and / or has a reduced brightness) after a period of time, even if box 562 is not identified.
[0263] exist Figure 5V 560 , when the partially completed three-dimensional model of room 520 in preview 568 is rotated to an orientation that corresponds to the user's current viewpoint and to the currently displayed portion of the physical environment, representation 532 ″ of wall 532 moves to a position that will visually occlude representations of other objects and / or surfaces that are beyond the interior of the three-dimensional model (e.g., representation 530 ″ of wall 530, representation 534 ″ of wall 534, representation 550 ″ of television cabinet 550, representation 560 ″ of television 560, representation 546 ″ of stool 546, representation 556 ″ of floor lamp 556, representation 556 ″ of sofa 552). 52", representation 554" of side table 554, representation 558" of lamp 558, representation 562" of chest 562, representation 548" of cabinet 548, representation 544" of entryway 547, and / or other representations of structural elements and / or non-structural elements. In some embodiments, representation 532" of wall 532 is made more translucent or removed entirely so that all or part of representations of other parts of the partially completed three-dimensional model become visible in preview 568, which representations would otherwise be visually obscured by representation 532" of wall 532. Figure 5V As shown, representation 530" of wall 530 and representation 534" of wall 534 are visible, while representation 532" of wall 532 is removed (optionally retaining an outline), or made transparent, partially transparent, or translucent. In some embodiments, the outline of representation 532" of wall 532 remains displayed, while representation 532" of wall 532 is removed, or made more translucent. Figure 5V In response to detecting that scanning and modeling of box 562 are completed, device 100 displays representation 562'' of box 562 in the partially completed model in preview 568, wherein positions of representation 562'' of box 562 in the partially completed model of room 520 correspond to positions of box 562 in room 520, respectively.
[0264] After scanning and modeling the fifth portion of the physical environment, the user turns the camera to a sixth portion of the physical environment, where the sixth portion of the physical environment includes at least a portion of the previously scanned and modeled first portion of the physical environment. Figure 5W , device 100 scans and models the sixth portion of the physical environment and determines that the user has completed a cycle of capturing all walls of room 520, and that the edge between wall 530 and wall 536 has been detected and modeled. Figure 5W , device 100 also detects and models the edge between wall 536 and floor 540 and the edge between wall 530 and floor 540. Once the corners of wall 530, wall 536, and floor 540 are detected and the positioning of the corners is consistent with the positioning of the intersecting edges of wall 536, wall 530, and floor 540, the amount of feathering applied to the graphical objects corresponding to the different edges is reduced to indicate that the scanning and modeling of the edges and / or surfaces of walls 530, 536, and floor 540 is complete.
[0265] exist Figure 5W , the partially completed three-dimensional model of room 520 in preview 568 is rotated to an orientation that corresponds to the user's current viewpoint and to the currently displayed portion of the physical environment. Representation 532″ of wall 532 and representation 534″ of wall 534 are moved to a position that will visually obscure representations of other objects and / or surfaces that are interior to the three-dimensional model (e.g., representation 530″ of wall 530, representation 536″ of wall 536, and / or other representations of structural elements and / or non-structural elements in room 520). In some embodiments, representation 532″ of wall 532 and representation 534″ of wall 534 are made more translucent or removed entirely so that all or part of representations of other portions of the partially completed three-dimensional model become visible in preview 568 that would otherwise be visually obscured by representation 532″ of wall 532 and representation 534″ of wall 534. In some embodiments, the outlines of representation 532" of wall 532 and representation 534" of wall 534 remain displayed, while representations 532" and 534" are removed, or made more translucent. Figure 5V , the edge between representations 532" and 534" remains displayed in preview 568 to indicate the location of the edge between walls 532 and 534 in the physical environment.
[0266] exist Figure 5XIn some embodiments, in response to detecting that scanning and modeling of the entire room (e.g., all four walls and the interior thereof, or another set of required structural elements and / or non-structural elements) is complete, device 100 stops displaying the partially completed three-dimensional model of room 520 and displays an enlarged three-dimensional model 634 of room 520 that has been generated based on the completed scanning and modeling of room 520. In some embodiments, completed three-dimensional model 634 of room 520 is displayed in user interface 636 that does not include camera view 524. In some embodiments, if user interface 522 includes a see-through view of the physical environment as seen through a transparent or translucent display, device 100 optionally displays an opaque or translucent background layer that blocks and / or obscures the view of the physical environment when three-dimensional model 634 is displayed in user interface 636. In some embodiments, user interface 522 includes an affordance (e.g., an "Exit" button 638 or another user interface object that can be selected to terminate or pause the scanning and modeling process) that, when selected, causes user interface 636 to be displayed before device 100 determines, based on predetermined rules and criteria, that scanning and modeling of room 520 is complete. In some embodiments, if the user selects the affordance (e.g., an "Exit" button 638 or another similar user interface object) to terminate the scanning before the device determines that scanning and modeling of the physical environment is complete, device 100 stops the scanning and modeling process and displays an enlarged version of the partially completed three-dimensional model available at that time in user interface 636, and device 100 stores and displays the partially completed three-dimensional model as the completed three-dimensional mode of room 520 at that time.
[0267] In some embodiments, such as Figure 5X As shown, the device 100 displays the completed three-dimensional model 634 (or a partially completed three-dimensional model if the scan is terminated prematurely by the user) in an orientation that does not necessarily correspond to the current position and facing direction of the camera (e.g., as indicated by the position and facing direction of the object 566 in the top view 564 of the room 520). For example, in some embodiments, the orientation of the three-dimensional model 634 is selected by the device 100 to enable better viewing of the object detected in the physical environment. In some embodiments, the orientation of the three-dimensional model 634 is selected based on the user's initial viewpoint when the scan is first started or based on the user's final viewpoint when the scan is ended (e.g., the representation of the first wall scanned by the user faces the user in the user interface 636, or the representation of the last wall scanned by the user faces the user in the user interface 636).
[0268] In some embodiments, while displaying the user interface 636 including the completed three-dimensional model 634 of the room 520, the device 100 detects the start of user input to the three-dimensional model 634. In some embodiments, such as Figure 5X As shown, detecting the start of input includes detecting contact 638 at a location on touch screen 220 that corresponds to a portion of three-dimensional model 634 in user interface 636. Figure 5X , device 100 further detects movement of contact 638 across touch screen 220 in a first direction (eg, a slide input or a drag input on completed model 634 in user interface 636 ).
[0269] exist Figure 5Y In the example, in response to detecting an input including movement in a first direction (e.g., in response to detecting a sliding input or dragging input on a completed model 634 in a user interface in a first direction), the device 100 moves the completed three-dimensional model 634 in a first manner according to the input (e.g., rotating and / or translating the completed model 634 in the first direction). In this example, in response to a rightward sliding on the completed model 634, the device 100 rotates the completed model around a vertical axis (e.g., an axis in the direction of gravity, and / or an axis pointing in a downward direction of the model 634 and / or the user interface 636). In some embodiments, the amount and / or rate of rotation of the completed model is based on the distance and / or rate of the sliding input detected on the completed model. In some embodiments, during the rotation of the completed model in the user interface 636, objects and / or surfaces within the completed model may become visually obscured by other objects and / or surfaces in the partially completed model due to the rotation.
[0270] exist Figure 5Z , after the completed model of room 520 in user interface 636 is rotated in accordance with a drag input by contact 638, and before the drag input is terminated (e.g., before contact 638 is lifted off, or before other types of termination are detected depending on the input type), the completed model 634 of room 520 in user interface 636 is shown having the orientation specified by the user input. Figure 5AA In the embodiment, when the drag input is detected to be terminated, the device 100 does not restore the orientation of the completed model 634 in the user interface 636, so that the orientation of the completed model continues to be in the orientation specified by the user input (e.g., Figure 5X This is the opposite of the behavior of the partially completed model in preview 568, as described above with respect to Figure 5K to Figure 5M described.
[0271] In some embodiments, if another user input is detected at the location of the completed model 634 in the user interface 636 (e.g., an expand gesture performed by moving two contacts away from each other after touching down on the completed model 634 in the user interface 636, or another zoom input of a different input type), the device 100 rescales the completed model in the user interface 636 according to the user input (e.g., increasing the scale of the completed model according to the movement of the contacts in the expand gesture, and / or decreasing the scale of the completed model according to the movement of the contacts in the pinch gesture). In some embodiments, the direction and magnitude of the rescaling of the completed model 634 are based on the direction and magnitude of the relative movement of the user input (e.g., separated contacts cause the model to be enlarged, contacts moving together cause the model to be contracted, and / or the centers of contacts moving in corresponding directions cause the model to be translated, being rescaled simultaneously). In some embodiments, in response to detecting a user input corresponding to a request to rescale the completed model 634 in the user interface 636, the completed model 634 of the room 520 is rescaled. In some embodiments, the changed scale of the completed model 634 in the user interface 636 is maintained (e.g., the rescaled model may even be partially outside the display area of the display generation component) until the user input is detected to be terminated (e.g., lifting off contact, or another type of termination depending on the input type). After detecting the termination of the user input, the device 100 displays the completed model at the final scale used before the user input was terminated.
[0272] like Figure 5AA As shown, the user interface 636 optionally includes a plurality of selectable user interface objects that correspond to different operations related to the scanning process and / or operations related to the models and / or data that have been generated. For example, in some embodiments, the user interface 636 includes an enable representation (e.g., a "Done" button 638, or another type of user interface object) that, when selected by a user input (e.g., a tap input, an air tap gesture, and / or another type of selection input), causes the device 100 to terminate the scanning and modeling process described herein and return to the application from which the scanning and modeling process was initiated. For example, in response to activation of button 638, the device 100 stops displaying the user interface 636 and returns to the application from which the scanning and modeling process was initiated. Figure 5B ) displays a user interface 644 of the browser application (e.g., Figure 5AB In another example, in response to activation of button 638, device 100 stops displaying user interface 636 and the scanning and modeling process begins from user interface 514 of the paint design application (e.g., in response to activation of button 638). Figure 5C ) displays a user interface 646 of the paint design application (e.g., Figure 5AC shown).
[0273] In some embodiments, the user interface 636 includes an affordance (e.g., a "rescan" button 640 or another type of user interface object) that, when selected by a user input (e.g., a tap input, an air tap gesture, and / or another type of selection input), causes the device 100 to return to the user interface 522 and allow the user to restart the scanning and modeling process and / or rescan one or more portions of the physical environment. For example, in response to activation of button 640, the device 100 stops displaying the user interface 636 and displays the user interface 522 with a preview 568 (e.g., including a currently completed three-dimensional model of the room 520 to be further updated, or including a completely new partially completed three-dimensional model to be built from scratch) and a camera view 524 (e.g., updated based on the current viewpoint). In some embodiments, the redisplayed user interface 522 includes one or more user interface objects for the user to specify which portion of the model needs to be updated and / or rescanned. In some embodiments, the redisplayed user interface 522 includes one or more visual guides to indicate which portion of the model has a lower prediction accuracy.
[0274] In some embodiments, user interface 636 includes an affordance (e.g., a “Share” button 642 or another type of user interface object) that, when selected by user input (e.g., a tap input, an air tap gesture, and / or another type of selection input), causes device 100 to display a user interface with selectable options to interact with the generated model and corresponding data, such as sharing, storing, and / or opening using one or more applications (e.g., the application from which the scanning and modeling process was initiated, and / or an application different from the application from which the scanning and modeling process was first initiated). For example, in some embodiments, in response to activation of button 642, device 100 ceases to display user interface 636 and displays user interface 648 (e.g., as shown in FIG. 1 ). Figure 5AD ), where a user can interact with one or more selectable user interface objects to review the model and / or corresponding data and perform one or more operations with respect to the model and / or corresponding data.
[0275] In some embodiments, such as Figure 5ABAs shown, the user interface 644 of the browser application includes a representation of the completed three-dimensional model 634, optionally enhanced with other information and graphical objects. For example, the three-dimensional model 634 of the room 520 is used to show how the AV equipment selected by the user can be placed inside the room 520. In some embodiments, the user interface 644 allows the user to drag the three-dimensional model and use various inputs (e.g., using a drag input, a pinch input, and / or an expand input) to rescale the three-dimensional model. In some embodiments, the user interface 644 includes an affordance 645-1 (e.g., a "back" button or other similar user interface object) that, when selected, causes the device 100 to stop displaying the user interface 636 and redisplay the user interface 644. In some embodiments, user interface 644 includes an affordance 645-2 (e.g., a "Share" button or other similar user interface object) that, when selected, causes device 100 to display a plurality of selectable options for sharing model 634, corresponding data for model 634, a layout of AV equipment (e.g., selected and / or recommended) generated based on model 634, a list of AV equipment that has been selected by the user and their placement in room 520, a list of recommended AV equipment generated based on the model of room 520, scan data of room 520, and / or a list of objects identified in room 520. In some embodiments, device 100 also provides different options for sharing the above data and information, such as options for selecting one or more recipients and / or using one or more applications to share the above data and information (e.g., about Figure 5AD Examples are provided). In some embodiments, the browser application's user interface 644 includes an affordance (e.g., a "Print" button 654-3 or another similar user interface object) that, when selected, causes the current view of the three-dimensional model 634 (optionally including enhancements applied to the model) to be printed to a file or printer. In some embodiments, the device optionally displays a plurality of selectable options to configure printing of the model 634 (e.g., selecting a printer, selecting a subject and data for printing, and / or selecting a format for printing). In some embodiments, the user interface 644 includes an affordance (e.g., a "Rescan" button 645-4 or another similar user interface object) that, when selected, causes the device 100 to stop displaying the user interface 644 and display the user interface 522 (e.g., as shown in FIG. 5 ). Figure 5D and Figure 5E or Figure 5W ) or user interface 636 (e.g., as shown Figure 5X634 or build a new model of room 520 from scratch). In some embodiments, user interface 644 includes an affordance (e.g., "Checkout" or another similar user interface object) that, when selected, causes device 100 to generate a payment interface for paying for AV equipment and services (e.g., scanning and modeling services, and / or layout and recommendation services) provided through the browser application's user interface.
[0276] In some embodiments, such as Figure 5AC As shown, the user interface 646 of the paint design application includes a representation of the completed three-dimensional model 634, optionally enhanced with other information and graphical objects. For example, the three-dimensional model 634 of the room 520 is used to show what the room 520 will look like when the paint and / or wallpaper selected by the user is applied. In some embodiments, the user interface 646 allows the user to drag the three-dimensional model and rotate the three-dimensional model, and use various inputs (e.g., using a drag input, a pinch input, and / or an expand input) to rescale the three-dimensional model. In some embodiments, the user interface 646 includes an enable representation 647-1 (e.g., a "back" button or other similar user interface object), which, when selected, causes the device 100 to stop displaying the user interface 646 and redisplay the user interface 636. In some embodiments, user interface 646 includes an affordance 647-2 (e.g., a "Share" button or other similar user interface object) that, when selected, causes device 100 to display a plurality of selectable options for sharing model 634, corresponding data for model 634, a rendered view of room 520 with selected or recommended paints and wallpapers generated based on model 634, a list of selected paints and wallpapers that have been selected by the user and their placement in room 520, a list of recommended paints and / or wallpapers generated based on the model of room 520, scanned data of room 520, and / or a list of objects identified in room 520. In some embodiments, device 100 also provides different options for sharing the above data and information, such as options for selecting one or more recipients and / or using one or more applications to share the above data and information (e.g., about Figure 5ADExamples are provided). In some embodiments, user interface 646 includes an affordance (e.g., a "Print" button 657-3 or another similar user interface object) that, when selected, causes the current view of three-dimensional model 634 (optionally including enhancements applied to the model) to be printed to a file or printer. In some embodiments, the device optionally displays a plurality of selectable options to configure printing of model 634 (e.g., selecting a printer, selecting a subject and data for printing, and / or selecting a format for printing). In some embodiments, user interface 644 includes an affordance (e.g., a "New Room" button 647-4 or another similar user interface object) that, when selected, causes device 100 to cease displaying user interface 646 and display user interface 522 (e.g., as shown in FIG. 5 ). Figure 5D and Figure 5E 5) for the user to scan another room from scratch or to rescan room 520. In some embodiments, user interface 646 includes a summary of paint selections for different walls of room 520, and includes affordances 647-5 for changing paint and / or wallpaper selections for different walls. In some embodiments, once affordances 647-5 are used to change paint / wallpaper selections, models 634 in user interface 646 are automatically updated by device 100 to show the newly selected paint / wallpaper on their respective surfaces.
[0277] Figure 5AD An example user interface 648 associated with user interface 636, user interface 644, and / or a "share" function of user interface 646 is shown. In some embodiments, user interface 648 is optionally a user interface of an operating system and / or a local application (e.g., an application of a vendor of an API or developer toolkit that provides scanning and modeling functions) that provides scanning and modeling functions described herein. In some embodiments, user interface 648 provides a list of shareable topics. For example, a representation of model 634 of room 520, a representation of top view 564 of room 520, and / or a list of identified objects in room 520 (e.g., list 649-1 or another similar user interface object) is displayed in user interface 648 along with corresponding selection controls (e.g., check boxes, radial buttons, and / or other selection controls). In some embodiments, subsequent sharing functions are applied to one or more of model 634, top view 564, and list 649-1 based on their respective selection states as specified by the selection controls. In some embodiments, user interface 648 includes an affordance (e.g., a “back” button 649-2 or other similar user interface object) that, when selected, causes device 100 to stop displaying user interface 648 and redisplay the user interface from which user interface 648 was triggered (e.g., Figure 5AAUser interface 636, Figure 5AB User interface 644 or Figure 5ACIn some embodiments, the user interface 648 displays multiple selectable representations of contacts or potential recipients 649-3 for sending the selected subject (e.g., model 634, top view 564, and / or list 649-1). In some embodiments, selection of one or more of the representations of contacts or potential recipients 649-3 causes display of a communication user interface (e.g., an instant messaging user interface, an email user interface, a network communication user interface (e.g., a WiFi, P2P, and / or Bluetooth transmission interface) and / or a shared network device user interface) for sending and / or sharing the selected subject (e.g., model 634, top view 564, and / or list 649-1). In some embodiments, the user interface 648 displays multiple selectable representations of an application 649-4 for opening and / or sending the selected subject (e.g., model 634, top view 564, and / or list 649-1). In some embodiments, selection of one or more of the representations of application 649-4 causes device 100 to display a corresponding user interface of the selected application, wherein the selected subject (e.g., model 634, top view 564, and / or list 649-1) can be viewed, stored, and / or shared with another user of the selected application. In some embodiments, user interface 648 includes an affordance (e.g., a "copy" button 649-5 or another similar user interface object) that, when selected, causes device 100 to make a copy of the selected subject (e.g., model 634, top view 564, and / or list 649-1) in a clipboard or memory so that the copy can be pasted into another application and / or user interface opened later. In some embodiments, user interface 648 includes an affordance (e.g., a “Publish” button 649-6 or another similar user interface object) that, when selected, causes device 100 to display a user interface for publishing the selected subject (e.g., model 634, top view 564, and / or list 649-1) to an online location (e.g., a website, an online bulletin board, a social networking platform, and / or a public and / or private sharing platform) so that other users can see the selected subject remotely from another device. In some embodiments, user interface 648 includes an affordance (e.g., an “Add to” button 649-8 or another similar user interface object) that, when selected, causes device 100 to display a user interface for inserting the selected subject (e.g., model 634, top view 564, and / or list 649-1) into an existing model of a physical environment (e.g., a model of a house including room 520 and other rooms, and / or an existing collection of models).In some embodiments, the user interface 648 includes an enable representation (e.g., a “save as” button 649-9 or another similar user interface object) that, when selected, causes the device 100 to display a user interface for saving the selected subject (e.g., the model 634, the top view 564, and / or the list 649-1) in a different format that is more suitable for sharing with another user or platform.
[0278] 6A to 6F 6 is a flowchart illustrating a method 650 for displaying a preview of a three-dimensional model of an environment during scanning and modeling of the environment according to some embodiments. The method 650 is performed on a computer system (e.g., a portable multifunction device 100 ( Figure 1A )、Device 300( Figure 3A ) or computer system 301( Figure 3B )) is executed at a computer system having a display device (e.g., a display (optionally a touch-sensitive display), a projector, a head-mounted display, a head-up display, etc.), such as a touch screen 112 ( Figure 1A )、Display 340( Figure 3A ) or display generation component 304 ( Figure 3B )), one or more cameras (e.g., optical sensor 164 ( Figure 1A ) or camera 305( Figure 3B )) and optionally one or more depth sensing devices, such as a depth sensor (e.g., one or more depth sensors, such as a time-of-flight sensor 220 ( Figure 2B )). Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0279] As described below, method 650 is a method for displaying a preview of a three-dimensional model of an environment during scanning and modeling of the environment and adding additional information to the preview of the three-dimensional model as the scan progresses. The preview of the three-dimensional model can be manipulated (e.g., rotated or otherwise oriented) independently of the field of view of one or more cameras of the computer system. Displaying a preview of the three-dimensional model and allowing manipulation independent of the field of view of the camera of the computer system improves the efficiency of the computer system by reducing the number of inputs a user needs to interact with the preview of the three-dimensional model. For example, a user can freely rotate the preview of the three-dimensional model to a desired orientation without having to constantly readjust the orientation of the preview (e.g., as would be required if the preview always tried to realign the orientation to match the field of view of one or more cameras of the computer system). This also provides improved visual feedback to the user (e.g., improved visual feedback about the progress of the scan) because the preview of the three-dimensional environment can be updated with additional information as the scan progresses.
[0280] In method 650, a computer system (e.g., device 100, device 300, or another computer system described herein) displays (652) via a display generation component a first user interface (e.g., a scanning user interface displayed to show progress of an initial scan of a physical environment to build a three-dimensional model of the physical environment, a camera user interface, and / or a user interface displayed in response to a user's request to perform a scan of the physical environment or to begin an augmented reality session in the physical environment), wherein the first user interface simultaneously includes (e.g., in an overlaid or adjacent manner): a representation of a field of view of one or more cameras (e.g., an image or video of a live feed from a camera, or a view of the physical environment through a transparent or semi-transparent display), the representation of the field of view including a first view of the physical environment corresponding to a first viewpoint of a user in the physical environment (e.g., the first viewpoint of the user corresponds to a user viewing the physical environment via a head-mounted XR device or via a handheld device (such as a handheld device)). The method also provides for providing a method of viewing a physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is convenient for the user to view the physical environment in a manner that is that ... FIG. 5E to FIG. 5H , a first user interface (e.g., user interface 520) includes a camera view 524 that captures a first view of the room 520, which corresponds to a first viewpoint of the user (e.g., as represented by object 566 in top view 564 of the room 520); and a preview of a three-dimensional model of the room 520 (e.g., preview 568 of a partially completed three-dimensional model of a first portion of the room 520).
[0281] In method 650, while displaying the first user interface (e.g., while scanning is in progress and / or not completed), the computer system detects (654) a first movement of one or more cameras in the physical environment that changes a current viewpoint of a user in the physical environment from a first viewpoint to a second viewpoint (e.g., the movement of the one or more cameras includes translation and / or rotation in three dimensions in the physical environment) (e.g., the movement of the one or more cameras includes a pan movement and / or a tilt movement that changes the direction the camera is facing; a horizontal movement and / or a vertical movement that changes the x, y, z positioning of the camera relative to the physical environment; and / or various combinations of the foregoing). For example, FIG. 5H to FIG. 5I As shown, while displaying a first user interface (e.g., user interface 522 including camera view 524 and preview 568), one or more cameras of device 100 (e.g., as indicated by Fig.5I The object 566 in the top view 564 is relative to Figure 5J movement and rotation of the ).
[0282] In method 650, in response to detecting a first movement of one or more cameras: the computer system updates (656) a preview of the three-dimensional model in a first user interface (and optionally updates a representation of the camera's field of view) based on the first movement of the one or more cameras, including: adding additional information to the partially completed three-dimensional model (e.g., based on depth information captured by the one or more cameras); and rotating the partially completed three-dimensional model from the first orientation corresponding to the first viewpoint of the user to a second orientation corresponding to the second viewpoint of the user. For example, the preview includes a view of the updated, partially completed three-dimensional model of the physical environment from the perspective of a virtual user located at or near a second viewpoint relative to the three-dimensional model. In some embodiments, the updated, partially completed model is oriented so that the model and the physical environment have the same or substantially similar orientation relative to the user's second viewpoint. In some embodiments, updating the preview of the three-dimensional model includes scaling the view of the three-dimensional model to accommodate more portions of the model in the same display area as portions are added to the model. For example, as Fig.5I As shown, in response to detecting movement of one or more cameras of device 100 (as indicated by movement and rotation of object 566 in top view 564 of room 520), device 100 updates camera view 524 to show a second portion of the physical environment and rotates preview 568 to a second orientation corresponding to the user's updated viewpoint.
[0283] In method 650, while displaying a first user interface (e.g., while scanning is in progress and / or incomplete), the computer system detects (658) a first input (e.g., a sliding input on a touch-sensitive surface and / or in air; and / or an air gesture specifying a direction of movement or rotation) directed to the preview of the three-dimensional model in the first user interface (e.g., the first input is determined to be directed to the preview because the preview has input focus and / or the location of the first input corresponds to the positioning of the preview in the first user interface). For example, as Figure 5K As shown, while displaying user interface 522 having camera view 524 showing a second portion of room 520 and preview 568 of a three-dimensional model of room 520, device 100 detects a slide input by contact 616 in a first direction on the partially completed three-dimensional model in preview 568, wherein the partially completed three-dimensional model in preview 568 is shown to have a first orientation that is consistent with room 520 relative to the current viewpoint (e.g., as indicated by Figure 5K The viewpoint indicated by the object 566 in the top view 564 of the room 520 in FIG. FIG. 5I to FIG. 5J The second orientation corresponds to the orientation of (the same as the situation shown).
[0284] In method 650, in response to detecting a first input for a preview of the three-dimensional model in a first user interface: the computer system updates (660) the preview of the three-dimensional model in the first user interface based on the first input, including: based on determining that the first input satisfies a first criterion (e.g., the first input includes a swipe input in a first direction, a pinch-drag air gesture, or another similar input of another input type, while the preview of the three-dimensional model has input focus), rotating the partially completed three-dimensional model from a second orientation corresponding to a second viewpoint of the user to a third orientation not corresponding to the second viewpoint of the user (e.g., while the representation of the field of view continues to show a second view of the physical environment corresponding to the second viewpoint of the user, or while continuing to update the representation of the field of view based on movement of one or more cameras performed during the first input). For example, Figure 5K to Figure 5L As shown, in response to the sliding input by contact 616, the computer system rotates the partially completed three-dimensional model of room 520 in preview 568 to a second orientation ( Figure 5K Different new orientations (such as Figure 5LAs shown). In some embodiments, in response to detecting a first input for a preview of the three-dimensional model in the first user interface, the computer system updates the partially completed three-dimensional model based on the depth information of the corresponding portion of the physical environment in the current field of view of one or more cameras (e.g., continuously updating the field of view based on the movement of one or more cameras, and continuously updating the model based on the newly acquired depth information of the portion of the physical environment in the field of view). For example, in some embodiments, at least the first depth information of the first portion of the physical environment corresponding to the first viewpoint and the second depth information of the second portion of the physical environment corresponding to the second viewpoint are used to generate the three-dimensional model. In some embodiments, the depth information includes data required to detect and / or determine the corresponding distances to various objects and / or surfaces in a portion of the physical environment in the field of view of the camera. In some embodiments, the depth information is used to determine the spatial relationships and spatial characteristics of physical features (e.g., objects, surfaces, edges and / or lines) in the physical environment. In some embodiments, the movement of the camera that changes the user's viewpoint is not a condition required to achieve the manual rotation of the preview of the three-dimensional model described above. For example, while a camera is capturing depth information of a corresponding portion of a physical environment in a field of view corresponding to a current viewpoint of a user (e.g., a first viewpoint, a second viewpoint, or another viewpoint different from the first viewpoint or the second viewpoint), a computer system detects another input for a preview of a three-dimensional model; and in response to detecting the new input for the preview of the three-dimensional model in a first user interface, the computer system updates the three-dimensional model based on the depth information and updates the preview of the three-dimensional model in the first user interface according to the new input, wherein updating the preview includes: rotating the partially completed three-dimensional model from a corresponding orientation corresponding to the user's current viewpoint to a new orientation that does not correspond to the user's current viewpoint (e.g., while the representation of the field of view continues to show a corresponding view of the physical environment corresponding to the user's current viewpoint, or while continuing to update the representation of the field of view according to movements of one or more cameras performed during the first input) based on determining that the new input satisfies a first criterion (e.g., the new input includes a sliding input in a first direction, a pinch-drag air gesture, or another similar input of a different input type, while the preview of the three-dimensional model has input focus). In some embodiments, the orientation of the partially completed three-dimensional model changes in direction and / or by an amount determined based on one or more characteristics (e.g., direction, duration, distance, rate, and / or speed of the input that satisfies the first criteria).
[0285] In some embodiments, while displaying a first user interface that includes a representation of the field of view and a preview of the three-dimensional model, the computer system adds (662) to the representation of the field of view a corresponding graphical object at a location (e.g., a location overlaying the representation of the field of view) that corresponds to one or more physical features (e.g., a physical object, a physical surface, a physical plane, a physical boundary, and / or a physical edge) that have been detected in a corresponding portion of the physical environment that is visible in the representation of the field of view. For example, Fig. 5F As shown, graphic objects 572, 578, 576, and 571 are added to the locations of various structural elements (such as the edges between the wall 530, the wall 532, the ceiling 538, and the floor 540 in the camera view 524). In addition, graphic objects 572, 578, 576, and 571 are added to the locations of non-structural elements (such as Fig. 5F In another example, a graphical object 580 is added to the location of the cabinet 548 in the camera view 524. Figure 5J As shown, graphical object 576 is added to the location of the edge between wall 532 and floor 540 in camera view 524, and graphical objects 598 and 594 are added to the locations of floor lamp 556 and TV cabinet 550 in the camera view. For example, in some embodiments, during a scan of a portion of a physical environment, when physical objects, surfaces, and / or planes are detected in the physical environment (e.g., when various spatial characteristics (e.g., length, size, width, shape, boundary, surface, and / or a combination of two or more of the above) and / or identity information (e.g., object type, category, grouping, ownership, class, and / or a combination of two or more of the above) have been estimated and determined to exceed a certain threshold level of accuracy based at least in part on information obtained during the scan (e.g., image and / or depth information), the computer system displays visual feedback to visually indicate the progress of the scan in the form of an outline or overlay that conveys the estimated spatial characteristics and identity information of the detected objects and, optionally, the predicted accuracy of the estimated spatial characteristics and identity information of the detected objects. In some embodiments, the visual feedback (with respect to the detected object) is dynamically updated based on changes in the predicted accuracy of the estimated spatial characteristics of the detected object. 9A to 9E and accompanying description for more details of this visual feedback). Adding respective graphical objects at locations corresponding to one or more physical features that have been detected in respective portions of the physical environment provides improved visual feedback to the user (e.g., improved visual feedback regarding the locations of the physical features in the physical environment and / or improved visual feedback regarding which physical features in the physical environment have been detected by the computer system).
[0286] In some embodiments, the one or more physical features include at least (664) a first physical object (e.g., a piece of furniture, an appliance, an item of equipment, an item of home decor, a person, a pet, etc.), and the corresponding graphical objects include at least the first graphical object displayed at a first location on the representation of the field of view corresponding to the first physical object. Figure 5H In , once one or more edges and surfaces of the cabinet 548 have been detected, a graphical object 580 is displayed at the location of the cabinet 548 in the camera view 524. In another example, in Figure 5J In the embodiment, once one or more edges and surfaces of the television 560 have been detected, a graphical object 592 is displayed at the location of the television 560 in the camera view 524. In some embodiments, the first graphical object is of a first type, the first type including an outline of a first physical object, a bounding box, and / or an overlay with the shape of the first physical object, wherein the first graphical object of the first type has a spatial characteristic indicating the spatial characteristic of the first physical object. In some embodiments, the first graphical object is of a second type, the second type including a label, an icon, and / or an avatar of the first physical object indicating the type, property, grouping, and / or category of the first physical object, but the spatial characteristic of the first graphical object (except for the display positioning of the first graphical object) does not necessarily correspond to the spatial characteristic of the first physical object. In some embodiments, when more information about the physical object is determined and the object type is identified from the physical characteristics of the physical object, the first graphical object is transformed from the first type to the second type during scanning. In some embodiments, the computer system displays both the first type of graphical object and the second type of graphical object for the corresponding physical object during scanning, for example, at least for a period of time. Adding to the representation of the field of view at least a first graphical object displayed at a first position on the representation of the field of view corresponding to the first physical object provides improved visual feedback to the user (e.g., improved visual feedback regarding the position of the first physical object and / or improved visual feedback that the computer system has detected the first physical object).
[0287] In some embodiments, the one or more physical features include at least (666) a first physical surface (e.g., a curved surface and / or a flat surface) (e.g., a wall, a window, a door, an entryway, a floor, a ceiling, and / or a tabletop), and the corresponding graphical object includes at least a second graphical object (e.g., an outline, a bounding box, a filled area, an overlay, a color filter, and / or a transparent filter) displayed at a second location on the representation of the field of view corresponding to the first physical surface. In some embodiments, as Figure 5HAs shown, once the surfaces of wall 530 and wall 532 are detected and characterized, an overlay is optionally displayed on the surfaces of wall 530 and wall 532 in camera view 524. In some embodiments, as shown in FIG. Figure 5H As shown, once the surface of the cabinet 524 is detected and characterized, an overlay is optionally displayed on the surface of the cabinet 548 in the camera view 524. In some embodiments, the second graphical object is of a first type, the first type including an outline of a first physical surface, a bounding box and / or an overlay with the shape of the first physical surface, wherein the second graphical object of the first type has a spatial characteristic indicating the spatial characteristic of the first physical surface. In some embodiments, the second graphical object is of a second type, the second type including a label, an icon and / or an avatar indicating the type, property, grouping and / or category of the first physical surface, but the spatial characteristic of the second graphical object (except for the display positioning of the second graphical object) does not necessarily correspond to the spatial characteristic of the first physical surface. In some embodiments, when more information about the physical surface is determined and the surface type is identified from the physical characteristics of the physical surface, the second graphical object is transformed from the first type to the second type during scanning. In some embodiments, the computer system simultaneously displays both the first type of graphical object and the second type of graphical object for the corresponding physical surface during scanning, for example, at least for a period of time. Adding to the representation of the field of view at least a second graphical object displayed at a second location on the representation of the field of view corresponding to the first physical surface provides improved visual feedback to the user (e.g., improved visual feedback about the location of the first physical surface and / or improved visual feedback that the computer system has detected the first physical surface).
[0288] In some embodiments, after rotating the partially completed three-dimensional model to a third orientation based on the first input (e.g., based on the magnitude, duration, and / or direction of the first input), the computer system detects (668) termination of the first input. In response to detecting termination of the first input: the computer system updates a preview of the three-dimensional model in the first user interface, including: rotating the partially completed three-dimensional model from the third orientation to a fourth orientation corresponding to the user's current viewpoint (e.g., the three-dimensional model is rotated so that a view of the three-dimensional model from the user's viewpoint is the same or similar to a view of the physical environment from the user's viewpoint relative to the physical environment) (e.g., after the effect of the first input terminates, the partially completed three-dimensional model automatically rotates to an orientation corresponding to the orientation of the physical environment relative to the user's current viewpoint) (e.g., the user's current viewpoint remains the user's second viewpoint and the representation of the field of view continues to show a second view of the physical environment corresponding to the user's second viewpoint, or while the representation of the field of view continues to be updated based on movements of one or more cameras performed during the first input and after the first input ends, the current viewpoint is the user's continuously updated viewpoint). For example, in Figure 5L to Figure 5M In the example, after the partially completed three-dimensional model in the preview 568 is rotated according to the sliding input performed by the contact 616 (as shown in FIG. Figure 5L In response to detecting the termination of the sliding input by contact 616, device 100 rotates the partially completed three-dimensional model to its original orientation, which corresponds to the orientation of room 200 relative to the current viewpoint (e.g., as indicated by object 566 in top view 564 of room 520) (e.g., as indicated by object 566 in top view 564 of room 520). Figure 5M After the partially completed three-dimensional model is rotated to the third orientation, rotating the partially completed three-dimensional model from the third orientation to a fourth orientation corresponding to the user's current viewpoint reduces the amount of input required to display the partially completed three-dimensional model in the proper orientation (e.g., the user does not need to perform additional user input to realign the partially completed three-dimensional model with the user's current viewpoint).
[0289] In some embodiments, while the first user interface is displayed (e.g., while a scan is in progress and / or is not complete), the computer system detects (670) a second input (e.g., a pinch input or a reverse pinch input on a touch-sensitive surface and / or in the air; and / or an air gesture specifying a type and amount of zoom) directed to the preview of the three-dimensional model in the first user interface (e.g., the second input is determined to be directed to the preview because the preview has input focus and / or the location of the second input corresponds to the positioning of the preview in the first user interface). In response to detecting a second input directed to the preview of the three-dimensional model in the first user interface: the computer system updates the preview of the three-dimensional model in the first user interface according to the second input, including: in accordance with determining that the second input satisfies a second criterion different from the first criterion (e.g., the second input includes a pinch or reverse pinch input on the touch-sensitive surface, a pinch and flick air gesture, or another similar input of a different input type, while the preview of the three-dimensional model has input focus), changing the scale of the representation of the partially completed three-dimensional model (e.g., an enlarged or shrunken partially completed three-dimensional model) relative to the field of view according to the second input (e.g., based on a direction and / or magnitude of the second input). For example, as Figure 5N to Figure 5OAs shown, in response to detecting an expand gesture performed by contacts 618-1 and 618-2, device 100 enlarges the partially completed three-dimensional model in preview 568 relative to camera view 524 in user interface 522. For example, in some embodiments, based on determining that the second input includes movement in a first direction (e.g., movement to the right, movement in a clockwise direction, and / or movement that reduces the gap between two fingers), the computer system reduces the scale of the partially completed three-dimensional model (e.g., by an amount corresponding to the magnitude of the movement of the second input); and based on determining that the second input includes movement in a second direction (e.g., movement to the left, movement in a counterclockwise direction, and / or movement that increases the gap between two fingers), the computer system increases the scale of the partially completed three-dimensional model (e.g., by an amount corresponding to the magnitude of the movement of the second input). In some embodiments, the first input that rotates the partially completed three-dimensional model and the second input that scales the partially completed three-dimensional model relative to the representation of the field of view are optionally detected as parts of the same gesture (e.g., a pinch or expand gesture that also includes a translational movement of the whole hand), and thus, the rotation and scaling of the partially completed three-dimensional model are performed simultaneously according to the gesture. Changing the scale of the representation of the partially completed three-dimensional model relative to the field of view based on a second input that satisfies a second criterion provides additional control options without cluttering the UI with additional display controls (e.g., additional display controls for rotating the partially completed three-dimensional model and / or additional display controls for changing the scale of the partially completed three-dimensional model).
[0290] In some embodiments, the preview of the three-dimensional model of the physical environment (e.g., a model generated and / or updated based on depth information being captured by one or more cameras during scanning) includes (672) corresponding three-dimensional representations of one or more surfaces that have been detected in the physical environment (e.g., the corresponding three-dimensional representations of the one or more surfaces include representations of surfaces of a floor, one or more walls, or one or more pieces of furniture arranged in three-dimensional space, where the spatial relationships and spatial characteristics correspond to their spatial relationships and spatial characteristics). For example, Figure 5H As shown, preview 568 of the three-dimensional model of room 520 includes three-dimensional representations 530″, 530″, and 540″ for wall 530, wall 532, and floor 540, and three-dimensional representation 548″ for cabinet 548, which includes a plurality of surfaces corresponding to the surfaces of cabinet 548. In another example, in Figure 5S, when an additional surface of wall 534 is detected, a representation 534 of wall 534 is added to the partially completed model in preview 568. In some embodiments, the corresponding representation of one or more surfaces that have been detected in the physical environment includes a virtual surface, bounding box, and / or wireframe in the three-dimensional model, which has spatial characteristics (e.g., size, orientation, shape, and / or spatial relationship) that correspond to (e.g., are reduced in scale relative to) the spatial characteristics (e.g., size, orientation, shape, and / or spatial relationship) of one or more surfaces that have been detected in the physical environment. Displaying a preview of the three-dimensional model (including the corresponding three-dimensional representation of one or more surfaces that have been detected in the physical environment) provides the user with improved visual feedback (e.g., improved visual feedback regarding the detected surfaces in the physical environment).
[0291] In some embodiments, the preview of the three-dimensional model of the physical environment (e.g., a model generated and / or updated based on depth information being captured by one or more cameras during scanning) includes (674) corresponding representations of one or more physical objects that have been detected in the physical environment (e.g., the corresponding representations of the one or more objects include representations of one or more pieces of furniture, physical objects, people, pets, windows, and / or doors in the physical environment). For example, Figure 5V As shown, preview 568 of the three-dimensional model of room 520 includes corresponding representation 548″ for cabinet 548, representation 546″ for stool 546, and / or representation 552″ for sofa 552, as well as other representations for other objects detected in room 520. In some embodiments, the representations of the objects are three-dimensional representations. In some embodiments, the corresponding representations of one or more objects that have been detected in the physical environment include outlines, wireframes, and / or virtual surfaces in the three-dimensional preview, which have spatial characteristics (e.g., size, orientation, shape, and / or spatial relationships) that correspond to (e.g., are reduced in scale relative to) the spatial characteristics (e.g., size, orientation, shape, and / or spatial relationships) of the one or more objects that have been detected in the physical environment. In some embodiments, the representations of the objects have reduced structure and visual detail in the three-dimensional model compared to their corresponding objects in the physical environment. Displaying a preview of the three-dimensional model of the physical environment (including corresponding representations of one or more physical objects that have been detected in the physical environment) provides improved visual feedback to the user (e.g., improved visual feedback about the detected physical objects in the physical environment).
[0292] In some embodiments, after adding additional information to the partially completed three-dimensional model in the preview of the three-dimensional model, based on determining that the partially completed three-dimensional model of the physical environment meets preset criteria (e.g., criteria for determining when a scan of the physical environment is complete, such as because sufficient information has been obtained from the scan and preset conditions regarding detecting surfaces and objects in the physical environment have been met; or because the user has requested that the scan be completed immediately), the computer system replaces (676) the display of the partially completed three-dimensional model in the preview of the three-dimensional model with a display of a first view of the completed three-dimensional model of the physical environment, wherein the first view of the completed three-dimensional model includes an enlarged copy of the partially completed three-dimensional model that meets the preset criteria (and, optionally, rotated to a preset orientation that does not correspond to the user's current viewpoint). For example, as Figure 5W to Figure 5X As shown, after completing the scan of room 520 (e.g., Figure 5W , all four walls of room 520 have been scanned and modeled), device 100 replaces the display of user interface 522 with user interface 636 (eg Figure 5X 520 ), wherein the user interface 636 includes an enlarged version of the completed three-dimensional model 634 of the room 520. For example, when the computer system determines that the scan is complete and the model of the physical environment meets preset criteria, the computer system replaces the preview of the three-dimensional model with a view of the completed three-dimensional model, wherein the view of the completed three-dimensional model is larger than the partially completed model shown in the preview. In some embodiments, the view of the completed three-dimensional model shows the three-dimensional model with a preset orientation (e.g., the orientation of the partially completed model shown when the scan was completed, a preset orientation independent of the orientation of the partially completed model shown when the scan was completed, and independent of the current viewpoint). Replacing the display of the partially completed three-dimensional model in the preview of the three-dimensional model with the display of the first view of the completed three-dimensional model of the physical environment including the enlarged copy of the partially completed three-dimensional model after the additional information is added to the partially completed three-dimensional model in the preview of the three-dimensional model reduces the number of inputs required to display the completed three-dimensional model of the physical environment at an appropriate size (e.g., after the computer system adds the additional information to the partially completed three-dimensional model in the preview of the three-dimensional model, the user does not need to perform additional user input to enlarge the completed three-dimensional model of the physical environment).
[0293] In some embodiments, while displaying a first view of the completed three-dimensional model in a first user interface (e.g., just after a scan is completed or some time after completion) (e.g., optionally, the representation of the field of view includes a corresponding view of the physical environment corresponding to the user's current viewpoint (e.g., the user's first viewpoint, second viewpoint, or another viewpoint different from the first viewpoint and the second viewpoint)), the computer system detects (678) a third input (e.g., a sliding input on a touch-sensitive surface or in the air; or an air gesture that specifies a direction of movement or rotation) directed to the first view of the completed three-dimensional model in the first user interface (e.g., the third input is determined to be directed to the completed three-dimensional model because the view of the three-dimensional model has input focus, or the location of the first input corresponds to the positioning of the view of the three-dimensional model in the first user interface). In response to detecting a third input for the first view of the completed three-dimensional model in the first user interface: the computer system updates the first view of the completed three-dimensional model in the first user interface according to the third input, including: according to determining that the third input satisfies the first criterion (for example, the third input includes a sliding input in the first direction, a pinch-drag air gesture, or another similar input of a different input type, and the view of the completed three-dimensional model has input focus), according to the third input, rotating the completed three-dimensional model from a fourth orientation (for example, a corresponding orientation corresponding to the user's current viewpoint and / or a preset orientation) to a fifth orientation different from the fourth orientation. For example, as Figure 5X to Figure 5Y As shown, after the completed three-dimensional model 634 is displayed in the user interface 636, the device 100 detects a sliding input (such as a sliding input made by a contact 638) for the completed three-dimensional model 634. Figure 5X In response to detecting the sliding input, the device 100 rotates the completed three-dimensional model 634 in the user interface 636 to a new orientation (as shown in FIG. 1 ) according to the sliding input. Figure 5Y In some embodiments, after the scan is complete and the completed three-dimensional model of the physical environment is displayed in the first user interface (e.g., with or without simultaneously displaying representations of the field of view of one or more cameras), the computer system allows the user to rotate the model (e.g., freely or under preset angular constraints) about one or more rotation axes (e.g., rotate about the x-axis, y-axis, z-axis, and / or tilt, yaw, pan, views of the model) to view the three-dimensional model from different angles. In response to detecting a third input for the first view of the completed three-dimensional model in the first user interface, rotating the completed three-dimensional model from a fourth orientation to a fifth orientation different from the fourth orientation in accordance with the third input provides improved visual feedback to the user (e.g., improved visual feedback about the appearance of the three-dimensional model as viewed in different orientations).
[0294] In some embodiments, after rotating the completed three-dimensional model to the fifth orientation based on the third input (e.g., based on the magnitude, duration, and / or direction of the third input), the computer system detects (680) termination of the third input. In response to detecting termination of the third input, the computer system abandons updating the first view of the completed three-dimensional model in the first user interface, including: maintaining the completed three-dimensional model in the fifth orientation (e.g., regardless of the current viewpoint, movement of the display generation component, and / or movement of one or more cameras). For example, Figure 5Y to Figure 5AA As shown, in accordance with the sliding input by the contact 638, the three-dimensional model 634 of the room 520 is rotated to a new orientation (e.g., as Figure 5Z After that, device 100 detects the termination of the sliding input. In response to detecting the termination of the sliding input, device 100 does not further rotate three-dimensional model 634, and does not rotate three-dimensional model 634 back to the orientation shown before the sliding input (for example, Figure 5X and Figure 5Y The orientation of the three-dimensional model 634 shown in FIG. 6A ) and the three-dimensional model 634 is maintained in the current orientation (as shown in FIG. 6B ). Figure 5Z and Figure 5AA ). In some embodiments, maintaining the changed orientation of the completed three-dimensional model after detecting termination of the third input to rotate the model allows the user time to examine the model from a desired viewing angle, decide whether to further rotate the model to examine the model from another viewing angle, and provide appropriate input to do so as needed. In response to detecting termination of the third input, abandoning updating the first view of the completed three-dimensional model in the first user interface (including maintaining the completed three-dimensional model in the fifth orientation) reduces the number of inputs required to interact with the completed three-dimensional model (e.g., the computer system does not change the orientation of the completed three-dimensional model (e.g., to reflect the user's current viewpoint), and therefore the user does not need to perform additional user input to continually readjust the orientation of the completed three-dimensional model back to the fifth orientation).
[0295] In some embodiments, the completed three-dimensional model includes (682) a corresponding graphical representation of a first structural element detected in the physical environment and corresponding graphical representations of one or more physical objects detected in the physical environment. Displaying a first view of the completed three-dimensional model includes: based on determining that a current orientation of the completed three-dimensional model in the first user interface (e.g., when the model is stationary and / or being rotated based on user input) will cause the corresponding graphical representation of the first structural element (e.g., a wall, floor, or another structural element in the physical environment) to occlude a view of corresponding graphical representations of one or more objects (e.g., physical objects in an interior portion of the physical environment such as furniture, physical objects, smart home appliances, people and / or pets), reducing (e.g., while still displaying at least a portion of the graphical representation of the first structural element) an opacity of the graphical representation of the first structural element or ceasing to display the graphical representation of the first structural element (e.g., abandoning display of the corresponding graphical representation of the first structural element in conjunction with the first view of the three-dimensional model). The method further comprises: displaying a corresponding graphical representation of one or more objects in a first view of the three-dimensional model (e.g., not displaying the graphical representation of the first structural element when the completed three-dimensional model is rotated according to the third input); and displaying the corresponding graphical representation of the first structural element together with the corresponding graphical representation of the one or more objects in the first view of the three-dimensional model (e.g., displaying the graphical representation of the first structural element when the completed three-dimensional model is rotated according to the third input) based on a determination that the current orientation of the completed three-dimensional model will not cause the corresponding graphical representation of the first structural element (e.g., a wall, floor, or other structural element in the physical environment) to obscure the view of the corresponding graphical representation of one or more objects (e.g., physical objects in an interior portion of the physical environment, such as furniture, physical objects, smart home appliances, people and / or pets). Figure 5X to Figure 5AA As shown, three-dimensional model 634 in user interface 636 includes representations of multiple structural elements, such as wall 530 , wall 532 , wall 534 , wall 536 , and floor 540 . Figure 5X 54" of wall 534 is not displayed in the view of three-dimensional model 634 in user interface 636 (e.g., optionally, the outline of the representation is displayed when the fill material of the representation is made transparent) because it would obscure representations of physical objects detected in the interior of room 520, such as representation 560" of television 560, representation 556" of floor lamp 556, representation 552" of sofa 552, representation 554" of side table 554, and representations of one or more other objects (e.g., box 562 and table lamp 558) that have been detected in room 520. In another example, after rotating completed three-dimensional model 634 in user interface 636, as shown in FIG. Figure 5YAs shown, representation 536" of wall 536 and representation 534" of wall 534 are removed or made transparent or partially transparent (optionally leaving an outline without fill material) because they would obscure representations of objects that have been detected in room 520 (e.g., representation 560" of television 560, representation 556" of floor lamp 556, representation 552" of sofa 552, representation 554" of side table 554, and representations of one or more other objects (e.g., box 562 and desk lamp 558)). Figure 5X , representation 530″ of wall 530 is displayed simultaneously with representations of objects detected in room 520 because representation 530″ will not obstruct any of the objects in user interface 636 having the current orientation of completed three-dimensional model 634. In another example, in Figure 5Z and Figure 5AA , representation 532″ of wall 532 is displayed simultaneously with representations of objects detected in room 520 because representation 532″ will not obstruct any of the objects in user interface 636 having the current orientation of completed three-dimensional model 634. Based on determining that the current orientation of the completed three-dimensional model in the first user interface will cause the corresponding graphical representation of the first structural element to obstruct the view of the corresponding graphical representation of one or more objects, abandoning display of the corresponding graphical representation of the first structural element and the corresponding representations of one or more objects in the first view of the three-dimensional model, and based on determining that the current orientation of the completed three-dimensional model will not cause the corresponding graphical representation of the first structural element to obstruct the view of the corresponding graphical representation of one or more objects, displaying the corresponding graphical representation of the first structural element and the corresponding representations of one or more objects in the first view of the three-dimensional model simultaneously reduces the number of inputs required to display an appropriate view of the completed three-dimensional model (e.g., if one or more objects in the three-dimensional model are obstructed, the user does not need to perform additional user input to adjust the orientation of the completed three-dimensional model).
[0296] In some embodiments, before displaying the first user interface, the computer system displays (684) a corresponding user interface of a third-party application (e.g., any of a plurality of third-party applications that implement an application interface for the room scanning capability described herein). While displaying the corresponding user interface of the third-party application, the computer system detects a corresponding input to the corresponding user interface of the third-party application, wherein the first user interface is displayed in response to detecting the corresponding input to the corresponding user interface of the third-party application and in accordance with determining that the corresponding input corresponds to a request to scan the physical environment (e.g., satisfies requirements of a system application programming interface (API) for scanning the physical environment). For example, Figure 5B and Figure 5C As shown, then Figure 5DAs shown, a user interface 522 for scanning and modeling a physical environment may be displayed in response to activation of a "start scanning" button 512 in any of the user interfaces of the browser application and the paint design application. In some embodiments, the same scanning process described herein is triggered in response to user input to a corresponding user interface of another different third-party application, wherein the user input corresponds to a request to scan the physical environment (e.g., satisfies the requirements of a system application programming interface (API) for scanning the physical environment). A preview of the three-dimensional model in the first user interface is updated according to a first movement of one or more cameras, and the partially completed three-dimensional model is rotated from a second orientation corresponding to the user's second viewpoint to a third orientation not corresponding to the user's second viewpoint according to a determination that the first input satisfies a first criterion and in response to detecting the first input to the preview of the three-dimensional model in the first user interface, wherein the first user interface is displayed in response to detecting the corresponding input to the corresponding user interface of the third-party application and in response to determining that the corresponding input corresponds to a request to scan the physical environment Provides improved visual feedback to the user (e.g., improved visual feedback about the progress of the partially completed three-dimensional model, and / or improved visual feedback about the appearance of the three-dimensional model, as viewed in different orientations).
[0297] In some embodiments, based on determining that generation of the three-dimensional model meets preset criteria (e.g., a complete three-dimensional model that meets the preset criteria has been obtained based on scanning the physical environment, and / or a user request to terminate scanning the physical environment has been detected), the computer system redisplays (686) the third-party application (e.g., based on spatial information contained in the three-dimensional model, displays the completed three-dimensional model in the user interface of the third-party application, and / or displays content from the third-party application together with at least a portion of the three-dimensional model (e.g., a corresponding set of user interface objects corresponding to a corresponding plurality of actions in the third-party application)). For example, Figure 5AA As shown, then Figure 5AB or Figure 5ACAs shown, after the three-dimensional model has been generated, selection of the "Finish" button 638 causes the device 100 to redisplay the user interface of the application that initiated the scanning and modeling process (e.g., the user interface 644 of the browser application or the user interface 646 of the paint design application). For example, in some embodiments, multiple different third-party applications may utilize the scanning user interfaces and processes described herein to obtain a three-dimensional model of the physical environment, and at the end of the scan, the computer system redisplays the third-party application that initiated the scanning process, and optionally, displays the user interface of the third-party application, which provides one or more options for interacting with the model and utilizing the model to complete one or more tasks of the third-party application. In some embodiments, the user interfaces and functions provided by different third-party applications are different from each other. Based on determining that the generation of the three-dimensional model meets the preset criteria, redisplaying the third-party application reduces the amount of user input required to redisplay the third-party application (e.g., the user does not need to perform additional user input to redisplay the third-party application).
[0298] In some embodiments, displaying a preview of the three-dimensional model including the partially completed three-dimensional model includes (688) displaying a graphical representation of a first structural element (e.g., a wall, floor, entryway, window, door, or ceiling) detected in the physical environment in a first orientation relative to corresponding graphical representations of one or more objects (e.g., physical objects in an interior portion of the physical environment, such as furniture, physical objects, people, and / or ...
Claims
1. A method, comprising: at a computer system in communication with a display generation component, one or more input devices, and one or more cameras: during a scan of a physical environment to obtain depth information of at least a portion of the physical environment: displaying, via the display generation component, a first user interface, wherein the first user interface includes a representation of the field of view of one or more cameras; and displaying a plurality of graphical objects that overlay the representation of the field of view of the one or more cameras, the display including: at least displaying a first graphical object at a first location and a second graphical object at a second location, the first graphical object representing one or more estimated spatial attributes of a first physical feature that has been detected in a corresponding portion of the physical environment within the field of view of the one or more cameras, and the second graphical object representing one or more estimated spatial attributes of a second physical feature that has been detected in the corresponding portion of the physical environment within the field of view of the one or more cameras; and when displaying the plurality of graphical objects that overlay the representation of the field of view of the one or more cameras, changing one or more visual attributes of the first graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the first physical feature, and changing one or more visual attributes of the second graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the second physical feature.
2. The method according to claim 1, wherein the first graphical object includes a first set of one or more lines representing the one or more estimated spatial attributes of the first physical feature, and the second graphical object includes a second set of one or more lines representing the one or more estimated spatial attributes of the second physical feature.
3. The method according to claim 2, wherein displaying the first graphical object comprises: extending a corresponding length of the first set of one or more lines according to the corresponding prediction accuracy of the one or more estimated spatial attributes of the first physical feature.
4. The method according to claim 1, wherein the first graphical object includes a first filled region representing the one or more estimated spatial attributes of the first physical feature, and the second graphical object includes a second filled region representing the one or more estimated spatial attributes of the second physical feature.
5. The method according to claim 1, wherein changing one or more visual attributes of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attributes of the first physical feature comprises: changing a corresponding opacity of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attributes of the first physical feature.
6. The method according to claim 1, wherein changing one or more visual attributes of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attributes of the first physical feature comprises: Changing the corresponding feathering amount applied to the edge of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature.
7. The method according to claim 6, wherein changing the corresponding feathering amount applied to the edge of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature comprises: Reducing the corresponding feathering amount applied to the edge of the first graphical object according to determining that the scan of the corner corresponding to the first graphical object satisfies a first criterion; and Increasing the corresponding feathering amount applied to the edge of the first graphical object according to determining that the scan of the corner corresponding to the first graphical object does not satisfy the first criterion.
8. The method according to claim 7, wherein increasing the corresponding feathering amount and reducing the corresponding feathering amount according to the first criterion are performed according to determining that the first graphical object includes a structural object and is not a non-structural object.
9. The method according to claim 1, wherein changing one or more visual attributes of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature comprises: Changing the corresponding sharpness of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature.
10. The method according to any one of claims 1 to 9, wherein changing one or more visual attributes of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature comprises: At a first time: According to determining that the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature is a first accuracy value of a first part of the first physical feature, displaying a first part of the first graphical object using a first attribute value of a first visual attribute among the one or more visual attributes, and At a second time later than the first time: According to determining that the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature is a second accuracy value of the first part of the first physical feature, displaying the first part of the first graphical object using a second attribute value of the first visual attribute among the one or more visual attributes, wherein the second accuracy value is different from the first accuracy value, and the second attribute value is different from the first attribute value.
11. The method according to any one of claims 1 to 9, wherein at a third time, Based on the corresponding prediction accuracy of the estimated spatial attribute that determines the first physical feature being the third accuracy value of the third part of the first physical feature and the fourth accuracy value of the fourth part of the first physical feature, the third part of the first graphical object is displayed using the third attribute value of the second visual attribute among the one or more visual attributes, and the fourth part of the first graphical object is displayed using the fourth attribute value of the second visual attribute among the one or more visual attributes, where the fourth part of the first physical feature is different from the third part of the first physical feature, the fourth accuracy value is different from the third accuracy value, and the fourth attribute value is different from the third attribute value.
12. The method according to any one of claims 1 to 9, wherein: the first physical feature includes a fifth part of the first physical feature and a sixth part of the first physical feature, the fifth part of the first physical feature is not visually occluded by another object in the field of view of the one or more cameras, and the sixth part of the first physical feature is visually occluded by another object in the field of view of the one or more cameras, and displaying the first graphical object includes: displaying the fifth part of the first graphical object corresponding to the fifth part of the first physical feature using a fifth attribute value corresponding to the fifth accuracy value of the corresponding prediction accuracy of the one or more estimated spatial attributes of the first physical feature, and displaying the sixth part of the first graphical object corresponding to the sixth part of the first physical feature using a sixth attribute value corresponding to the sixth accuracy value of the corresponding prediction accuracy of the one or more estimated spatial attributes of the first physical feature, wherein, in the first user interface, the sixth attribute value corresponds to a lower visibility compared to the visibility corresponding to the fifth attribute value.
13. The method according to any one of claims 1 to 9, comprising: Based on determining that the scanning of the first physical feature is completed: displaying a corresponding change in the one or more visual attributes of the first graphical object to indicate the completion of the scanning for the first physical feature; and stopping changing the one or more visual attributes of the first graphical object according to the change in the corresponding prediction accuracy of the estimated spatial attribute of the first physical feature.
14. The method according to claim 13, wherein displaying the corresponding change in the one or more visual attributes of the first graphical object comprises: Based on determining that the first physical feature has a first feature type, displaying a first type of change in the one or more visual attributes of the first graphical object to indicate the completion of the scanning of the first physical feature; and In response to determining that the first physical feature is a second feature type different from the first feature type, display a change of a second type of the one or more visual attributes of the first graphical object that is different from a change of the first type to indicate completion of scanning of the first physical feature.
15. The method according to claim 14, wherein the first graphical object includes a set of one or more lines, and display a change of the first type of the one or more visual attributes of the first graphical object to indicate completion of scanning of the first physical feature comprising: Reduce the amount of feathering.
16. The method according to claim 14, wherein the first graphical object includes a surface, and display a change of the second type of the one or more visual attributes of the first graphical object to indicate completion of scanning of the first physical feature comprising: Display a preset change sequence of one or more visual attributes in the surface.
17. The method according to any one of claims 1 to 9, comprising: Detect completion of scanning of the first physical feature; and, In response to detecting completion of scanning of the first physical feature, reduce the visual saliency of the first graphical object from a first visibility level to a second visibility level lower than the first visibility level.
18. A computer system communicatively coupled to a display generation component, one or more input devices, and one or more cameras, comprising: One or more processors; and A memory storing one or more programs, wherein the one or more programs are configured to be executed by the one or more processors, and the one or more programs include instructions for performing the following operations: During scanning of a physical environment to obtain depth information of at least a portion of the physical environment: Display a first user interface via the display generation component, wherein the first user interface includes a representation of the fields of view of one or more cameras; and Display a plurality of graphical objects covering the representation of the fields of view of the one or more cameras, the display including: at least display a first graphical object at a first location and a second graphical object at a second location, the first graphical object representing one or more estimated spatial attributes of a first physical feature that has been detected in a corresponding portion of the physical environment in the fields of view of the one or more cameras, and the second graphical object representing one or more estimated spatial attributes of a second physical feature that has been detected in the corresponding portion of the physical environment in the fields of view of the one or more cameras; and When displaying the plurality of graphical objects covering the representation of the fields of view of the one or more cameras, change one or more visual attributes of the first graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the first physical feature, and change one or more visual attributes of the second graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the second physical feature.
19. The computer system according to claim 18, wherein the one or more programs include instructions for performing the method according to any one of claims 2 to 17.
20. A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computer system in communication with a display generation component, one or more input devices, and one or more cameras, cause the computer system to perform the following operations: During scanning of a physical environment to obtain depth information of at least a portion of the physical environment: Display, via the display generation component, a first user interface, wherein the first user interface includes a representation of the fields of view of the one or more cameras ; and Display a plurality of graphical objects that overlay the representation of the fields of view of the one or more cameras, the display including: At least display a first graphical object at a first location and a second graphical object at a second location, the first graphical object representing one or more estimated spatial attributes of a first physical feature that has been detected in a corresponding portion of the physical environment in the fields of view of the one or more cameras, and the second graphical object representing one or more estimated spatial attributes of a second physical feature that has been detected in the corresponding portion of the physical environment in the fields of view of the one or more cameras; and When displaying the plurality of graphical objects that overlay the representation of the fields of view of the one or more cameras, change one or more visual attributes of the first graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the first physical feature, and change one or more visual attributes of the second graphical object according to a change in a corresponding prediction accuracy of the estimated spatial attributes of the second physical feature.
21. The computer-readable storage medium according to claim 20, wherein the one or more programs include instructions that, when executed by the computer system, cause the computer system to perform the method according to any one of claims 2 to 17.