Interior / exterior building walkthrough image interface

The system integrates interior and exterior image frames with a 3D model using LIDAR data to provide a seamless virtual walkthrough, addressing the challenge of aligning and switching between building views.

JP2026503083APending Publication Date: 2026-01-27OPEN SPACE LABS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025540361
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2024-01-08
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing systems lack an efficient method to integrate and align interior and exterior image frames with a 3D model of a building, limiting the ability to provide a comprehensive virtual walkthrough of environments.

Method used

A system that captures interior image frames using a mobile device and exterior frames using a UAV, aligns them with a 3D model generated from LIDAR data, and provides an interface to seamlessly switch between interior and exterior views based on user interaction.

Benefits of technology

Enables a comprehensive virtual walkthrough of buildings by accurately integrating and switching between interior and exterior views, enhancing user interaction and navigation within the 3D model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503083000001_ABST
    Figure 2026503083000001_ABST
Patent Text Reader

Abstract

The computing device accesses interior image frames captured by the mobile device while the mobile device is moving inside the building. The computing device accesses exterior image frames captured by the UAV while the UAV is navigating the exterior perimeter of the building. The computing device generates a 3D model representing the building based on the image frames. The computing device generates an interface that displays the 3D model in a first interface portion. The computing device identifies displayed portions of the 3D model that correspond to the one or more accessed exterior image frames. The computing device modifies the first interface portion to display an interface element at a position corresponding to the identified portion of the 3D model. In response to a selection of the displayed interface element, the computing device modifies a second interface portion to display the one or more accessed exterior image frames that correspond to the identified portion of the 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0002] The present disclosure relates to generating a model of an environment. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 438,182, filed January 10, 2023, which is incorporated herein by reference in its entirety.

[0003]

[0003] Images of an environment can be useful for ascertaining details associated with the environment without actually visiting that environment. For example, a real estate agent may want to create a virtual tour of a home by capturing a series of photographs of the home's rooms so that parties can virtually tour the home. Similarly, a construction company may want to monitor progress at a construction site by capturing images of the construction site at various points during construction and comparing images captured at different times. Summary of the Invention

[0004] A system accesses interior image frames captured by a mobile device while the mobile device is moving through the interior of a building, and exterior image frames captured by an unmanned aerial vehicle (UAV) while the UAV is navigating around the exterior of the building. The system generates a 3D model representing the building based on these image frames. The system interfaces to display the 3D model in a first interface portion. The system identifies portions of the displayed 3D model that correspond to one or more of the accessed exterior image frames. The system modifies the first interface portion to display an interface element in a position corresponding to the identified portion of the 3D model. In response to selection of the displayed interface element, the system modifies a second interface portion to display one or more of the accessed exterior image frames that correspond to the identified portion of the 3D model. [Brief explanation of the drawings]

[0005] [Figure 1]

[0005] FIG. 1 illustrates a system environment for a spatial indexing system, according to one embodiment. [Figure 2A]

[0006] FIG. 2 is a block diagram of a pass module, according to one embodiment. [Figure 2B]

[0007] FIG. 2 is a block diagram of a model generation module, according to one embodiment. [Figure 3]

[0008] 4 is a flowchart illustrating an exemplary method for automatic spatial indexing of frames using features in a floorplan, according to one embodiment. [Figure 4]

[0009] 1 is a flowchart illustrating an exemplary method for generating an interface, according to one embodiment. [Figure 5]

[0010] FIG. 1 illustrates an exemplary interface, according to one embodiment. [Figure 6]

[0011] FIG. 6 illustrates a variation of the interface of FIG. 5, according to one embodiment. [Figure 7]

[0012] FIG. 6 illustrates a variation of the interface of FIG. 5, according to one embodiment. [Figure 8]

[0013] FIG. 1 illustrates an exemplary interface according to one embodiment. [Figure 9]

[0014] FIG. 1 illustrates a computer system for implementing embodiments herein, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0006] I. Overview

[0015] The spatial indexing system receives a video including a series of image frames representing an environment and aligns the image frames with a 3D model of the environment generated using light detection and ranging (LIDAR) data. The image frames are captured by a video capture system moving through the environment along a path. The LIDAR data is collected by the LIDAR system, and the spatial indexing system generates a 3D model of the environment based on the LIDAR data received from the LIDAR system. The spatial indexing system aligns the images with the 3D model. In some embodiments, the LIDAR system is integrated with the video capture system such that the image frames and the LIDAR data are captured simultaneously and synchronized in time. Based on the time synchronization, the spatial indexing system can determine the location where each of the image frames was captured and the portion of the 3D model to which the image frame corresponds. In other embodiments, the LIDAR system is separate from the video capture system, and the spatial indexing system can use feature vectors associated with the LIDAR data and feature vectors associated with the image frames for alignment.

[0007]

[0016] The spatial indexing system generates an interface having a first interface portion for displaying the 3D model and a second interface portion for displaying image frames. The spatial indexing system may receive an interaction from a user indicating a portion of the 3D model to be displayed. For example, the interaction may include selecting a waypoint icon associated with a location within the 3D model or selecting an object within the 3D model. The spatial indexing system identifies image frames associated with the selected portion of the 3D model and displays the corresponding image frames in the second interface portion. When the spatial indexing system receives another interaction indicating a different portion of the 3D model to be displayed, the interface is updated to display the other portion of the 3D model in the first interface and to display different image frames associated with the other portion of the 3D model.

[0008]

[0017] In some embodiments, the spatial indexing system accesses interior image frames captured by a mobile device while the mobile device is moving inside a building. The spatial indexing system accesses exterior image frames captured by a UAV (or other external image capture system, although for simplicity's sake we will refer to a UAV herein) while the UAV is navigating the exterior perimeter of the building. The spatial indexing system generates a 3D model representing the building based on the image frames. The spatial indexing generates an interface in a first interface portion that displays the 3D model. The spatial indexing system identifies displayed portions of the 3D model that correspond to one or more of the accessed exterior image frames. The spatial indexing system modifies the first interface portion to display an interface element at a position corresponding to the identified portion of the 3D model. When the displayed interface element is selected, the spatial indexing system modifies the second interface portion to display the accessed exterior image frames that correspond to the identified portion of the 3D model.

[0009]

[0018] In some embodiments, the spatial indexing system accesses interior image frames and / or depth information captured by a mobile device while the mobile device is moving inside a building. The spatial indexing system accesses exterior image frames captured by a UAV while the UAV is navigating the exterior perimeter of the building. The spatial indexing system accesses a floor plan of the building. The spatial indexing system aligns the interior and exterior image frames to the accessed floor plan. The spatial indexing system generates an interface that displays one or more interior image frames in a first interface portion. The spatial indexing system uses the floor plan to identify displayed interior image frames that correspond to one or more of the accessed exterior image frames. The spatial indexing system modifies the first interface portion to display an interface element at a position corresponding to the identified displayed interior frame. In response to a selection of the displayed interface element, the spatial indexing system modifies the second interface portion to display the accessed exterior image frame that corresponds to the identified displayed interior frame.

[0010]

[0019] In some embodiments, the spatial indexing system accesses interior image frames and / or depth information captured by a mobile device while the mobile device is moving inside the building. The spatial indexing system accesses exterior image frames captured by a UAV while the UAV is navigating the exterior perimeter of the building. The spatial indexing system aligns the interior image frames and the exterior image frames to a coordinate system. The spatial indexing system generates an interface that displays one or more interior image frames in a first interface portion. The spatial indexing system uses the coordinate system to identify displayed interior image frames that correspond to one or more of the accessed exterior frames. The spatial indexing system modifies the first interface portion to display an interface element at a position corresponding to the identified and displayed interior frame. In response to a selection of the displayed interface element, the spatial indexing system modifies the second interface portion to display the accessed exterior image frame that corresponds to the identified and displayed interior frame.

[0011] II. System Environment

[0020] Figure 1 is a diagram illustrating a system environment 100 for a spatial indexing system, according to one embodiment. In the embodiment shown in Figure 1, the system environment 100 includes a video capture system 110, a UAV 118, a network 120, a spatial indexing system 130, a LIDAR system 150, and a client device 160. Although a single video capture system 110, a single LIDAR system 150, and a single client device 160 are shown in Figure 1, in some implementations, the spatial indexing system 130 interacts with multiple video capture systems 110, multiple LIDAR systems 150, and / or multiple client devices 160.

[0012]

[0021] Video capture system 110 collects one or more of frame data, motion data, and position data as video capture system 110 moves along a path. In the embodiment shown in FIG. 1 , video capture system 110 includes camera 112, motion sensor 114, and position sensor 116. Video capture system 110 may be implemented as a device having a form factor suitable for moving along a path. In one embodiment, video capture system 110 is a portable device that a user physically moves along a path, such as a wheeled cart or a device attached to or integrated with an object worn on the user's body (e.g., a backpack or helmet). In another embodiment, video capture system 110 is attached to or integrated into a vehicle. The vehicle may be, for example, a wheeled vehicle (e.g., a wheeled robot) or an aerial vehicle (e.g., a UAV 118, a quadcopter drone, etc.) and may be configured to navigate autonomously along a pre-determined route or to be controlled by a human user in real time. In some embodiments, video capture system 110 is part of a mobile computing device, such as a smartphone, tablet computer, or laptop computer, and can be carried by a user and used to capture video as the user moves through an environment along a path.

[0013]

[0022] Camera 112 collects video including a series of image frames as video capture system 110 moves along a path. In some embodiments, camera 112 is a 360-degree camera that captures 360-degree frames. Camera 112 may be implemented by positioning multiple non-360-degree cameras within video capture system 110 so that the non-360-degree cameras are oriented at various angles relative to each other and configuring the multiple non-360-degree cameras to capture frames of the environment substantially simultaneously from each angle of the non-360-degree cameras. The image frames can then be combined to form a single 360-degree frame. Camera 112 may then be implemented by capturing frames substantially simultaneously from two 180-degree panoramic cameras oriented in opposite directions. In other embodiments, camera 112 has a narrow field of view and is configured to capture typical 2D images instead of 360-degree frames.

[0014]

[0023] The frame data captured by video capture system 110 may further include frame timestamps, which are data corresponding to the time at which each of the frames was captured by video capture system 110. As used herein, frames are captured substantially simultaneously if they are captured within a threshold time interval (e.g., within 1 second, within 100 milliseconds, etc.) of each other.

[0015]

[0024] In one embodiment, the camera 112 captures walk-through video as the video capture system 110 moves through the environment. The walk-through video includes a series of video frames, which can be captured at any frame rate, such as a high frame rate (e.g., 60 frames per second) or a low frame rate (e.g., 1 frame per second). Capturing a sequence of image frames at a higher frame rate generally produces more stable results, while capturing a sequence of image frames at a lower frame rate allows for reduced data storage and transmission. In another embodiment, the camera 112 captures a sequence of still frames separated by a fixed time interval. In yet another embodiment, the camera 112 captures a single image frame. The motion sensor 114 and the position sensor 116 collect motion data and position data, respectively, while the camera 112 captures frame data. The motion sensor 114 may include, for example, an accelerometer and a gyroscope. The motion sensor 114 may also include a magnetometer, which measures the direction of the magnetic field surrounding the video capture system 110.

[0016]

[0025] The position sensor 116 may include a global navigation satellite system (e.g., a GPS receiver) receiver that determines the latitude and longitude coordinates of the video capture system 110. In some embodiments, the position sensor 116 additionally or alternatively includes an indoor positioning system (IPS) receiver that determines the location of the video capture system based on signals received from transmitters installed at known locations within the environment. For example, multiple radio frequency (RF) transmitters that transmit RF fingerprints are placed throughout the environment, and the position sensor 116 also includes a receiver that detects the RF fingerprints and estimates the location of the video capture system 110 within the environment based on the relative strength of the RF fingerprints.

[0017]

[0026] 1 includes a camera 112, a motion sensor 114, and a position sensor 116, in other embodiments, some of the components 112, 114, 116 may be omitted from the video capture system 110. For example, one or both of the motion sensor 114 and the position sensor 116 may be omitted from the video capture system.

[0018]

[0027] In some embodiments, video capture system 110 is implemented as part of a computing device (e.g., computer system 600 shown in FIG. 6 ) that also includes a storage device for storing captured data and a communication interface for transmitting the captured data over network 120 to spatial indexing system 130. In one embodiment, video capture system 110 stores the captured data locally as video capture system 110 moves along a path, and after data collection is complete, the data is transmitted to spatial indexing system 130. In another embodiment, video capture system 110 transmits the captured data to spatial indexing system 130 in real time as system 110 moves along a path.

[0019]

[0028] Video capture system 110 communicates with other systems via network 120. Network 120 may include any combination of local area and / or wide area networks using both wired and / or wireless communication systems. In one embodiment, network 120 uses standard communication technologies and / or protocols. For example, network 120 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of network protocols used to communicate over network 120 include multiprotocol label switching (MPLS), transmission control protocol / internet protocol (TCP / IP), hypertext transfer protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Network 120 may also be used to deliver push notifications via various push notification services such as APPLE Push Notification Service (APNS) and GOOGLE Cloud Messaging (GCM). Data exchanged over network 110 may be represented using any suitable format, such as Hypertext Markup Language (HTML), Extensible Markup Language (XML), or JavaScript Object Notation (JSON). In some embodiments, all or part of the communication links of network 120 may be encrypted using any suitable technique or techniques.

[0020]

[0029] Continuing with reference to Figure 1, the UAV 118 may interact with external systems, such as a spatial indexing system 130, via a network 120. The UAV 118 may be equipped with or integrated with a video capture system 110 that captures aerial image frames of an environment, such as a building. For example, the UAV 118 may capture images of a building from outside the building.

[0021]

[0030] In some embodiments, a camera 112 is mounted on the UAV and is responsible for capturing exterior image frames of the building from different angles. For example, the camera 112 may be a multi-lens camera system that provides multiple viewpoints and covers a wide field of view. In some embodiments, the UAV may be equipped with a depth-sensing system, such as a LIDAR sensor, a structured light sensor, or a time-of-flight sensor. These depth-sensing systems may capture depth information in the form of a depth map, which may then be integrated with the exterior image frames to construct a detailed 3D model.

[0022]

[0031] In some embodiments, the UAV 118 has built-in motion sensors 114, such as an accelerometer and a gyroscope, that measure linear acceleration and rotational motion, respectively. Data obtained from these sensors helps estimate and correct the UAV's position and attitude during flight. In some embodiments, the UAV 118 may have a position sensor 116, such as a GPS, that provides precise position information during flight. This data can be used to help the UAV estimate its position relative to a building and to precisely align external image frames with a 3D model.

[0023]

[0032] The UAV 118 may include a propulsion system consisting of an electric motor, propellers, and a battery. This system provides the thrust necessary to keep the UAV airborne, guide its flight path, and acquire external image frames and other relevant data while maneuvering around a building. The UAV 118 may also have a flight controller that serves as the UAV's central processing and control unit. It processes data from various sensors, manages the propulsion and stabilization systems, and transmits data, such as captured image frames and other sensor data, through communication with an external system. The UAV 118 may also have a communications module that enables wireless data transmission between the UAV and the external system. Communication may occur via Wi-Fi, radio frequency, or other wireless communication protocols. This module may transmit captured image frames, depth maps, and sensor data to a system for further processing and / or 3D model generation.

[0024]

[0033] In some embodiments, the UAV 118 may use a camera and depth sensing system to collect image frames and depth information (if available) while flying around a building. This information, along with the UAV's position and attitude data obtained from the UAV's motion and position sensors, may be transmitted over the network 120 to the spatial indexing system 130.

[0025]

[0034] Continuing with reference to FIG. 1 , a light detection and ranging (LIDAR) system 150 uses a laser 152 and a detector 154 to collect three-dimensional data representative of an environment as the LIDAR system 150 moves through the environment. The laser 152 emits a laser pulse, and the detector 154 detects when the laser pulse returns to the LIDAR system 150 after being reflected by multiple points on objects or surfaces in the environment. The LIDAR system 150 also includes a motion sensor 156 and a position sensor 158 that indicate the movement and position of the LIDAR system 150, which can be used to determine the direction from which the laser pulse is emitted. The LIDAR system 150 generates LIDAR data associated with the laser pulse detected after being reflected from the surface of an object or water in the environment. The LIDAR data may include a known direction from which the laser pulse was emitted and a set of (x, y, z) coordinates determined based on the duration between the emission of the laser 152 and its detection by the detector 154. The LIDAR data may also include other attribute data, such as the intensity of the detected laser pulse. In other embodiments, LIDAR system 150 may be replaced with another depth-sensing system. Exemplary depth-sensing systems include radar systems, 3D camera systems, and the like.

[0026]

[0035] In some embodiments, LIDAR system 150 is integrated with video capture system 110. For example, LIDAR system 150 and video capture system 110 may be components of a smartphone configured to capture video and LIDAR data. Video capture system 110 and LIDAR system 150 may be operated simultaneously, such that video capture system 110 captures video of the environment while LIDAR system 150 collects LIDAR data. When video capture system 110 and LIDAR system 150 are integrated, motion sensor 114 may be the same as motion sensor 156, and position sensor 116 may be the same as position sensor 158. LIDAR system 150 and video capture system 110 may be aligned, and points in the LIDAR data may be mapped to pixels in an image frame captured simultaneously with the points, such that the points are associated with image data (e.g., RGB values).

[0027]

[0036] LIDAR system 150 may also collect timestamps associated with the points. Thus, the image frames and LIDAR data may be associated with each other based on the timestamps. As used herein, a timestamp in the LIDAR data may correspond to the time a laser pulse is emitted toward a point or the time a laser pulse is detected by detector 154. That is, for a timestamp associated with an image frame indicating the time the image frame was captured, one or more points in the LIDAR data may be associated with the same timestamp. In some embodiments, LIDAR system 150 may be used while video capture system 110 is not in use, or vice versa. In some embodiments, LIDAR system 150 is a separate system from video capture system 110. In such embodiments, the path of video capture system 110 may be different from the path of LIDAR system 150.

[0028]

[0037] Continuing with reference to FIG. 1 , spatial indexing system 130 receives image frames captured by video capture system 110 and LIDAR collected by LIDAR system 150, performs a spatial indexing process to automatically identify the spatial location where each of the image frames and LIDAR data was captured, and aligns the image frames to a 3D model generated using the LIDAR data. After aligning the image frames to the 3D model, spatial indexing system 130 provides a visualization interface that allows client device 160 to select portions of the 3D model and view corresponding image frames side-by-side together. In the embodiment shown in FIG. 1 , spatial indexing system 130 includes a path module 132, a path storage 134, a floor plan storage 136, a model generation module 138, a model storage 140, a model integration module 142, an interface module 144, and a query module 146. In other embodiments, spatial indexing system 130 may include fewer, different, or additional modules.

[0029]

[0038] The path module 132 receives image frames and other locations in the walk-through video and motion data collected by the video capture system 110 and determines a path for the video capture system 110 based on the received frames and data. In one embodiment, the path is defined as a 6D camera pose for each frame of a walk-through video that includes a series of frames. The 6D camera pose for each frame is an estimate of the relative position and orientation of the camera 112 when the image frame was captured. The path module 132 may store the path in a path storage device 134.

[0030]

[0039] In one embodiment, path module 132 uses a SLAM (simultaneous localization and mapping) algorithm to (1) simultaneously determine a path estimate by inferring the position and orientation of camera 112, and (2) model the environment using direct methods or landmark features extracted from a walkthrough video, which is a sequence of frames (e.g., oriented FAST and rotated BRIEF (ORB), scale-invariant feature transform (SIFT), speeded up robust features (SURF), etc.). Path module 132 outputs vectors of six-dimensional (6D) camera poses over time, with one 6D vector (three dimensions of position, three dimensions of orientation) for each frame in the sequence, and the 6D vectors may be stored in path storage 134.

[0031]

[0040] The spatial indexing system 130 may also include a floor plan storage device 136 that stores one or more floor plans, such as a floor plan of the environment captured by the video capture system 110. As referred to herein, a floor plan is a to-scale, two-dimensional (2D) diagrammatic representation of an environment (e.g., a building or portion of a structure) from a top-down perspective. In an alternative embodiment, the floor plan may be a 3D model of the expected completed building instead of a 2D diagram (e.g., a Building Information Modeling (BIM) model). The floor plan may be annotated to specify the locations, dimensions, and types of physical objects expected to be present in the environment. In some embodiments, the floor plan is manually annotated by a user associated with the client device 160 and provided to the spatial indexing system 130. In other embodiments, the floor plan is annotated by the spatial indexing system 130 using a machine learning model trained using a training dataset of annotated floor plans to identify the locations, dimensions, and object types of physical objects expected to be present in the environment. Different portions of a building or structure may be represented by separate floor plans. For example, spatial indexing system 130 may store a separate floor plan for each floor of a building, unit, or substructure.

[0032]

[0041] Model generation module 138 generates a 3D model of the environment. In some embodiments, the 3D model is based on image frames captured by video capture system 110. To generate the 3D model of the environment based on the image frames, model generation module 138 may use methods such as structure from motion (SfM), simultaneous localization and mapping (SLAM), monocular depth map generation, or other methods. The 3D model may be generated using image frames from a walk-through video of the environment, the relative position of each image frame (as indicated by the 6D pose of the image frames), and (optionally) the absolute position of each image frame with respect to the floor plan of the environment. Image frames from video capture system 110 may be stereo images that can be combined to generate the 3D model. In some embodiments, model generation module 138 generates a 3D point cloud based on the image frames using photogrammetry. In some embodiments, model generation module 138 generates the 3D model based on LIDAR data from system 150. The model generation module 138 may process the LIDAR data to generate a point cloud, which may have higher resolution compared to a 3D model generated using the image frames. After generating the 3D model, the model generation module 138 stores the 3D model in the model storage device 140.

[0033]

[0042] In one embodiment, model generation module 136 receives a frame sequence and its corresponding path (e.g., 6D pose vectors defining a 6D pose for each frame in a walkthrough video, which is a sequence of frames) from path module 132 or path storage device 134 and extracts a subset of image frames in the sequence and their corresponding 6D poses for inclusion in the 3D model. For example, if the walkthrough video, which is a sequence of frames, is a frame in a video captured at 30 frames per second, model generation module 136 subsamples the image frames by extracting frames and their corresponding 6D poses at 0.5 second intervals. An embodiment of model generation module 136 is described in more detail below with respect to FIG. 2B.

[0034]

[0043] 1 , the 3D model is generated by model generation module 138 within spatial indexing system 130. However, in alternative embodiments, model generation module 138 may be generated by a third-party application (e.g., an application installed on a mobile device that includes video capture system 110 and / or LIDAR system 150). Image frames captured by video capture system 110 and / or LIDAR data collected by LIDAR system 150 may be transmitted over network 120 to a server associated with the application, which processes the data to generate the 3D model. Spatial indexing system 130 may then access the generated 3D model, align the 3D model with other data associated with the environment, and present the aligned representation to one or more users.

[0035]

[0044] Model integration module 142 integrates the 3D model with other data describing the environment. The other types of data may include one or more images (e.g., image frames from video capture system 110), 2D floor plans, diagrams, and annotations describing characteristics of the environment. Model integration module 142 determines the similarity between the 3D model and the other data and aligns the other data with associated portions of the 3D model. Model integration module 142 may determine to which portions of the 3D model the other data corresponds and may store identifiers associated with the determined 3D portions associated with the other data.

[0036]

[0045] In some embodiments, model integration module 142 may align a 3D model generated based on the LIDAR data to one or more image frames based on time synchronization. As described above, video capture system 110 and LIDAR system 150 may be integrated into a single system that simultaneously captures image frames and LIDAR data. For each image frame, model integration module 142 may determine the timestamp at which the image frame was captured and identify a set of points in the LIDAR data associated with the same timestamp. Model integration module 142 may then determine which portion of the 3D model contains the identified set of points and align the image frame with that portion. Furthermore, model integration module 142 may map pixels in the image frame to the set of points in the LIDAR data.

[0037]

[0046] In some embodiments, the model integration module 142 may align a point cloud generated using LIDAR data (hereinafter referred to as the “LIDAR point cloud”) to another point cloud generated based on image frames (hereinafter referred to as the “low-resolution point cloud”). This method may be used when the LIDAR system 150 and the video capture system 110 are separate systems. The model integration module 142 may generate feature vectors for each point in the LIDAR point cloud and each point in the low-resolution point cloud (e.g., using ORB, SIFT, or HardNET). The model integration module 142 may determine feature distances between the feature vectors and matching point pairs between the feature distance-based LIDAR point cloud and the feature distance-based low-resolution point cloud. The 3D poses between the LIDAR point cloud and the low-resolution point cloud are determined to generate a larger number of geometric inliers for the point pairs, for example, using random sample consensus (RANSAC) or nonlinear optimization. As the low-resolution point clouds are generated for each image frame, the LIDAR point clouds are also aligned to each image frame.

[0038]

[0047] In some embodiments, the model integration module 142 may align the 3D model to the diagram or one or more image frames based on annotations associated with the diagram or one or more image frames. The annotations may be provided by a user or determined by the spatial indexing system 130 using an image recognition or machine learning model. The annotations may describe characteristics of objects or surfaces in the environment, such as dimensions or object type. The model integration module 142 may extract features in the 3D model and compare the extracted features to the annotations. For example, if the 3D model represents a room in a building, features extracted from the 3D model may be used to determine the dimensions of the room. The determined dimensions may be compared to a floor plan of a construction site annotated with the dimensions of various rooms in the building, and the model integration module 142 may identify rooms within the floor plan that match the determined dimensions. In some embodiments, the model integration module 142 may perform 3D object detection on the 3D model and compare the output of the 3D object detection to the output from the image recognition or machine learning model based on the diagram or one or more images.

[0039]

[0048] In some embodiments, model integration module 142 integrates the 3D model with other data, such as external image frames received and / or stored in spatial indexing system 130 or external image frames captured at UAV 118 of FIG. 1 . For example, by processing the external image frames, model integration module 142 may identify portions of the 3D model that correspond to the external image frames. This identification process includes feature extraction, feature matching, alignment based on matched features and estimated camera poses, and mapping of the external image frames to the corresponding portions of the 3D model. These steps may provide a consistent and accurate spatial representation of the integrated 3D model and the external image frames.

[0040]

[0049] In some embodiments, the model integration module 142 processes external image frames captured by the UAV to extract characteristic features such as points, edges, or object boundaries. Feature extraction algorithms such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), or Oriented FAST and Rotated BRIEF (ORB) can be employed for this purpose.

[0041]

[0050] In some embodiments, model integration module 142 may identify corresponding features in the 3D model by searching for similarities between features extracted from the external image frames and features in the 3D model. This may be achieved by using feature matching algorithms such as k-Nearest Neighbors (KNN), Fast Approximate Nearest Neighbors (FLANN), or Bag-of-Words-based techniques. By finding these correspondences, model integration module 142 may associate specific portions of the 3D model with the captured external image frames.

[0042]

[0051] In some embodiments, using the corresponding feature points and estimated camera pose of the UAV, the model integration module 142 aligns the external image frames with the 3D model. This alignment accurately maintains the spatial relationship between the external image frames and the 3D model. A bundle adjustment algorithm can be used to optimize and refine the alignment by minimizing reprojection error and ensure that feature points are consistently positioned in both the image frames and the 3D model. In some cases, manual alignment or an iterative closest point (ICP) algorithm can also be used to further refine the placement and orientation of the 3D model based on the image frames.

[0043]

[0052] In some embodiments, after aligning the external image frames with the 3D model, the model integration module 142 may identify displayed portions of the 3D model that correspond to one or more external image frames. During this step, the model integration module 142 may generate a mapping or index that associates the external image frames with the corresponding portions of the 3D model.

[0044]

[0053] In some embodiments, the 3D model may be manually aligned with the diagram based on input from a user. The 3D model and diagram may be presented on a client device 160 associated with the user, and the user may select a location within the diagram that indicates the location corresponding to the 3D model. For example, the user may place a pin at a location in a floor plan that corresponds to the LIDAR data.

[0045]

[0054] The interface module 144 provides a visualization interface to the client device 160 to present information associated with the environment. The interface module 144 may generate the visualization interface in response to receiving a request from the client device 160 to show one or more models representing the environment. The interface module 144 may initially generate the visualization interface to include a 2D overhead map interface representing a floor plan of the environment from the floor plan storage device 136. The 2D overhead map may be an interactive interface such that clicking a point on the map moves to a portion of the 3D model corresponding to the selected point in space. The visualization interface provides a first-person perspective of portions of the 3D model, allowing the user to pan and zoom around the 3D model and to navigate to other portions of the 3D model by selecting waypoint icons representing the relative locations of the other portions.

[0046]

[0055] The visualization interface also allows a user to select an object within the 3D model, causing the visualization interface to display an image frame corresponding to the selected object. The user may select an object by interacting with a point on the object (e.g., clicking a point on the object). When interface module 144 detects an interaction from the user, interface module 144 sends a signal to query module 146 indicating the location of the point within the 3D model. Query module 146 identifies an image frame aligned with the selected point, and interface module 144 updates the visualization interface to display the image frame. The visualization interface may include a first interface portion for displaying the 3D model and a second interface portion for displaying the image frame.

[0047]

[0056] In some embodiments, interface module 144 may receive a request to measure the distance between selected endpoints on a 3D model or image frame. Interface module 144 may provide identification of the endpoints to query module 146, and query module 146 may determine the (x, y, z) coordinates associated with the endpoints. Query module 146 may calculate the distance between the two coordinates and return the distance to interface module 144. Interface module 144 may update an interface portion to display the requested distance to the user. Similarly, interface module 144 may receive additional endpoints with a request to determine the area or volume of an object.

[0048]

[0057] In some embodiments, the interface module 144 may modify a first interface portion of the interface to provide user interaction and display of an exterior view by displaying interface elements at locations corresponding to portions of the 3D model. The interface module may create interface elements that visually indicate the availability of one or more exterior views associated with portions of the 3D model. The interface elements may take the form of icons, buttons, highlights, shading, tooltips, hotspots, arrows, lines, text labels, or overlay blends. The selection and design of the interface elements may be tailored to a particular building structure, layout, or user preferences. To accurately position the interface elements within the first interface portion, the interface module 144 may use location information associated with both the interior and exterior images. The location information may include GPS coordinates, a universal coordinate system, or building floor plan coordinates. Using the location information, the interface module may place the interface elements at a desired location within the first interface portion. This location corresponds to an identified portion of the 3D model, ensuring that the interface elements are accurately positioned to visually represent that portion of the model that has an available exterior view. After placing the interface elements, interface module 144 may update the content of the first interface portion to include the newly generated interface elements, which may include, for example, rendering the interface elements using an appropriate rendering technique (such as a 2D or 3D graphics library) or updating the DOM (Document Object Model) of the web-based interface to include the new interface elements.

[0049]

[0058] In some embodiments, the interface module 144 attaches event listeners or input handlers to the newly created interface elements to monitor user interactions (e.g., clicks or taps). These listeners or handlers trigger responses when the user interacts with the interface elements, allowing the system to update a second interface portion of the interface with an image frame corresponding to the building's exterior.

[0050]

[0059] In response to the selection of the interface element, the interface module 144 may modify a second interface portion of the interface to display an image frame of the building's exterior that corresponds to the portion of the 3D model. For example, when a user selects an interface element, the interface module 144 looks up location information that corresponds to the portion of the building that the interface element represents. This information may include GPS coordinates, a universal coordinate system, or building floor plan coordinates. Using the location information, the interface module 144 may identify the external image that corresponds to the interface element. For example, this process may include comparing the interface element's location information with the location information of each image in the external image frame. Based on this comparison, the system may identify the relevant external image that is linked to the interface element's location.

[0051]

[0060] After identifying the external images corresponding to the positions of the interface elements, the interface module 144 may display them in a second interface portion of the interface. Updating the content of the interface or displaying the selected image within the second interface portion may be accomplished using appropriate rendering techniques and graphics libraries, such as 2D or 3D graphics libraries. In some embodiments, the interface module 144 may generate additional interface elements (e.g., buttons, icons, or sliders) within the second interface portion to provide controls for the user to switch or navigate through the image frames. Event listeners or input handlers may continuously monitor user interaction with these additional interface elements. When the user interacts with these elements, the system may modify the content of the second interface portion to switch or control the image frames in response to the user input. In some embodiments, the interface module 144 may modify the second interface portion to display a corresponding interior view of the building. This may be accomplished by updating the content of the second interface portion to add the relevant interior image based on location information associated with the exterior view.

[0052]

[0061] Client device 160 may be any mobile computing device, such as a smartphone, tablet computer, or laptop computer, or a non-mobile computing device, such as a desktop computer, that can connect to network 120 and be used to access spatial indexing system 130. Client device 160 displays an interface to a user on a display device, such as a screen, and receives user input that allows the user to interact with the interface. An exemplary implementation of a client device is described below with reference to computer system 900 of FIG. 9.

[0053] III. Route Generation Overview

[0062] 2A shows a block diagram of the path module 132 of the spatial indexing system 130 shown in FIG. 1, according to one embodiment. The path module 132 receives input data (e.g., a sequence of frames 212, motion data 214, position data 223, floor plan 257) captured by the video capture system 110 and the LIDAR system 150 and generates a path 226. In the embodiment shown in FIG. 2A, the path module 132 includes a simultaneous localization and mapping (SLAM) module 216, a motion processing module 220, and a path generation and alignment module 224.

[0054]

[0063] The SLAM module 216 receives the sequence of frames 212 and runs a SLAM algorithm to generate a first path estimate 218. Before running the SLAM algorithm, the SLAM module 216 may perform one or more preprocessing steps on the image frames 212. In one embodiment, the preprocessing steps include extracting features from the image frames 212 by converting the sequence of frames 212 into a sequence of vectors, each of which is a feature representation of a respective frame. In particular, the SLAM module may extract SIFT features, SURF features, or ORB features.

[0055]

[0064] After extracting features, the preprocessing step may also include a segmentation process. The segmentation process divides the walk-through video, which is a sequence of frames, into segments based on the quality of each feature of the image frame. In one embodiment, the feature quality of a frame is defined as the number of features extracted from the image frame. In this embodiment, the segmentation step classifies each of the frames as having high or low feature quality based on whether the feature quality of the image frame is above or below a threshold. (That is, frames with feature quality above the threshold are classified as high quality, and frames with feature quality below the threshold are classified as low quality.) Low feature quality may be caused, for example, by excessive subject motion or poor lighting conditions.

[0056]

[0065] After classifying the image frames, the segmentation process divides the sequence such that consecutive frames with high feature quality are combined into segments and frames with low feature quality are not included in any segment. For example, consider a path that follows a dimly lit hallway, entering and exiting a series of brightly lit rooms. In this example, image frames captured in each of the rooms are likely to have high feature quality, while image frames captured in the hallway are likely to have low feature quality. As a result, the segmentation process divides the walk-through video, which is a sequence of frames, such that each sequence of consecutive frames captured in the same room is divided into a single segment (resulting in a separate segment for each room), while image frames captured in the hallway are not included in any segment.

[0057]

[0066] After the pre-processing step, the SLAM module 216 runs a SLAM algorithm to generate a first estimate 218 of the path. In one embodiment, the first estimate 218 is also a vector of 6D camera pose over time, with one 6D vector for each frame in the sequence. In an embodiment where the pre-processing step includes segmenting the walk-through video, which is a sequence of frames, the SLAM algorithm is run separately for each of the segments to generate path segments for each of the segments of frames.

[0058]

[0067] Motion processing module 220 receives motion data 214 collected as video capture system 110 moves along a path and generates a second estimate 222 of the path. Similar to first estimate 218 of the path, second estimate 222 may also be represented as a 6D vector of camera pose over time. In one embodiment, motion data 214 includes acceleration and gyroscope data collected by an accelerometer and a gyroscope, respectively, and motion processing module 220 generates second estimate 222 by performing a dead reckoning process on the motion data. In embodiments in which motion data 214 also includes data from a magnetometer, the magnetometer data may be used in addition to or instead of the gyroscope data to determine changes in orientation of video capture system 110.

[0059]

[0068] The data generated by many consumer-grade gyroscopes contains a time-varying bias (also called drift) that can affect the accuracy of the second path estimate 222 if the bias is not corrected. In embodiments in which the motion data 214 includes all three types of data described above (accelerometer, gyroscope, and magnetometer data), the motion processing module 220 may use the accelerometer and magnetometer data to detect and correct this bias in the gyroscope data. In particular, the motion processing module 220 determines the direction of the gravity vector (which would typically point in the direction of gravity) from the accelerometer data and uses the gravity vector to estimate the two-dimensional tilt of the video capture system 110. Meanwhile, the magnetometer data is used to estimate the gyroscope's heading bias. Because magnetometer data can be noisy, especially when used in buildings whose interior structures include steel frames, the motion processing module 220 may calculate and use a rolling average of the magnetometer data to estimate the heading bias. In various embodiments, the rolling average may be calculated over a time window of 1 minute, 5 minutes, 10 minutes, or some other duration.

[0060]

[0069] The path generation and alignment module 224 combines the first 218 and second 222 path estimates into a combined estimate of the path 226. In embodiments in which the video capture system 110 also collects position data 223 while moving along the path, the path generation module 224 may also use the position data 223 when generating the path 226. If a floor plan of the environment is available, the path generation and alignment module 224 may also receive the floor plan 257 as an input and align the combined estimate of the path 216 to the floor plan 257.

[0061] IV. Model Generation Overview

[0070] 2B shows a block diagram of the model generation module 138 of the spatial indexing system 130 shown in FIG. 1 , according to one embodiment. FIG. 2B shows a 3D model 266 generated based on image frames. The model generation module 138 receives the path 226 generated by the path module 132, along with a sequence of frames 212 captured by the video capture system 110, a floor plan 257 of the environment, and information 254 about the camera. The output of the model generation module 138 is the 3D model 266 of the environment. In the described embodiment, the model generation module 138 includes a route generation module 252, a route filtering module 258, and a frame extraction module 262.

[0062]

[0071] The route generation module 252 receives the path 226 and the camera information 254 and generates one or more candidate route vectors 256 for each extracted frame. The camera information 254 includes a camera model 254A and a camera height 254B. The camera model 254A is a model that maps each 2D point in a frame (i.e., as defined by a pair of coordinates identifying a pixel in the image frame) to a 3D ray representing the line of sight direction from the camera to that 2D point. In one embodiment, the spatial indexing system 130 stores a separate camera model for each type of camera supported by the system 130. The camera height 254B is the height of the camera relative to the floor of the environment while the walkthrough video, which is a sequence of frames, is being captured. In one embodiment, the camera height is assumed to have a constant value during the image frame capture process. For example, if the camera is mounted on a helmet worn on the user's body, then the height has a constant value equal to the sum of the user's height and the height of the camera relative to the top of the user's head (both numbers may be received as user input).

[0063]

[0072] As referred to herein, a root vector of an extracted frame is a vector that represents the spatial distance between the extracted frame and one of the other extracted frames. For example, a root vector associated with an extracted frame has its end point in the extracted frame and its start point in the other extracted frame, such that adding the root vector to the spatial position of the frame associated with the root vector results in the spatial position of the other extracted frame. In one embodiment, the root vector is calculated by performing vector subtraction to calculate the difference between the three-dimensional positions of the two extracted frames, as indicated by the 6D pose vectors of each of the two extracted frames.

[0064]

[0073] With reference to interface module 144, the route vector for the extracted frame is used after interface module 144 receives 3D model 266 and displays the first-person view of the extracted frame. When displaying the first-person view, interface module 144 renders a waypoint icon at a location within the image frame that represents the location of another frame (e.g., the image frame of the origin of the route vector). In one embodiment, interface module 144 determines the location within the image frame at which to render the waypoint icon that corresponds to the route vector using the following equation:

[0065]

number

[0066]

[0074] In this formula, Mproj is the projection matrix containing the parameters of the camera projection function used for rendering, Mview is an isometry matrix representing the user's position and orientation relative to the current frame, Mdelta is the root vector, Gring is the geometry (list of 3D coordinates) representing the mesh model of the waypoint icon being rendered, and Picon is the geometry of the icon within the first person view of the image frame.

[0067]

[0075] Referring again to the route generation module 138, the route generation module 252 may calculate a candidate route vector 256 between each pair of extracted frames. However, displaying a separate waypoint icon for each candidate route vector associated with a frame may result in a large number of waypoint icons (e.g., dozens) being displayed within a frame, which may be burdensome to the user and make it difficult to distinguish between the individual waypoint icons.

[0068]

[0076] To avoid displaying too many waypoint icons, the route filtering module 258 receives the candidate route vectors 256 and selects a subset of the route vectors, which are displayed route vectors 260, represented from a first-person perspective with corresponding waypoint icons. The route filtering module 256 may select the displayed route vectors 256 based on various criteria. For example, the candidate route vectors 256 may be filtered based on distance (e.g., only route vectors having a length less than a threshold length are selected).

[0069]

[0077] In some embodiments, the route filtering module 256 also receives a floor plan 257 of the environment and filters for candidate route vectors 256 based on features in the floor plan. In one embodiment, the route filtering module 256 uses the features of the floor plan to eliminate any candidate route vectors 256 that pass through walls, resulting in a set of displayed route vectors 260 that point only to locations that are visible in the image frame. This may be done, for example, by extracting a floor plan frame patch from the area of ​​the floor plan surrounding the candidate route vector 256 and inputting the image frame patch into a frame classifier (e.g., a feedforward, deep convolutional neural network) to determine whether a wall is present within the patch. If a wall is present within the patch, then the candidate route vector 256 does not pass through a wall and is not selected as one of the displayed route vectors 260. If a wall is not present, then the candidate route vector does not pass through a wall and may be selected as one of the displayed route vectors 260, subject to any other selection criteria (such as distance) configured by the filtering module 258.

[0070]

[0078] The image frame extraction module 262 receives the sequence of 360-degree frames and extracts some or all of the image frames to generate extracted frames 264. In one embodiment, the sequence of 360-degree frames is captured as frames of a 360-degree walk-through video, and the image frame extraction module 262 generates a separate extracted frame for each frame. As described above with respect to FIG. 1 , the image frame extraction module 262 may also extract a subset of image frames from the walk-through video. For example, if the walk-through video, which is a sequence of frames 212, was captured at a relatively high frame rate (e.g., 30 or 60 frames per second), the image frame extraction module 262 may extract a subset of image frames at regular intervals (e.g., 2 frames per second for the video), thereby providing a more manageable number of extracted frames 264 to be displayed to the user as part of the 3D model.

[0071]

[0079] The floor plan 257, the displayed route vectors 260, the path 226, and the extracted frames 264 are integrated into a 3D model 266. As described above, the 3D model 266 is a representation of an environment that comprises a set of extracted frames 264 for the environment, the relative positions of each of the image frames (as indicated by their 6D poses within the path 226). In the embodiment shown in Figure 2B, the 3D model also includes the floor plan 257, the absolute positions of each of the image frames on the floor plan, and the displayed route vectors 260 for some or all of the extracted frames 264.

[0072] V. Spatial Indexing of Frames Based on Floorplan Features

[0080] As described above, the visualization interface may provide a 2D overhead view map that displays the location of each frame within the floor plan of the environment. In addition to being displayed in the overhead view, the floor plan of the environment may also be used as part of a spatial indexing process to determine the location of each frame.

[0073]

[0081] 3 is a flowchart illustrating an exemplary method 300 for automatic spatial indexing of frames using features in a floor plan, according to one embodiment. In other embodiments, method 300 may include additional, fewer, or different steps, and the steps illustrated in FIG. 3 may be performed in a different order. For example, method 300 may be performed without acquiring a floor plan 330, in which case an integrated estimate of the path is generated 340 without using features in the floor plan.

[0074]

[0082] The spatial indexing system 130 receives 310 a walkthrough video, which is a sequence of frames output from the video capture system 110. The image frames in the sequence are captured as the video capture system 110 moves along a particular path within an environment (e.g., a construction site floor). In one embodiment, the individual image frames are frames captured by a camera on the video capture system (e.g., camera 112 described with respect to FIG. 1). In another embodiment, the viewing angle of the individual image frames is narrower, e.g., 90 degrees.

[0075]

[0083] The spatial indexing system 130 generates a first estimate of a path based on the walkthrough video, which is a sequence of frames. This first estimate of the path is represented, for example, as a 6D vector that specifies a 6D camera pose for each frame in the sequence. In one embodiment, a component of the spatial indexing system 130 (e.g., the SLAM module 216 described with reference to FIG. 2A ) runs a SLAM algorithm on the walkthrough video, which is a sequence of frames, to simultaneously determine the 6D camera pose for each frame and generate a three-dimensional virtual model of the environment.

[0076]

[0084] The spatial indexing system 130 obtains 330 a floor plan of the environment. For example, multiple floor plans (including a floor plan of the environment depicted in a walkthrough video, which is a sequence of received frames) may be stored in floor plan storage 136, and the spatial indexing system 130 accesses floor plan storage 136 to obtain the floor plan of the environment. The floor plan of the environment may not be stored in floor plan storage 136, but may instead be received from a user via video capture system 110 or client device 160.

[0077]

[0085] The spatial indexing system 130 generates 340 an integrated estimate of the path based on the first estimate of the path and the physical objects in the floor plan. After generating the integrated estimate 340, the spatial indexing system 130 generates 350 a 3D model of the environment. For example, the model generation module 138 generates the 3D model by integrating the floor plan, the multiple route vectors, the integrated estimate of the path, and frames extracted from the walkthrough video, which is a sequence of frames, as described above in FIG. 2B .

[0078]

[0086] In some embodiments, spatial indexing system 130 may also receive additional data (separate from the walk-through video, which is a sequence of frames) captured while the video capture system is moving along the path. For example, the spatial indexing system may receive motion data or position data, as described above with reference to FIG. 1. In embodiments in which spatial indexing system 130 receives additional data, spatial indexing system 130 may use the additional data in addition to the floor plan in generating the integrated estimate of the path.

[0079]

[0087] In an embodiment in which spatial indexing system 130 receives motion data along with a walk-through video, which is a sequence of frames, spatial indexing system 130 may perform a dead-reckoning process on the motion data to generate a second estimate of the path, as described above with respect to FIG. 2A. In this embodiment, generating a consolidated estimate of the path 340 includes using portions of the second estimate to fill gaps in the first estimate of the path. For example, the first estimate of the path may be split into path segments due to poor feature quality in some captured frames (resulting in gaps for which the SLAM algorithm cannot generate a reliable 6D pose, as described above with respect to FIG. 2A). In this case, the 6D poses of the second estimate of the path may be used to connect the segments of the first estimate of the path by filling the gaps between the segments of the first estimate of the path.

[0080]

[0088] As noted above, in some embodiments, the method 300 may be performed without obtaining a floorplan 330, and the integrated estimate of the path may be generated 340 without using features of the floorplan. In one of these embodiments, the first estimate of the path is used as the integrated estimate of the path without additional data processing or analysis.

[0081]

[0089] In another form of these embodiments, a combined estimate of the path is generated 340 by generating one or more additional estimates of the path, calculating confidence scores for the 6D poses in each path estimate, and selecting the 6D pose with the highest confidence score at each spatial location on the path. For example, the additional path estimates may include one or more of a second estimate using motion data, a third estimate using data from a GPS receiver, and a fourth estimate using data from an IPS receiver, as described above. As described above, each path estimate is a vector of 6D poses that describes the relative position and orientation for each frame in the sequence.

[0082]

[0090] The 6D pose confidence score is calculated differently for each path estimate. For example, the confidence scores for the path estimates described above may be calculated as follows: the 6D pose confidence score for the first estimate (generated using a SLAM algorithm) represents the quality of the features in the image frame corresponding to the 6D pose (e.g., the number of features detected in the image frame); the 6D pose confidence score for the second estimate (generated using motion data) represents the noise level in the accelerometer, gyroscope, and / or magnetometer data for a time interval before, after, or around the time of the 6D pose; the 6D pose confidence score for the third estimate (generated using GPS data) represents the GPS signal strength of the GPS data used to generate the 6D pose; and the 6D pose confidence score for the fourth estimate (generated using IPS data) represents the IPS signal strength (e.g., RF signal strength) of the IPS data used to generate the 6D pose.

[0083]

[0091] After generating the confidence scores, the spatial indexing system 130 iteratively scans each estimate of the path estimate and selects the 6D pose with the highest confidence score for each frame in the sequence, and the selected 6D pose is output as the 6D pose of the image frame in the integrated estimate of the path. Because the confidence scores of the individual path estimates are calculated differently, the confidence scores of the individual path estimates may be normalized to a common scale (e.g., a scalar value between 0 and 1, where 0 represents the lowest confidence and 1 represents the highest confidence) before the iterative scanning process is performed.

[0084] VI. Interface Generation Overview

[0092] 4 is a flowchart illustrating an example method 400 for generating an interface that integrates image frames of both the interior and exterior of a building, according to one embodiment. In other embodiments, method 400 may include additional, fewer, or different steps, and the steps illustrated in FIG. 4 may be performed in a different order. In some embodiments, method 400 may be performed by a computer system, such as spatial indexing system 130 of FIG. 1. Method 400 may be performed by any suitable system.

[0085]

[0093] Continuing with reference to FIG. 4, the system accesses 410 interior image frames captured by a mobile device as the mobile device moves within a building. In some embodiments, a mobile device equipped with a video capture system (e.g., video capture system 110 shown in FIG. 1) and a depth sensing system (e.g., LIDAR system 150 shown in FIG. 1) captures interior image frames as it moves within a building (e.g., a floor of a construction site) along a camera path. The individual frames may be 360-degree frames captured by a 360-degree camera in the video capture system, such as 360-degree camera 112 described above with respect to FIG. 1.

[0086]

[0094] The depth sensing system may generate a depth map corresponding to each captured image frame. In some embodiments, along with the image frame and depth data, the mobile device collects data from built-in motion sensors, such as motion sensor 114 of FIG. 1, and position sensors, such as position sensor 116 of FIG. 1, to estimate the mobile device's location and orientation within the building.

[0087]

[0095] The captured image frames, depth information, and associated sensor data may be stored in the mobile device's local storage or transmitted directly to an external storage system, such as a cloud storage service, via a network connection. The system may access the stored image frames and depth information from the storage location. For example, this may be done by retrieving the data from cloud storage or other external storage systems via a direct connection with the mobile device or a network connection (e.g., Wi-Fi, cellular, or wired connection).

[0088]

[0096] Continuing with reference to FIG. 4, the system accesses 420 exterior image frames obtained by a UAV (e.g., UAV 118 of FIG. 1) capturing images of a building from outside the building. The UAV may be equipped with a camera that captures the exterior image frames while flying around the building. The captured images may cover various angles of the building's exterior and provide a complete view of the building and its surroundings. Along with capturing the image frames, the UAV may collect data from its built-in motion sensors (e.g., accelerometer, gyroscope) and position sensors (e.g., GPS) to estimate its position and attitude relative to the building during its flight.

[0089]

[0097] The external image frames and associated sensor data stored in the UAV's local storage during flight may be transmitted to an external storage system, such as a cloud storage service or a remote server, via a network connection after the flight is completed. The system may access the stored external image frames from a designated storage location through a direct connection with the UAV or by retrieving the data from cloud storage or other external storage system via a network connection.

[0090]

[0098] In some embodiments, for the lower portion of the building, images may be captured using a mobile device while a user walks near the ground around the perimeter of the building. Similarly, for the upper portion of the building, images may be captured using a UAV while the UAV flies around the perimeter of the building.

[0091]

[0099] Continuing with reference to FIG. 4 , the system generates a 3D model representing a building based on the image frames. In some embodiments, the system integrates depth information with the interior image frames so that the interior image frames and depth information more comprehensively represent the building's interior. The depth data may provide a third dimension that planar image frames lack. In some embodiments, the system may process the interior image frames and, optionally, corresponding depth information (e.g., LIDAR data) to create an interior 3D model. Structure-from-Motion (SFM), Simultaneous Localization and Mapping (SLAM), or other depth estimation techniques may be employed to generate the 3D model. These techniques may utilize depth maps and position and orientation data from capture device sensors to construct an accurate three-dimensional representation of the building's interior.

[0092]

[0100] In some embodiments, the external image frames and depth information captured by the UAV can be processed by the system to generate an external 3D model. This step can be similar to the internal 3D model generation, but can use the external image frames and any corresponding depth information, as well as motion and position data captured by the UAV or other suitable capture device. Techniques such as SFM, SLAM, or other depth estimation techniques can be employed to construct the external 3D model.

[0093]

[0101] After generating the 3D models, the system may map any one of the 3D models to a common coordinate system or floor plan of the building. This step may include transforming the corresponding 3D model to maintain consistency of scale, orientation, and position within the coordinate system of the floor plan or building. The floor plan may be a 2D representation or a 3D model, such as a Building Information Model (BIM).

[0094]

[0102] Alignment of the 3D model and image frames can be achieved using various algorithms that ensure consistency and accuracy of the spatial representation. For example, feature matching techniques can be used to identify points or features common to both the image frames and the 3D model, and then used to accurately position and orient the image frames relative to the 3D model. Bundle adjustment algorithms can optimize the alignment by minimizing reprojection error, thereby ensuring that feature points are consistently positioned in both the image frames and the 3D model. In some cases, manual alignment or iterative closest point (ICP) algorithms can be used to refine the placement and orientation of the 3D model based on the image frames.

[0095]

[0103] In some embodiments, the system may integrate interior and exterior 3D models of a building to create a single, integrated 3D model representing the building. This comprehensive model incorporates both interior and exterior data, allowing users to more effectively interact with and visualize the building. In such embodiments, the interior and exterior 3D models of the building are aligned to identify locations in the interior 3D model that correspond to the same portion of the building's exterior wall in the exterior 3D model. In other embodiments, the system may process only the interior or exterior 3D models.

[0096]

[0104] Continuing with reference to Figure 4, the system generates 440 an interface that displays the 3D model in a first interface portion. In some embodiments, the system's walkthrough interface can be modified to display both interior and exterior representations of the building.

[0097]

[0105] In some embodiments, the system may create an interface design featuring two primary portions: a first portion may be designed for displaying the 3D model, while a second portion may be designed for displaying image frames corresponding to particular regions of the 3D model. Advantageously, this interface structure allows a user to interactively explore the 3D model along with the contextual image frames.

[0098]

[0106] Additionally, the system may configure various interface components, such as buttons, sliders, menus, and display panels, that allow users to interact with the 3D models and image frames. These components are organized and arranged within the interface layout to create an intuitive and user-friendly experience.

[0099]

[0107] In some embodiments, the system may incorporate the 3D model and its corresponding image frame into the interface structure by embedding a graphical representation of the 3D model in a first portion of the interface and displaying the corresponding image frame in a second portion of the interface. This allows the user to seamlessly navigate and visualize the 3D model and image frame within the interface. The system may implement interactive features such as zoom, pan, and rotate options for viewing the 3D model, as well as click or tap events for selecting a portion of the model displayed in the first portion of the interface and displaying the corresponding image frame in the second portion of the interface. The user may also interact with other interface components, such as buttons or menus, to change display settings, view additional information, or navigate between different areas of the 3D model and image frame.

[0100]

[0108] In some embodiments, the system may deploy an interface to a user device (e.g., a desktop computer, laptop, tablet, or smartphone) for visualization and interaction. The interface may be presented through a web browser, a standalone application, or a platform-specific app. The 3D models and image frames may be rendered using an appropriate rendering engine and API (e.g., OpenGL, WebGL, DirectX, or Vulkan) for smooth and responsive visualization and user interaction.

[0101]

[0109] In some embodiments, portions corresponding to the interior and exterior of a building may be identified in interior and exterior images of the building. In some embodiments, location information (e.g., GPS coordinates) may be used to identify interior and exterior images corresponding to the same portion. For example, an interior view of a building's exterior wall may be identified in an interior image of the building using a set of GPS coordinates captured by the device that captured the interior image. In some embodiments, images corresponding to the exterior of the building may be identified by using the GPS coordinates of the interior image to query GPS coordinates associated with the exterior image and identifying the exterior image that is closest to the GPS coordinates of the interior image. In some embodiments, the interior and exterior images may be mapped to a common coordinate system (e.g., using GPS or other localization / alignment techniques). In some embodiments, either the interior or exterior image may be mapped to a building floor plan.

[0102]

[0110] Continuing with reference to Figure 4, the system identifies 450 displayed portions of the 3D model that correspond to the accessed one or more external image frames. In some embodiments, to identify the displayed portions of the 3D model that correspond to the external image frames, the system may extract features from the external image frames, match the features of the external image frames with corresponding features in the 3D model, align the external image frames to the 3D model based on the matched features, and determine the displayed portions of the 3D model that correspond to the external image frames based on the alignment.

[0103]

[0111] The system may process the external image frames and extract characteristic features (e.g., points, edges, or object boundaries) from the images. Feature extraction algorithms such as SIFT, SURF, or ORB may be used for this purpose. For feature matching with the 3D model, the system may identify corresponding features in the 3D model by searching for similarities between features extracted from the external image frames and features in the model. This may be done using feature matching algorithms such as KNN, FLANN, or bag-of-words-based methods. By finding these correspondences, the system may associate specific portions of the 3D model with the external image frames of the building. For example, using the matched features and the estimated camera pose, the system may determine which displayed portions of the 3D model correspond to one or more external image frames. During this step, the system may create a mapping or index that links the external image frames to their associated model portions.

[0104]

[0112] Continuing with reference to FIG. 4 , the system may modify 460 a first portion of the interface to display interface elements at locations corresponding to the identified portions of the 3D model. In some embodiments, the system may generate interface elements that indicate the availability of one or more exterior views associated with the identified portions of the 3D model. The interface elements may act as visual hints or indicators to help the user access the corresponding exterior views of the building. Examples of interface elements may include icons, buttons, highlights, shading, pop-up dialogs, tooltips, hotspots, arrows, lines, text labels, and overlay blends. For example, icons may take the form of a camera, magnifying glass, or other symbol indicating that an exterior view is available. The selection and design of interface elements may be tailored to the particular building structure, layout, or user preferences.

[0105]

[0113] In some embodiments, interface elements may be provided as buttons with clickable or tappable areas labeled to indicate the availability of an exterior view when pressed. Examples of buttons may include rectangles, rounded corners, or circles containing a label or icon. Portions of the 3D model, such as exterior walls, may be highlighted or shaded to indicate that an exterior view is available at that location. The highlighting or shading may change when the user hovers or clicks over the location.

[0106]

[0114] In some cases, when a user hovers or clicks on a particular area, a small dialog or tooltip may appear showing thumbnails or brief descriptions of available external views. Alternatively, hotspots, such as interactive areas within the 3D model, may be provided. These hotspots may change color, glow, or present animations when hovered over to indicate the availability of external views.

[0107]

[0115] Directional indicators, such as arrows or lines, may be used to connect corresponding exterior views to interior portions of the 3D model, guiding the user as to where to click or tap. Text labels may be placed near building exterior walls or at specific locations within the 3D model to inform the user that an exterior view is available in that area. In some cases, the system may composite or overlay an exterior image over an interior image with adjustable transparency, allowing the user to see a view that combines both interior and exterior perspectives. Combinations of elements, such as icons contained within buttons, may also be employed to create an intuitive and user-friendly interface.

[0108]

[0116] Once the interface element is generated, the system may place the element at a location within the first interface portion that corresponds to the identified portion of the 3D model. To achieve accurate placement, the system may use location information associated with both the interior and exterior images (e.g., GPS coordinates, a common coordinate system, or a floor plan).

[0109]

[0117] Continuing with reference to FIG. 4 , the system may modify 470 the second interface portion to display an external image frame corresponding to the identified portion of the 3D model. In some embodiments, the system may continuously monitor user interaction with interface elements within the first interface portion. This may be accomplished using event listeners or input handlers, depending on the programming language or framework employed for the interface. When a user interacts with an interface element (e.g., clicks or taps), the system may detect the activity and trigger a response. This detection may be accomplished through event handlers or callbacks programmed to respond to specific input events associated with the user interaction. If user interaction with an interface element is detected, the system may obtain location information corresponding to the building portion represented by the selected interface element. This location information may include GPS coordinates, a universal coordinate system, or building floor plan coordinates.

[0110]

[0118] Using the location information, the system may identify external images that correspond to the selected interface element. This process may include obtaining location information for the selected interface element and comparing the location information of each external image in the external image frame with the location information for the selected interface element. Based on this comparison, the system may identify relevant external images that match, are close to, or are within a predetermined distance from the location of the interface element. In some embodiments, the predetermined distance may be less than 1 meter, 1 meter, 2 meters, 3 meters, 4 meters, or 5 meters.

[0111]

[0119] After identifying the corresponding external images, the system may display them in the second portion of the interface, which may be accomplished by updating the content of the interface or by rendering the selected images within the second portion of the interface using an appropriate rendering technique, such as a 2D or 3D graphics library.

[0112]

[0120] In some embodiments, the system may generate an additional interface element (e.g., a button, icon, or slider) in the second portion of the interface. This additional interface element may provide the user with controls for switching or navigating the external image frame. The system may place the additional interface element in the second portion of the interface in a location that is easily accessible and visible to the user. The system may continuously monitor the user's interaction with the additional interface element in the second portion of the interface using an event listener or input handler, depending on the framework employed by the interface. Upon detecting that the user has interacted with the additional interface element, the system may modify the second interface portion to provide switching or control over the accessed external image frame. For example, this may be achieved by updating the content of the second interface portion and adjusting the display of the external image in response to user input.

[0113]

[0121] In some embodiments, the system may modify the second interface portion to display a corresponding interior view of the building. This may be achieved by updating the content of the second interface portion based on location information associated with the exterior view and adding relevant interior imagery. These features may allow a user to quickly navigate between exterior and interior views of a particular area of ​​the 3D model, offering a comprehensive understanding and visualization of the building's structure and / or environment.

[0114]

[0122] In some embodiments, the system may identify displayed portions of the 3D model that correspond to one or more internal image frames. This process may include using an algorithm to match features present in the 3D model and the internal image frames. Based on the identification, the system may modify a first interface portion to display an interface element at a location corresponding to the identified portion of the 3D model. For example, the system may modify a first portion of the interface to display an interface element, such that the element appears at a location corresponding to the identified portion of the 3D model. In other words, the interface element may act as a marker or indicator that an internal image corresponding to that portion of the model is available. In response to selecting a displayed interface element, the system may also modify a second interface portion to display an internal image frame corresponding to the identified portion of the 3D model. For example, when a user selects an interface element (engages the element with a mouse click or touch), the system modifies another portion of the interface to display an internal image frame associated with the selected area in the 3D model. This process allows the user to visually associate real-world imagery with the 3D model. Advantageously, this sequence of operations provides the user with a deeper understanding of spatial relationships within the building, as the user can simultaneously see the correspondence between the real image and the 3D spatial model.

[0115]

[0123] 5 illustrates an interface 502 displaying an image of a building interior 510. Within the interface 502, a portion 520 of the exterior wall of a building under construction is shown as seen from the building interior 510. The portion 520 of the exterior wall of the building may correspond to one or more of the exterior images of the building. The interface 502 may then be modified to add an interface element at a location within the image of the exterior wall of the building. The interface element may be in any suitable form, such as, for example, an icon or a button. The interface element may indicate that one or more exterior views of the identified portion of the exterior wall are available for viewing.

[0116]

[0124] 6 shows the interface 502 of FIG. 5 modified to add an interface element 530 at the location of a portion of an exterior wall. The interface element 530 is positioned adjacent to the portion of the exterior wall 520. In some embodiments, the location of the identified portion of the exterior wall may be determined within an interior 3D model or image of the building so that the location of the interface element in the displayed interface does not change significantly as the user "navigates" between different views, positions, or viewpoints within the building's interior.

[0117]

[0125] 7 shows interface 502 of FIG. 6 modified to include interface element 530 from a different perspective in building interior 510. In response to selecting interface element 530, the interface may be modified to include one or more of the building exterior images corresponding to the location of the portion of the building exterior indicated by interface element 530. Interface 502 may be modified to include one or more of the corresponding building exterior images. For example, the interface may be modified so that the building interior is displayed in a first interface portion and the building exterior is displayed in a second interface portion.

[0118]

[0126] FIG. 8 illustrates interface 502 modified to display an exterior image of a building in a location corresponding to interface element 530 shown in FIGS. 6 and 7. For example, the image shown in FIG. 8 may be captured by a UAV. The image shows the exterior wall of a building corresponding to a portion of the exterior wall displayed in the example building interior interface of FIGS. 6 and 7. In the displayed image, additional interface elements 810 and 820 are displayed. The additional interface elements 810 and 820 are selectable (e.g., a user may click on the interface elements). When selected, interface 502 may be modified to include a representation of an interior view of the building in a location corresponding to the selected interface element. For example, a selected interface element may be modified to display in the interface a representation of the building floor corresponding to that element, thereby allowing a user to quickly navigate between interior views of different floors of the building based on the exterior view of the building.

[0119]

[0127] Although this example shows an interface that includes only an exterior view of a building, in practice, a first portion of the interface (e.g., the left half of the interface) may show an interior view of the building, and in response to selection of an interface element corresponding to the exterior wall of the building shown in the first portion, a second portion of the interface (e.g., the right half of the interface) may show an exterior view of the building corresponding to the exterior wall of the building.

[0120]

[0128] If a different interface element displayed in the second portion of the interface is selected (e.g., an interface element displayed on an exterior image of a building), the interior view of the building displayed in the first portion of the interface may be modified to include a representation of the floor corresponding to the selected interface element. Similarly, if an interface element corresponding to a different exterior wall of a building is displayed in the first interface portion and is newly selected, the exterior portion of the building displayed in the second portion of the interface may change to display an image of the different exterior wall corresponding to the newly selected interface element.

[0121]

[0129] It should also be noted that a change in the interior view of a building shown in a first interface portion may also result in a change in the exterior view of the building shown in a second interface portion. For example, if a user modifies the perspective of the interior of the building to the left, the perspective of the exterior of the building may shift to the right, while the portion of the exterior wall of the building shown in each interface portion remains constant. The amount of perspective shift in each interface portion may depend on the relative distance between the image capture device and the exterior wall. For example, if the distance between a first device capturing an interior image of the building and the exterior wall is approximately half the distance between a second device capturing an exterior image of the building and the exterior wall, the angle corresponding to the change in perspective of the exterior image displayed in the interface may be approximately half the angle corresponding to the change in perspective of the interior image displayed in the interface.

[0122] VII. Hardware Components

[0130] Figure 9 is a block diagram illustrating a computer system 900 on which embodiments described herein may be implemented. For example, in the context of Figure 1, video capture system 110, LIDAR system 150, spatial indexing system 130, or client device 160 may be implemented using computer system 900 as described in Figure 9. Video capture system 110, LIDAR system 150, spatial indexing system 130, or client device 160 may also be implemented using a combination of multiple computer systems 900 as described in Figure 9. Computer system 900 may be, for example, a laptop computer, a desktop computer, a tablet computer, or a smartphone.

[0123]

[0131] In one implementation, system 900 includes processing resources 901, main memory 903, read-only memory (ROM) 905, storage device 907, and communication interface 909. System 900 includes at least one processor 901 for processing information and main memory 903, such as random access memory (RAM) or other dynamic storage device, for storing information and instructions executed by processor 901. Main memory 903 may also be used for storing temporary variables or other intermediate information during execution of instructions executed by processor 901. System 900 may also include ROM 905 or other static storage device for storing static information and instructions for processor 901. A storage device 907, such as a magnetic or optical disk, is provided for storing information and instructions.

[0124]

[0132] The communications interface 909 may enable the system 900 to communicate with one or more networks (e.g., network 140) through the use of a network link (wireless or wired). Using the network link, the system 900 may communicate with one or more computing devices and one or more servers. The system 900 may also include a display device 911, such as a cathode ray tube (CRT), LCD monitor, or television set, for displaying graphics and information to a user. An input mechanism 913, such as a keyboard including alphanumeric and other keys, may be coupled to the system 900 for communicating information and command selections to the processor 901. Other non-limiting illustrative examples of the input mechanism 913 include a mouse, a trackball, a touch screen, or cursor direction keys for communicating directional information and command selections to the processor 901 and for controlling cursor movement on the display device 911. Additional examples of the input mechanism 913 include a radio frequency identification (RFID) reader, a barcode reader, a three-dimensional scanner, and a three-dimensional camera.

[0125]

[0133] According to one embodiment, the techniques described herein are performed by system 900 in response to processor 901 executing one or more sequences of one or more instructions contained in main memory 903. Such instructions may be read into main memory 903 from another machine-readable medium, such as storage device 907. Execution of the sequences of instructions contained in main memory 903 causes processor 901 to perform the process steps described herein. In alternative implementations, hardwired circuitry may be used in place of or in combination with software instructions to implement the examples described herein. Thus, the described examples are not limited to any specific combination of hardware circuitry and software.

[0126] VIII. Internal / External Interface Generation

[0134] In some embodiments, the walk-through interfaces described herein may be modified to display representations of both the interior and exterior of a building. For example, we have described above generating a 3D model of the interior of a building based (at least in part) on images and depth information captured by a device as the device moves around the interior of the building. For example, we have described generating a 3D model of the exterior of a building based on images and depth information captured by a UAV as the UAV moves around the exterior of the building. Note that images and / or depth information representing the exterior of the building may be captured using other devices. For example, images of the lower part of a building may be captured by a user with a user mobile device while the user is walking around the exterior of the building, at or near ground level. Similarly, images of the upper part of a building may be captured by a UAV as the UAV flies around the exterior of the building.

[0127]

[0135] Corresponding interior and exterior portions may be identified within the interior and exterior images of a building. In some embodiments, location information (e.g., GPS coordinates) may be used to identify interior and exterior images that correspond to the same portion of a building. For example, an interior view of a building's exterior wall may be identified within the interior image using a set of GPS coordinates captured by the mobile device that captured the interior image. In one embodiment, corresponding exterior images may be identified by querying GPS coordinates associated with the exterior image with the interior GPS coordinate set and identifying the exterior image that is closest to the interior GPS coordinate set. In other embodiments, the interior and exterior images may be mapped to a common coordinate system (e.g., using GPS or other location / alignment techniques). In yet other embodiments, both the interior and exterior images are mapped to a building floor plan.

[0128]

[0136] In some embodiments, in addition to generating an interior 3D model of a building, an exterior 3D model of the building may be generated, for example, using exterior imagery and depth information captured by a UAV flying (e.g., at one or more altitudes) over the exterior of the building. In such embodiments, the interior and exterior 3D models of the building may be aligned, allowing for identification of locations in the interior 3D model and locations in the exterior 3D model that correspond to the same portion of the building's exterior wall.

[0129]

[0137] By identifying images, 3D models, floor plans, or interior portions of a building that correspond to exterior portions of the building within a coordinate system common to the interior and exterior, an interface can be generated that allows a user to switch between or simultaneously view the interior and exterior of a building.

[0130]

[0138] In an interface that displays the exterior wall of a building under construction from inside the building, portions of the building's exterior wall that correspond to one or more exterior images of the building may be identified. The interface may then be modified to include an interface element at a location within the image of the building's exterior wall. The interface element may be in any suitable form, such as an icon or button, that indicates that one or more exterior images of the identified portion of the exterior wall are available for viewing. The interface may be modified to include an interface element at the location of the identified portion of the exterior wall.

[0131]

[0139] In some embodiments, the location of the identified portion of the exterior wall may be determined within a 3D model of the building interior or an image of the building interior such that the location of interface elements within the displayed interface does not change significantly as the user "moves" between different views, positions, or viewpoints within the building interior. The interface may be modified to include interface elements from different viewpoints within the building interior.

[0132]

[0140] In response to selecting the interface element, the interface may be modified to include one or more exterior images of the building corresponding to the location of the identified portion of the building exterior indicated by the interface element. The entire interface may be modified to include the corresponding one or more exterior images of the building. The interface may be modified so that the interior of the building is displayed in a first interface portion and the exterior of the building is displayed in a second interface portion.

[0133]

[0141] The interface may be modified to display an exterior image of the building at a location corresponding to the interface element. The image may display the exterior wall of the building corresponding to the location of the interface element. Additional interface elements may be displayed in the displayed image that, when selected, modify the interface to include a representation of the interior of the building at a location corresponding to the selected interface element. For example, a selected interface element would modify the interface to show a representation of the floor of the building corresponding to the interface element. This allows a user to quickly navigate between interior views of different floors of a building based on an exterior view of the building.

[0134]

[0142] Although the interface may include only an exterior view of the building, in reality a first portion of the interface (e.g., the left half of the interface) may show an interior view of the building, and in response to selection of an interface element corresponding to an exterior wall of the building shown in the first portion of the interface, a second portion of the interface (e.g., the right half of the interface) may show an exterior view of the building corresponding to the exterior wall of the building.

[0135]

[0143] If another interface element displayed in the second portion of the interface (e.g., an interface element displayed on an exterior image of a building) is selected, the interior view of the building shown in the first portion of the interface may be modified to include a representation of the floor corresponding to the selected interface element. Similarly, if an interface element corresponding to another exterior wall of the building shown in the first portion is newly selected, the exterior portion of the building shown in the second portion of the interface may change to show an image of the other exterior wall corresponding to the newly selected interface element.

[0136]

[0144] It should also be noted that a change in the interior view of the building shown in a first interface portion may result in a change in the exterior view of the building shown in a second interface portion. For example, if a user changes the viewpoint of the interior of the building to the left, the viewpoint of the exterior of the building may shift to the right, such that the portion of the exterior wall of the building shown in the respective interface portion remains constant. The amount of viewpoint shift in the respective interface portion may depend on the relative distance between the video capture device and the exterior wall. For example, if the distance between a first device capturing an interior image of the building and the exterior wall is approximately half the distance between a second device capturing an exterior image of the building and the exterior wall, the angle corresponding to the change in viewpoint of the exterior image displayed in the interface may be approximately half the angle corresponding to the change in viewpoint of the interior image displayed in the interface.

[0137] IX. Additional Considerations

[0145] As used herein, the term "comprising," followed by one or more elements, does not exclude the presence of one or more additional elements. The term "or" should be interpreted as an inclusive "or" rather than an exclusive "or" (e.g., "A or B" can refer to "A," "B," or "A and B"). The article "a" or "an" refers to one or more instances of the next following element, unless a single instance is expressly specified.

Claims

1. accessing interior image frames captured by a mobile device while the mobile device moves through the interior of the building; accessing exterior image frames captured by an unmanned aerial vehicle ("UAV") while the UAV navigates around the exterior of the building; generating a 3D model representing the building based on the image frames; generating an interface that displays the 3D model in a first interface portion; identifying a displayed portion of the 3D model that corresponds to one or more accessed external image frames; modifying the first interface portion to display an interface element at a location corresponding to the identified portion of the 3D model; modifying a second interface portion to display one or more accessed external image frames corresponding to the identified portion of the 3D model in response to selection of the displayed interface element; A method comprising:

2. generating the 3D model representing the building based on the image frames, Integrating depth information into the interior image frames such that both the interior image frames and the depth information provide a more comprehensive representation of the interior of the building; generating an internal 3D model based on the internal image frames and the depth information; and mapping the interior 3D model to a coordinate system or floor plan of the building.

3. generating the 3D model representing the building based on the image frames, The method of claim 2 , further comprising aligning the 3D model to the image frame for consistency and accuracy of spatial representation.

4. generating the interface for displaying the 3D model in the first interface portion, defining a structure of an interface, the interface including two parts, one part configured to display the 3D model and another part configured to display the image frame; Integrating the 3D model and the image frames into an interface structure; and deploying the interface to a user device for visualization and interaction by a user.

5. Identifying the displayed portion of the 3D model that corresponds to the accessed one or more external image frames includes: extracting features from the external image frames; matching features of the external image frames with corresponding features in the 3D model; aligning the external image frame to the 3D model based on the matching; and determining the displayed portion of the 3D model that corresponds to the external image frame based on the alignment.

6. Modifying the first interface portion to display the interface element at the location corresponding to the identified portion of the 3D model includes: generating the interface element, the interface element indicating that one or more external views of the identified portion of the 3D model are available for viewing; and placing the interface element at the location corresponding to the identified portion of the 3D model.

7. Modifying the second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model may include: listening for user interactions with the interface element and detecting when a user selects the interface element; In response to detecting the user interaction with the interface element, obtaining location information corresponding to the portion of the building represented by the interface element; using the location information to identify an external image corresponding to the selected interface element; and displaying the identified external image in the second interface portion.

8. Identifying the external image corresponding to the selected interface element includes: obtaining the location information corresponding to the portion of the building represented by the interface element, the location information including GPS coordinates, a coordinate system, or building floor plan coordinates; comparing the position of the interface element with the position information of each of the exterior images in the external image frame; and identifying the external image that corresponds to the selected interface element based on a comparison of the location information.

9. Modifying the second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model may include:

10. The method of claim 1, comprising providing an additional interface element within the second interface portion, selection of the additional interface element causing the second interface portion to switch or control the one or more accessed external image frames.

10. Modifying the second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model may include:

2. The method of claim 1, comprising providing an additional interface element in the second interface portion, selection of the additional interface element causing the second interface portion to display a corresponding interior view of the building.

11. identifying the displayed portion of the 3D model corresponding to one or more of the accessed internal image frames; modifying the first interface portion to display the interface element at the location corresponding to the identified portion of the 3D model; 10. The method of claim 1, further comprising: in response to a selection of the displayed interface element, modifying the second interface portion to display the one or more accessed internal image frames corresponding to the identified portion of the 3D model.

12. a hardware processor; a non-transitory computer-readable storage medium having executable instructions stored thereon that, when executed by the hardware processor, cause the hardware processor to perform steps, the steps comprising: accessing interior image frames captured by a mobile device while the mobile device moves through the interior of the building; accessing exterior image frames captured by an unmanned aerial vehicle ("UAV") while the UAV navigates around the exterior of the building; generating a 3D model representing the building based on the image frames; generating an interface that displays the 3D model in a first interface portion; identifying a displayed portion of the 3D model that corresponds to one or more of the accessed external image frames; modifying the first interface portion to display an interface element at a location corresponding to the identified portion of the 3D model; and in response to a selection of the displayed interface element, modifying a second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model.

13. generating the 3D model representing the building based on the image frames, Integrating depth information into the interior image frames such that both the interior image frames and the depth information provide a more comprehensive representation of the interior of the building; generating an internal 3D model based on the internal image frames and the depth information; and mapping the interior 3D model to a coordinate system or floor plan of the building.

14. generating the 3D model representing the building based on the image frames, The system of claim 13 , further comprising aligning the 3D model to the image frames for consistency and accuracy of spatial representation.

15. generating the interface for displaying the 3D model in the first interface portion, defining a structure of an interface, the interface including two parts, one part configured to display the 3D model and another part configured to display the image frame; Integrating the 3D model and the image frames into an interface structure; and deploying the interface to a user device for visualization and interaction by a user.

16. Identifying the displayed portion of the 3D model that corresponds to the accessed one or more external image frames includes: extracting features from the external image frames; matching features of the external image frames with corresponding features in the 3D model; aligning the external image frame to the 3D model based on the matching; and determining the displayed portion of the 3D model that corresponds to the external image frame based on the alignment.

17. Modifying the first interface portion to display the interface element at the location corresponding to the identified portion of the 3D model includes: generating the interface element, the interface element indicating that one or more external views of the identified portion of the 3D model are available for viewing; and placing the interface element at the location corresponding to the identified portion of the 3D model.

18. Modifying the second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model may include: listening for user interactions with the interface element and detecting when a user selects the interface element; In response to detecting the user interaction with the interface element, obtaining location information corresponding to the portion of the building represented by the interface element; using the location information to identify an external image corresponding to the selected interface element; and displaying the identified external image in the second interface portion.

19. Identifying the external image corresponding to the selected interface element includes: obtaining the location information corresponding to the portion of the building represented by the interface element, the location information including GPS coordinates, a coordinate system, or building floor plan coordinates; comparing the position of the interface element with the position information of each of the exterior images in the external image frame; and identifying the external image that corresponds to the selected interface element based on a comparison of the location information.

20. A non-transitory computer-readable storage medium having stored thereon executable instructions that, when executed by a hardware processor, cause the hardware processor to perform steps, the steps comprising: accessing interior image frames captured by a mobile device while the mobile device moves through the interior of the building; accessing exterior image frames captured by an unmanned aerial vehicle ("UAV") while the UAV navigates around the exterior of the building; generating a 3D model representing the building based on the image frames; generating an interface that displays the 3D model in a first interface portion; identifying a displayed portion of the 3D model that corresponds to one or more of the accessed external image frames; modifying the first interface portion to display an interface element at a location corresponding to the identified portion of the 3D model; and in response to a selection of the displayed interface element, modifying a second interface portion to display the one or more accessed external image frames corresponding to the identified portion of the 3D model.

21. accessing interior image frames captured by a mobile device while the mobile device moves through the interior of the building; accessing exterior image frames captured by an unmanned aerial vehicle ("UAV") while the UAV navigates around the exterior of the building; accessing a floor plan of the building; aligning the internal image frame and the external image frame to the accessed floorplan; generating an interface that displays one or more of the internal image frames in a first interface portion; using the floor plan to identify displayed internal image frames corresponding to one or more of the accessed external image frames; modifying the first interface portion to display an interface element in a position corresponding to the identified and displayed internal frame; and in response to a selection of the displayed interface element, modifying a second interface portion to display the one or more accessed external image frames that correspond to the identified and displayed internal frame.

22. accessing interior image frames captured by a mobile device while the mobile device moves through the interior of the building; accessing exterior image frames captured by an unmanned aerial vehicle ("UAV") while the UAV navigates around the exterior of the building; aligning the internal image frame and the external image frame to a coordinate system; generating an interface that displays one or more internal image frames in a first interface portion; using a coordinate system to identify a displayed internal image frame that corresponds to one or more of said accessed external image frames; modifying the first interface portion to display an interface element at a location corresponding to the identified and displayed internal frame; and in response to a selection of the displayed interface element, modifying a second interface portion to display the one or more accessed external image frames that correspond to the identified and displayed internal frame.