Depth-Based 3D Model Rendering Using Integrated Image Frames

The spatial indexing system efficiently aligns image frames with 3D models generated from LIDAR data, addressing the challenge of manual matching by integrating synchronized data to enhance environmental reviews.

JP7769693B2Active Publication Date: 2025-11-13OPEN SPACE LABS INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023521454
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-02
Filing Date
2021-09-21
Publication Date
2025-11-13
Estimated Expiration
2041-09-21

AI Technical Summary

Technical Problem

Generating three-dimensional models from two-dimensional images and LIDAR data is time-consuming and difficult due to the manual matching of large numbers of images and separate 3D models, which limits the usefulness of combined environmental reviews.

Method used

A spatial indexing system that aligns image frames with a 3D model generated from LIDAR data by using time synchronization or feature vectors, allowing for an interface that displays selected portions of the 3D model alongside corresponding image frames.

Benefits of technology

Facilitates efficient alignment and visualization of 3D models with image frames, providing more insightful environmental reviews by automating the association process and enhancing the understanding of environments through synchronized data integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007769693000002
    Figure 0007769693000002
  • Figure 0007769693000003
    Figure 0007769693000003
  • Figure 0007769693000004
    Figure 0007769693000004
Patent Text Reader

Abstract

The system aligns a 3D model of the environment with image frames of the environment to generate a visualization interface that displays a portion of the 3D model and a portion of the corresponding image frames. The system receives LIDAR data collected in the environment and generates a 3D model based on the LIDAR data. For each image frame, the system aligns the image frame with the 3D model. After aligning the image frame with the 3D model, when the system presents the portion of the 3D model in the interface, the system also presents the image frame that corresponds to the portion of the 3D model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field The present disclosure relates to generating models of an environment, and in particular to coordinating a three-dimensional model of an environment generated based on depth information (e.g., light detection and ranging, or LIDAR, data) with image frames of the environment to present portions of the three-dimensional model with corresponding image frames. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 090,095, filed October 9, 2020, and U.S. Provisional Application No. 63 / 165,682, filed March 24, 2021, which claims the benefit of U.S. Patent Application No. 17 / 367,204, filed July 2, 2021, each of which is incorporated herein by reference in its entirety.

[0003] Images of an environment can be useful for reviewing details associated with the environment without having to visit the environment in person. For example, a real estate agent may want to create a virtual tour of a home by capturing a series of photographs of the rooms in the home so that parties can virtually view the home. Similarly, a contractor may want to monitor progress at a construction site by capturing images of the construction site at various points during construction and comparing images captured at different times.

[0004] However, because images are limited to two dimensions (2D), a three-dimensional (3D) model of the environment can be generated using a LIDAR system to provide further details about the environment. When multiple views of the environment are presented simultaneously, it can provide more useful insight into the environment compared to when the images and 3D models are considered separately. However, when there are a large number of images and separate 3D models, manually reviewing the images and matching them to corresponding portions of the 3D models can be difficult and time-consuming. Summary of the Invention

[0005] The spatial indexing system receives image frames captured in an environment and LIDAR data collected in the same environment and associates a 3D model generated based on the LIDAR data with the image frames. The spatial indexing system associates the 3D model with the image frames by mapping each of the image frames to a portion of the 3D model. In some embodiments, the image frames and LIDAR data are captured simultaneously by a mobile device as the mobile device moves through the environment, and the image frames are mapped to the LIDAR data based on timestamps. In some embodiments, the video capture system and the LIDAR system are separate systems, and the image frames are mapped to the 3D model based on feature vectors. After association, the spatial indexing system generates an interface presenting a selected portion of the 3D model and the image frames corresponding to the selected portion of the 3D model. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 illustrates a system environment for a spatial indexing system, according to one embodiment. [Figure 2A] FIG. 2 is a block diagram of a pass module, according to one embodiment. [Figure 2B] FIG. 2 is a block diagram of a model generation module, according to one embodiment. [Figure 3A] 1 illustrates an example of a model visualization interface displaying a first interface portion including a 3D model and a second interface portion including an image associated with the 3D model, according to one embodiment. [Figure 3B] 1 illustrates an example of a model visualization interface displaying a first interface portion including a 3D model and a second interface portion including an image associated with the 3D model, according to one embodiment. [Figure 3C] 1 illustrates an example of a model visualization interface displaying a first interface portion including a 3D model and a second interface portion including an image associated with the 3D model, according to one embodiment. [Figure 3D] 1 illustrates an example of a model visualization interface displaying a first interface portion including a 3D model and a second interface portion including an image associated with the 3D model, according to one embodiment. [Figure 4] 1 is a flowchart illustrating an exemplary method for automatic spatial indexing of frames using features in a floor plan, according to one embodiment. [Figure 5] 1 is a flowchart illustrating an exemplary method for generating an interface for displaying a 3D model in conjunction with image frames, according to one embodiment. [Figure 6] FIG. 1 illustrates a computer system for implementing embodiments herein, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] I. Overview The spatial indexing system receives a video including a series of image frames representing an environment and aligns the image frames with a 3D model of the environment generated using LIDAR data. The image frames are captured by a video capture system moving through the environment along a path. The LIDAR data is collected by the LIDAR system, and the spatial indexing system generates a 3D model of the environment based on the LIDAR data received from the LIDAR system. The spatial indexing system aligns the images with the 3D model. In some embodiments, the LIDAR system is integrated with the video capture system such that the image frames and the LIDAR data are captured simultaneously and synchronized in time. Based on the time synchronization, the spatial indexing system can determine the location where each of the image frames was captured and the portion of the 3D model to which the image frame corresponds. In other embodiments, the LIDAR system is separate from the video capture system, and the spatial indexing system can use feature vectors associated with the LIDAR data and feature vectors associated with the image frames for alignment.

[0008] The spatial indexing system generates an interface having a first interface portion for displaying the 3D model and a second interface portion for displaying image frames. The spatial indexing system may receive an interaction from a user indicating a portion of the 3D model to be displayed. For example, the interaction may include selecting a waypoint icon associated with a location within the 3D model or selecting an object within the 3D model. The spatial indexing system identifies image frames associated with the selected portion of the 3D model and displays the corresponding image frames in the second interface portion. When the spatial indexing system receives another interaction indicating another portion of the 3D model to be displayed, the interface is updated to display the other portion of the 3D model in the first interface and to display different image frames associated with the other portion of the 3D model.

[0009] II. System Environment Figure 1 illustrates a system environment 100 for a spatial indexing system, according to one embodiment. In the embodiment shown in Figure 1, system environment 100 includes a video capture system 110, a network 120, a spatial indexing system 130, a LIDAR system 150, and a client device 160. Although a single video capture system 110, a single LIDAR system 150, and a single client device 160 are shown in Figure 1, in some implementations, spatial indexing system 130 interacts with multiple video capture systems 110, multiple LIDAR systems 150, and / or multiple client devices 160.

[0010] Video capture system 110 collects one or more of frame data, motion data, and position data as video capture system 110 moves along a path. In the embodiment shown in FIG. 1 , video capture system 110 includes camera 112, motion sensor 114, and position sensor 116. Video capture system 110 is implemented as a device having a form factor suitable for moving along a path. In one embodiment, video capture system 110 is a portable device that a user physically moves along a path, such as, for example, a wheeled cart or a device attached to or integrated with an object worn on the user's body (e.g., a backpack or helmet). In another embodiment, video capture system 110 is attached to or integrated with a vehicle. The vehicle may be, for example, a wheeled vehicle (e.g., a wheeled robot) or an aerial vehicle (e.g., a quadcopter drone) and can be configured to navigate autonomously along a pre-set route or to be controlled by a human user in real time. In some embodiments, video capture system 110 is part of a mobile computing device, such as a smartphone, tablet computer, or laptop computer, and can be carried by a user and used to capture video as the user moves through an environment along a path.

[0011] Camera 112 collects video including a series of image frames as video capture system 110 moves along a path. In some embodiments, camera 112 is a 360-degree camera that captures 360-degree frames. Camera 112 can be implemented by positioning multiple non-360-degree cameras within video capture system 110 so that the non-360-degree cameras are oriented at various angles relative to each other and configuring the multiple non-360-degree cameras to capture frames of the environment from each angle of the non-360-degree cameras approximately simultaneously. The image frames can then be combined to form a single 360-degree frame. For example, camera 112 can be implemented by capturing frames substantially simultaneously from two 180-degree panoramic cameras oriented in opposite directions. In other embodiments, camera 112 has a narrow field of view and is configured to capture typical 2D images instead of 360-degree frames.

[0012] The frame data captured by video capture system 110 may further include frame timestamps, which are data corresponding to the time each of the frames was captured by video capture system 110. As used herein, frames are captured substantially simultaneously if they are captured within a threshold time interval of each other (e.g., within 1 second, within 100 milliseconds, etc.).

[0013] In one embodiment, the camera 112 captures a walk-through video as the video capture system 110 moves through the environment. The walk-through video includes a series of image frames that can be captured at any frame rate, such as a high frame rate (e.g., 60 frames per second) or a low frame rate (e.g., 1 frame per second). Capturing a sequence of image frames at a higher frame rate generally produces more stable results, while capturing a sequence of image frames at a lower frame rate allows for reduced data storage and transmission. In another embodiment, the camera 112 captures a sequence of still frames separated by a fixed time interval. In yet another embodiment, the camera 112 captures a single image frame. The motion sensor 114 and the position sensor 116 collect motion data and position data, respectively, while the camera 112 captures frame data. The motion sensor 114 can include, for example, an accelerometer and a gyroscope. The motion sensor 114 can also include a magnetometer that measures the direction of the magnetic field surrounding the video capture system 110.

[0014] The location sensor 116 may include a global navigation satellite system (e.g., a GPS receiver) receiver that determines the latitude and longitude coordinates of the video capture system 110. In some embodiments, the location sensor 116 additionally or alternatively includes an indoor positioning system (IPS) receiver that determines the location of the video capture system based on signals received from transmitters installed at known locations within the environment. For example, multiple radio frequency (RF) transmitters that transmit RF fingerprints are placed throughout the environment, and the location sensor 116 also includes a receiver that detects the RF fingerprints and estimates the location of the video capture system 110 within the environment based on the relative strength of the RF fingerprints.

[0015] 1 includes a camera 112, a motion sensor 114, and a position sensor 116, in other embodiments, some of the components 112, 114, 116 may be omitted from the video capture system 110. For example, one or both of the motion sensor 114 and the position sensor 116 may be omitted from the video capture system.

[0016] In some embodiments, video capture system 110 is implemented as part of a computing device (e.g., computer system 600 shown in FIG. 6 ) that also includes a storage device for storing captured data and a communication interface for transmitting the captured data to spatial indexing system 130 over network 120. In one embodiment, video capture system 110 stores the captured data locally as video capture system 110 moves along a path, and after data collection is complete, the data is transmitted to spatial indexing system 130. In another embodiment, video capture system 110 transmits the captured data to spatial indexing system 130 in real time as system 110 moves along a path.

[0017] Video capture system 110 communicates with other systems via network 120. Network 120 may include any combination of local area and / or wide area networks using both wired and / or wireless communication systems. In one embodiment, network 120 uses standard communication technologies and / or protocols. For example, network 120 includes communication links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 3G, 4G, code division multiple access (CDMA), digital subscriber line (DSL), etc. Examples of network protocols used to communicate over network 120 include multiprotocol label switching (MPLS), transmission control protocol / internet protocol (TCP / IP), hypertext transfer protocol (HTTP), simple mail transfer protocol (SMTP), and file transfer protocol (FTP). Network 120 may also be used to deliver push notifications via various push notification services such as APPLE Push Notification Service (APNS) and GOOGLE Cloud Messaging (GCM). Network 120 Data exchanged over may be represented using any suitable format, such as Hypertext Markup Language (HTML), Extensible Markup Language (XML), or JavaScript Object Notation (JSON). In some embodiments, all or part of the communication links of network 120 may be encrypted using any suitable technique or techniques.

[0018] Light detection and ranging (LIDAR) system 150 uses laser 152 and detector 154 to collect three-dimensional data representing the environment as LIDAR system 150 moves through the environment. Laser 152 emits a laser pulse, and detector 154 detects when the laser pulse returns to LIDAR system 150 after being reflected by multiple points on objects or surfaces in the environment. LIDAR system 150 also includes motion sensor 156 and position sensor 158, which indicate the movement and position of LIDAR system 150 and can be used to determine the direction from which the laser pulse is emitted. LIDAR system 150 generates LIDAR data associated with the laser pulse detected after being reflected from the surface of an object or water in the environment. The LIDAR data may include a set of (x, y, z) coordinates determined based on the known direction from which the laser pulse was emitted and the duration between the emission of laser 152 and its detection by detector 154. The LIDAR data may also include other attribute data, such as the intensity of the detected laser pulse. In other embodiments, the LIDAR system 150 may be replaced by another depth-sensing system. Examples of depth-sensing systems include radar systems, 3D camera systems, etc.

[0019] In some embodiments, LIDAR system 150 is integrated with video capture system 110. For example, LIDAR system 150 and video capture system 110 may be components of a smartphone configured to capture video and LIDAR data. Video capture system 110 and LIDAR system 150 may be operated simultaneously, such that video capture system 110 captures video of the environment while LIDAR system 150 collects LIDAR data. When video capture system 110 and LIDAR system 150 are integrated, motion sensor 114 may be the same as motion sensor 156, and position sensor 116 may be the same as position sensor 158. LIDAR system 150 and video capture system 110 may work together, and points in the LIDAR data may be mapped to pixels in image frames captured simultaneously with the points, such that the points are associated with image data (e.g., RGB values). LIDAR system 150 may also collect timestamps associated with the points. Thus, the image frames and LIDAR data may be associated with each other based on the timestamps. As used herein, a timestamp in the LIDAR data may correspond to the time a laser pulse was emitted toward a point or the time a laser pulse was detected by detector 154. That is, for a timestamp associated with an image frame indicating the time the image frame was captured, one or more points in the LIDAR data may be associated with the same timestamp. In some embodiments, LIDAR system 150 may be used while video capture system 110 is not, or vice versa. In some embodiments, LIDAR system 150 is a separate system from video capture system 110. In such embodiments, the path of video capture system 110 may be different from the path of LIDAR system 150.

[0020] The spatial indexing system 130 aggregates the image frames captured by the video capture system 110 and the LIDAR data collected by the LIDAR system 150. dataand performs a spatial indexing process to automatically identify the spatial locations where each of the image frames and LIDAR data was captured and align the image frames to a 3D model generated using the LIDAR data. After aligning the image frames to the 3D model, spatial indexing system 130 provides a visualization interface that allows client device 160 to select portions of the 3D model and view corresponding image frames side-by-side together. In the embodiment shown in FIG. 1 , spatial indexing system 130 includes a route module 132, a route storage 134, a floorplan storage 136, a model generation module 138, a model storage 140, a model integration module 142, an interface module 144, and a query module 146. In other embodiments, spatial indexing system 130 may include fewer, different, or additional modules.

[0021] The path module 132 receives image frames in the walkthrough video and other location and motion data collected by the video capture system 110 and determines a path for the video capture system 110 based on the received frames and the received data. In one embodiment, the path is defined as the 6D camera pose for each frame of the walkthrough video, which includes a series of frames. The 6D camera pose for each frame is an estimate of the relative position and orientation of the camera 112 when the image frame was captured. The path module 132 can store the path in a path storage device 134.

[0022] In one embodiment, path module 132 uses a SLAM (simultaneous localization and mapping) algorithm to (1) simultaneously determine a path estimate by inferring the position and orientation of camera 112, and (2) model the environment using direct methods or landmark features extracted from a walkthrough video, which is a sequence of frames (e.g., oriented FAST and rotated BRIEF (ORB), scale-invariant feature transform (SIFT), speeded up robust features (SURF), etc.). Path module 132 outputs vectors of six-dimensional (6D) camera pose over time, with one 6D vector (three dimensions of position, three dimensions of orientation) for each frame in the sequence, and the 6D vectors can be stored in path storage 134.

[0023] The spatial indexing system 130 may also include a floor plan storage device 136 that stores one or more floor plans, such as floor plans of the environment captured by the video capture system 110. As referred to herein, a floor plan is a to-scale, two-dimensional (2D) diagrammatic representation of an environment (e.g., a building or portion of a structure) from a top-down perspective. In an alternative embodiment, the floor plan may be a 3D model of the expected completed building instead of a 2D drawing (e.g., a Building Information Modeling (BIM) model). The floor plan may be annotated to specify the locations, dimensions, and types of physical objects expected to be in the environment. In some embodiments, the floor plan is manually annotated by a user associated with the client device 160 and provided to the spatial indexing system 130. In other embodiments, the floor plan is annotated by the spatial indexing system 130 using a machine learning model trained using a training dataset of annotated floor plans to identify the locations, dimensions, and object types of physical objects expected to be in the environment. Different portions of a building or structure may be represented by separate floor plans. For example, the spatial indexing system 130 may store a separate floor plan for each floor of a building, unit, or substructure.

[0024] Model generation module 138 generates a 3D model of the environment. In some embodiments, the 3D model is based on image frames captured by video capture system 110. To generate the 3D model of the environment based on the image frames, model generation module 138 may use methods such as structure from motion (SfM), simultaneous localization and mapping (SLAM), monocular depth map generation, or other methods. The 3D model may be generated using image frames from a walkthrough video of the environment, the relative positions of each of the image frames (as indicated by the 6D pose of the image frames), and (optionally) the absolute positions of each of the image frames with respect to the floor plan of the environment. Image frames from video capture system 110 may be stereo images that can be combined to generate the 3D model. In some embodiments, model generation module 138 generates a 3D point cloud based on the image frames using photogrammetry. In some embodiments, model generation module 138 generates the 3D model based on LIDAR data from system 150. The model generation module 138 may process the LIDAR data to generate a point cloud, which may have higher resolution compared to a 3D model generated using the image frames. After generating the 3D model, the model generation module 138 stores the 3D model in the model storage device 140.

[0025] In one embodiment, the model generation module 138 receives a frame sequence and its corresponding path (e.g., 6D pose vectors defining a 6D pose for each of the frames in a walkthrough video, which is a sequence of frames) from the path module 132 or path storage device 134, and extracts a subset of the image frames in the sequence and their corresponding 6D poses for inclusion in the 3D model. For example, if the walkthrough video, which is a sequence of frames, is a frame in a video captured at 30 frames per second, the model generation module 138The model generation module subsamples the image frames by extracting the frames and their corresponding 6D poses at 0.5 second intervals. 138 This embodiment is described in more detail below with respect to FIG. 2B.

[0026] 1 , the 3D model is generated by model generation module 138 in spatial indexing system 130. However, in alternative embodiments, model generation module 138 may be generated by a third-party application (e.g., an application installed on a mobile device that includes video capture system 110 and / or LIDAR system 150). Image frames captured by video capture system 110 and / or LIDAR data collected by LIDAR system 150 may be transmitted over network 120 to a server associated with the application, which processes the data to generate the 3D model. Spatial indexing system 130 may then access the generated 3D model, associate the 3D model with other data associated with the environment, and present the associated representation to one or more users.

[0027] The model integration module 142 integrates the 3D model with other data that describes the environment. The other types of data may include one or more images (e.g., image frames from the video capture system 110), 2D floor plans, diagrams, and annotations that describe characteristics of the environment. The model integration module 142 determines the similarities between the 3D model and other data in order to align the other data with relevant portions of the 3D model. The model integration module 142 、3 Which part of the D model corresponds to other data The 3D portion may be determined and an identifier associated with the determined 3D portion may be stored in relation to other data.

[0028] In some embodiments, the model integration module 142 may align a 3D model generated based on the LIDAR data with one or more image frames based on time synchronization. As described above, the video capture system 110 and the LIDAR system 150 may be integrated into a single system that simultaneously captures image frames and LIDAR data. For each image frame, the model integration module 142 may determine the timestamp at which the image frame was captured and identify a set of points in the LIDAR data associated with the same timestamp. The model integration module 142 may then determine which portion of the 3D model contains the identified set of points and align the image frame with that portion. Furthermore, the model integration module 142 may map pixels in the image frame to sets of points in the LIDAR data.

[0029] In some embodiments, the model integration module 142 may align a point cloud generated using LIDAR data (hereinafter referred to as the “LIDAR point cloud”) with another point cloud generated based on image frames (hereinafter referred to as the “low-resolution point cloud”). This method may be used when the LIDAR system 150 and the video capture system 110 are separate systems. The model integration module 142 may generate feature vectors for each point in the LIDAR point cloud and each point in the low-resolution point cloud (e.g., using ORB, SIFT, or HardNET). The model integration module 142 may determine feature distances between the feature vectors and matching point pairs between the feature distance-based LIDAR point cloud and the feature distance-based low-resolution point cloud. The 3D pose between the LIDAR point cloud and the low-resolution point cloud is determined to generate a larger number of geometric inliers for the point pairs, for example, using random sample consensus (RANSAC) or nonlinear optimization. Since the low-resolution point cloud is generated with the image frame, the LIDAR point cloud is also aligned with the image frame itself.

[0030] In some embodiments, the model integration module 142 may associate a 3D model with a diagram or one or more image frames based on annotations associated with the diagram or one or more image frames. The annotations may be provided by a user or determined by the spatial indexing system 130 using an image recognition or machine learning model. The annotations may describe characteristics of objects or surfaces in the environment, such as dimensions or object type. The model integration module 142 may extract features in the 3D model and compare the extracted features to the annotations. For example, if the 3D model represents a room in a building, features extracted from the 3D model may be used to determine the dimensions of the room. The determined dimensions may be compared to a floor plan of a construction site annotated with the dimensions of various rooms in the building, and the model integration module 142 may identify rooms within the floor plan that match the determined dimensions. In some embodiments, the model integration module 142 may perform 3D object detection on the 3D model and compare the output of the 3D object detection with the output from the image recognition or machine learning model based on the diagram or one or more images.

[0031] In some embodiments, the 3D model may be manually associated with the diagram based on input from a user. The 3D model and diagram may be presented on a client device 160 associated with the user, and the user may select a location within the diagram that indicates the location corresponding to the 3D model. For example, the user may place a pin in a floor plan that corresponds to the LIDAR data.

[0032] The interface module 144 provides a visualization interface to the client device 160 to present information associated with the environment. The interface module 144 may generate the visualization interface in response to receiving a request from the client device 160 to view one or more models representing the environment. The interface module 144 may initially generate the visualization interface to include a 2D overhead map interface representing a floor plan of the environment from the floor plan storage 136. The 2D overhead map may be an interactive interface such that clicking a point on the map navigates to a portion of the 3D model corresponding to the selected point in space. The visualization interface provides a first-person perspective of a portion of the 3D model, allowing the user to pan and zoom around the 3D model and navigate to other portions of the 3D model by selecting waypoint icons representing the relative locations of other portions.

[0033] The visualization interface also allows a user to select an object within the 3D model, causing the visualization interface to display an image frame corresponding to the selected object. The user may select an object by interacting with a point on the object (e.g., clicking a point on the object). When interface module 144 detects an interaction from the user, interface module 144 sends a signal indicating the location of the point within the 3D model to query module 146. Query module 146 identifies an image frame associated with the selected point, and interface module 144 updates the visualization interface to display the image frame. The visualization interface may include a first interface portion for displaying the 3D model and a second interface portion for displaying the image frame. Exemplary visualization interfaces are described with respect to FIGS. 3A-3D .

[0034] In some embodiments, interface module 144 may receive a request to measure the distance between selected endpoints on a 3D model or image frame. Interface module 144 may provide identification of the endpoints to query module 146, and query module 146 may determine the (x, y, z) coordinates associated with the endpoints. Query module 146 may calculate the distance between the two coordinates and return the distance to interface module 144. Interface module 144 may update an interface portion to display the requested distance to the user. Similarly, interface module 144 may receive additional endpoints with a request to identify the area or volume of an object.

[0035] Client device 160 is any mobile computing device, such as a smartphone, tablet computer, or laptop computer, or a non-mobile computing device, such as a desktop computer, that can connect to network 120 and be used to access spatial indexing system 130. Client device 160 displays an interface to a user, for example, on a display device such as a screen, and receives user input to allow the user to interact with the interface. An exemplary implementation of a client device is described below with reference to computer system 600 of FIG. 6.

[0036] III. Route Generation Overview 2A illustrates a block diagram of the path module 132 of the spatial indexing system 130 shown in FIG. 1, according to one embodiment. The path module 132 receives input data (e.g., a sequence of frames 212, motion data 214, position data 223, floor plan 257) captured by the video capture system 110 and the LIDAR system 150 to generate a path 226. In the embodiment shown in FIG. 2A, the path module 132 includes a simultaneous localization and mapping (SLAM) module 216, a motion processing module 220, and a path generation and coordination module 224.

[0037] The SLAM module 216 receives the sequence of frames 212 and runs a SLAM algorithm to generate a first estimate of the path 218. Before running the SLAM algorithm, the SLAM module 216 may perform one or more preprocessing steps on the image frames 212. In one embodiment, the preprocessing step includes extracting features from the image frames 212 by converting the sequence of frames 212 into a sequence of vectors, each of which is a feature representation of a respective frame. In particular, the SLAM module may extract SIFT features, SURF features, or ORB features.

[0038] After extracting features, the preprocessing step can also include a segmentation process. The segmentation process divides the walk-through video, which is a sequence of frames, into segments based on the quality of each feature of the image frame. In one embodiment, the feature quality of a frame is defined as the number of features extracted from the image frame. In this embodiment, the segmentation step classifies each of the frames as having high or low feature quality based on whether the feature quality of the image frame is above or below a threshold. (That is, frames with feature quality above the threshold are classified as high quality, and frames with feature quality below the threshold are classified as low quality.) Low feature quality can be caused by, for example, excessive subject blur or poor lighting conditions.

[0039] After classifying the image frames, the segmentation process divides the sequence such that consecutive frames with high feature quality are combined into segments and frames with low feature quality are not included in any segment. For example, imagine a path follows a dimly lit hallway, entering and exiting a series of brightly lit rooms. In this example, image frames captured in each of the rooms are likely to have high feature quality, while image frames captured in the hallway are likely to have low feature quality. As a result, the segmentation process divides the walkthrough video, which is a sequence of frames, such that each sequence of consecutive frames captured in the same room is divided into a single segment (resulting in a separate segment for each room), while image frames captured in the hallway are not included in any segment.

[0040] After the pre-processing step, the SLAM module 216 runs a SLAM algorithm to generate a first estimate 218 of the path. In one embodiment, the first estimate 218 is also a vector of the 6D camera pose over time, with one 6D vector for each frame in the sequence. In an embodiment where the pre-processing step includes segmenting the walk-through video, which is a sequence of frames, the SLAM algorithm is run separately for each of the segments of frames to generate path segments for each of the segments.

[0041] Motion processing module 220 receives motion data 214 collected as video capture system 110 moves along a path and generates a second estimate 222 of the path. Similar to first estimate 218 of the path, second estimate 222 can also be represented as a 6D vector of the camera pose over time. In one embodiment, motion data 214 includes acceleration and gyroscope data collected by an accelerometer and a gyroscope, respectively, and motion processing module 220 generates second estimate 222 by performing a dead reckoning process on the motion data. In embodiments in which motion data 214 also includes data from a magnetometer, the magnetometer data may be used in addition to or instead of the gyroscope data to determine changes in orientation of video capture system 110.

[0042] The data generated by many consumer-grade gyroscopes contains a time-varying bias (also called drift) that, if not corrected for, can affect the accuracy of the second estimate of path 222. In embodiments where motion data 214 includes all three types of data described above (accelerometer, gyroscope, and magnetometer data), , MoThe motion processing module 220 can use accelerometer and magnetometer data to detect and correct for this bias in the gyroscope data. In particular, the motion processing module 220 determines the direction of the gravity vector (which would typically point in the direction of gravity) from the accelerometer data and uses the gravity vector to estimate the two-dimensional tilt of the video capture system 110. Meanwhile, the magnetometer data is used to estimate the gyroscope's heading bias. Because magnetometer data can be noisy, especially when used in buildings whose interior structures include steel frames, the motion processing module 220 can calculate and use a rolling average of the magnetometer data to estimate the heading bias. In various embodiments, the rolling average can be calculated over a time window of 1 minute, 5 minutes, 10 minutes, or some other period.

[0043] A path generation and association module 224 combines the first 218 and second 222 path estimates into a combined estimate of the path 226. In embodiments where the video capture system 110 also collects position data 223 while moving along the path, the path generation module 224 and collaboration Module 224 can also use location data 223 when generating route 226. If a floor plan of the environment is available, route generation and coordination module 224 also receives the floor plan 257 as input and generates the route. 226 The integrated estimate can be linked to the floor plan 257.

[0044] IV. Model Generation Overview 2B illustrates a block diagram of the model generation module 138 of the spatial indexing system 130 shown in FIG. 1 , according to one embodiment. FIG. 2B illustrates a 3D model 266 generated based on image frames. The model generation module 138 receives the path 226 generated by the path module 132, along with the sequence of frames 212 captured by the video capture system 110, a floor plan 257 of the environment, and information 254 about the camera. The output of the model generation module 138 is the 3D model 266 of the environment. In the described embodiment, the model generation module 138 includes a route generation module 252, a route filtering module 258, and image The frame extraction module 262 is included.

[0045] The route generation module 252 receives the path 226 and the camera information 254 and generates one or more candidate route vectors 256 for each extracted frame. The camera information 254 includes a camera model 254A and a camera height 254B. The camera model 254A is a model that maps each 2D point in a frame (i.e., as defined by a pair of coordinates that identify a pixel in the image frame) to a 3D ray that represents the line of sight direction from the camera to that 2D point. In one embodiment, the spatial indexing system 130 stores a separate camera model for each type of camera supported by the system 130. The camera height 254B is the height of the camera relative to the floor of the environment while the walkthrough video, which is a sequence of frames, is being captured. In one embodiment, the camera height is assumed to have a constant value during the image frame capture process. For example, if the camera is mounted on a helmet worn on the user's body, then the height has a constant value equal to the sum of the user's height and the height of the camera relative to the top of the user's head (both numbers can be received as user input).

[0046] As referred to herein, a root vector of an extracted frame is a vector that represents the spatial distance between the extracted frame and one of the other extracted frames. For example, a root vector associated with an extracted frame has its end point at the extracted frame and its start point at the other extracted frame, such that adding the root vector to the spatial position of the frame associated with the root vector results in the spatial position of the other extracted frame. In one embodiment, the root vector is calculated by performing vector subtraction to calculate the difference between the three-dimensional positions of the two extracted frames, as indicated by their respective 6D pose vectors.

[0047] Referring to interface module 144, the route vector for the extracted frame is used after interface module 144 receives 3D model 266 and displays the first-person perspective of the extracted frame. When displaying the first-person perspective, interface module 144 renders a waypoint icon (shown as a circle in FIG. 3B ) at a location within the image frame that represents the location of another frame (e.g., the image frame of the origin of the route vector). In one embodiment, interface module 144 determines the location within the image frame at which to render the waypoint icon that corresponds to the route vector using the following equation:

[0048]

number

[0049] In this formula, M proj is the projection matrix containing the parameters of the camera projection function used for rendering, and M view is an isometric matrix representing the user's position and orientation relative to his or her current frame, and M delta is the root vector, and G ringis the geometry (list of 3D coordinates) representing the mesh model of the waypoint icon being rendered, and P icon is the geometry of the icon inside the first person view of the image frame.

[0050] Route Generation Module 252 Referring again to , the route generation module 252 can calculate a candidate route vector 256 between each pair of extracted frames. However, displaying a separate waypoint icon for each candidate route vector associated with a frame can result in a large number of waypoint icons (e.g., dozens) being displayed within a frame, which can be overwhelming for a user and make it difficult to distinguish between individual waypoint icons.

[0051] To avoid displaying too many waypoint icons, a route filtering module 258 receives the candidate route vectors 256 and selects a subset of the route vectors, which are displayed route vectors 260, represented from a first-person perspective with corresponding waypoint icons. 258 Based on various criteria, candidate A root vector 256 may be selected. For example, the candidate root vectors 256 may be filtered based on distance (e.g., only root vectors having a length less than a threshold length are selected).

[0052] In some embodiments, the route filtering module 258 also receives a floor plan 257 of the environment and filters for candidate route vectors 256 based on features within the floor plan. 258uses the floor plan features to eliminate any candidate route vectors 256 that pass through walls, resulting in a set of displayed route vectors 260 that point only to locations that are visible in the image frame. This can be done, for example, by extracting floor plan frame patches from areas of the floor plan surrounding the candidate route vectors 256 and inputting the image frame patches into a frame classifier (e.g., a feedforward, deep convolutional neural network) to determine whether a wall is present within the patch. If a wall is present within the patch, then the candidate route vector 256 passes through the wall and is not selected as one of the displayed route vectors 260. If a wall is not present, then the candidate route vector does not pass through the wall, and Route Filtering It may be selected as one of the displayed route vectors 260 subject to any other selection criteria (such as distance) configured by the module 258 .

[0053] The image frame extraction module 262 receives the sequence of 360-degree frames and extracts some or all of the image frames to generate extracted frames 264. In one embodiment, the sequence of 360-degree frames is captured as frames of a 360-degree walk-through video, and the image frame extraction module 262 generates a separate extracted frame for each frame. As described above with respect to FIG. 1 , the image frame extraction module 262 can also extract a subset of image frames from the walk-through video. For example, if the walk-through video, which is a sequence of frames 212, was captured at a relatively high frame rate (e.g., 30 or 60 frames per second), the image frame extraction module 262 can extract a subset of image frames at regular intervals (e.g., 2 frames per second for the video), thereby displaying a more manageable number of extracted frames 264 to the user as part of the 3D model.

[0054] The floor plan 257, the displayed route vector 260, the path 226, and the extracted frames 264 are combined into a 3D model 266. As described above, the 3D model 266 is a representation of an environment that comprises a set of extracted frames 264 for the environment, the relative positions of each of the image frames (as indicated by the 6D pose within the path 226). In the embodiment shown in Figure 2B, the 3D model also includes the floor plan 257, the absolute positions of each of the image frames on the floor plan, and the displayed route vectors 260 for some or all of the extracted frames 264.

[0055] V. Model Visualization Interface 3A-3D show an example model visualization interface 300 displaying a first interface portion 310 including a 3D model and a second interface portion 320 including an image associated with the 3D model, according to one embodiment. The environment illustrated in FIGS. 3A-3D is a portion of a building (e.g., the back of the building). A user uses a mobile device to capture video and simultaneously collect LIDAR data while walking around the building. The video and LIDAR data are provided to a spatial indexing system 130, which generates a 3D model based on the LIDAR data and associates image frames in the video with corresponding portions of the 3D model. The interface module 144 of the spatial indexing system 130 generates the model visualization interface 300 to display the 3D model and image frames.

[0056] The 3D model shown in the first interface portion 310 may be a point cloud generated based on LIDAR data. While the 3D model helps visualize a building in three dimensions, it may miss details or have incorrect parts. Therefore, to compensate for the deficiencies of the 3D model, it is advantageous to display image frames with high-resolution 2D data alongside the 3D model. The 3D model coordinates with the image frames so that when the first interface portion 310 displays a portion of the 3D model, the second interface portion 320 displays an image frame corresponding to the portion of the 3D model displayed in the first interface portion 310. As described above with respect to FIG. 2B , waypoint icons are associated with the path taken to capture the image frames and represent the relative location of the frames within the environment. Waypoint icons are provided within the 3D model to indicate where one or more image frames were captured.

[0057] 3A illustrates a model visualization interface 300 presented in response to a user interacting with a waypoint icon 330. A first interface portion 310 allows the user Gau Waypoint Icon 33 to 0 Display a first-person view of a portion of a 3D model that matches what the user would see if standing at the corresponding location in the real environment .cormorant Waypoint Icon 330 is Way The point icon 330 is associated with a first image frame 340 that was captured at a location corresponding to the point icon 330. The first image frame 340 is overlaid on the 3D model and at an angle perpendicular to the angle at which the first image frame 340 was captured. As described above, each of the image frames is associated with a 6D vector (three dimensions for position, three dimensions for direction), and the first image frame 340 is overlaid with respect to the 3D model. frame The angle at which to tilt 340 is determined based on the third dimension relative to the direction in the 6D vector. The second interface portion 320 displays the same first image frame 340.

[0058] The interface module 144 receives an interaction with point A on the 3D model (e.g., clicking point A) and updates the model visualization interface 300 to display a different portion of the 3D model. The interface module 144 may also update the model visualization interface 300 after receiving other types of interactions within the 3D model indicating requests to reveal different portions of the 3D model and image frames by zooming in and out, rotating, and shifting. As the first interface portion 310 is updated to display a different portion of the 3D model, the second interface portion 320 is simultaneously updated to display image frames corresponding to the different portion of the 3D model.

[0059] In some embodiments, the model visualization interface 300 may include a measurement tool that can be used to measure dimensions of an object or surface of interest. The measurement tool allows a user to identify precise dimensions from the 3D model without having to revisit the building in person. The interface module 144 may receive two endpoints of the dimension of interest from the user and determine the distance 350 between the endpoints. In the example shown in FIG. 3B , the measurement tool is used to measure how far a portion of a wall extends outward. Because the 3D model in the first interface portion 310 and the image frame in the second interface portion 320 are aligned, endpoints can be selected from either interface portion. To determine the distance, the interface module 144 may provide an identification of the endpoint selected by the user to the query module 146, which retrieves the (x, y, z) coordinates of the endpoint. The query module 146 may calculate the distance 350 between the coordinates and may provide the distance 350 to the interface module 144, which displays the distance 350 within at least one of the first interface portion 310 and the second interface portion 320.

[0060] FIG. 3C illustrates another display of a 3D model associated with location B and a second image frame 360 ​​associated with location B. The interface module 144 may receive an interaction with location B and update the model visualization interface 300, as shown in FIG. 3C. For example, a user may want to see more of a window at location B and click on location B. In response to receiving the interaction, the first interface portion 310 is overlaid on a second image frame 360 ​​that is positioned at an angle perpendicular to the capture angle of the second image frame 360 ​​(e.g., tilted downward toward the ground). The second interface portion 320 is also updated to display the second image frame 360. The interface module 144 may receive a request for measurements that include endpoints that match the width of the window. As illustrated in FIG. 3D, the first interface portion 310 and the second interface portion 320 are updated to display the distance between the endpoints. 370 Includes.

[0061] In the example illustrated in FIGS. 3A-3D , a split-screen mode of the model visualization interface 300 having a first interface portion 310 and a second interface portion 320 is illustrated, simultaneously showing a 3D model and an image frame. However, the model visualization interface 300 may be presented in other viewing modes. For example, the model visualization interface 300 may initially show one of the first interface portion 310 and the second interface portion 320 and change to split-screen mode in response to receiving a request from a user to display both. In another example, the user interface may initially display a floor plan of an area including one or more graphic elements at locations within the floor plan where a 3D model or image frame is available. In response to receiving an interaction with a graphic element, the user interface may update to display the 3D model or image frame.

[0062] In other embodiments, a different pair of models may be displayed in the model visualization interface 300. That is, instead of the 3D model and image frames from the LIDAR database, one of the models may be replaced with a diagram, another 3D model (e.g., a BIM model, an image-based 3D model), or other representation of the building.

[0063] VI. Spatial Indexing of Frames Based on Floor Plan Features As described above, the visualization interface can provide a 2D overhead view map that displays the location of each of the frames within the floor plan of the environment. In addition to being displayed in an overhead view, the floor plan of the environment can also be used as part of a spatial indexing process to determine the location of each of the frames.

[0064] 4 is a flowchart illustrating an exemplary method 400 for automated spatial indexing of frames using features in a floor plan, according to one embodiment. In other embodiments, method 400 may include additional, fewer, or different steps, and the steps shown in FIG. 4 may be performed in a different order. For example, method 400 may be performed without obtaining a floor plan 430, in which case an integrated path estimate is generated 440 without using floor plan features.

[0065] The spatial indexing system 130 receives 410 a walkthrough video, which is a series of frames, from the video capture system 110. The image frames in the sequence are captured as the video capture system 110 moves along a path through an environment (e.g., a construction site floor). In one embodiment, each of the image frames is a frame captured by a camera on the video capture system (e.g., camera 112 described above with respect to FIG. 1). In another embodiment, each of the image frames has a narrower field of view, such as 90 degrees.

[0066] The spatial indexing system 130 generates 420 a first estimate of a path based on the walkthrough video, which is a sequence of frames. The first estimate of the path can be represented, for example, as a six-dimensional vector that defines a 6D camera pose for each of the frames in the sequence. In one embodiment, a component of the spatial indexing system 130 (e.g., the SLAM module 216 described above with reference to FIG. 2A ) runs a SLAM algorithm on the walkthrough video, which is a sequence of frames, to simultaneously determine the 6D camera pose for each of the frames and generate a three-dimensional virtual model of the environment.

[0067] The spatial indexing system 130 obtains 430 a floor plan of the environment. For example, multiple floor plans (including a floor plan for the environment depicted in the received walkthrough video, which is a sequence of frames) may be stored in floor plan storage 136, and the spatial indexing system 130 accesses floor plan storage 136 to obtain the floor plan of the environment. The floor plan of the environment may also be received from a user via video capture system 110 or via client device 160 without being stored in floor plan storage 136.

[0068] The spatial indexing system 130 generates 440 an integrated path estimate based on the first estimate of the path and the physical objects in the floor plan. After generating 440 the integrated path estimate, the spatial indexing system 130 generates 450 a 3D model of the environment. For example, the model generation module 138 generates the 3D model by combining the floor plan, the multiple route vectors, the integrated path estimate, and frames extracted from the walkthrough video, which is a sequence of frames, as described above with respect to FIG. 2B .

[0069] In some embodiments, the spatial indexing system 130 may also receive additional data (other than walk-through video, which is a sequence of frames) captured while the video capture system is moving along the path. For example, the spatial indexing system may also receive motion data or position data, as described above with respect to FIG. 1. In embodiments in which the spatial indexing system 130 receives additional data, the spatial indexing system 130 may use 440 the additional data in addition to the floor plan when generating the integrated path estimate.

[0070] In an embodiment in which spatial indexing system 130 receives motion data along with a walkthrough video, which is a sequence of frames, spatial indexing system 130 can perform a dead-reckoning process on the motion data to generate a second estimate of the path, as described above with respect to FIG. 2A. In this embodiment, generating a combined path estimate 440 includes using portions of the second estimate to fill gaps in the first estimate of the path. For example, the first estimate of the path may be split into path segments due to insufficient feature quality in some of the captured frames (which causes gaps for which the SLAM algorithm cannot generate a reliable 6D pose, as described above with respect to FIG. 2A). In this case, the 6D pose from the second path estimate can be used to combine the segments of the first path estimate by filling the gaps between them.

[0071] As noted above, in some embodiments, the method 400 may be performed without obtaining a floor plan 430, and the integrated path estimate is generated without using features in the floor plan 440. In one of these embodiments, the first estimate of the path is used as the integrated path estimate without additional data processing or additional data analysis.

[0072] In another of these embodiments, the integrated path estimate is generated by generating one or more additional path estimates, calculating a confidence score for each 6D pose in each of the path estimates, and selecting the 6D pose with the highest confidence score for each spatial location along the path 440. For example, the additional path estimates may include one or more of a second estimate using motion data, a third estimate using data from a GPS receiver, and a fourth estimate using data from an IPS receiver, as described above. As described above, each path estimate is a vector of 6D poses that describes the relative position and orientation for each of the frames in the sequence.

[0073] A confidence score for the 6D pose is calculated separately for each of the path estimates. For example, the confidence scores for the above-described path estimates may be calculated in the following manner: the confidence score for the 6D pose of a first estimate (generated using a SLAM algorithm) represents the feature quality (e.g., the number of features detected in the image frame) of the image frame corresponding to the 6D pose; the confidence score for the 6D pose of a second estimate (generated using motion data) represents the level of noise in the accelerometer, gyroscope, and / or magnetometer data in a time interval centered around, preceding, or following the time of the 6D pose; the confidence score for the 6D pose of a third estimate (generated using GPS data) represents the GPS signal strength for the GPS data used to generate the 6D pose; and the confidence score for the 6D pose of a fourth estimate (generated using IPS data) represents the IPS signal strength (e.g., RF signal strength) for the IPS data used to generate the 6D pose.

[0074] After generating the confidence scores, the spatial indexing system 130 iteratively scans each of the path estimates and selects the 6D pose with the highest confidence score for each frame in the sequence, and the selected 6D pose is output as the 6D pose for the image frame in the integrated path estimate. Because the confidence scores for each of the path estimates are calculated separately, the confidence scores for each of the path estimates can be normalized to a common scale (e.g., a scalar value between 0 and 1, where 0 represents the lowest possible confidence and 1 represents the highest possible confidence) before the iterative scanning process occurs.

[0075] VII. Overview of Interface Generation 5 is a flowchart 500 illustrating an exemplary method for generating an interface displaying a 3D model associated with image frames, according to one embodiment. A spatial indexing system receives image frames and LIDAR data collected by a mobile device as the mobile device moves through an environment 510. Based on the LIDAR data, the spatial indexing system generates a 3D model representing the environment 520. The spatial indexing system associates the image frames with the 3D model 530. The spatial indexing system generates an interface comprising a first interface portion and a second interface portion 540. The spatial indexing system displays a portion of the 3D model within the first interface portion 550. The spatial indexing system receives a selection of an object within the 3D model displayed within the first interface portion 560. The spatial indexing system identifies an image frame corresponding to the selected object 570. The spatial indexing system displays an image frame corresponding to the selected object within the second interface portion 580.

[0076] VIII. Hardware Components Figure 6 is a block diagram illustrating a computer system 600 on which embodiments described herein may be implemented. For example, in the context of Figure 1, video capture system 110, LIDAR system 150, spatial indexing system 130, or client device 160 may be implemented using computer system 600 as described in Figure 6. Video capture system 110, LIDAR system 150, spatial indexing system 130, or client device 160 may also be implemented using a combination of multiple computer systems 600 as described in Figure 6. Computer system 600 may be, for example, a laptop computer, a desktop computer, a tablet computer, or a smartphone.

[0077] In one implementation, the system 600 includes: Processor System 600 includes at least one processor 601 for processing information, a main memory 603, a read-only memory (ROM) 605, a storage device 607, and a communication interface 609. System 600 includes at least one processor 601 for processing information and a main memory 603, such as a random access memory (RAM) or other dynamic storage device, for storing information and instructions executed by processor 601. Main memory 603 may also be used for storing temporary variables or other intermediate information during execution of instructions executed by processor 601. System 600 may also include a ROM 605 or other static storage device for storing static information and instructions for processor 601. Storage device 607, such as a magnetic or optical disk, is provided for storing information and instructions.

[0078] The communication interface 609 allows the system 600 to communicate with one or more networks (e.g., networks) through the use of network links (wireless or wired). 120) using a network link, system 600 can communicate with one or more computing devices and one or more servers. System 600 can also include a display device 611, such as a cathode ray tube (CRT), LCD monitor, or television set, for displaying graphics and information to a user. An input mechanism 613, such as a keyboard including alphanumeric and other keys, can be coupled to system 600 for communicating information and command selections to processor 601. Other non-limiting illustrative examples of input mechanism 613 include a mouse, trackball, touch screen, or cursor direction keys for communicating directional information and command selections to processor 601 and for controlling cursor movement on display device 611. Additional examples of input mechanism 613 include a radio frequency identification (RFID) reader, a barcode reader, a three-dimensional scanner, and a three-dimensional camera.

[0079] According to one embodiment, the techniques described herein are performed by system 600 in response to processor 601 executing one or more sequences of one or more instructions contained in main memory 603. Such instructions may be read into main memory 603 from another machine-readable medium, such as storage device 607. Execution of the sequences of instructions contained in main memory 603 causes processor 601 to perform the process steps described herein. In alternative implementations, hardwired circuitry may be used in place of or in combination with software instructions to implement the examples described herein. Thus, the described examples are not limited to any specific combination of hardware circuitry and software.

[0080] IX. Additional Considerations As used herein, the term "comprising," followed by one or more elements, does not exclude the presence of one or more additional elements. The term "or" should be interpreted as an inclusive "or" rather than an exclusive "or" (e.g., "A or B" may refer to "A," "B," or "A and B"). The article "a" or "an" refers to one or more instances of the next following element, unless a single instance is expressly specified.

[0081] The drawings and written description illustrate exemplary embodiments of the present disclosure and should not be construed as reciting essential features of the present disclosure. The scope of the present invention should be interpreted from any claims published in any patent containing this specification.

Claims

1. receiving image frames and light detection and ranging (LIDAR) data collected by a mobile device as the mobile device moves through an environment; generating a 3D model representing the environment based on the LIDAR data; Aligning the image frames with the 3D model; generating an interface comprising a first interface portion and a second interface portion; displaying a portion of the 3D model within the first interface portion; receiving a selection of an object within the 3D model displayed within the first interface portion; identifying an image frame corresponding to the selected object; displaying within the second interface portion the image frame corresponding to the selected object; Equipped with modifying the first interface portion to include the image frame overlaid within the portion of the 3D model from which the object was selected, at a location corresponding to the selected object; A method further comprising:

2. The method of claim 1 , wherein the image frames are rendered by the mobile device within the 3D model at an angle perpendicular to the capture angle of the image frames.

3. receiving a selection of two endpoints of the object within a 3D model displayed within the first interface portion; determining the distance between the two end points; displaying the determined distance within the first interface portion; The method of claim 1 further comprising:

4. receiving a selection of two endpoints of the object within the image frame within the second interface portion; determining the distance between the two end points; displaying the determined distance within the second interface portion; The method of claim 1 further comprising:

5. The method of claim 1 , wherein the 3D model is generated by performing a co-registration cartography process on the LIDAR data.

6. The step of associating the image frames with the 3D model comprises: determining a first set of feature vectors associated with a plurality of points of the 3D model, the first set of feature vectors being generated based on the LIDAR data; generating a second 3D model representing the environment based on the image frames; determining a second set of feature vectors associated with a plurality of points of the second 3D model generated based on the image frames; mapping the points of the 3D model to the points of the second 3D model based on the first set of feature vectors and the second set of feature vectors; The method of claim 1 further comprising:

7. The step of associating the image frames with the 3D model comprises: For each image frame determining a time period associated with the image frame; identifying a portion of the LIDAR data associated with the time period, the portion of the LIDAR data being associated with the 3D model; storing an identification of the image frame associated with the identified set of points; The method of claim 1 further comprising:

8. The step of associating the image frames with the 3D model comprises: extracting features associated with the 3D model; comparing the extracted features to notes associated with one or more image frames; storing, based on the comparison, an identification of an image frame associated with a portion of the 3D model, wherein one or more annotations associated with the image frame match one or more characteristics of the portion of the 3D model; The method of claim 1 further comprising:

9. receiving an interaction with the first interface portion; updating the first interface portion to display a different portion of the 3D model according to the interaction; updating the second interface portion to display a different image frame according to the interaction; The method of claim 1 further comprising:

10. The method of claim 9 , wherein the first interface portion and the second interface portion are updated simultaneously.

11. The method of claim 9 , wherein the interaction includes at least one of zooming in, zooming out, rotating, and shifting.

12. A non-transitory computer-readable storage medium storing executable instructions that, when executed by a hardware processor, cause the hardware processor to perform the method of any one of claims 1 to 11.

13. 13. A system comprising: a processor; and the non-transitory computer-readable storage medium of claim 12.

Citation Information

Patent Citations

  • Methods and apparatus for generating and using reduced resolution images and / or communicating such images to a playback or content distribution device

    CN107615338A

  • Image processing device, image processing method, and program

    JP2016139199A

  • Three-dimensional measuring device

    JP2018031747A

  • Methods and apparatus for supporting content generation, transmission and / or playback

    JP2018507650A

  • Image display system and method

    JP2020005186A