Structure scan using unmanned aerial vehicle

The system enhances UAV scanning capabilities by using stereoscopic computer vision and automatic docking for extended operation, addressing battery limitations and enabling efficient, autonomous scanning of large structures with high-resolution imaging and obstacle detection.

JP2025170309APending Publication Date: 2025-11-18SKYDIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025135633
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-08
Filing Date
2025-08-18
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing unmanned aerial vehicles (UAVs) face limitations in performing robust, fully autonomous scanning of large structures like roofs and bridges due to battery life constraints and the need for human intervention during charging, and lack efficient methods for consistent image capture and obstacle avoidance.

Method used

The system employs UAVs equipped with stereoscopic computer vision and automatic docking stations for extended operation, enabling autonomous scanning with visual inertial odometry for localization, and automatic charging and landing, allowing for continuous scanning across multiple battery cycles.

Benefits of technology

Enables consistent, efficient scanning of large structures with reduced human intervention, providing high-resolution images and obstacle detection, and supports sustained operation through automated battery management and precise landing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170309000001_ABST
    Figure 2025170309000001_ABST
Patent Text Reader

Abstract

To provide a system and a method for a structure scan using an unmanned aerial vehicle.SOLUTION: Some methods include: accessing a three-dimensional map of a structure; generating facets based on the three-dimensional map, the facets each being a polygon on a plane in a three-dimensional space that fits to a subset of the points in the three-dimensional map; generating a scan plan based on the facets, the scan plan including a sequence of poses for an unmanned aerial vehicle to assume to enable capture, using image sensors of the unmanned aerial vehicle, of images of the structure; causing the unmanned aerial vehicle to fly to assume a pose corresponding to one of the sequence of poses of the scan plan; and capturing one or more images of the structure from the pose.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates to structural scanning using unmanned aerial vehicles. [Background technology]

[0002] Unmanned air vehicles (e.g., drones) can be used to capture images from vantage points that are otherwise difficult to reach. Drones are typically operated by humans using dedicated controllers to remotely control the movement and image capture functions of the unmanned air vehicle. Several automated image capture modes have been implemented, such as recording video while tracking a recognized user or a user carrying a beacon device as the user moves through an environment. Summary of the Invention [Means for solving the problem]

[0003] The present disclosure can be best understood from the following detailed description when read in conjunction with the accompanying drawings, in which: It is emphasized that, according to common practice, the various features of the drawings are not to scale. On the contrary, the dimensions of the various features have been suitably increased or reduced for clarity. [Brief explanation of the drawings]

[0004] [Figure 1] FIG. 1 illustrates an example system for scanning structures using unmanned air vehicles. [Figure 2A] FIG. 1 illustrates an example of an unmanned air vehicle configured for overhead structure scanning. [Figure 2B] FIG. 1 illustrates an example of an unmanned air vehicle configured for scanning a structure from below. [Figure 2C] FIG. 1 illustrates an example of a controller for an unmanned aerial vehicle. [Figure 3] FIG. 1 illustrates an example dock for facilitating autonomous landing of an unmanned air vehicle. [Figure 4]A block diagram showing an example of the hardware configuration of an unmanned aerial vehicle. [Figure 5A] FIG. 10 illustrates an example of a graphical user interface of an unmanned aerial vehicle used to present a two-dimensional polygonal projection of facets overlaid on an overview image of a structure to enable editing of the facets to facilitate scanning of the structure. [Figure 5B] FIG. 10 illustrates an example of a graphical user interface of an unmanned aerial vehicle used to present a scan plan overlaid on an overview image of a structure to allow user review to facilitate scanning of the structure. [Figure 6] 1 is a flowchart illustrating an example process for scanning a structure using an unmanned air vehicle. [Figure 7] 10 is a flowchart illustrating an example process for enabling user editing of facets. [Figure 8] 1 is a flowchart illustrating an example of a process that attempts to simplify polygons representing facets by removing convex edges. [Figure 9] 1 is a flowchart illustrating an example of a process for presenting coverage information for a scan of a structure. [Figure 10] 1 is a flowchart illustrating an example of a process for generating a three-dimensional map of a structure. [Figure 11] 1 is a flowchart illustrating an example process for generating a three-dimensional map of a roof. [Figure 12] 10 is a flowchart illustrating an example process for presenting context information for a roof scan. [Figure 13A] FIG. 10 illustrates an example of a graphical user interface of an unmanned aerial vehicle used to present a proposed bounding polygon overlaid on an overview image of a roof to enable editing of the bounding polygon to facilitate scanning of the roof. [Figure 13B]FIG. 10 illustrates an example of a graphical user interface of an unmanned aerial vehicle used to present a proposed bounding polygon overlaid on an overview image of a roof to enable editing of the bounding polygon to facilitate scanning of the roof. [Figure 14A] FIG. 10 illustrates an example of an input polygon that can be associated with a facet. [Figure 14B] FIG. 14B is a diagram showing an example of a simplified polygon determined based on the input polygon of FIG. 14A. DETAILED DESCRIPTION OF THE INVENTION

[0005] Much of the value and challenge of autonomous unmanned aerial vehicles lies in enabling robust, fully autonomous missions. Disclosed herein is technology for scanning structures (e.g., roofs, bridges, or construction sites) in a thorough and repeatable manner using unmanned aerial vehicles (UAVs). Some implementations can provide advantages over conventional systems, such as: providing more consistent framing of structure scan images than is achievable by manual control of the unmanned aerial vehicle by maintaining a consistent distance and orientation relative to the section of the structure's surface being imaged, which can facilitate more robust detection of structure maintenance issues using machine learning or human review of the scan data; reducing the need for human operator attention; and / or faster, broader scanning of large structures.

[0006] In some implementations, an initial coarse scan is performed with a range sensor (e.g., an array of image sensors configured for stereoscopic computer vision) to obtain a three-dimensional map of the structure at a first resolution based on a user-provided rough bounding box of the structure of interest. A set of facets is then generated based on the three-dimensional map. In some implementations, user feedback on the set of facets is solicited by presenting the facets as two-dimensional polygonal projections of the facets in an overview image (e.g., a still image) of the structure. The user may be able to edit the two-dimensional polygons to make corresponding changes to the facets as they exist in three dimensions. A scanplan is generated based on the set of facets, the scanplan including a series of poses of the unmanned air vehicle near the surface to be scanned and modeled by the facets. For example, the scanplan poses may be orthographic poses and at a constant distance relative to the surface to be scanned. The scanplan is then executed by maneuvering the UAV to these poses and capturing relatively high-resolution images of the facets, which can then be stitched together. The captured images can be examined in real time or offline by humans or trained machine learning modules.

[0007] For large structures, scan planning can be performed over many charging cycles of the UAV's battery. This capability is greatly enhanced using fully automatic docking and charging in specially marked docks. Automatic docking and charging may be used in conjunction with the ability to pause scan planning after a series of attitudes and, after the charging session is complete, perform large-scale scanning with human intervention, robustly localizing at the next attitude in the series. For example, localization at the next attitude may be facilitated by using robust visual inertial odometry (VIO) for high-resolution localization and obstacle detection and avoidance.

[0008] In some implementations, the setup phase begins with a user on the ground pointing the unmanned aerial vehicle toward the structure to be scanned (e.g., a building with a roof). The user may hit "take off" on the unmanned aerial vehicle's user interface. The unmanned aerial vehicle takes off, moves diagonally above the target house of interest, and flies high enough to look directly down onto the roof of the building below and capture all of the area of ​​interest within its field of view.

[0009] The user interface presents a polygon, the user can drag the vertices of which to identify the rooftop area of ​​interest for scanning. The user can then select an approximate height (e.g., relative to the ground) that defines the rooftop volume of interest in 3D space. This specifies the 3D space where the scan will occur. A camera image can also be taken from this overview vantage point, which is used as a static viewpoint in the user interface. As the unmanned air vehicle continues to fly and approach the rooftop, the image on the screen freezes in the overview view, but a 3D render of the unmanned air vehicle is drawn in the user interface at the correct perspective for where the physical drone will be. This allows the user to see the unmanned air vehicle in the image and the status of the shape estimation and path planning for future steps.

[0010] For example, an unmanned air vehicle may be enabled to load data stored on the vehicle or on a user device to continue progress from a previously incomplete scan or to repeat a previously performed scan. In this case, after the vehicle reaches an overhead view, the unmanned air vehicle can skip the search phase and relocalize itself based on visual and inertial data. Relocalization may be enabled without the need for any global positioning services or visual fiducials / data.

[0011] In the initial search phase, after a three-dimensional bounding box is defined, a small number of points of interest are generated and flown from an oblique view at the corners of the roof. The unmanned aerial vehicle may then fly a flight path (e.g., a flight path relative to a dynamic surface) to obtain an initial three-dimensional map of the roof. This may be done by flying a lawnmower-style back-and-forth pattern while using a dynamic local obstacle map to fly at a fixed altitude above the roof surface. Stereo imaging may be used to accumulate distance information into a single three-dimensional map of the entire roof. The size of the lawnmower-pattern grid and the height above the surface may be selected to trade off between obtaining a high-quality three-dimensional map (e.g., close to the surface, multiple passes, slow flight) and quickly acquiring a map (e.g., farther from the surface, fewer passes, fast flight). These techniques may enable autonomous flying of a surface-relative pattern to generate mapping data.

[0012] Software running on a processing unit within the unmanned air vehicle and / or on a controller for the UAV may be used to implement the structure scanning techniques described herein.

[0013] FIG. 1 illustrates an example of a system 100 for structure scanning using an unmanned air vehicle 110. System 100 includes unmanned air vehicle 110, controller 120, and docking station 130. Controller 120 can communicate with unmanned air vehicle 110 via a wireless communications link (e.g., via a WiFi network or a Bluetooth link) to receive video or images and to issue commands (e.g., commands related to takeoff, landing, following, manual control, and / or performing an autonomous or semi-autonomous scan of a structure (e.g., a roof, bridge, or building under construction)). For example, controller 120 may be controller 250 of FIG. 2C . In some implementations, the controller includes a smartphone, tablet, or laptop running software configured to communicate with and control unmanned air vehicle 110. For example, system 100 can be used to implement process 600 of FIG. 6 . For example, system 100 can be used to implement process 700 of FIG. 7 . For example, system 100 can be used to implement process 800 of FIG. 8 . For example, system 100 can be used to implement process 900 of Figure 9. For example, system 100 can be used to implement process 1000 of Figure 10.

[0014] Unmanned air vehicle 110 includes a propulsion mechanism (e.g., including a propeller and a motor), one or more image sensors, and a processing unit. For example, unmanned air vehicle 110 may be unmanned air vehicle 200 of Figures 2A-2B. For example, unmanned air vehicle 110 may include hardware configuration 400 of Figure 4. The processing device (e.g., processing device 410) can be configured to: access a three-dimensional map of the structure encoding a set of points in three-dimensional space on the surface of the structure; generate one or more facets based on the three-dimensional map, where a given facet of the one or more facets is a polygon on the surface in three-dimensional space that fits a subset of points in the three-dimensional map; generate a scan plan based on the one or more facets, the scan plan including a series of poses for unmanned air vehicle 110 to assume that enable capture of an image of the structure at a fixed distance from each of the one or more facets using one or more image sensors; control the propulsion mechanism to fly unmanned air vehicle 110 at a pose corresponding to one of the series of poses in the scan plan; and capture one or more images of the structure from that pose using the one or more image sensors. The processing device can be further configured to continue executing the scan plan by controlling the propulsion mechanism to fly unmanned air vehicle 110 at a pose corresponding to each of the series of poses in the scan plan; and capture one or more images of the structure from each of these poses using the one or more image sensors until an image covering all of the one or more facets has been captured. In some implementations, the processing device may be configured to stitch the captured images together to obtain a composite image of one or more surfaces of the structure. For example, the image stitching may be performed based in part on out-of-band information associated with the images through each facet, such as three-dimensional map points associated with the facet or the boundary of one or more facets. For example, the set of poses in the scan plan may be poses for orthogonal imaging of each of one or more facets such that the unmanned air vehicle's image sensor (e.g., image sensor 220) is oriented toward the facet along a normal to the surface of the facet. For example, the structure may be the roof of a building.For example, the structure may be a bridge. For example, the structure may be a building under construction.

[0015] In some implementations, unmanned air vehicle 110 is configured to generate facets, in part, by soliciting user feedback and edits of proposed facets generated based on an automated analysis of a three-dimensional map of the structure. For example, a processing unit of unmanned air vehicle 110 can be configured to: capture an overview image of the structure using one or more image sensors; generate facet proposals based on the three-dimensional map; determine a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the two-dimensional polygon corresponding to the proposed facet when the subset of points is projected into an image plane of the overview image; present the two-dimensional polygon overlaid on the overview image; determine an edited two-dimensional polygon in the image plane of the overview image based on data indicating user edits of the two-dimensional polygon; and determine one of the one or more facets based on the edited two-dimensional image. In some implementations, the processing device is configured to: simplify the 2D polygon by removing convex edges from the 2D polygon and extending edges of the 2D polygon adjacent to the convex edges to points where the extended edges intersect with each other before presenting the 2D polygon overlaid on the overview image. For example, the processing device may be configured to verify that removing the convex edges increases the area of ​​the 2D polygon by an amount less than a threshold. For example, the processing device may be configured to verify that removing the convex edges increases the perimeter of the 2D polygon by an amount less than a threshold.

[0016] In some implementations, the unmanned air vehicle 110 is also used to generate a three-dimensional map of a structure by performing an initial coarse scan of the structure using range sensors (e.g., an array of image sensors configured for stereoscopic computer vision, a radar sensor, and / or a lidar sensor). For example, the unmanned air vehicle 110 may include one or more image sensors configured to support stereoscopic imaging used to provide distance data. For example, the processing unit may be configured to: control the propulsion mechanism to fly the unmanned air vehicle 110 near the structure; and scan the structure using one or more image sensors to generate the three-dimensional map. In some implementations, the structure is scanned to generate the three-dimensional map from a distance greater than the fixed distance used for facet imaging.

[0017] For example, the generated facet-based scanplan may be presented to a user for approval before execution of the scanplan begins. In some implementations, the processing device is configured to: capture a schematic image of the structure using one or more image sensors; present to the user a graphical representation of the scanplan superimposed on the schematic image; and receive an indication from the user of approval of the scanplan.

[0018] In some implementations, the scan plan may be dynamically updated during execution of the scan plan to accommodate dynamically detected obstacles or occlusions and to take advantage of higher resolution sensor data that becomes available as the unmanned air vehicle 110 approaches the surface of the structure represented by the facet. For example, the processing unit may be configured to: detect an obstacle while flying between poses in the series of poses in the scan plan based on images captured using one or more image sensors; and dynamically adjust poses in the series of poses in the scan plan to avoid the obstacle.

[0019] A facet is a polygon oriented in 3D space that approximates a surface of a structure (e.g., a roof). The true surface does not necessarily match this planar model. Deviation is the distance of a point on the actual surface from the facet that corresponds to the true surface. For example, deviations can arise from aggregation inherent in the facet estimation process, which cannot model smaller features such as a vent cap on a roof or a small skylight. Deviations can also arise from errors in the 3D scanning process. Deviations are detected by analyzing images captured from close-up during scanplan execution (e.g., two or more images providing stereoscopic vision). Adjustments are made to maintain a consistent distance from the actual surface, taking into account higher-resolution data about deviations that becomes available as the scanplan approaches a normal pose for image capture. For example, the processing device may be configured to: detect deviations of points on the surface of the structure from one of the one or more facets based on images captured using the one or more image sensors while flying between poses in the series of poses in the scan plan; and dynamically adjust an pose in the series of poses in the scan plan to accommodate the deviations and maintain a constant distance for image capture.

[0020] The unmanned air vehicle 110 may output image data and / or other sensor data captured during execution of the scanplan to controller 120 for user viewing, storage, and / or further offline analysis. For example, the processing device may be configured to: determine an area estimate for each of the one or more facets; and present a data structure including the one or more facets, the area estimates for each of the one or more facets, and an image of the structure captured during execution of the scanplan. For example, the area estimates may be converted to or accompanied by a corresponding cost estimate for maintenance work on the portion of the structure corresponding to the facet. Output from the unmanned air vehicle 110 may also include an indication of the coverage of the structure achieved by execution of the scanplan. For example, the processing device may be configured to: generate a coverage map of the one or more facets indicating which of the one or more facets were successfully imaged during execution of the scanplan; and present the coverage map (e.g., via sending data encoding the coverage map to controller 120).

[0021] Some structures may be too large to complete execution of a scan plan on a single charge of the unmanned air vehicle's 110 battery. It may be useful to pause execution of a scan plan while the unmanned air vehicle 110 lands and recharges, and then continue execution of the scan plan where it was interrupted. For example, a docking station 130 may facilitate safe landing and charging of the unmanned air vehicle 110 while execution of the scan plan is paused. In some implementations, the processing unit is configured to: store a scan plan state indicating a next attitude in the series of attitudes of the scan plan after initiation and before completion of the scan plan; control the propulsion mechanism to fly the unmanned air vehicle to land after storing the scan plan state; control the propulsion mechanism to fly the unmanned air vehicle to take off after landing; access the scan plan state; and control the propulsion mechanism to fly the unmanned air vehicle to the next attitude to continue execution of the scan plan based on the scan plan state. For example, the scan plan state may include a copy of the flight plan and an indication of the next attitude, such as a pointer to the next attitude in the series of attitudes of the scan plan. In some implementations, docking station 130 is configured to enable automatic landing, charging, and takeoff of unmanned air vehicle 110. For example, docking station 130 may be dock 300 of FIG.

[0022] FIG. 2A illustrates an example of an unmanned aerial vehicle 200 configured for such structure scanning. The unmanned aerial vehicle 200 includes a propulsion mechanism 210 including four propellers and a motor configured to rotate the propellers. For example, the unmanned aerial vehicle 200 may be a quadcopter drone. The unmanned aerial vehicle 200 includes image sensors, including a high-resolution image sensor 220 mounted on a gimbal to support stable, low-blur image capture and target tracking. For example, the image sensor 220 can be used to perform high-resolution scans of the surface of a structure during execution of a scan plan. The unmanned aerial vehicle 200 also includes lower-resolution image sensors 221, 222, and 223 spaced around the top of the unmanned aerial vehicle 200, each capped by a fisheye lens to provide a wide field of view and support stereoscopic computer vision. The unmanned aerial vehicle 200 also includes an on-board processing unit (not shown in FIG. 2A ). For example, unmanned air vehicle 200 may include hardware configuration 400 of Figure 4. In some implementations, the processing unit is configured to automatically fold the propellers when entering a docking station (e.g., dock 300 of Figure 3), thereby allowing the dock to have a footprint smaller than the area swept by the propellers of propulsion mechanism 210.

[0023] FIG. 2B illustrates an example of an unmanned air vehicle 200 configured for structure scanning. From this vantage point, three additional image sensors are visible: image sensor 224, image sensor 225, and image sensor 226, located at the bottom of the unmanned air vehicle 200. These image sensors (224-226) may also be covered by respective fisheye lenses to provide a wide field of view and support stereoscopic computer vision. This array of image sensors (220-226) may enable visual inertial odometry (VIO) for high-resolution localization and obstacle detection and avoidance. For example, the array of image sensors (220-226) may be used to scan a structure to obtain range data and generate a three-dimensional map of the structure.

[0024] Unmanned air vehicle 200 may be configured to autonomously land on landing surface 310. Unmanned air vehicle 200 also includes a battery in battery pack 240 attached to the bottom of unmanned air vehicle 200 with conductive contacts 230 to enable battery charging. For example, the techniques described in connection with FIG. 3 may be used to land unmanned air vehicle 200 on landing surface 310 of dock 300.

[0025] The bottom of battery pack 240 is the bottom of unmanned air vehicle 200. Battery pack 240 is shaped to fit onto landing surface 310 at the bottom of the funnel. As unmanned air vehicle 200 makes final approach to landing surface 310, the bottom of battery pack 240 contacts landing surface 310 and is mechanically guided to a central position at the bottom of the funnel by the tapered sides of the funnel. Once landing is complete, conductive contacts of battery pack 240 can contact conductive contacts 330 on landing surface 310, forming an electrical connection to enable charging of the battery of unmanned air vehicle 200. Dock 300 may include a charger configured to charge the battery while unmanned air vehicle 200 is on landing surface 310.

[0026] 2C illustrates an example controller 250 for an unmanned air vehicle. Controller 250 can provide a user interface for controlling the unmanned air vehicle and reviewing data (e.g., images) received from the unmanned air vehicle. Controller 250 includes touchscreen 260; left joystick 270; and right joystick 272. In this example, touchscreen 260 is part of a smartphone 280 that connects to controller mount 282, which can provide additional control surfaces including left joystick 270 and right joystick 272, as well as extended-range communication capabilities for longer-range communication with the unmanned air vehicle.

[0027] In some implementations, processing (e.g., image processing and control functions) may be performed by an application running on a processor of a remote controller device (e.g., controller 250 or a smartphone) for an unmanned air vehicle controlled using the remote controller device. Such a remote controller device may provide interactive functionality, with the application providing all of the functionality using video content provided by the unmanned air vehicle. For example, the various steps of processes 600, 700, 800, 900, 1000, 1100, and 1200 of Figures 6-12 may be implemented using a processor of a remote controller device (e.g., controller 250 or a smartphone) that communicates with the unmanned air vehicle to control it.

[0028] Much of the value and challenge of autonomous unmanned aerial vehicles lies in enabling robust, fully autonomous missions. Disclosed herein is a docking platform that enables unmanned charging, takeoff, landing, and mission planning for unmanned aerial vehicles (UAVs). Some implementations enable reliable operation of such platforms and associated application programming interface designs that make the system accessible to a variety of consumer and commercial applications.

[0029] One of the biggest limiting factors for operating a drone is the battery. A typical drone can operate for 20–30 minutes before needing a new battery pack. This limits the amount of time an autonomous drone can operate without human intervention. When a battery pack depletes, the operator must land the drone and replace it with a fully charged one. While battery technology continues to improve, enabling higher energy densities, these improvements are incremental, and no clear roadmap for sustained autonomous operation can be drawn. One approach to alleviating the need for periodic human intervention is to automate battery management operations through some type of automated base station.

[0030] Some methods disclosed herein utilize visual tracking and control software capable of pinpoint landings on much smaller targets. By using visual fiducial marks to assist in absolute position tracking relative to a base station, a UAV (e.g., a drone) can reliably hit a 5 cm x 5 cm target in a variety of environmental conditions. This results in extremely accurate positioning of the UAV with the aid of a small, passive funnel-shaped geometry that serves to guide the UAV's battery, extending to a set of charging contacts below the rest of the UAV's structure, without requiring any complex actuation or bulky structures. This may enable a basic implementation of the base station to simply consist of a funnel-shaped nest with a set of spring contacts and a visual tag within it. To reduce the turbulent ground effects that a UAV typically experiences during landing, this nest can be positioned high above the ground, and the profile of the nest itself can be small enough to remain centered between the UAV's propeller wash during landing. Prop wash, or propeller wash, is a turbulent mass of air pushed by an airplane's propellers. To enable reliable operation in GPS-denied environments, the reference marks (e.g., small visual tags) within the nest may be supplemented with larger reference marks (e.g., large visual tags) located somewhere outside the landing nest, such as a flexible mat that can be spread on the ground near the base station or attached to a nearby wall. The additional visual tag can be easily found by the UAV from a significant distance, allowing the UAV to reacquire its absolute position relative to the landing nest in GPS-denied environments, regardless of visual inertial odometry (VIO) flight drift that may increase during the UAV's mission. Finally, a reliable communications link with the UAV can be maintained so that the UAV can cover a large area. Because ideal landing recharge locations are often not suitable for transmitter installation, communications circuitry may be placed in a separate range-extension module, which can be placed in a central location in the intended mission space, ideally at an elevated position for maximum coverage.

[0031] The simplicity and low cost of such a system, when compared to more complex and expensive battery-swapping systems, makes up for the time the UAV is unavailable while the battery is recharged. For many use cases, intermittent operation is sufficient, and users requiring greater UAV coverage can simply increase UAV availability by adding another UAV and base station system. This cheaper approach can be cost-competitive with larger, more expensive battery-swapping systems and can also significantly improve system reliability by eliminating the possibility of a single point of failure that could bring down the entire system.

[0032] For use cases where UAVs (e.g., drones) need to be protected from the elements but no existing UAV-accessible structures are available, the UAV nest can be housed in a small, custom hangar. This hangar may consist of a covered section attached to an uncovered access area beneath which the UAV can land and act as a windbreak, allowing the UAV to approach and precision land even in high winds. One useful feature of such a shelter would be open or ventilated sections along the entire perimeter of the bottom of the walls to direct the drone's downdraft away from the structure rather than circulating turbulently inside and adversely affecting stable flight.

[0033] For use cases where the UAV (e.g., drone) needs to be more robustly secured against dust, cold, theft, etc., a mechanized boxed drone enclosure may be used. For example, a drawer-like box slightly larger than the UAV itself may be used as a dock for the UAV. In some implementations, a motorized door on the side of the box can open 180 degrees to avoid being caught in the UAV's downdraft. For example, a telescoping linear slide may be attached to the charging nest within the box, holding the UAV far enough away from the box that it does not touch the box when the UAV takes off or lands. In some implementations, after the UAV lands, the slide folds its propellers into a small space, and then pulls the UAV back into the box while the UAV slowly rotates its propellers backward to move them out of the way of the door. This allows the box's footprint to be smaller than the area swept by the UAV's propellers. In some implementations, the two-bar coupling connecting the door to the door motor is designed to rotate through a center when closed so that the motor cannot be backdriven by pulling on the door from the outside, thereby effectively locking the door. For example, the UAV may be physically secured within the nest by a coupling mechanism that utilizes the last few centimeters of sliding movement to firmly press the UAV into the nest with soft rollers. Once secured, the box can be safely transported and even turned over without the UAV being removed from the box.

[0034] This operating enclosure design may be shelf-mounted or freestanding on an elevated base that ensures the UAV is positioned high enough above the ground to avoid ground effect when landing. The box's square profile makes it easy to stack multiple boxes on top of each other for a multi-drone hive configuration, where each box is rotated 90 degrees relative to the box below it, allowing multiple drones to take off and land simultaneously without interfering with each other. Because the UAV is physically secured within the enclosure when the box is closed, the box can be mounted to a car or truck, avoiding experiencing charging interruptions while the air vehicle is in motion. For example, in implementations where the UAV is positioned sideways from the box, the box can be recessed into a wall to ensure it is completely out of the way when not landing or taking off.

[0035] The box can be manufactured to have a very high Ingress Protection (IP) rating when closed and can include a basic cooling and heating system to allow the system to function in many outdoor environments. For example, a High-Efficiency Particulate Absorbing (HEPA) filter over the intake cooling fan may be used to protect the interior of the enclosure from environmental dust. A heater built into the top of the box can melt snow accumulation on site during winter.

[0036] For example, the top and sides of the box can be formed from a material that does not block radio frequencies, so that a version of the range extender can be built into the box itself for mobile use cases. In this way, a UAV (e.g., a drone) can maintain GPS lock while charging and deploy instantly. In some implementations, the door can incorporate a window, or the door and side panels of the box can be transparent, so that the UAV can see its surroundings before deploying and can act as its own surveillance camera to prevent theft or vandalism.

[0037] In some implementations, a spring-loaded microfiber wiper can be placed inside the box so that the navigation camera lens is wiped clean every time the drone slides in or out of the box. In some implementations, a small diaphragm pump inside the box can fill a small pressure vessel that can be used to clean all of the drone's lenses by blowing air at the lenses through small nozzles inside the box.

[0038] For example, the box can be mounted to a car using three linear actuators hidden within the mounting base that can lift and tilt the box during launch or landing to accommodate the vehicle standing on uneven streets or uneven terrain.

[0039] In some implementations, the box may have a single or double door at the top of the box that slides or rotates open, allowing the landing nest to extend into the open air rather than to the side. This would also take advantage of the UAV's ability to land on small targets, away from obstacles or surfaces that would interfere with the UAV's propeller wash (which would make a stable landing more difficult), and after the UAV has landed, the UAV and nest can be stored in a secure enclosure.

[0040] Software running on a processing unit within the unmanned air vehicle and / or a processing unit in a dock for the UAV may be used to implement the autonomous landing techniques described herein.

[0041] For example, a robust estimation and relocalization procedure may include visual relocalization of a dock with a multi-scale landing surface. For example, the UAV software may support GPS visual localization migration. Depending on the implementation, arbitrary fiducial mark (e.g., visual tag) designs, sizes, and orientations around the dock may be supported. For example, the software may enable detection and elimination of false detections.

[0042] For example, a UAV's takeoff and landing procedure may include robust planning and control in wind using model-based wind estimation and / or model-based wind compensation. For example, a UAV's takeoff and landing procedure may include a landing honing procedure that allows a short stop above a dock's landing surface. Because state estimation and visual detection are more accurate than control in high-wind environments, the UAV waits until the position, velocity, and angle errors between the actual air vehicle and a reference mark on the landing surface are low before committing to landing. For example, a UAV's takeoff and landing procedure may include a dock-specific landing detection and abort procedure. For example, actual contact with the dock may be detected, and the system may distinguish between a successful landing and a near miss. For example, a UAV's takeoff and landing procedure may include employing slow reverse motor rotation to enable self-retracting propellers.

[0043] Depending on the implementation, the UAV's takeoff and landing procedures may include support and fallback actions in case of failure, such as setting a predetermined landing location in case of failure; going to another box; or the option to land on a dock if the box fails.

[0044] For example, an application programming interface design for single drone, single dock operation may be provided. For example, engineering work may be performed on a schedule or as often as possible given battery life or recharge amounts.

[0045] For example, an application programming interface design may be provided for operation of N drones with M docks. In some implementations, mission parameters may be defined such that UAVs (e.g., drones) are automatically deployed and retrieved with overlap to always meet the mission parameters.

[0046] An unmanned air vehicle (UAV) may be configured to automatically fold its propellers to fit into a dock. For example, the dock may be smaller than the entire UAV. Multiple UAVs may be docked, charged, missioned, docked on standby, and / or cooperatively charged for sustained operation. In some implementations, the UAV may be automatically serviced while in place in the dock. For example, automatic servicing of a UAV may include: charging batteries, cleaning sensors, more general cleaning and / or drying of the UAV, propeller replacement, and / or battery replacement.

[0047] To be robust against drift, the UAV may track its state (e.g., attitude, including position and orientation) using a combination of sensing modalities (e.g., visual inertial odometry (VIO) and global positioning system (GPS)-based motion).

[0048] In some implementations, during takeoff and landing, the UAV continuously aims at the landing spot as it approaches the dock. This aiming process can make the takeoff and landing procedure robust to wind, ground effects, and other disturbances. For example, intelligent aiming can use position, heading, and trajectory to arrive within very tight tolerances. In some implementations, the rear motor may reverse for approach.

[0049] Some implementations can offer advantages over conventional systems, such as: a small, inexpensive, and simple dock; a retraction mechanism that allows for a circling hold and can reduce aerodynamic turbulence issues with landing; robust visual landings that can be more accurate; charging the UAV; automatic retraction of the propellers to allow for compact storage during maintenance and storage; the ability to service the UAV while docked without human intervention; and sustained autonomous operation of multiple vehicles via docks, SDKs, air vehicles, and services (hardware and software).

[0050] FIG. 3 illustrates an example dock 300 for facilitating autonomous landing of an unmanned air vehicle. Dock 300 includes a landing surface 310 with reference marks 320 and charging contacts 330 for a battery charger. Dock 300 includes a box 340 in the shape of a rectangular box with a door 342. To facilitate takeoff and landing of the unmanned air vehicle, dock 300 includes a retractable arm 350 that supports landing surface 310 and enables positioning of box 340 either externally or internally for storage and / or servicing of the unmanned air vehicle. Dock 300 includes a second auxiliary reference mark 322 on the upper exterior surface of box 340. Route reference marks 320 and auxiliary reference marks 322 may be detected and used for visual localization of the unmanned air vehicle in relation to dock 300 to enable precise landing on narrow landing surface 310. To land an unmanned air vehicle on landing surface 310 of dock 300, techniques described, for example, in US Patent Application No. 62 / 915,639, which is incorporated herein by reference, may be used.

[0051] Dock 300 includes landing surface 310 configured to hold an unmanned air vehicle (e.g., unmanned air vehicle 200 of FIG. 2 ) and fiducial marks 320 on landing surface 310. Landing surface 310 has a funnel-shaped geometry that fits onto the bottom of the unmanned air vehicle at the base of the funnel. The tapered sides of the funnel may help mechanically guide the bottom of the unmanned air vehicle to a centered position on the base of the funnel during landing. For example, after the bottom of the air vehicle fits into the base of the funnel-shaped shape of landing surface 310, corners at the base of the funnel may function to prevent the air vehicle from rotating on landing surface 310. For example, fiducial marks 320 may include an asymmetric pattern that enables robust detection and determination of the attitude (i.e., position and orientation) of fiducial marks 320 relative to the unmanned air vehicle based on an image of fiducial marks 320 captured by an image sensor of the unmanned air vehicle. For example, fiducial marks 320 may include a visual tag from the AprilTag family.

[0052] Dock 300 is positioned at the bottom of the funnel and includes conductive contacts 330 for a battery charger on landing surface 310. Dock 300 includes a charger configured to charge the battery while the unmanned air vehicle is on landing surface 310.

[0053] Dock 300 includes a box 340 configured to enclose landing surface 310 in a first configuration (shown in FIG. 4 ) and expose landing surface 310 in a second configuration (shown in FIGS. 3 and 3 ). Dock 300 may be configured to automatically transition from the first configuration to the second configuration by performing steps including opening a door 342 of box 340 and extending a retractable arm 350 to move landing surface 310 from the interior of box 340 to the exterior of box 340. Auxiliary reference marks 322 are located on the exterior surface of box 340.

[0054] Dock 300 includes retractable arms 350, with landing surface 310 positioned at the end of retractable arms 350. When retractable arms 350 are extended, landing surface 310 is positioned away from box 340 of dock 300, which can reduce or prevent propeller wash from the unmanned air vehicle's propellers during landing, thus simplifying the landing operation. Retractable arms 350 may include an aerodynamic cowling to redirect propeller wash to further mitigate the problem of propeller wash during landing.

[0055] For example, fiducial mark 320 may be a route fiducial mark, and auxiliary fiducial mark 322 may be larger than route fiducial mark 320 to facilitate visual localization from a greater distance as the unmanned air vehicle approaches dock 300. For example, the area of ​​auxiliary fiducial mark 322 may be 25 times the area of ​​route fiducial mark 320. For example, auxiliary fiducial mark 322 may include an asymmetric pattern that enables robust detection and determination of the attitude (i.e., position and orientation) of auxiliary fiducial mark 322 relative to the unmanned air vehicle based on an image of auxiliary fiducial mark 322 captured by an image sensor of the unmanned air vehicle. For example, auxiliary fiducial mark 322 may include a visual tag from the AprilTag family. For example, a processing device (e.g., processing device 410) of the unmanned air vehicle may be configured to detect auxiliary fiducial mark 322 in at least one of one or more images captured using an image sensor of the unmanned air vehicle; determine an attitude of auxiliary fiducial mark 322 based on the one or more images; and control a propulsion mechanism to fly the unmanned air vehicle to a first location near landing surface 310 based on the attitude of auxiliary fiducial mark 322. Thus, auxiliary fiducial mark 322 can facilitate the unmanned air vehicle getting sufficiently close to landing surface 310 to enable detection of route fiducial mark 320.

[0056] Dock 300 can enable automatic landing and recharging of the unmanned air vehicle, further enabling automatic scanning of large structures (e.g., large roofs, bridges, or large construction sites) that require multiple battery pack charges for scanning so that they can be scanned automatically without user intervention. For example, the unmanned air vehicle may be configured to: after initiating and before completing a scan plan, store a scan plan state indicating the next attitude in a series of attitudes for the scan plan; after storing the scan plan state, control the propulsion mechanism to fly the unmanned air vehicle for landing; after landing, control the propulsion mechanism to fly the unmanned air vehicle for takeoff; access the scan plan state; and, based on the scan plan state, control the propulsion mechanism to fly the unmanned air vehicle to the next attitude and continue executing the scan plan. In some implementations, controlling the propulsion mechanism to fly the unmanned air vehicle for landing includes: controlling the propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to a first location near a dock (e.g., dock 300) that includes a landing surface (e.g., landing surface 310) configured to hold the unmanned air vehicle and a reference mark on the landing surface; accessing one or more images captured using an image sensor of the unmanned air vehicle; detecting the reference mark in at least one of the one or more images; determining an attitude of the reference mark based on the one or more images; and controlling the propulsion mechanism to land the unmanned air vehicle on the landing surface based on the attitude of the reference mark. For example, this technique of automatic landing may include automatically charging a battery of the unmanned air vehicle using a charger included in the dock while the unmanned air vehicle is on the landing surface.

[0057] FIG. 4 is a block diagram illustrating an example hardware configuration 400 of an unmanned air vehicle. The hardware configuration may include a processing unit 410, a data storage device 420, a sensor interface 430, a communication interface 440, a propulsion control interface 442, a user interface 444, and an interconnect 450 through which the processing unit 410 can access other components. For example, the hardware configuration 400 may be or be part of an unmanned air vehicle (e.g., unmanned air vehicle 200). For example, the unmanned air vehicle can be configured to scan a structure (e.g., a roof, a bridge, or a construction site). For example, the unmanned air vehicle may be configured to perform process 600 of FIG. 6. In some implementations, the unmanned air vehicle may be configured to detect one or more fiducial marks on a dock (e.g., dock 300) to use an estimate of the attitude of the one or more fiducial marks for landing on a narrow landing surface to facilitate automated maintenance of the unmanned air vehicle.

[0058] The processing unit 410 is operable to execute instructions stored in the data storage device 420. In some implementations, the processing unit 410 is a processor having a random access memory for temporarily storing instructions read from the data storage device 420 during execution of the instructions. The processing unit 410 may include a single or multiple processors, each processor having a single or multiple processing cores. Alternatively, the processing unit 410 may include another type of device, or multiple devices, capable of manipulating or processing data. For example, the data storage device 420 may be a non-volatile information storage device, such as a solid-state device, a read-only memory device (ROM), an optical disk, a magnetic disk, or any other suitable type of storage device, such as non-transitory computer-readable memory. The data storage device 420 may also include another type of device, or multiple devices, capable of storing data for retrieval or processing by the processing unit 410. The processing unit 410 may access and manipulate data stored in the data storage device 420 via the interconnect 450. For example, data storage device 420 may store instructions executable by processing unit 410 that, when executed by processing unit 410, cause processing unit 410 to perform operations (e.g., operations to implement process 600 of FIG. 6, process 700 of FIG. 7, process 800 of FIG. 8, process 900 of FIG. 9, and / or process 1000 of FIG. 10).

[0059] The sensor interface 430 can be configured to control and / or receive data (e.g., temperature measurements, pressure measurements, global positioning system (GPS) data, acceleration measurements, angular velocity measurements, magnetic flux measurements, and / or visible spectrum imagery) from one or more sensors (e.g., including the image sensor 220). In some implementations, the sensor interface 430 may implement a serial port protocol (e.g., I2C or SPI) for communication with one or more sensor devices over wires. In some implementations, the sensor interface 430 may include a wireless interface for communicating with one or more sensor groups via low-power, short-range communications (e.g., a vehicle area network protocol).

[0060] The communication interface 440 facilitates communication with other devices, such as a mating dock (e.g., dock 300), a dedicated controller, or a user computing device (e.g., a smartphone or tablet). For example, the communication interface 440 may include a wireless interface that can facilitate communication over a Wi-Fi network, a Bluetooth link, or a ZigBee link. For example, the communication interface 440 may include a wired interface that can facilitate communication over a serial port (e.g., RS-232 or USB). The communication interface 440 facilitates communication over a network.

[0061] The propulsion control interface 442 can be used by the processing unit to control the propulsion system (e.g., including one or more propellers driven by electric motors). For example, the propulsion control interface 442 may include circuitry for converting digital control signals from the processing unit 410 into analog control signals for the actuators (e.g., the electric motors that drive each propeller). In some implementations, the propulsion control interface 442 may implement a serial port protocol (e.g., I2C or SPI) for communication with the processing unit 410. In some implementations, the propulsion control interface 442 may include a wireless interface for communication with the one or more motors via low-power, short-range communications (e.g., a vehicle area network protocol).

[0062] The user interface 444 allows for the input and output of information to and from a user. In some implementations, the user interface 444 may include a display, which may be a liquid crystal display (LCD), a light-emitting diode (LED) display (e.g., an OLED display), or other suitable display. For example, the user interface 444 may include a touchscreen. For example, the user interface 444 may include buttons. For example, the user interface 444 may include a positional input device, such as a touchpad, a touchscreen, or other suitable human or machine interface device.

[0063] For example, interconnect 450 may be a system bus or a wired or wireless network (e.g., a vehicle area network). In some implementations (not shown in FIG. 4), some components of the unmanned air vehicle, such as user interface 444, may be omitted.

[0064] FIG. 5A illustrates an example of a graphical user interface 500 associated with an unmanned air vehicle used to present two-dimensional polygonal projections of facets overlaid on an overview image of a structure to enable editing of the facets to facilitate structure scanning. Graphical user interface 500 includes an overview image 510 of the structure (e.g., a static image of a roof as shown in FIG. 5A ). Graphical user interface 500 includes a graphical representation of a two-dimensional polygon 520 corresponding to the facet proposal, which is a projection of a convex hull of points of a three-dimensional map of the structure. This two-dimensional polygon 520 includes four vertices, including vertex 522. A user can edit 2D polygon 520 by interacting with vertex 522 (e.g., using a touchscreen display interface) to move vertex 522 within the plane of overview image 510. Once the user is satisfied with the visual coverage of 2D polygon 520, the user can interact with confirm icon 530 to cause data indicating the user edits of 2D polygon 520 to be returned to the unmanned aerial vehicle, which can then determine facets based on the facet suggestions and the user edits. Graphical user interface 550 may then be updated to present the final facets with a final 2D polygon overlaid on overview image 510, similar to final 2D polygon 540 shown for a nearby section of the roof structure without interactive vertices. This process may continue with the user reviewing and / or editing the facet suggestions until the structure is satisfactorily covered by facets. For example, the user interface may be displayed on a computing device remote from the unmanned aerial vehicle, such as controller 120. For example, the unmanned aerial vehicle may be configured to present graphical user interface 500 to the user by transmitting data encoding this graphical user interface 500 to a computing device (e.g., controller 250 for display on touchscreen 260).

[0065] FIG. 5B illustrates an example of an unmanned air vehicle graphical user interface 550 used to present a scan plan overlaid on an overview image of the structure for user review to facilitate structure scanning. Graphical user interface 550 includes an overview image 510 of the structure (e.g., a still image of a roof as shown in FIG. 5B ). Graphical user interface 550 includes a graphical representation of an image sensor's field of view 570 from a given pose among a series of poses in the scan plan. Field of view 570 may be projected within the plane of overview image 510. To facilitate user review and approval of the scan plan, a collection of fields of view corresponding to each pose in the scan plan provides a graphical representation of the scan plan. In some implementations, a user can adjust scan plan parameters, such as vertical overlap, horizontal overlap, and distance from the surface, to recreate the scan plan and resulting field of view for those poses. Once the user is satisfied with the scan plan, the user can approve the scan plan by interacting with approval icon 580, causing the unmanned air vehicle to begin executing the scan plan.

[0066] 6 is a flowchart illustrating an example process 600 for scanning a structure using an unmanned air vehicle. Process 600 includes accessing 610 a three-dimensional map of the structure, where the three-dimensional map encodes a set of points in three-dimensional space on a surface of the structure; generating 620 one or more facets based on the three-dimensional map, where a given facet of the one or more facets is a polygon on a surface in three-dimensional space that fits a subset of the points in the three-dimensional map; generating 630 a scanplan based on the one or more facets, where the scanplan includes a series of poses for the unmanned air vehicle to assume to enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors of the unmanned air vehicle; controlling 640 a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle in an pose corresponding to one of the poses of the scanplan; and capturing 650 one or more images of the structure from the pose using the one or more image sensors. For example, process 600 may be performed by unmanned aerial vehicle 110 in Figure 1. For example, process 600 may be performed by unmanned aerial vehicle 200 in Figures 2A-2B. For example, process 600 may be implemented using hardware configuration 400 in Figure 4.

[0067] Process 600 includes accessing 610 a three-dimensional map of the structure. The three-dimensional map encodes a set of points in three-dimensional space on the surface of the structure. For example, the structure may be a building roof, a bridge, or a building under construction. In some embodiments, the three-dimensional map may include a voxel occupancy map or a signed distance map. For example, the three-dimensional map may have been generated based on sensor data collected by a range sensor (e.g., an array of image sensors configured for stereoscopic computer vision, a radar sensor, and / or a lidar sensor). In some implementations, the unmanned air vehicle accessing 610 the three-dimensional map itself has recently generated the three-dimensional map by performing a relatively low-resolution scan using a range sensor while operating at a safe distance from the structure. In some implementations, the structure is scanned to generate the three-dimensional map from a distance greater than the fixed distance used for facet imaging. For example, process 1000 of FIG. 10 may be performed to generate the three-dimensional map. For example, process 1100 of FIG. 11 may be performed to generate the three-dimensional map. The three-dimensional map can be accessed 610 in a variety of ways. For example, the three-dimensional map may be accessed 610 directly from a range sensor via a sensor interface (e.g., sensor interface 430) or by reading it from a memory (e.g., data storage device 420) via an interconnect (e.g., interconnect 450).

[0068] The process 600 includes generating 620 one or more facets based on the three-dimensional map. A given facet of the one or more facets is a polygon on a surface in three-dimensional space that fits a subset of points in the three-dimensional map. For example, a facet may be identified by searching the three-dimensional map for the largest expanse of coplanar points with a low proportion of outlier points, and then fitting a surface to these subsets of points. In some implementations, isolated outlier points may be filtered out.

[0069] In some implementations, user input may be used to identify portions of facets and / or refine the boundaries of facets. For example, a graphical user interface (e.g., graphical user interface 500) may provide a user with an overview image (e.g., a static viewpoint) of the structure. The user may click on the center of a facet as it appears in the overview image. One or more points in the overview image at the location of the click interaction are projected onto a point in the three-dimensional map, or equivalently, points from the top surface of the three-dimensional map are projected onto the overview image and associated with the location of the click interaction. Once a mapping from the click interaction location to a small subset of points in the three-dimensional map is established, a surface is fitted to this small subset of points (e.g., using random sample consensus (RANSAC)). The entire three-dimensional map surface may then be considered to select points that are coplanar with and adjacent to the points in the small subset, and this subset may be iteratively refined. When the iterations converge, the resulting subset of points in the three-dimensional map becomes the basis for facet suggestions. The convex hull of these points when projected onto the image may be calculated to obtain a 2D polygon in the image plane of the overview image. In some implementations, a user click over the top of the structure is simulated, and the proposed facet boundary is used as the final facet boundary to determine the facets more quickly. In some implementations, these locations of the 3D facets are jointly optimized for cleaner boundaries between facets. In some implementations, image-based machine learning is used to detect facets in image space (e.g., the plane of the overview image) instead of 3D space.

[0070] The resulting two-dimensional polygon (or convex hull) may be simplified by removing edges and extending neighboring edges, as long as the area or edge length of the resulting polygon is not excessively increased. More specifically, for an input polygon, each edge may be considered. If an edge is “convex” such that two adjacent edges would meet outside the polygon, the polygon that would result from removing the convex edge and intersecting the corresponding adjacent edge is considered. The increase in area and increase in edge length that would result from using this alternative polygon may be considered. For example, the “convex” edge that would have the smallest area increase may be removed as long as both the increase in area and edge length are below specified thresholds. For example, input polygon 1400 of FIG. 14A may be simplified to obtain simplified polygon 1450 of FIG. 14B. To simplify a two-dimensional polygon representing a facet proposal, process 800 of FIG. 8 may be performed, for example. This simplified polygon may be presented to a user. The user can then move, add, or remove vertices of the polygons in the surface of the overview image to better fit the desired facet as it appears to the user in the overview image. To fit the final facet, all of the surface points are projected onto the image and RANSAC is run on them to find the 3D surface on which the facet lies. The 2D vertices of the polygon in the image may then be intersected with this surface to determine a polygon within this facet surface, which is the final facet. Surface points in the 3D map that belong to this facet may be ignored when suggesting or fitting subsequent facets. For example, generating 620 one or more facets based on the 3D map may include performing process 700 of FIG. 7 to solicit user feedback on the proposed facets.

[0071] Process 600 includes generating 630 a scanplan based on the one or more facets. The scanplan includes a series of poses for the unmanned air vehicle to assume to enable capture of images of the structure at a fixed distance (e.g., 1 meter) from each of the one or more facets using one or more image sensors (e.g., including image sensor 220) of the unmanned air vehicle. A pose in the series of poses may include a position of the unmanned air vehicle (e.g., a tuple of coordinates x, y, and z), an orientation of the unmanned air vehicle (e.g., Euler angles or a quaternion). In some implementations, a pose may include an orientation of the unmanned air vehicle's image sensor with respect to the unmanned air vehicle or with respect to another coordinate system. After the set of facets is generated, the unmanned air vehicle can plan a path to capture images of all facets at a desired ground sampling distance (GSD). After the path is generated, it may be presented to a user via a graphical user interface (e.g., a live augmented reality (AR) display) for the user to approve or reject. For example, graphical user interface 550 of FIG. 5B may be used to present the scanplan to the user for approval. In some implementations, process 600 includes capturing a general image of the structure using one or more image sensors; presenting a graphical representation of the scan plan overlaid on the general image; and receiving an indication from a user of approval of the scan plan.

[0072] A scanplan may be generated 630 based on one or more facets and several scanplan configuration parameters such as distance from the surface, vertical and horizontal overlap between fields of view of one or more image sensors at different poses in the scanplan series, etc. For example, the scanplan series may be for orthogonal imaging of each of the one or more facets.

[0073] Process 600 includes controlling 640 a propulsion mechanism of the unmanned air vehicle (e.g., unmanned air vehicle 200) to fly the unmanned air vehicle in an attitude corresponding to one of a series of attitudes in the scan plan; and capturing 650 one or more images of the structure from that attitude using one or more image sensors (e.g., image sensor 220). For example, steps 640 and 650 may be repeated for each of the scan attitudes until images covering all of one or more facets have been captured. In some implementations, a processing device may be configured to stitch the captured images together to obtain a composite image of one or more surfaces of the structure. For example, image stitching may be performed in part based on out-of-band information associated with the images via each facet, such as three-dimensional map points associated with the facet or the boundaries of one or more facets. For example, a processing device (e.g., processing device 410) may use a propulsion controller interface (e.g., propulsion controller interface 442) to control 640 a propulsion mechanism (e.g., one or more propellers driven by electric motors).

[0074] During scanning, the vehicle may fly the calculated path while capturing images. In addition to avoiding obstacles, the vehicle may dynamically update the path for things like obstacle avoidance or improved image registration. For example, while flying between attitudes in the series of attitudes in the scan plan, process 600 may include detecting an obstacle based on images captured using one or more image sensors; and dynamically adjusting an attitude in the series of attitudes in the scan plan to avoid the obstacle. For example, while flying between attitudes in the series of attitudes in the scan plan, the air vehicle may detect a deviation of a point on the surface of the structure from one of the one or more facets based on images captured using one or more image sensors; and dynamically adjust an attitude in the series of attitudes in the scan plan to accommodate the deviation and maintain a constant distance for image capture.

[0075] During scanning, the operator can monitor the drone via a static viewpoint (e.g., an overview image of the structure) or from a live feed from the vehicle's camera. The operator also has the controls to intervene manually at this stage.

[0076] When the unmanned air vehicle completes the scan or needs to abort the scan (e.g., due to a low battery or vehicle malfunction), the vehicle can automatically return to the takeoff point and land. To the extent that the unmanned air vehicle must land before completing the scan plan, it may be useful to save the scan progress state so that the unmanned air vehicle can resume scanning where it left off after any conditions that caused the unmanned air vehicle to land are resolved. For example, process 600 may include storing a scan plan state after initiating and before completing the scan plan that indicates a next attitude in the series of attitudes of the scan plan; controlling the propulsion mechanism to fly the unmanned air vehicle for landing after storing the scan plan state; controlling the propulsion mechanism to fly the unmanned air vehicle for takeoff after landing; accessing the scan plan state; and controlling the propulsion mechanism to fly the unmanned air vehicle to assume the next attitude and continue executing the scan plan based on the scan plan state. The scan plan includes at least a series of attitudes of the unmanned air vehicle and may also include additional information. The pose may be encoded in various coordinate systems (e.g., a global coordinate system, or a coordinate system relative to a dock for the unmanned air vehicle or relative to the structure being scanned). The scan plan state can be used in conjunction with a visual inertial odometry (VIO) system to assume the next pose in the scan plan after recharging. For example, the unmanned air vehicle can automatically land and take off from dock 300 of FIG. 3 after automatically charging its battery.In some implementations, controlling the propulsion mechanism to fly the unmanned air vehicle for landing includes: controlling the propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to a first location near a dock (e.g., dock 300) that includes a landing surface (e.g., landing surface 310) configured to hold the unmanned air vehicle and a reference mark on the landing surface; accessing one or more images captured using an image sensor of the unmanned air vehicle; detecting the reference mark in at least one of the one or more images; determining an attitude of the reference mark based on the one or more images; and controlling the propulsion mechanism to land the unmanned air vehicle on the landing surface based on the attitude of the reference mark. For example, process 600 may include automatically charging a battery of the unmanned air vehicle using a charger included in the dock while the unmanned air vehicle is on the landing surface.

[0077] Upon completion of the scanplan execution, collected data (e.g., high-resolution images of the surface of the structure (e.g., a roof, bridge, or construction site) and associated metadata) may be transmitted to another device (e.g., controller 120 or a cloud server) for viewing or offline analysis. Estimates of the area of ​​the facets and / or cost estimates for repairing the facets may be useful. In some implementations, process 600 includes determining an area estimate for each of one or more facets; and presenting (e.g., transmitting, storing, or displaying) a data structure including the one or more facets, the area estimates for each of the one or more facets, and images of the structure captured during the scanplan execution. In some implementations, a status report summarizing the progress or effectiveness of the scanplan execution may be presented. For example, process 900 of FIG. 9 may be implemented to generate and present a scanplan coverage report.

[0078] Once the unmanned air vehicle has landed, it can begin transmitting data to the operator device, which may include stitched composite images of each facet, photographs taken, and metadata including camera pose and flight summary data (number of facets, photo capture, percentage of flight completed, flight time, etc.).

[0079] FIG. 7 is a flowchart illustrating an example of a process 700 for enabling user editing of facets. Process 700 includes capturing an overview image of a structure using one or more image sensors 710; generating facet proposals based on a three-dimensional map 720; determining a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the two-dimensional polygon corresponding to the facet proposal when the subset of points is projected onto an image plane of the overview image 730; presenting the two-dimensional polygon overlaid on the overview image 740; determining an edited two-dimensional polygon in the image plane of the overview image based on data indicating user editing of the two-dimensional polygon 750; and determining one of one or more facets based on the edited two-dimensional polygon 760. For example, process 700 may be performed by unmanned aerial vehicle 110 of FIG. 1 . For example, process 700 may be performed by unmanned aerial vehicle 200 of FIGS. 2A-2B . For example, process 700 may be performed using hardware configuration 400 of FIG. 4 .

[0080] The process 700 includes capturing 710 an overview image of the structure using one or more image sensors (e.g., image sensors (220-226)). The overview image can be used as a still picture of the structure that can form part of a graphical user interface to enable a user to track the progress of the execution of the scan plan and provide user feedback at various stages of the structure scanning process. Incorporating the overview image into the graphical user interface can facilitate localization of the user's intent relative to the structure being scanned by relating pixels of the graphical user interface to points on the three-dimensional surface of the structure in a three-dimensional map. For example, the overview image may be captured 710 from a position sufficiently far away from the structure such that the entire structure appears within the field of view of the image sensor used to capture 710 the overview image.

[0081] The process 700 includes generating 720 facet proposals based on the three-dimensional map. For example, facet proposals may be generated 720 by searching the three-dimensional map for the largest spread of coplanar points with a low proportion of outliers, and then fitting a surface to this subset of points. Depending on the implementation, isolated outliers may be filtered out. User input may be used to identify a portion of a facet of interest. For example, an overview image (e.g., a static viewpoint) of the structure may be presented to the user in a graphical user interface (e.g., graphical user interface 500). The user may click on the center of a facet that appears in the overview image. One or more points at the location of the click interaction in the overview image may be projected onto a point on the three-dimensional map, or equivalently, points from the top surface of the three-dimensional map are projected onto the overview image and associated with the location of the click interaction. Once a mapping from the click interaction location to a small subset of points on the three-dimensional map is established, a surface is fitted to this small subset of points (e.g., using random sample consensus (RANSAC)). The entire 3D map surface is then considered to select points that are coplanar and neighboring with the points in that smaller subset, and this subset can be iteratively refined. Once the iterations converge, the resulting subset of 3D map points becomes the basis for facet proposals.

[0082] The process 700 includes determining 730 a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the subset of points corresponding to the facet proposal when projected onto the image plane of the overview image. The convex hull of these points when projected onto the image may be calculated to obtain a two-dimensional polygon in the image plane of the overview image. In some implementations, the two-dimensional polygon is simplified before being presented 740. For example, the process 800 of FIG. 8 may be performed to simplify the two-dimensional polygon.

[0083] Process 700 includes presenting 740 a two-dimensional polygon overlaid on the overview image. For example, the two-dimensional polygon overlaid on the overview image may be presented 740 as part of a graphical user interface (e.g., graphical user interface 500 of FIG. 5A ). For example, a processing unit of the unmanned air vehicle may present 740 the two-dimensional polygon overlaid on the overview image by transmitting data encoding the two-dimensional polygon overlaid on the overview image to a user computing device (e.g., controller 120) (e.g., via a wireless communications network).

[0084] Process 700 includes determining 750 an edited two-dimensional polygon in the image plane of the overview image based on data indicative of user edits of the two-dimensional polygon. For example, the data indicative of user edits of the two-dimensional polygon may have been generated by a user interacting with a graphical user interface (e.g., graphical user interface 500), such as by dragging vertex icons to move vertices of the two-dimensional polygon within the plane of the overview image (e.g., using touchscreen 260). For example, the data indicative of the user edits may be received by the unmanned air vehicle via a network communications interface (e.g., communications interface 440).

[0085] The process 700 includes determining 760 one of the one or more facets based on the edited two-dimensional polygon. The edited two-dimensional polygon may be mapped to a new subset of points of the three-dimensional map. In some implementations, all points of the three-dimensional map are projected onto a surface of the overview image, and those points that have a projection within the edited two-dimensional polygon are selected as members of a new subset of points underlying the determined 760 new facet. In some implementations, an inverse projection of the edited two-dimensional polygon is used to select the new subset of points underlying the determined 760 new facet. For example, determining 760 one of the one or more facets may include fitting a surface to the new subset of points and calculating a convex hull of the points in the new subset when projected onto the surface of the new facet.

[0086] 8 is a flowchart illustrating an example of a process 800 for attempting to simplify a polygon representing a facet by removing convex edges. Process 800 includes identifying 810 convex edges of a two-dimensional polygon; determining 820 the area increase resulting from removing the convex edges; verifying in step 825 that removing the convex edges increases the area of ​​the two-dimensional polygon by an amount less than a threshold; and if (in step 825) the area increase is equal to or greater than the threshold (e.g., a 10% increase), retaining 830 the convex edges in the two-dimensional polygon and repeating process 800 as necessary for any other convex edges in the two-dimensional polygon. If (step 825) the area increase does not exceed a threshold (e.g., a 10% increase), determine 840 the increase in perimeter resulting from removing the convex edge; verify (step 845) that removing the convex edge increases the perimeter of the 2D polygon by an amount less than the threshold (e.g., a 10% increase); if (step 845) the perimeter increase is equal to or greater than the threshold (e.g., a 10% increase), retain 830 the convex edge in the 2D polygon and repeat process 800 as necessary for any other convex edges in the 2D polygon. If (step 845) the increase is less than the threshold (e.g., a 10% increase), simplify 850 the 2D polygon by removing the convex edge from the 2D polygon and extending edges of the 2D polygon adjacent to the convex edge to the points where the extended edges intersect each other. Process 800 may be repeated for any other convex edges in the 2D polygon. In some implementations, only the increase in perimeter resulting from removing the convex edge is verified. In some implementations, only the area increase resulting from the removal of convex edges is identified. For example, process 800 may be performed to simplify input polygon 1400 of FIG. 14A to obtain simplified polygon 1450 of FIG. 14B. For example, process 800 may be performed by unmanned aerial vehicle 110 of FIG. 1. For example, process 800 may be performed by unmanned aerial vehicle 200 of FIGS. 2A-2B. For example, process 800 may be performed using hardware configuration 400 of FIG. 4.

[0087] FIG. 9 is a flowchart illustrating an example of a process 900 for presenting coverage information for a scan of a structure. Process 900 includes generating 910 a coverage map of one or more facets that indicates which of the one or more facets were successfully imaged during execution of the scan plan; and presenting 920 the coverage map. The unmanned aerial vehicle may calculate the image coverage of selected facets on-board so that the operator can verify that all data was captured. If a facet is determined to lack sufficient coverage, an application on the operator device (e.g., controller 120) may indicate the location of the coverage deficiency and prompt an action to obtain coverage (e.g., generate an automated path to capture the missing image or prompt the operator to manually fly the unmanned aerial vehicle to capture the image). For example, process 900 may be performed by unmanned aerial vehicle 110 of FIG. 1. For example, process 900 may be performed by unmanned aerial vehicle 200 of FIGS. 2A-2B. For example, the process 900 may be implemented using the hardware configuration 400 of FIG.

[0088] FIG. 10 is a flowchart illustrating an example of a process 1000 for generating a three-dimensional map of a structure. Process 1000 includes controlling 1010 a propulsion mechanism to fly an unmanned air vehicle adjacent to the structure; and scanning 1020 the structure using one or more image sensors configured to support stereoscopic imaging, which are used to provide distance data, to generate a three-dimensional map of the structure. For example, the three-dimensional map may include a voxel occupancy map or a signed distance map. For example, the structure may be the roof of a building. For example, process 1100 of FIG. 11 may be implemented to scan 1020 the roof. For example, the structure may be a bridge. In some implementations, scanning is performed from a single orientation far enough from the structure that the entire structure is within the field of view of one or more image sensors (e.g., image sensors 224, 225, and 226). For example, the structure may be a building under construction. For example, process 1000 may be implemented by unmanned air vehicle 110 of FIG. 1. For example, process 1000 may be implemented by unmanned air vehicle 200 of FIGS. 2A-2B. For example, the process 1000 may be implemented using the hardware configuration of FIG.

[0089] 11 is a flowchart illustrating an example of a process 1100 for generating a three-dimensional roof map. Process 1100 includes capturing 1110 an overview image of a building roof from a first pose of the unmanned air vehicle positioned above the roof; presenting 1120 a user with a graphical representation of a proposed bounding polygon overlaid on the overview image; accessing 1130 data encoding user edits of one or more vertices of the proposed bounding polygon; determining 1140 a bounding polygon based on the proposed bounding polygon and the data encoding the user edits; determining 1150 a flight path based on the bounding polygon; controlling 1160 a propulsion mechanism to fly the unmanned air vehicle through a series of scan poses with a horizontal position aligned with each pose of the flight path and a vertical position determined to maintain a constant distance above the roof; and scanning 1170 the roof from the series of scan poses to generate a three-dimensional map of the roof. For example, process 1100 may be performed by unmanned air vehicle 110 of FIG. 1. For example, process 1100 may be performed by unmanned air vehicle 200 of Figures 2A-2B. For example, process 1100 may be performed using hardware configuration 400 of Figure 4.

[0090] Process 1100 includes capturing 1110 an overview image of a building roof using one or more image sensors (e.g., image sensor 220) of the unmanned air vehicle (e.g., unmanned air vehicle 200) from a first attitude of the unmanned air vehicle positioned above the roof. The overview image may be used as a still image of the structure, which may form part of a graphical user interface for enabling a user to track the progress of the unmanned air vehicle along a flight path (e.g., a flight path relative to a dynamic surface), which will be used to generate a three-dimensional map of the roof and to provide user feedback at various stages of the scanning procedure. Incorporating the overview image into the graphical user interface can facilitate localization of user intent relative to the roof being scanned by associating pixels of the graphical user interface with portions of the roof. For example, the overview image can be captured 1110 from an attitude sufficiently far from the roof so that the entire roof appears within the field of view of the image sensor used to capture 1110 the overview image.

[0091] In some implementations, the unmanned air vehicle may be configured to automatically fly to the attitude used to capture 1110 the overview image of the roof. For example, a user may initially point the vehicle on the ground toward a building whose roof is to be scanned. The user may operate a takeoff icon in the unmanned air vehicle's user interface, causing the unmanned air vehicle to take off and ascend diagonally above the target building of interest, looking directly down at the roof of the building below and flying high enough to capture 1110 all of the area of ​​interest within the field of view of one or more image sensors (e.g., image sensor 220). In some implementations, the unmanned air vehicle may be manually controlled to the attitude used to capture 1110 the overview image of the roof, and process 1100 may begin once the unmanned air vehicle is so positioned.

[0092] Process 1100 includes presenting 1120 to the user a graphical representation of the proposed bounding polygon overlaid on the overview image. The proposed bounding polygon includes vertices corresponding to respective vertex icons in the graphical representation that allow the user to move the vertices within the plane. For example, the proposed bounding polygon may be a rectangle in a horizontal plane. In some implementations, the proposed bounding polygon (e.g., a triangle, rectangle, pentagon, or hexagon) is overlaid at the center of the overview image and has a fixed default size. In some implementations, the proposed bounding polygon is generated using computer vision processing to identify the perimeter of the roof as it appears in the overview image and generate a proposed bounding polygon that closely corresponds to the identified perimeter of the roof. In some implementations (e.g., the overview image is captured from an oblique perspective), the proposed bounding polygon is projected from a horizontal plane onto the surface of the overview image before being overlaid on the overview image. For example, a graphical representation of the proposed bounding polygon may be presented 1120 as part of a graphical user interface (e.g., graphical user interface 1300 of FIGS. 13A-13B ). For example, a processing unit (e.g., processing unit 410) of the unmanned air vehicle may present 1120 the graphical representation of the proposed bounding polygon overlaid on the overview image by transmitting (e.g., via a wireless communication network) to a user computing device (e.g., controller 120) data encoding the graphical representation of the proposed bounding polygon overlaid on the overview image.

[0093] Process 1100 includes accessing 1130 data encoding user edits of one or more of the vertices of the proposed bounding polygon. A user may use a computing device (e.g., controller 120, tablet, laptop, or smartphone) to receive, interpret, and / or interact with the graphical user interface in which the proposed bounding polygon is presented 1120. For example, a user may use a touchscreen to interact with one or more of the vertex icons to move the vertices of the proposed bounding polygon to edit the proposed bounding polygon corresponding to the perimeter of the roof that is scanned as it appears in the overview image. The user may use their computing device to encode these edits to one or more vertices of the proposed bounding polygon as data, which may be transmitted to a device (e.g., unmanned air vehicle 200) performing process 1100 and further processed therein. The device receives the data. For example, this data may include changed coordinates in the plane of the vertices of the proposed bounding polygon. The data encoding user edits of one or more of the vertices of the proposed bounding polygon may be accessed 1130 in a variety of ways. For example, the data encoding user edits of one or more of the vertices of the proposed bounding polygon may be accessed 1130 by receiving it from a remote computing device (e.g., controller 120) via a communications interface (e.g., communications interface 440). For example, the data encoding user edits of one or more of the vertices of the proposed bounding polygon may be accessed 1130 by being read from a memory (e.g., data storage device 420) via an interconnect (e.g., interconnect 450).

[0094] Process 1100 includes determining 1140 a bounding polygon based on the proposed bounding polygon and data encoding the user edits. The data encoding the user edits may be incorporated to update one or more vertices of the proposed bounding polygon to determine 1140 the bounding polygon. In some implementations (e.g., the overview image is captured from an oblique perspective), the bounding polygon is projected from the plane of the overview image onto a horizontal plane. For example, the bounding polygon may be a geofence of an unmanned air vehicle.

[0095] Process 1100 includes determining 1150 a flight path (e.g., a flight path relative to the dynamic surface) based on the bounding polygon. The flight path includes a series of poses of the unmanned air vehicle having respective fields of view at fixed heights that collectively cover the bounding polygon. For example, the flight path may be determined as a lawnmower pattern. In some implementations, the user (e.g., using a user interface used to edit the proposed bounding polygon) also inputs or selects an approximate height (e.g., above ground) that, together with the bounding polygon, defines a volume in three-dimensional space within which the roof is expected to reside. For example, a bounding box in three-dimensional space may be determined based on this height parameter and the bounding polygon, and a flight path may be determined 1150 based on the bounding box. In some implementations, additional parameters of the three-dimensional scanning operation may be specified or adjusted by the user. For example, the flight path may be determined 1150 based on one or more scan parameters presented for selection by the user, including one or more parameters from a set of parameters including a grid size, a nominal height above the roof surface, and a maximum flight speed.

[0096] Process 1100 includes controlling 1160 the propulsion mechanism to fly the unmanned air vehicle through a series of scanning attitudes having horizontal positions consistent with each attitude of the flight path (e.g., the flight path relative to the dynamic surface) and vertical positions determined to maintain a constant distance (e.g., 3 meters or 5 meters) above the roof. The unmanned air vehicle may be configured to automatically detect and avoid obstacles (e.g., chimneys or tree branches) encountered during the scanning procedure. For example, while flying between attitudes in the series of scanning attitudes, an obstacle may be detected based on images captured using one or more image sensors (e.g., image sensors 220-226) of the unmanned air vehicle; the attitude of the flight path may be dynamically adjusted to avoid the obstacle. For example, a roof may be scanned to generate a three-dimensional map from a distance greater than the constant distance used for facet imaging (e.g., using process 600 of FIG. 6 ), which may make scanning 1170 to generate a three-dimensional map of the roof safer and faster. For example, a processing unit (eg, processing unit 410) may control 1160 a propulsion mechanism (eg, one or more propellers driven by an electric motor) using a propulsion controller interface (eg, propulsion control interface 442).

[0097] In some implementations, after the 3D bounding box is defined, a small number of points of interest, such as an oblique viewpoint (e.g., from an elevated position, looking down) at a corner of the roof, are generated and flown over, and the unmanned air vehicle then flies that flight path (e.g., a flight path relative to the dynamic surface) to generate a 3D map of the roof.

[0098] Process 1100 includes scanning 1170 the roof from a series of scan poses to generate a three-dimensional map of the roof. For example, the three-dimensional map may include a voxel occupancy map or a signed distance map. For example, one or more image sensors may be configured to support stereoscopic imaging used to provide distance data, and the roof may be scanned 1170 using the one or more image sensors to generate a three-dimensional map of the roof. Depending on the implementation, the unmanned air vehicle may include other types of range or distance sensors (e.g., a lidar sensor or a radar sensor). For example, the roof may be scanned 1170 using a radar sensor to generate a three-dimensional map of the roof. For example, the roof may be scanned 1170 using a lidar sensor to generate a three-dimensional map of the roof.

[0099] A 3D map (e.g., a voxel map) may be created by fusing stereo range images from an airborne image sensor. For example, voxels in the 3D map may be marked as occupied or empty space. Surface voxels may be a subset of occupied voxels that are adjacent to empty space. In some implementations, surface voxels may simply be the highest occupied voxel at each horizontal (x,y) location.

[0100] For example, the 3D map may be a signed distance map. The 3D map may be created by fusing stereo range images from an onboard image sensor. The 3D map may be represented as a dense voxel grid of signed distance values. For example, the signed distance map may be a truncated signed distance field (TSDF). Values ​​can be updated by projecting voxel centers onto the range image and updating a weighted average of the signed distance values. The top surface of the signed distance map may be calculated by ray marching using rays selected at the desired resolution. Depending on the implementation, implicit surface locations (e.g., zero crossings) of the signed distance function along the rays may be interpolated for improved accuracy.

[0101] In some cases, it may be advantageous to interrupt the scanning procedure, such as when the unmanned air vehicle needs to be recharged. Maintaining a low-drift visual inertial odometry (VIO) estimate of the unmanned air vehicle's position as it moves can enable the interruption with a relatively seamless continuation of the scanning procedure after performing an intervening task, such as recharging. For example, process 1100 may include storing a scan state indicating a next attitude in a series of attitudes for a flight path; controlling a propulsion mechanism to fly the unmanned air vehicle for landing (e.g., on dock 300) after storing the scan state; controlling the propulsion mechanism to take off the unmanned air vehicle after landing; accessing the scan state; and controlling the propulsion mechanism, based on the scan state, to fly the unmanned air vehicle to an attitude corresponding to the next attitude in the series of scan attitudes and continue scanning 1170 the roof to generate a three-dimensional map.

[0102] FIG. 12 is a flowchart illustrating an example of a process 1200 for presenting progress information for a roof scan. For example, a scan may be performed (e.g., using process 1100 of FIG. 11 ) to generate a three-dimensional map of the roof. Process 1200 includes presenting 1210 a graphical representation of the unmanned air vehicle overlaid on an overview image; and presenting 1220 an indication of progress along a flight path (e.g., a flight path relative to a dynamic surface) overlaid on the overview image. The overview image is used as a static viewpoint in the unmanned air vehicle's user interface. As the unmanned air vehicle continues to fly and approach the roof, a background image displayed in the user interface may be static in the overview image, but additional situational information regarding the scanning procedure being performed may be updated and overlaid on this background image to provide spatial context for the situational information. For example, process 1200 may be performed by unmanned air vehicle 110 of FIG. 1 . For example, process 1200 may be performed by unmanned air vehicle 200 of FIGS. 2A-2B . For example, process 1200 may be performed using hardware configuration 400 of FIG. 4 .

[0103] Process 1200 includes presenting 1210 a graphical representation of the unmanned air vehicle overlaid on the overview image. The graphical representation of the unmanned air vehicle corresponds to the current horizontal position of the unmanned air vehicle. In some implementations, the graphical representation of the unmanned air vehicle includes a three-dimensional rendering of the unmanned air vehicle. For example, the three-dimensional rendering of the unmanned air vehicle can be depicted in a user interface in perspective with respect to the unmanned air vehicle's physical location (e.g., current location or planned location). For example, the unmanned air vehicle's physical location relative to the roof, as viewed from the overview image, can be determined by maintaining a low-drift visual inertial odometry (VIO) estimate of the unmanned air vehicle's position as it moves. Presenting 1210 the graphical representation (e.g., three-dimensional rendering) of the unmanned air vehicle can enable a user to view the unmanned air vehicle in the context of the overview image to better understand where the unmanned air vehicle is located relative to the roof and the scanning procedure at hand. For example, the graphical representation of the unmanned air vehicle may be presented 1210 as part of a graphical user interface (e.g., graphical user interface 1300 of FIGS. 13A-13B). For example, a processing unit (e.g., processing unit 410) of the unmanned air vehicle may present 1210 the graphical representation of the unmanned air vehicle overlaid on the overview image by transmitting (e.g., via a wireless communication network) data encoding the graphical representation of the unmanned air vehicle overlaid on the overview image to a user computing device (e.g., controller 120).

[0104] Process 1200 includes presenting 1220 an indication of progress along the flight path (e.g., the flight path relative to the dynamic surface) superimposed on the overview image. For example, the indication of progress along the flight path may include color-coding sections of the roof that were successfully scanned from an attitude corresponding to the attitude of the flight path. Presenting 1220 the indication of progress along the flight path may enable a user to understand the status and / or shape estimation of the 3D scanning procedure and the path plan for future steps. For example, the indication of progress along the flight path may be presented 1220 as part of a graphical user interface (e.g., graphical user interface 1300 of FIGS. 13A-13B ). For example, a processing unit (e.g., processing unit 410) of the unmanned air vehicle may present 1220 an indication of progress along the flight path superimposed on the overview image by transmitting (e.g., via a wireless communication network) data encoding the indication of progress along the flight path superimposed on the overview image to a user computing device (e.g., controller 120).

[0105] 13A illustrates an example of a graphical user interface 1300 of an unmanned air vehicle (e.g., unmanned air vehicle 200) used to present a proposed bounding polygon overlaid on an overview image of a roof to enable editing of the bounding polygon to facilitate rooftop scanning. Graphical user interface 1300 includes overview image 1310 including a view of building roof 1320. Graphical user interface 1300 also includes a graphical representation of proposed bounding polygon 1330 overlaid on overview image 1310. The graphical representation of the proposed bounding polygon includes vertex icons 1340, 1342, 1344, and 1346 corresponding to respective vertices of the proposed bounding polygon. A user can interact with one or more of vertex icons 1340, 1342, 1344, and 1346 (e.g., using a touchscreen of their computing device) to move the corresponding vertex of the proposed bounding polygon.

[0106] 13B illustrates an example of an unmanned air vehicle graphical user interface 1300 used to present a proposed bounding polygon overlaid on an overview image of a roof to enable editing of the bounding polygon to facilitate roof scanning. FIG. 13B shows graphical user interface 1300 after a user interacts with vertex icons 1340, 1342, 1344, and 1346 to edit the proposed bounding polygon to correspond to the entire perimeter of the roof being scanned. In this example, the user is using the zoom functionality of graphical user interface 1300 to zoom in on a portion of overview image 1310 to facilitate fine-tuning the position of vertex icon 1340 and vertex icon 1342. When the user has completed editing the proposed bounding polygon, the user can indicate completion by interacting with bounding polygon approval icon 1360.

[0107] 14A is a diagram illustrating an example of an input polygon 1400 that can be associated with a facet. The input polygon 1400 has a convex edge 1410 with adjacent edges 1420 and 1422 that would intersect if extended outside the input polygon 1400. The input polygon 1400 can be simplified by removing the convex edges and extending the adjacent edges to reduce the number of edges and vertices.

[0108] Figure 14B illustrates an example of a simplified polygon 1450 determined based on the input polygon 1400 of Figure 14A. For example, process 800 of Figure 8 may be performed to simplify the input polygon 1400 to obtain simplified polygon 1450. Convex edge 1410 is identified and removed, and adjacent edges 1420 and 1422 are extended to a point 1460 where they intersect outside the input polygon 1400. If the resulting increase in perimeter and area of ​​simplified polygon 1450 relative to the input polygon is sufficiently small (e.g., less than a threshold), simplified polygon 1450 can be used in place of input polygon 1400.

[0109] Disclosed herein is an implementation of structure scanning using unmanned aerial vehicles.

[0110] In a first aspect, the subject matter described herein can be embodied in a system including an unmanned aerial vehicle including a propulsion mechanism, one or more image sensors, and a processing device, wherein the processing device is configured to: access a three-dimensional map of the structure encoding a set of points in three-dimensional space on the surface of the structure; based on the three-dimensional map, generate one or more facets, which are polygons on the surface in three-dimensional space that each fit a subset of the points in the three-dimensional map; based on the one or more facets, generate a scan plan including a series of poses of the unmanned aerial vehicle that enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors; control the propulsion mechanism to fly the unmanned aerial vehicle to assume a pose corresponding to one of the series of poses in the scan plan; and capture one or more images of the structure from that pose using the one or more image sensors.

[0111] In a second aspect, the subject matter described in this specification can be implemented in a method that includes accessing a three-dimensional map of a structure that encodes a set of points in three-dimensional space on the surface of the structure; generating, based on the three-dimensional map, one or more facets that are polygons on the surface in three-dimensional space that each fit a subset of the points in the three-dimensional map; generating, based on the one or more facets, a scan plan that includes a series of poses of the unmanned aerial vehicle that enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors of the unmanned aerial vehicle; controlling a propulsion mechanism to fly the unmanned aerial vehicle to assume an pose corresponding to one of the series of poses in the scan plan; and capturing one or more images of the structure from that pose using the one or more image sensors.

[0112] In a third aspect, the subject matter described herein can be embodied in a non-transitory computer-readable storage medium including instructions that, when executed by a processor, prompt the performance of an operation including: accessing a three-dimensional map of a structure encoding a set of points in three-dimensional space on the surface of the structure; generating, based on the three-dimensional map, one or more facets that are polygons on the surface in three-dimensional space, each of which fits a subset of points in the three-dimensional map; generating, based on the one or more facets, a scan plan including a series of poses of the unmanned aerial vehicle that enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors of the unmanned aerial vehicle; controlling a propulsion mechanism to fly the unmanned aerial vehicle to assume an pose corresponding to one of the series of poses in the scan plan; and capturing one or more images of the structure from that pose using the one or more image sensors.

[0113] In a fourth aspect, the subject matter described herein may be embodied in an unmanned air vehicle that includes a propulsion mechanism, one or more sensors, and a processing device, the processing device: capturing an overview image of a building roof from a first attitude of the unmanned air vehicle positioned above the roof using one or more image sensors; presenting to a user a graphical representation of a proposed bounding polygon overlaid on the overview image, the proposed bounding polygon including vertices corresponding to respective vertex icons of the graphical representation that allow a user to move the vertices within a plane; and indicating user editing of one or more of the vertices of the proposed bounding polygon. accessing data encoding the proposed bounding polygon; determining a bounding polygon based on the proposed bounding polygon and the data encoding the user edits; determining a flight path based on the bounding polygon, the flight path including a series of poses of the unmanned air vehicle having respective fields of view at fixed heights that collectively cover the bounding polygon; controlling a propulsion mechanism to fly the unmanned air vehicle to assume a series of scan poses having horizontal positions consistent with each pose of the flight path and vertical positions determined to maintain a constant distance above the roof; and configured to scan the roof from the series of scan poses to generate a three-dimensional map of the roof.

[0114] In a fifth aspect, the subject matter described herein includes a method for providing a user with a method for creating a building rooftop using one or more image sensors of the unmanned air vehicle to capture an overview image of a building rooftop from a first pose of the unmanned air vehicle positioned above the rooftop; presenting to a user a graphical representation of a proposed bounding polygon overlaid on the overview image, the proposed bounding polygon including vertices corresponding to respective vertex icons of the graphical representation that enable the user to move the vertices within a plane; accessing data encoding user edits of one or more of the vertices of the proposed bounding polygon; and The present invention can be embodied in a method including: determining a bounding polygon based on data encoding user edits; determining a flight path based on the bounding polygon, the flight path including a series of attitudes of the unmanned aerial vehicle having respective fields of view at fixed heights that collectively cover the bounding polygon; controlling a propulsion mechanism to fly the unmanned aerial vehicle in a series of attitudes having horizontal positions consistent with each attitude of the flight path and vertical positions determined to maintain a constant distance above the roof; and scanning the roof from the series of scan attitudes to generate a three-dimensional map of the roof.

[0115] In a sixth aspect, the subject matter described herein may be embodied in a non-transitory computer-readable storage medium including instructions that, when executed by a processor, perform the following actions: capture, using one or more image sensors of the unmanned air vehicle, an overview image of a building roof from a first attitude of the unmanned air vehicle positioned above the roof; present to a user a graphical representation of a proposed bounding polygon overlaid on the overview image, the proposed bounding polygon including vertices corresponding to respective vertex icons of the graphical representation that enable a user to move the vertices within a plane; and encode user edits of one or more of the vertices of the proposed bounding polygon. The method facilitates performance of operations including: accessing data encoding the proposed bounding polygon; determining a bounding polygon based on the proposed bounding polygon and data encoding user edits; determining a flight path based on the bounding polygon, the flight path including a series of poses of the unmanned aerial vehicle having respective fields of view at fixed heights that collectively cover the bounding polygon; controlling a propulsion mechanism to fly the unmanned aerial vehicle in a series of poses having horizontal positions consistent with each pose of the flight path and vertical positions determined to maintain a constant distance above the roof; and scanning the roof from the series of scan poses to generate a three-dimensional map of the roof.

[0116] In some implementations, the unmanned air vehicle includes a propulsion mechanism, one or more image sensors, and a processing unit, where the processing unit is configured to: access a three-dimensional map of the structure encoding a set of points in three-dimensional space on a surface of the structure; generate one or more facets based on the three-dimensional map, where a given facet of the one or more facets is a polygon on the surface in three-dimensional space that fits a subset of points in the three-dimensional map; generate a scan plan based on the one or more facets, the scan plan including a series of poses for the unmanned air vehicle to assume to capture images of the structure using the one or more image sensors; control the propulsion mechanism to fly the unmanned air vehicle to assume an pose corresponding to one of the series of poses in the scan plan; and capture one or more images of the structure from the pose using the one or more image sensors. In some implementations of the unmanned air vehicle, the processing unit is configured to: capture an overview image of the structure using the one or more image sensors and generate a facet proposal based on the three-dimensional map; and determine a two-dimensional polygon as a convex hull of a subset of points in the three-dimensional map that corresponds to the facet proposal when the subset of points is projected onto the image plane of the overview image. In some such implementations, the processing device is configured to: present a two-dimensional polygon overlaid on the overview image; determine an edited two-dimensional polygon in the image plane of the overview image based on data indicating user edits of the two-dimensional polygon; and determine one of the one or more facets based on the edited two-dimensional polygon. In some such implementations, the processing device is configured to: before presenting the two-dimensional polygon overlaid on the overview image, remove convex edges from the two-dimensional polygon; and simplify the two-dimensional polygon by extending edges of the two-dimensional polygon adjacent to the convex edges to points where the extended edges intersect with each other. In some such implementations, the processing device is configured to: verify that removing the convex edges increases the area of ​​the two-dimensional polygon by an amount less than a threshold. In some such implementations, the processing device is configured to verify that removing the convex edges increases the perimeter of the two-dimensional polygon by an amount less than a threshold.In some implementations of the unmanned air vehicle, the scan plan series of poses are poses for orthogonal imaging of each of the one or more facets. In some implementations of the unmanned air vehicle, the one or more image sensors are configured to support stereoscopic imaging used to provide distance data, and the processing unit is configured to: control a propulsion mechanism to fly the unmanned air vehicle to a vicinity of the structure; and scan the structure using the one or more image sensors to generate a three-dimensional map. In some implementations of the unmanned air vehicle, the processing unit is configured to: capture an overview image of the structure using the one or more image sensors; present to a user a graphical representation of the scan plan overlaid on the overview image; and receive an indication from the user that the scan plan is approved. In some implementations of the unmanned air vehicle, the processing unit is configured to: detect obstacles based on images captured using the one or more image sensors while flying between poses in the scan plan series of poses; and dynamically adjust poses in the scan plan series to avoid the obstacles. In some implementations of the unmanned air vehicle, the processing unit is configured to: detect deviations of points on the surface of the structure from one of the one or more facets based on images captured using the one or more image sensors while flying between attitudes in the series of attitudes in the scan plan; and dynamically adjust the attitude in the series of attitudes in the scan plan to accommodate the deviations and maintain a constant distance for image capture. In some implementations of the unmanned air vehicle, the processing unit is configured to: generate a coverage map of the one or more facets indicating which facets in the one or more facets were successfully imaged during execution of the scan plan; and present the coverage map. In some implementations of the unmanned air vehicle, the processing unit is configured to: determine area estimates for each of the one or more facets; and present a data structure including the one or more facets, the area estimates for each of the one or more facets, and images of the structure captured during execution of the scan plan. In some implementations of the unmanned air vehicle, the structure is a building roof, a bridge, or a building under construction.

[0117] In some implementations, the method includes: accessing a three-dimensional map encoding a set of points in three-dimensional space on a surface of the structure; generating one or more facets based on the three-dimensional map, where a given facet of the one or more facets is a polygon on a surface in three-dimensional space that fits a subset of the points in the three-dimensional map; and generating a scanplan based on the one or more facets, where the scanplan includes a series of poses for the unmanned air vehicle to assume to enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors of the unmanned air vehicle. In some implementations of the method, the method includes: controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle in a pose corresponding to one of the series of poses in the scanplan; and capturing one or more images of the structure from the poses using the one or more image sensors. In some implementations of the method, generating the one or more facets includes: capturing an overview image of the structure using one or more image sensors; generating facet proposals based on a three-dimensional map; and determining a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the subset of points corresponding to the facet proposal when projected onto an image plane of the overview image. In some implementations of the method, the method includes: presenting a two-dimensional polygon overlaid on the overview image; determining an edited two-dimensional polygon in the image plane of the overview image based on data indicating user edits of the two-dimensional polygon; and determining one of the one or more facets based on the edited two-dimensional polygon. In some such implementations, the method includes: before presenting the two-dimensional polygon overlaid on the overview image, removing convex edges from the two-dimensional polygon and simplifying edges of the two-dimensional polygon adjacent to the convex edges by extending the extended edges to points where the extended edges intersect each other. In some such implementations, the method includes: verifying that removing the convex edge increases the area of ​​the two-dimensional polygon by an amount less than a threshold.In some such implementations, the method includes: verifying that removal of convex edges increases the perimeter of the two-dimensional polygon by an amount less than a threshold. In some implementations of this method, the series of scanplan poses are poses for orthogonal imaging of each of the one or more facets. In some such implementations, the one or more image sensors are configured to support stereoscopic imaging used to provide distance data, and the method includes: controlling a propulsion mechanism to fly the unmanned air vehicle to a vicinity of the structure; and scanning the structure using the one or more image sensors to generate a three-dimensional map. In some implementations of this method, the method further includes: capturing an overview image of the structure using the one or more image sensors; presenting a graphical representation of the scanplan overlaid on the overview image; and receiving an indication of approval of the scanplan from a user. In some implementations of this method, the method further includes: detecting an obstacle while flying between poses in the series of scanplan poses, the detection being based on images captured using the one or more image sensors; and dynamically adjusting poses in the series of scanplan poses to avoid the obstacle. In some implementations of the method, the method further includes: detecting deviations of points on a surface of the structure from one of the one or more facets while flying between poses of the series of poses of the scan plan, the detection being based on images captured using the one or more image sensors; and dynamically adjusting poses of the series of poses of the scan plan to accommodate the deviations and maintain a constant distance for image capture. In some implementations of the method, the method further includes: generating a coverage map of the one or more facets indicating which facets of the plurality of facets were successfully imaged during execution of the scan plan; and presenting the coverage map.In some implementations of the method, the method further includes: determining an area estimate for each of the one or more facets; and presenting a data structure including the one or more facets, the area estimate for each of the one or more facets, and an image of the structure captured during execution of the scan plan. In some implementations of the method, the method further includes: controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to a first location in proximity to a dock including a landing surface configured to hold the unmanned air vehicle and a fiducial mark on the landing surface; accessing one or more images captured using an image sensor of the unmanned air vehicle; detecting the fiducial mark in at least one of the one or more images; determining an attitude of the fiducial mark based on the one or more images; and controlling the propulsion mechanism to land the unmanned air vehicle on the landing surface based on the attitude of the fiducial mark. In some implementations of the method, the method further includes: automatically charging a battery of the unmanned air vehicle using a charger included in the dock while the unmanned air vehicle is on the landing surface.

[0118] In some implementations, a non-transitory computer-readable storage medium includes instructions that, when executed by a processor, cause the performance of an operation including: accessing a three-dimensional map encoding a set of points in three-dimensional space on a surface of a structure; generating one or more facets based on the three-dimensional map, where a given one of the facets is a polygon on the surface in three-dimensional space that fits a subset of the points in the three-dimensional map; and generating a scanplan based on the one or more facets, the scanplan including a series of poses for the unmanned air vehicle to assume to enable capture of images of the structure at a fixed distance from each of the one or more facets using one or more image sensors of the unmanned air vehicle. In some implementations of the non-transitory computer-readable storage medium, the instructions, when executed by a processor, cause the performance of an operation including: controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle in an pose corresponding to one of the series of poses in the scanplan; and capturing one or more images of the structure from the poses using the one or more image sensors. In some implementations of the non-transitory computer-readable storage medium, the instructions, when executed by a processor, facilitate performance of operations including: capturing an overview image of the structure using one or more image sensors; generating facet proposals based on the three-dimensional map; and determining a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the subset of points corresponding to the facet proposal when projected onto an image plane of the overview image. In some such implementations, the instructions, when executed by a processor, facilitate performance of operations including: presenting the two-dimensional polygon overlaid on the overview image; determining an edited two-dimensional polygon in the image plane of the overview image based on data indicating user edits of the two-dimensional polygon; and determining one of the one or more facets based on the edited two-dimensional polygon.In some such implementations, the instructions, when executed by a processor, facilitate the performance of an operation including: prior to presenting the 2D polygon overlaid on the overview image, simplifying by removing convex edges from the 2D polygon and extending edges of the 2D polygon adjacent to the convex edges to a point where the extended edges intersect one another. In some such implementations, the instructions, when executed by a processor, facilitate the performance of an operation including: verifying that removing the convex edges increases the area of ​​the 2D polygon by an amount less than a threshold. In some such implementations, the instructions, when executed by a processor, facilitate the execution of an operation including: verifying that removing the convex edges increases the perimeter of the 2D polygon by an amount less than a threshold. In some implementations of the non-transitory computer-readable storage medium, the set of poses of the scanplan are poses for orthogonal imaging of each of the one or more facets. In some such implementations, the one or more image sensors are configured to support stereoscopic imaging used to provide distance data, and the non-transitory computer-readable storage medium includes instructions that, when executed by a processor, prompt the performance of operations including: controlling a propulsion mechanism to fly the unmanned air vehicle to a vicinity of the structure; and scanning the structure using the one or more image sensors to generate a three-dimensional map. In some implementations of the non-transitory computer-readable storage medium, the instructions, when executed by a processor, prompt the performance of operations including: capturing an overview image of the structure using the one or more image sensors; presenting a graphical representation of a scan plan overlaid on the overview image; and receiving an indication from a user that the scan plan is approved. In some such implementations, the instructions, when executed by a processor, prompt the performance of operations including: generating a coverage map of one or more facets indicating which of the one or more facets were successfully imaged during execution of the scan plan; and presenting the coverage map.In some such implementations, the instructions, when executed by a processor, facilitate performance of operations including: determining an area estimate for each of the one or more facets; and presenting a data structure including the one or more facets, the area estimate for each of the one or more facets, and an image of the structure captured during execution of the scan plan.

[0119] In some implementations, the method includes: capturing an overview image of a building roof from a first attitude of the unmanned air vehicle using one or more image sensors of the unmanned air vehicle, where the unmanned air vehicle is positioned above the roof at the first attitude; presenting to a user a graphical representation of a proposed bounding polygon superimposed on the overview image, where the proposed bounding polygon includes vertices corresponding to respective vertex icons of the graphical representation that enable a user to move the vertices within a plane; determining a bounding polygon based on the proposed bounding polygon and data encoding user edits of one or more of the vertices of the proposed bounding polygon; and determining a flight path based on the bounding polygon, where the flight path includes a series of attitudes of the unmanned air vehicle having respective fields of view at fixed heights that collectively cover the bounding polygon. In some implementations of the method, the method further includes presenting a graphical representation of the unmanned air vehicle superimposed on the overview image, where the graphical representation of the unmanned air vehicle corresponds to a current horizontal position of the unmanned air vehicle. In some such implementations, the graphical representation of the unmanned air vehicle includes a three-dimensional rendering of the unmanned air vehicle. In some implementations of the method, the flight path is determined based on one or more scan parameters presented for selection by a user, the one or more scan parameters including one or more parameters from a set of parameters including a grid size, a nominal height above the roof surface, and a maximum flight speed. In some implementations of the method, the flight path is determined as a lawnmower pattern. In some implementations of the method, the method further includes: detecting an obstacle while the unmanned air vehicle flies between attitudes of the series of scan poses based on images captured using the one or more image sensors; and dynamically adjusting the attitude of the flight path to avoid the obstacle. In some implementations of the method, the method further includes: presenting an indication of progress along the flight path superimposed on the overview image. In some implementations of the method, the method further includes: scanning the roof to generate a three-dimensional map of the roof.In some such implementations, the method further includes: storing a scan state indicating a next attitude in the series of attitudes of the flight path; controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to a landing site and land the unmanned air vehicle; controlling the propulsion mechanism to cause the unmanned air vehicle to take off from the landing site after landing; accessing the scan state; and controlling the propulsion mechanism, based on the scan state, to fly the unmanned air vehicle to assume an attitude in the series of scan attitudes corresponding to the next attitude and continue scanning the roof to generate a three-dimensional map. In some such implementations, the method further includes: controlling the propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to assume a series of scan attitudes having horizontal positions consistent with each attitude of the flight path and a vertical position determined to maintain a constant distance above the roof, and scanning the roof to generate a three-dimensional map of the roof includes: scanning the roof from the series of scan attitudes to generate a three-dimensional map of the roof. In some such implementations, the roof is scanned to generate the three-dimensional map from a distance greater than the constant distance later used for facet imaging. In some such implementations, the one or more image sensors are configured to support stereoscopic imaging, and scanning the roof to generate the three-dimensional map of the roof includes: scanning the roof using stereoscopic imaging to generate the three-dimensional map of the roof. In some such implementations, scanning the roof to generate the three-dimensional map of the roof includes: scanning the roof using a radar sensor on the unmanned air vehicle to generate the three-dimensional map of the roof. In some such implementations, scanning the roof to generate the three-dimensional map of the roof includes: scanning the roof using a lidar sensor on the unmanned air vehicle to generate the three-dimensional map of the roof.

[0120] In some implementations, a non-transitory computer-readable storage medium includes instructions that, when executed by a processor, facilitate the performance of acts to perform the methods described herein.

[0121] In some implementations, an unmanned air vehicle including a processing device is configured to perform a method as disclosed herein, the unmanned air vehicle further including: a propulsion mechanism; and one or more image sensors. In some implementations of the unmanned air vehicle, the processing device is configured to: control the propulsion mechanism to fly the unmanned air vehicle through a series of scan attitudes having horizontal positions aligned with respective attitudes of the flight path and vertical positions determined to maintain a constant distance above the roof; and scan the roof from the series of scan attitudes to generate a three-dimensional map of the roof. In some such implementations, the unmanned air vehicle further includes a radar sensor, and the processing device is configured to: scan the roof using the radar sensor to generate a three-dimensional map of the roof. In some such implementations, the unmanned air vehicle further includes a lidar sensor, and the processing device is configured to: scan the roof using the lidar sensor to generate a three-dimensional map of the roof. In some implementations of the unmanned aerial vehicle, the processing device is configured to: store a scan state indicating a next attitude in a series of attitudes along a flight path; control the propulsion mechanism to fly the unmanned aerial vehicle to a landing site and land the unmanned aerial vehicle; control the propulsion mechanism to take off the unmanned aerial vehicle from the landing site after landing, and access the scan state; and control the propulsion mechanism to fly the unmanned aerial vehicle to assume an attitude in the series of scan attitudes corresponding to the next attitude and continue scanning the roof to generate a three-dimensional map of the roof.

[0122] Although the present disclosure has been described in connection with particular embodiments, it should be understood that the present disclosure should not be limited to the embodiments of the present disclosure, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, the scope of which is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures.

Claims

1. accessing a three-dimensional map of the structure, the three-dimensional map encoding a set of points in three-dimensional space on a surface of the structure; generating one or more facets based on the three-dimensional map, a given facet of the one or more facets being a polygon on a surface in three-dimensional space that fits a subset of points in the three-dimensional map; generating a scanplan based on the one or more facets, the scanplan including a series of poses for the unmanned air vehicle to assume to enable capture of images of the structure at a consistent distance from each of the one or more facets using one or more image sensors of the unmanned air vehicle; A method comprising:

2. controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to an attitude corresponding to one of a series of attitudes in the scan plan; capturing one or more images of the structure from that pose using one or more image sensors; The method of claim 1 , comprising:

3. Creating one or more facets capturing a general image of the structure using one or more image sensors; generating facet suggestions based on the three-dimensional map; determining a two-dimensional polygon as a convex hull of a subset of points of the three-dimensional map, the subset of points corresponding to the facet proposal when projected onto an image plane of the overview image; The method of claim 1 , comprising:

4. presenting a two-dimensional polygon superimposed on the overview image; determining an edited two-dimensional polygon in the image plane of the overview image based on data indicative of the user edit of the submitted two-dimensional polygon; determining one of the one or more facets based on the edited two-dimensional polygon; The method of claim 3, comprising:

5. 5. The method of claim 4, including simplifying the two-dimensional polygon before presenting the two-dimensional polygon superimposed on the overview image by removing convex edges from the two-dimensional polygon and extending edges of the two-dimensional polygon adjacent to the convex edges to points where the extended edges intersect each other.

6. The method of claim 5 , including verifying that removing convex edges increases the area of ​​the two-dimensional polygon by an amount less than a threshold.

7. The method of claim 5 , including verifying that removing convex edges increases the perimeter of the two-dimensional polygon by an amount less than a threshold.

8. determining an area estimate for each of the one or more facets; presenting a data structure including one or more facets, area estimates for each of the one or more facets, and an image of the structure captured during execution of the scan plan; 8. The method of claim 2, comprising:

9. The one or more image sensors are configured to support stereoscopic imaging used to provide distance data, and the method includes: Controlling the propulsion mechanism to fly the unmanned aerial vehicle to the vicinity of the structure; scanning the structure using one or more image sensors to generate a three-dimensional map; 8. The method of claim 2, comprising:

10. generating a coverage map of the one or more facets indicating which facets of the one or more facets were successfully imaged during execution of the scan plan; Presenting a coverage map 8. The method of claim 2, comprising:

11. Detecting obstacles while flying between postures in a series of postures in a scan plan, the detection being based on images captured using one or more image sensors; Dynamically adjusting a pose among a series of poses in a scan plan to avoid obstacles; 8. The method of claim 2, comprising:

12. Detecting obstacles while flying between postures in a series of postures in a scan plan, the detection being based on images captured using one or more image sensors; Dynamically adjusting a pose among a series of poses in a scan plan to avoid obstacles; 8. The method of claim 2, comprising:

13. detecting deviations of points on a surface of the structure from one of the one or more facets while flying between poses of the series of poses of the scan plan, the detection being based on images captured using the one or more image sensors; Dynamically adjusting poses among a series of scan planning poses to accommodate deviations and maintain a consistent distance for image capture; 8. The method of claim 2, comprising:

14. capturing a general image of the structure using one or more image sensors; presenting a graphical representation of the scan plan superimposed on the overview image; receiving an indication from the user of approval of the scan plan; 8. The method of claim 1, comprising:

15. The method of claim 1 , wherein the series of poses of the scan plan are poses for orthogonal imaging of each of the one or more facets.

16. 8. The method of any one of claims 1 to 7, wherein the structure is a roof of a building, a bridge, or a building under construction.

17. controlling a propulsion mechanism of the unmanned air vehicle to fly the unmanned air vehicle to a first location proximate a dock including a landing surface configured to hold the unmanned air vehicle and a reference mark on the landing surface; accessing one or more images captured using an image sensor of the unmanned air vehicle; Detecting a fiducial mark in at least one of the one or more images; determining a pose of the fiducial mark based on the one or more images; Controlling the propulsion mechanism to land the unmanned aerial vehicle on the landing surface based on the attitude of the reference mark; 8. The method of claim 1, comprising:

18. 20. The method of claim 17, comprising automatically charging a battery of the unmanned air vehicle using a charger included in the dock while the unmanned air vehicle is on the landing surface.

19. a propulsion mechanism; one or more image sensors; a processing device configured to perform the method of any one of claims 1 to 18; Unmanned aerial vehicles, including

20. 19. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor, facilitate performance of the method of any one of claims 1 to 18.