Smart landing of aircraft

UAVs use sensory inputs and navigation systems to select safe landing locations, addressing unsafe flight conditions by identifying flat areas and avoiding obstacles, ensuring reliable and safe landings.

JP2026012676APending Publication Date: 2026-01-27SKYDIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025153531
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-11
Filing Date
2025-09-16
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Unmanned aerial vehicles (UAVs) face unsafe or undesirable flight conditions due to low battery charge or malfunctioning systems, necessitating autonomous landing that avoids contact with people and assets while being convenient for the user.

Method used

UAVs utilize sensory inputs to select a safe landing location by incorporating geometric and semantic knowledge of the surrounding environment, using image capture devices and navigation systems to identify the flattest area and avoid obstacles, with mechanisms to adjust image capture orientation and generate planned trajectories.

Benefits of technology

Enables safe, autonomous landing by selecting suitable landing sites that minimize damage and inconvenience, ensuring reliable operation under varying conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012676000001_ABST
    Figure 2026012676000001_ABST
Patent Text Reader

Abstract

Techniques for autonomous landing by an aircraft are presented.SOLUTION: The presented techniques include processing sensor data, such as images captured by an on-board camera, to generate a ground map including a plurality of cells. An appropriate footprint comprising a subset of the plurality of cells in the ground map that satisfies the one or more landing criteria is selected, and control commands are generated to autonomously land the aircraft in an area corresponding to the footprint. In some embodiments, the presented technology involves a geometric smart landing process for selecting a relatively flat area on the ground for landing. In some implementations, the techniques presented include a semantic smart landing process in which semantic information about detected objects is incorporated into the ground map.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit and / or priority of U.S. Provisional Application No. 62 / 628,871 (Attorney Docket No. 113391-8012.US00), entitled "Unmanned Aerial Vehicle Smart Landing," filed February 9, 2018, the contents of which are incorporated herein by reference in their entirety for all purposes. Accordingly, this application is entitled to a priority date of February 9, 2018.

[0002] Technical Field The present disclosure relates generally to autonomous transport vehicle technology. [Background technology]

[0003] Transport vehicles can be configured to autonomously navigate a physical environment. For example, an autonomous transport vehicle equipped with various onboard sensors can be configured to generate sensory inputs based on the surrounding physical environment and use those sensory inputs to estimate the position and / or orientation of the autonomous transport vehicle within the physical environment. An autonomous navigation system can then utilize those position and / or orientation estimates to guide the autonomous transport vehicle through the physical environment. Summary of the Invention

[0004] Overview There are a variety of circumstances that may make continued flight by an unmanned aerial vehicle (UAV) unsafe or undesirable. For example, the battery charge on board the UAV may be low, or certain systems (such as those for visual navigation) may not be functioning properly. Even in a nominal use case, the UAV may reach a point during normal flight where there is only enough charge left in the battery to complete a controlled descent and landing before power runs out. Whatever the circumstances, an autonomous landing by the UAV may be necessary.

[0005] Techniques are presented to enable autonomous smart landing by UAVs. In some embodiments, the presented techniques enable a UAV (or an associated autonomous navigation system) to utilize available information to select a safe location for landing. Such a safe location for landing 1) avoids contact with and / or injury to people, 2) avoids contact with or damage to all assets, including the UAV itself, and / or 3) is convenient or desirable for a user and their particular use case. In some embodiments, a UAV may be configured to perform a geometric smart landing, selecting the flattest area of ​​ground currently visible to the UAV for landing. Some embodiments may involve incorporating semantic knowledge of the surrounding physical environment to select a landing site (as close as possible) that meets the above requirements. In some embodiments, a UAV may be configured to collect information during a user-guided or controlled landing to train the landing selection process. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 illustrates an exemplary configuration of an autonomous vehicle in the form of an unmanned aerial vehicle (UAV) to which certain techniques described herein may be applied.

[0007] [Figure 2] FIG. 2 shows a block diagram of an exemplary navigation system that may be implemented by the UAV of FIG.

[0008] [Figure 3A] FIG. 3A shows a block diagram of an exemplary exercise planning system that may be part of the navigation system of FIG.

[0009] [Figure 3B] FIG. 3B shows a block diagram representing exemplary goals that may be incorporated into the motion planning system shown in FIG. 3A.

[0010] [Figure 4]FIG. 4 is a flowchart of an exemplary process for autonomous landing by a UAV.

[0011] [Figure 5A] FIG. 5A shows an example ground map generated by a UAV in flight.

[0012] [Figure 5B] FIG. 5B shows an example footprint included in the ground map of FIG. 5A.

[0013] [Figure 6] Figure 6 shows an elevation view of a UAV flying above the ground.

[0014] [Figure 7A] FIG. 7A is a flowchart of an exemplary process for autonomous landing by a UAV using geometric smart landing techniques.

[0015] [Figure 7B] FIG. 7B is a flowchart of an exemplary process for downweighting older points in a given cell when updating the statistics of the cell based on new points.

[0016] [Figure 8A] FIG. 8A shows a UAV flying over a physical environment populated with various objects.

[0017] [Figure 8B] FIG. 8B shows an example ground map with semantic information added based on the various objects shown in FIG. 8A.

[0018] [Figure 9] FIG. 9 is a flowchart of an exemplary process for autonomous landing by a UAV using semantic smart landing techniques.

[0019] [Figure 10]FIG. 10 is a flowchart of an exemplary process for landing a UAV based on an operating state of the UAV.

[0020] [Figure 11A] FIG. 11A is a flowchart of an exemplary process for initiating a controlled descent in response to determining that a smart landing is not possible. [Figure 11B] FIG. 11B is a flowchart of an exemplary process for initiating a controlled descent in response to determining that a smart landing is not possible.

[0021] [Figure 12] FIG. 12 is a diagram of an example location system in which at least some operations described in this disclosure may be implemented.

[0022] [Figure 13] FIG. 13 illustrates the concept of visual odometry based on captured images.

[0023] [Figure 14] FIG. 14 is an exemplary illustration of a three-dimensional (3D) occupancy map of a physical environment.

[0024] [Figure 15] FIG. 15 is an exemplary image captured by a UAV flying through a physical environment with an associated visualization of data regarding the object being tracked based on processing of the captured image.

[0025] [Figure 16] FIG. 16 illustrates an exemplary process for estimating an object's trajectory based on multiple images captured by a UAV.

[0026] [Figure 17] FIG. 17 is a diagrammatic representation of an exemplary spatiotemporal factor graph.

[0027] [Figure 18] FIG. 18 illustrates an exemplary process for generating an intelligent initial estimate of where a tracked object will appear in subsequently captured images.

[0028] [Figure 19] FIG. 19 shows a visualization representing a dense pixel-by-pixel segmentation of the captured image.

[0029] [Figure 20] Figure 20 shows a visualization representing the instance segmentation of the captured image.

[0030] [Figure 21] FIG. 21 is a block diagram of an example UAV system including various functional system components in which at least some operations described in this disclosure may be implemented.

[0031] [Figure 22] FIG. 22 is a block diagram of an example processing system in which at least some of the operations described in this disclosure may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0032] Unmanned aerial vehicle implementation example FIG. 1 illustrates an exemplary configuration of a UAV 100 to which certain techniques described herein may be applied. As illustrated in FIG. 1, the UAV 100 may be configured as a rotor-based aircraft (e.g., a "quadcopter"), although other presented techniques may similarly be applied in other types of UAVs, such as fixed-wing aircraft. The exemplary UAV 100 includes a control actuator 110 for maintaining controlled flight. The control actuator 110 may include or be associated with a propulsion system (e.g., rotors) and / or one or more control surfaces (e.g., flaps, ailerons, rudder, etc.), depending on the configuration of the UAV. The exemplary UAV 100 illustrated in FIG. 1 includes a control actuator 110 in the form of an electronic rotor, which includes the propulsion system of the UAV 100. The UAV 100 also includes various sensors for automatic navigation and flight control 112 and one or more image capture devices 114 and 115 for capturing images of the surrounding physical environment during flight. "Image" in this context includes both still images and captured video. Although not shown in FIG. 1, UAV 100 may also include other sensors (e.g., a sensor for capturing audio) and systems for communicating with other devices, such as mobile device 104, via wireless communication channel 116.

[0033] 1 , image capture devices 114 and / or 115 are shown capturing an object 102 in a physical environment that happens to be human. In some cases, the image capture devices may be configured to capture images for display to a user (e.g., as an aerial video platform) and / or for use in autonomous navigation, as described above. In other words, UAV 100 may navigate a physical environment autonomously (i.e., without direct human control), for example, by processing images captured by any one or more image capture devices. While flying autonomously, UAV 100 may also capture images using any one or more image capture devices, which may be displayed in real time or recorded for later display on another device (e.g., mobile device 104).

[0034] FIG. 1 illustrates an exemplary configuration of a UAV 100 with multiple image capture devices, each configured for a different purpose. In the exemplary configuration illustrated in FIG. 1, the UAV 100 includes multiple image capture devices 114 arranged around the periphery of the UAV 100. The image capture devices 114 may be configured to capture images for use by a visual navigation system in guiding autonomous flight by the UAV 100 and / or for use by a tracking system for tracking other objects in the physical environment (e.g., as described with respect to FIG. 2). Specifically, the exemplary configuration of the UAV 100 illustrated in FIG. 1 includes an array of multiple stereo image capture devices 114 arranged around the periphery of the UAV 100 to provide maximum stereo image capture in 360 degrees around the UAV 100.

[0035] In addition to the array of image capture devices 114, the UAV 100 shown in FIG. 1 also includes another image capture device 115 configured to capture images. Images captured by this image capture device 115 are displayed for navigation purposes, but are not necessarily used for navigation purposes. In some embodiments, the image capture device 115 may be similar to the image capture device 114, except for the manner in which the captured images are utilized. However, in other embodiments, the image capture devices 115 and 114 may be configured differently to suit their respective roles.

[0036] In many cases, given certain hardware and software constraints, it is generally preferable to capture images intended to be viewed at the highest possible resolution. However, for use in visual navigation or object tracking, lower resolution images may be desirable in certain situations to reduce processing load and provide more robust motion planning performance. Thus, in some embodiments, image capture device 115 may be configured to capture relatively high resolution (e.g., 3840×2160) color images, while image capture device 114 may be configured to capture relatively low resolution (e.g., 320×240) grayscale images.

[0037] UAV 100 may be configured to track one or more objects, such as human subject 102, through a physical environment based on images received via image capture devices 114 and / or 115. Furthermore, UAV 100 may be configured to track image capture of such objects, e.g., for filming purposes. In some embodiments, image capture device 115 is coupled to the body of UAV 100 via an adjustable mechanism that provides one or more degrees of freedom of motion relative to the body of UAV 100. UAV 100 may be configured to automatically adjust the orientation of image capture device 115 to track image capture of an object (e.g., human subject 102) as both UAV 100 and the object move through the physical environment. In some embodiments, this adjustable mechanism may include a mechanical gimbal mechanism that rotates the attached image capture device about one or more axes. In some embodiments, the gimbal mechanism may be configured as a hybrid mechanical-digital gimbal system that couples image capture device 115 to the body of UAV 100. In a hybrid mechanical-digital gimbal system, the orientation of the image capture device 115 about one or more axes may be adjusted by mechanical means, while the orientation about other axes may be adjusted by digital means. For example, a mechanical gimbal mechanism may handle adjustments in pitch of the image capture device 115, while roll and yaw adjustments are achieved digitally by transforming (e.g., rotating, panning, etc.) the captured image to effectively provide at least three degrees of freedom of movement of the image capture device 115 relative to the UAV 100.

[0038] The mobile device 104 may include any type of mobile device, such as a laptop computer, a table computer (e.g., Apple iPad®), a mobile phone, a smartphone (e.g., Apple iPhone®), a handheld gaming device (e.g., Nintendo Switch®), a single-function remote control device, or any other type of device capable of receiving user input, sending signals to the UAV 100 for transmission (e.g., based on the user input), and / or presenting information to a user (e.g., based on sensor data collected by the UAV 100). In some embodiments, the mobile device 104 may include a touchscreen display and an associated graphical user interface (GUI) for receiving user input and presenting information. In some embodiments, the mobile device 104 may include various sensors (e.g., image capture devices, accelerometers, gyroscopes, GPS receivers, etc.) capable of collecting sensor data. In some embodiments, such sensor data may be communicated to the UAV 100, for use by, for example, an onboard navigation system of the UAV 100.

[0039] Figure 2 is a block diagram illustrating an example navigation system 120 that may be implemented as part of the example UAV 100 described with respect to Figure 1. Navigation system 120 may include any combination of hardware and / or software. For example, in some embodiments, navigation system 120 and associated subsystems may be implemented as instructions stored in memory and executable by one or more processors.

[0040] As shown in FIG. 2 , the exemplary navigation system 120 includes a motion planner 130 (also referred to herein as a “motion planning system”) for autonomously maneuvering the UAV 100 through a physical environment, a tracking system 140 for tracking one or more objects within the physical environment, and a landing system 150 for performing the smart landing techniques described herein. Note that the system configuration shown in FIG. 2 is an example provided for illustrative purposes and should not be construed as limiting. For example, in some embodiments, the tracking system 140 and / or the landing system 150 may be separate from the navigation system 120. Furthermore, the subsystems that make up the navigation system 120 may not be logically separated as shown in FIG. 2 , but instead may effectively operate as a single, integrated navigation system.

[0041] In some embodiments, motion planner 130 operates separately or in conjunction with tracking system 140 and is configured to generate a planned trajectory in three-dimensional (3D) space of the physical environment based on, for example, images received from image capture devices 114 and / or 115, data from other sensors 112 (e.g., IMU, GPS, proximity sensors, etc.), and / or one or more control inputs 170. Control inputs 170 may come from an external source, such as a mobile device operated by a user, or may come from other systems onboard the UAV.

[0042] In some embodiments, navigation system 120 may generate control commands configured to steer UAV 100 along a planned trajectory generated by motion planner 130. For example, the control commands may be configured to control one or more control actuators 110 (e.g., rotors and / or control surfaces) to steer UAV 100 along the planned 3D trajectory. Alternatively, the planned trajectory generated by motion planner 130 may be output to a separate flight controller 160 configured to process the trajectory information and generate appropriate control commands configured to control one or more control actuators 110.

[0043] Tracking system 140 may operate separately or in conjunction with motion planner 130 and may be configured to track one or more objects within the physical environment based, for example, on images received from image capture devices 114 and / or 115, data from other sensors 112 (e.g., IMU, GPS, proximity sensors, etc.), one or more control inputs 170 from external sources (e.g., remote user, navigation application, etc.), and / or specific one or more designated tracking targets. Tracking targets may include, for example, a user designation to track specific detected objects within the physical environment, or a persistent target to track specific classes of objects (e.g., people).

[0044] As mentioned above, tracking system 140 may communicate to motion planner 130, for example, to steer UAV 100 based on measured, estimated, and / or predicted positions, orientations, and / or trajectories of objects in the physical environment. For example, tracking system 140 may communicate navigational objectives to motion planner 130, such as maintaining a particular separation distance to a tracked object during movement.

[0045] In some embodiments, the tracking system 140, operating separately or in conjunction with the motion planner 130, is further configured to generate control commands configured to cause a mechanism to adjust the orientation of any image capture device 114 / 115 relative to the body of the UAV 100 based on the tracking of one or more objects. Such a mechanism may include a mechanical gimbal or a hybrid digital-mechanical gimbal, as described above. For example, while tracking an object moving relative to the UAV 100, the tracking system 140 may generate control commands configured to adjust the orientation of the image capture device 115 to keep the tracked object centered in the field of view (FOV) of the image capture device 115 while the UAV 100 is moving. Similarly, the tracking system 140 may generate commands or output data to a digital image processor (e.g., part of a hybrid digital-mechanical gimbal) to transform images captured by the image capture device 115 to keep the tracked object centered in the FOV of the image capture device 115 while the UAV 100 is moving.

[0046] Landing system 150 may operate separately or in conjunction with motion planner 130 to determine when to initiate a landing procedure (e.g., in response to a user command or in response to a detected event such as a low battery), when to identify a landing location (e.g., based on images received from image capture devices 114 and / or 115 and / or data from other sensors 112 (e.g., IMU, GPS, proximity sensors, etc.)), and when to generate control commands configured to cause the UAV to land at the selected location. Note that in some embodiments, landing system 150 is configured to generate output in the form of landing targets and input the landing targets to motion planner 130, which utilizes the landing targets in conjunction with other goals (e.g., avoiding collisions with other objects) to autonomously land the UAV.

[0047] In some embodiments, the navigation system 120 (e.g., specifically the motion planning component 130) is configured to incorporate multiple objectives to generate an output, such as a planned trajectory, that can be used to guide the autonomous behavior of the UAV 100 at any given time. For example, certain incorporated objectives, such as obstacle avoidance and vehicle dynamics constraints, may be combined with other input objectives (e.g., landing objectives) as part of the trajectory generation process. In some embodiments, the trajectory generation process may include gradient-based optimization, gradient-free optimization, sampling, end-to-end learning, or any combination thereof. The output of this trajectory generation process may be a planned trajectory over a time horizon (e.g., 10 seconds), which is configured to be interpreted and utilized by the flight controller 160 to generate control commands to steer the UAV 100 according to the planned trajectory. The motion planner 130 may continuously perform the trajectory generation process as new sensory inputs (e.g., imagery or other sensor data) and target inputs are received. Thus, the planned trajectory may be continuously updated over a time horizon, thereby enabling the UAV 100 to dynamically and autonomously respond to changing conditions.

[0048] 3A shows a block diagram representing an exemplary system for goal-based motion planning. As shown in FIG. 3A, a motion planner 130 (e.g., as discussed with respect to FIG. 2) may generate and continuously update a planned trajectory 320 based on a trajectory generation process that involves one or more goals (e.g., goals as described above) and / or more sensory inputs 306. The sensory inputs 306 may include images received from one or more image capture devices 114 / 115, processed results of such images (e.g., disparity images or depth values), and / or sensor data from one or more other sensors 112 onboard the UAV 100 or associated with other computing devices (e.g., mobile device 104) that communicate with the UAV 100. The one or more goals 302 utilized in the motion planning process may include built-in goals that govern high-level behaviors (e.g., avoiding collisions with other objects, smart landing techniques described herein, etc.) and goals based on control input(s) 308 (e.g., from a user). Each of the goals 302 may be encoded as one or more equations for incorporation into one or more motion planning equations utilized by the motion planner 130 in generating a planned trajectory to satisfy the one or more goals. The control input 308 may be in the form of control commands from a user or from other components of the navigation system 120, such as the tracking system 140 and / or the smart landing system 150. In some embodiments, such input is received in the form of a call to an application programming interface (API) associated with the navigation system 120. In some embodiments, the control input 308 may include predetermined goals generated by other components of the navigation system 120, such as the tracking system 140 or the landing system 150.

[0049] Each given goal in the set of one or more goals 302 utilized in the motion planning process may include one or more defined parameterizations exposed through the API. For example, FIG. 3B illustrates an example goal 332, which includes a target 334, a dead zone 336, a weighting factor 338, and other parameters 340.

[0050] Targets 344 define the goals of a particular objective that motion planner 130 attempts to meet when generating planned trajectory 320. For example, targets 344 for a given objective may be to maintain line of sight to one or more detected objects or to fly to a particular location within the physical environment.

[0051] A dead zone defines an area around a target 334 within which the motion planner 130 may not take corrective action. This dead zone 336 may be considered a tolerance level for meeting a given target 334. For example, the target in the goal for an exemplary image may be to maintain image capture of the tracked object so that it appears at a specific location (e.g., the center) in image space of the captured image. To avoid continuous adjustments based on slight deviations from this target, a dead zone is defined to allow for some tolerance. For example, dead zones may be defined in the y and x directions around the target location in image space. In other words, as long as the tracked object appears within the area enclosed by the target in the image and within the respective dead zone, the goal is considered met.

[0052] The weighting factor 336 (also referred to as the “aggressiveness” factor) defines the relative level of influence a particular goal 332 will have on the overall trajectory generation process performed by the motion planner 130. Recall that a particular goal 332 may be one of several goals 302, which may include competing targets. In an ideal situation, the motion planner 130 would generate a planner trajectory 320 that perfectly satisfies all of the associated goals at any given moment. For example, the motion planner 130 may generate a planned trajectory that steers the UAV 100 to specific GPS coordinates while tracking the tracked object, capturing an image of the tracked object, maintaining line of sight with the tracked object, and avoiding collisions with other objects. In practice, such ideal situations may be rare. Thus, the motion planner system 130 may need to prioritize one goal over another when satisfying both is impossible or impractical (for any number of reasons). The weighting factor of each of the goals 302 defines how they are considered by the motion planner 130.

[0053] In an exemplary embodiment, the weighting factor is a numeric value on a scale of 0.0 to 1.0. A value of 0.0 for a particular goal indicates that the motion planner 130 can completely ignore the goal (if necessary), while a value of 1.0 indicates that the motion planner 130 will make every effort to meet the goal while maintaining safe flight. A value of 0.0 may similarly be associated with an inactive goal and may be set to zero, for example, in response to being switched from an active state to an inactive state by the goal application 1210. A low weighting factor value (e.g., 0.0 to 0.4) may be set for a given goal based on subjective or aesthetic targets, such as maintaining visual salience in a captured image. Conversely, a higher weighting factor value (e.g., 0.5 to 1.0) may be set for a more important goal, such as avoiding a collision with another object.

[0054] In some embodiments, the weighting factor values ​​338 may remain static as the planned trajectory is continuously updated while the UAV 100 is flying. Alternatively, or in addition, the weighting factors for a given goal may change dynamically based on changing conditions while the UAV 100 is flying. For example, a goal to avoid regions in the captured image associated with uncertain depth value calculations (e.g., due to low light conditions) may have a variable weighting factor that increases or decreases based on other perceived threats to the safe operation of the UAV 100. In some embodiments, a goal may be associated with multiple weighting factor values ​​that change depending on how the goal is applied. For example, a collision avoidance goal may utilize different weighting factors depending on the class of detected object to be avoided. As an illustrative example, the system may be configured to place more emphasis on avoiding collisions with people or animals versus avoiding collisions with buildings or trees.

[0055] The UAV 100 shown in Figure 1 and the associated navigation system 120 shown in Figure 2 are examples provided for illustrative purposes. A UAV 100 according to the present teachings may include more or fewer components than those shown. Furthermore, the example UAV 100 shown in Figure 1 and the associated navigation system 120 shown in Figure 2 may include or be part of one or more components of the example UAV system 2100 described with reference to Figure 21 and / or the example computer processing system 2200 described with reference to Figure 22. For example, the navigation system 120 and associated motion planner 130, tracking system 140, and landing system 150 described above may include or be part of the UAV system 2100 and / or computer processing system 2200.

[0056] The presented techniques for smart landing are described in the context of an unmanned aerial vehicle, such as UAV 100 shown in FIG. 1 , for ease of illustration, although the presented approaches are not limited to this context. The presented techniques may similarly be applied to guide the landing of other types of aircraft, such as manned rotorcraft, such as helicopters, and manned or unmanned fixed-wing aircraft. For example, a manned aircraft may include an autonomous navigation component (e.g., navigation system 120) in addition to a component of manual control (direct or indirect). During the landing sequence, control of the aircraft may switch from the manual control component to the automatic control component, whereupon the presented touchdown detection techniques are implemented. The switch from manual control to automatic control occurs in response to pilot input and / or automatically in response to detected events, such as remote signals, environmental conditions, aircraft operating conditions, etc.

[0057] Smart Landing Process FIG. 4 shows a flowchart of an exemplary process 400 for autonomous landing by a UAV in accordance with the present technology. One or more steps of this exemplary process may be performed by any one or more of the components of the exemplary navigation system 120 shown in FIG. 2. For example, process 400 may be performed by the landing system 150 component of the navigation system 120. Furthermore, execution of process 400 may include any of the computing components of the exemplary computer systems of FIGS. 21 or 22. For example, process 400 shown in FIG. 4 may be represented in instructions stored in a memory, which are subsequently executed by a processing unit. Process 400 described with respect to FIG. 4 is an example provided for illustrative purposes and should not be construed as limiting. Other processes may include more or fewer steps than shown while remaining within the scope of the present disclosure. Furthermore, the steps shown in the exemplary process may be performed in an order different from that shown.

[0058] The example process 400 begins by receiving sensor data (i.e., sensory input) from one or more sensors onboard or otherwise associated with the UAV while the UAV is flying through a physical environment, at step 402. As previously mentioned, the sensors may include visual sensors, such as image capture devices 114 and / or 115, as well as other types of sensors 112 configured to recognize certain aspects within the physical environment.

[0059] The example process 400 continues at step 404 by processing the received sensor data to generate a “ground map” of the physical environment. To effectively evaluate multiple potential landing areas, landing system 150 may maintain information about smaller areas of the ground map, referred to herein as “cells.” In other words, the ground map may include an arrangement of multiple cells representing specific portions of a surface in the physical environment. Each cell may be associated with data (also referred to as “characteristic data”) indicative of perceived characteristics of the portion of the physical environment corresponding to the cell. For example, as described in more detail below, the characteristic data for a given cell may include statistics regarding the heights of various points along the surface within the area corresponding to the cell. The characteristic data may also include semantic information associated with physical objects detected within the region corresponding to the cell. In either case, the characteristic data may be based primarily on sensor data received in step 402 and collected by the UAV while it is in flight, but may also be supplemented with information from other data sources, such as remote sensors in the physical environment, sensors on other UAVs, sensors in the mobile device 104, a database of predetermined elevation values, object locations, etc.

[0060] FIG. 5A illustrates an exemplary UAV 100 in flight through a physical environment 500 that includes a physical ground surface 502. As shown in FIG. 5A, a landing system 150 associated with the UAV 100 may maintain and continuously update a two-dimensional (2D) ground map 510 that includes a plurality of cells. In the example shown in FIG. 5A, the cells of the ground map 510 are rectangular in shape and arranged in a rectangular 2D grid that is M cells wide by M cells long. This exemplary arrangement of cells in the ground map 510 is provided for illustrative purposes and should not be construed as limiting. For example, the cells of the ground map may be a different shape than that shown in FIG. 5A. Furthermore, the cells of the ground map 510 may be arranged differently than that shown in FIG. 5A. For example, an alternative ground map may be M cells wide by M cells long (where M is not equal to M).

[0061] In some embodiments, steps 402 and 404 are performed continuously as the UAV 100 flies over the physical environment 502. In other words, as new sensor data is received during UAV flight, the characteristic data associated with each cell of the ground map 510 is continuously updated to reflect the characteristics of the portion of the surface of the physical environment 502 associated with that cell. In this context, "continuously" may refer to performing a given step at short intervals, such as every millisecond. The actual interval utilized depends on the operational requirements of the UAV 100 and / or the computing capabilities of the landing system 150.

[0062] In some embodiments, the ground map 510 is UAV-centered. In other words, the ground map 510 remains centered or otherwise stationary relative to the position of the UAV 100 while the UAV 100 is flying. This means that at any given moment, the ground map 510 may only include characteristic data associated with the portion of the physical environment within a particular range of the UAV 100 based on the dimensions of the ground map 510.

[0063] Associating characteristic data (e.g., height statistics) with portions (i.e., cells) of the ground map simplifies the calculations required to subsequently select an appropriate region on the surface for landing, which may reduce the computational resources required to store and calculate the characteristic data and may decrease latency in selecting an appropriate landing area. Furthermore, the placement of cells in the ground map 510 may be adjusted to meet the operational requirements of the UAV 100 while remaining within certain constraints, such as available computing resources, power usage, etc. For example, increasing the “resolution” of the ground map 510 by increasing the value M may provide the landing system 150 with a more accurate perception of the characteristics of the portion of the physical environment 502 corresponding to the ground map 510. However, this increase in resolution may increase the computational resources required to calculate and maintain the characteristic data associated with the cells, which may exceed the capabilities of the onboard system and / or increase power consumption. In either case, the specific configuration of the ground map 510 may be adjusted (either during design or while the UAV 100 is in flight) to balance the operational requirements with the given constraints.

[0064] 4, the example process 400 continues at step 406 with identifying a footprint including a subset of cells in the ground map 510 associated with characteristic data that meets one or more prescribed landing criteria. As used herein, a "footprint" refers to the arrangement of any subset of cells included in the ground map 510. For example, a footprint may include a single cell in the ground map 510 or may include an arrangement of multiple cells. For example, FIG. 5B shows a diagram of the ground map 510 with a highlighted subset of cells representing footprint 512.

[0065] While the exemplary footprint 512 shown in FIG. 5B includes a rectangular cell arrangement that is N cells wide by N cells long (M is greater than N), the actual cell arrangement that comprises the footprint may vary depending on several factors. For example, the size and / or shape of the UAV 100, along with the accuracy of its navigation system, will affect the type of ground area required to perform a safe landing. As a specific illustrative example, if the UAV 100 measures 2 feet wide by 2 feet long, the ground area must be greater than 4 square feet due to a safety factor. The safety factor may vary based on, for example, the capabilities of the landing system 150, weather conditions, other navigation objectives, user preferences, etc. For example, in response to sensing high winds, the landing system 150 may increase the safety factor, thereby increasing the size of the footprint required to safely land on the surface.

[0066] Additionally, although the arrangement of cells comprising the exemplary footprint 512 is shown as a rectangle, other embodiments employing higher resolution ground maps may identify footprints with more irregular shapes that correspond more closely to the actual shape of the UAV 100. This may increase the computing resources required to evaluate multiple footprint candidates, but may enable landing in tight spaces or uneven terrain.

[0067] As described above, identifying a footprint may include evaluating characteristic data associated with the footprint against one or more prescribed landing criteria. For example, if the characteristic data is associated with height statistics, evaluating the footprint may include aggregating height values ​​for multiple cells of the ground map 510 to identify locations of cells having height statistics within a threshold level of variance (e.g., indicating relatively flat ground). Similarly, if the characteristic data is associated with semantic information about detected objects, evaluating the footprint may include identifying locations of cells that do not include semantic information indicative of dangerous objects, such as bodies of water, trees, rocks, etc. Evaluating the footprint may also include identifying locations of cells that include semantic information indicative of safe landing locations, such as prescribed landing pads, paved areas, or grass.

[0068] In some embodiments, identifying a footprint may include evaluating multiple footprint candidates (e.g., in parallel or serially) and selecting a particular footprint candidate from the multiple footprint candidates, where each of the multiple footprint candidates includes a different subset of cells in ground map 510 and is offset from one another (in the x and / or y directions) by at least one cell. In some embodiments, the selected footprint may be the first footprint candidate evaluated by landing system 150 that meets one or more prescribed landing criteria. Alternatively, landing system 150 may evaluate a certain number of footprint candidates and select the best footprint from those candidates as long as it meets one or more prescribed landing criteria. For example, in the case of height variance, landing system 150 may select the footprint candidate with the lowest height variance from the multiple candidates as long as the height variance is below a threshold variance level.

[0069] Similar to steps 402 and 404, the process of identifying footprints in step 406 may be performed continuously as UAV 100 flies over physical environment 502. In other words, as the ground map is updated based on new sensor data during UAV flight, landing system 150 may continuously evaluate and re-evaluate various footprint candidates against one or more prescribed landing criteria. Again, in this context, "continuously" may refer to re-performing a given step at short intervals, such as every millisecond.

[0070] Once the footprint is identified, the exemplary process 400 continues at step 408 by designating a landing area on a surface within the physical environment based on the identified footprint, and at step 410, autonomously landing the UAV 100 at the designated landing area.

[0071] Landing UAV 100 on the designated landing area may include, for example, generating control commands configured to maneuver UAV 100 to autonomously land on the designated landing area. In some embodiments, this step may include landing system 150 generating landing targets that set predetermined parameters (e.g., landing position, descent rate, etc.) and inputting or otherwise communicating the generated landing targets to motion planner 130, where motion planner 130 processes the landing targets in conjunction with one or more other behavioral targets to generate a planned trajectory, which may then be utilized by flight controller 160 to generate control commands.

[0072] Geometric Smart Landing As mentioned above, in certain embodiments, landing system 150 may be configured to select the flattest area of ​​the ground in response to determining that landing is requested or necessary. The technique described below is referred to as geometric smart landing and relies on the premise that the variance in height (z) for all points observed in a given area is inversely proportional to the flatness of that area. For example, FIG. 6 shows an elevation view of UAV 100 flying above ground 602. As shown in FIG. 6, the variance in height (z) of various points along the ground in area A is greater than the variance in height (z) of various points along the ground in area B due to the slope of the ground in area A. Based on this observation, it may be inferred that area B is more preferable than area A for landing UAV 100. The geometric smart landing technique described below may be implemented instead of or in addition to the semantic smart landing technique discussed below.

[0073] Geometric Smart Landing - Calculating and Maintaining Distributed Data As discussed above, characteristic data such as height statistics may be generated and maintained for multiple cells in a ground map that is continuously updated based on sensor data received from one or more sensors, for example, as described with respect to Figures 5A-5B. Using the height statistics for various cells in the ground map, landing system 150 may identify footprints with height variances that meet certain specified criteria (e.g., below a threshold level).

[0074] FIG. 7A shows a flowchart of an example process 700a for autonomous landing by a UAV using geometric smart landing techniques. Similar to the example process 400 of FIG. 4, one or more steps of the example process 700a may be performed by any one or more of the components of the example navigation system 120 shown in FIG. 2. For example, the process 700a may be performed by the landing system 150 component of the navigation system 120. Furthermore, execution of the process 700a may include any of the computing components of the example computer systems of FIGS. 21 or 22. For example, the process 700a may be represented by instructions stored in a memory that are subsequently executed by a processing unit. The process 700a described with respect to FIG. 7A is an example provided for illustrative purposes and should not be construed as limiting. Other processes may include more or fewer steps than shown while remaining within the scope of the present disclosure. Furthermore, the steps shown in the example process may be performed in an order different from that shown.

[0075] The example process 700a begins by receiving sensor data at step 702 from sensors onboard the UAV 100, for example, as described in step 402 of the example process 400. In the illustrated example, the sensors onboard the UAV 100 include a downward-facing stereo camera configured to capture images of the ground while the UAV 100 flies through the physical environment. The downward-facing stereo camera may be part of an array of stereoscopic navigation cameras that includes, for example, the image capture device 114 described above.

[0076] In step 704, the received sensor data is processed to determine (i.e., generate) data points indicative of height values ​​at a plurality of points along a surface (e.g., the ground) in the physical environment. In this context, a height value may indicate the relative vertical distance between a point along the surface and the position of the UAV 100. The height value may also indicate the vertical position of the point along the surface relative to another frame of reference, such as sea level (i.e., above sea level).

[0077] Steps 704a and 704b describe an exemplary subprocess for determining data points based on images captured from a downward-facing stereoscopic camera. In step 704, images captured by an image capture device (e.g., a stereoscopic camera) are processed to generate a disparity image. As used herein, a "disparity image" is an image representing the disparity between two or more corresponding images. For example, a stereo pair of images (e.g., a left image and a right image) captured by a stereoscopic image capture device will exhibit an inherent offset due to slight differences in the positions of two or more cameras associated with the stereoscopic image capture device. Despite this offset, at least some of the objects displayed in one image will also appear in the other image. However, the image locations of pixels corresponding to such objects will be different. By matching pixels in one image with corresponding pixels in the other and calculating the distance between these corresponding pixels, a disparity image with pixel values ​​based on the distance calculation can be generated. In other words, each pixel in the disparity image may be associated with a value (e.g., a height value) indicating the distance from the image capture device to a captured physical object, such as a surface in a physical environment.

[0078] In some embodiments, the subprocess may continue at step 706b by mapping pixels in the generated disparity image to 3D points in space corresponding to points along a surface (e.g., the ground) in the physical environment, resolving a height (z) value associated with that point. While preferred embodiments rely on capturing height information using a visual sensor such as a stereo camera, other types of sensors may be used as well. For example, UAV 100 may be equipped with a downward-facing range sensor, such as a LIDAR, to continuously scan the ground below the UAV to collect height data.

[0079] The example process 700a continues with generating a ground map at step 706, e.g., similar to that described with respect to step 404 of process 400. Steps 706a and 706b describe an example sub-process for adding height values ​​to particular cells to generate and continuously update the ground map while the UAV 100 flies through the physical environment. For example, in step 706a, height values ​​based on the generated data points are "added" to corresponding cells in the ground map. In some embodiments, this step of adding height values ​​may include determining which cells in the ground map a particular data point is associated with, for example, by determining which cells overlap a 3D point in space corresponding to the data point (i.e., binning the data in the x and y directions).

[0080] Once a particular cell is identified, the height value of the particular point is added to that cell. Specifically, in some embodiments, this process of adding height data to a cell may include updating a height statistic for the cell in step 706b, such as the number of data points collected, the average height value, the median height value, the minimum height value, the maximum height value, the sum of the squared height value differences, or any other type of statistic based on an aggregation of the data points for the particular cell.

[0081] This aggregation of data points through height statistic updates may reduce overall memory usage, as landing system 150 is only required to store information about maintained statistics instead of an entry for each of multiple collected data points. In some embodiments, similar height statistics may be continually updated for each of multiple footprint candidates based on height statistics for cells contained in each of the footprints. Alternatively, height statistics for each footprint candidate may be calculated based on the individual points contained in each footprint, which can be computationally expensive. In either case, the variance of height values ​​in a given footprint at any time may be calculated as the sum of the squared differences in height values ​​divided by the number of points in that area.

[0082] As previously mentioned, the ground map may be continuously updated based on the sensor data as it is received during the flight of the UAV 100. In other words, at least steps 702, 704, and 706 (including associated substeps) may be performed continuously to generate and continuously update the ground map. Again, "continuously" in this context may refer to re-executing a given step at short intervals, such as every millisecond.

[0083] The exemplary process 700a continues, e.g., at step 708, similar to step 406 of exemplary process 400, by identifying footprints that meet one or more indicated criteria. In some embodiments, the landing criteria are met if the height statistics associated with the footprint candidates have a variance below a threshold variance level. Again, the variance may be calculated based on the height statistics for one or more cells included in a given footprint candidate. In some embodiments, the landing criteria are met if the height statistics associated with a particular footprint candidate have the lowest variance compared to the other footprint candidates and if the height statistics are below a threshold variance level.

[0084] Once the footprint is identified, the exemplary process 700a continues by, for example, designating a landing area on a surface within the physical environment based on the identified footprint in step 710, similar to steps 408 and 410 of process 400, and autonomously landing the UAV 100 in the designated landing area in step 712.

[0085] Geometric Smart Landing - Accounting for Uncertainties in Stereo When stereo vision is applied, the geometric smart landing technique may be configured to account for the uncertainty in the estimated / measured height values. In stereo vision, the larger the parallax, the larger the uncertainty in the range (distance from the camera) of that point. The expected variance of a point is scaled by the fourth power of the range of that point. To account for this, landing system 150 may adjust the predetermined height value by an appropriate correction factor. For example, in some embodiments, landing system 150 may divide the increment of the sum of the squared height differences for this given point by the fourth power of the range of that point.

[0086] Geometric Smart Landing - Downweighting the Old Point To make the ground map more responsive to moving objects, landing system 150 may be configured to limit the number of data points contributing to the statistics for each cell or otherwise emphasize more recently collected data points. For example, if a data point based on height measured at a particular point is added to a cell that already has a maximum number of points, the system may be configured to calculate the average of the sum of squared height value differences and then subtract the calculated average from the sum of squared height value differences before updating the statistics for that cell. In this manner, the number of points collected for the cell does not increase. Downweighting of older data may also be performed based on the average height by fixing the number of data points. For example, a typical average height increment may be calculated based on the following: mean[n+1] = (x-mean[n]) / (n+1), where "n" represents both the number of points and the number of updates. This then becomes mean[n+1] = (x-mean[n]) / min(n+1, n_max), where the maximum number of points, "n_max," is set for a given cell. This has a similar effect to downweighting old data.

[0087] FIG. 7B shows a flowchart of an example process 700b for downweighting old points in a cell when updating statistics for a given cell based on new points. In step 720, a cell in the ground map corresponding to a given data point with a height value (i.e., generated in step 704 of process 700a) is identified. As previously described, landing system 150 maintains statistics for each of the cells in the ground map. One of these statistics may include the total number of data points added to that cell. If the total number of data points has not yet reached a designated maximum threshold, the process proceeds to step 724, where a data point is added to the cell and the height statistics for the cell are updated. If the maximum number of data points has already been added to the cell, process 700b instead proceeds to step 722, where another height value data is downweighted, e.g., as described in the previous paragraph, before updating the height statistics for the cell in step 724.

[0088] The maximum number of data points for each cell may be static or may change dynamically over time based on various factors such as the operational requirements of the UAV 100, the computing resources available at any given moment, environmental conditions, etc. The maximum number of points for each of the cells may be uniform across the ground map or may vary between cells. In some embodiments, the maximum number of points may be user-directed.

[0089] Geometric Smart Landing - Landing Spot Selection Assuming that UAV 100 needs to land (in response to a user command or in response to an event such as low battery), landing system 150 can use the calculated variance for each footprint in the ground map to identify footprints that meet certain prescribed criteria (i.e., as described with respect to steps 708 and 710 of process 700a) to select an area on the physical ground for landing. The criteria may dictate the selection of the footprint with the lowest variance in the ground map for landing. Alternatively or additionally, the criteria may dictate some maximum allowable variance. In some embodiments, cells of footprints containing too little data may be avoided regardless of the calculated height variance, because their height statistics may be unreliable. Such criteria may be user-instructed, fixed, or dynamically changed in response to predetermined conditions. For example, if UAV 100's battery is low and an immediate landing is required, landing system 150 may simply select the footprint in the immediate vicinity with the lowest variance in height values. Conversely, if landing is at the request of a user command and there is no immediate risk of power loss, landing system 150 may be comfortable being more selective in its footprint. In such a situation, UAV 100 may continue flying until a footprint is identified that results in a height variance lower than the maximum allowable variance.

[0090] In some embodiments, UAV 100 may be configured to more effectively position itself before selecting a footprint for landing. For example, in one embodiment, in response to determining that it needs to land, landing system 150 may autonomously maneuver UAV 100 to a commanded height above the ground to better assess the situation before landing. For example, a height of at least about 2 meters above the lowest observed point and up to about 4 meters above the highest observed point has been found to effectively balance being able to see important ground areas, not wasting battery power by climbing too high, and being close enough to detect relatively small obstacles.

[0091] One important edge case observed during testing is when two identically flat areas have different ground elevations. For example, if UAV 100 is flying near a cliff, both the top and bottom of the cliff may be flat (i.e., exhibit similar elevation variance). However, one landing area may still be preferable to the other, at least from the perspective of a user accessing UAV 100 after landing. To account for this, landing system 150 may be configured with a heuristic that prefers a footprint that is closest to the last known location (e.g., altitude) of an object (e.g., the user at the time of selecting the footprint for landing). For example, in the cliff scenario described above, landing system 150 may determine (e.g., based on information from tracking system 140) that the tracked user's last known location is at an altitude that matches the identified cliff-top footprint. Thus, landing system 150 may land UAV 100 on the cliff-top footprint rather than the cliff-bottom footprint to avoid the user having to climb down the cliff to retrieve UAV 100 after landing.

[0092] Other heuristics, which may be added manually or learned based on use cases, may ensure that the UAV 100 is configured to land in the most intuitive and convenient (i.e., preferred) location. For example, the UAV 100 may be configured to prefer landing at the UAV's original takeoff height when performing roof inspections. As another example, the UAV 100 may be configured to prefer landing on the path of a tracked object (rather than landing on a nearby bush or tree). As another example, the UAV 100 may be configured to prefer landing in an area of ​​ground seldom traversed by people or machinery to avoid unexpected contact. Conversely, the UAV 100 may also be configured to prefer landing in areas with more traffic to avoid landing in areas that are too far or inaccessible to the user. As another example, the UAV 100 may be configured to prefer landing near a building or other structure (rather than landing in the middle of a field) when operating in an agricultural / farmwork setting. These are just a few examples of heuristics that can be implemented to guide the landing system 150 in selecting an appropriate landing area. Other heuristics may be implemented as well based on the specific usage requirements of the UAV 100.

[0093] Semantic Smart Landing While geometric smart landing can be effective in some situations, it is limited in its ability to distinguish between certain real-world conditions. For example, using geometric smart landing techniques, landing system 150 may identify the surface of a body of water, the flat top of a tree, or a road as effective for landing because the height variance is sufficiently low. Such surfaces may not represent desirable landing areas for a variety of reasons. Furthermore, some relatively small obstacles (e.g., small bushes, thin tree branches, small rocks, etc.) may not have enough impact on the variance of height values ​​in a given cell to prevent the landing system from selecting a footprint containing that cell. Overall, this may result in the selection of an undesirable landing area.

[0094] To avoid such situations, in some embodiments, semantic information associated with objects in the surrounding environment may be extracted (e.g., based on an analysis of captured images of the surrounding physical environment) and incorporated into the aforementioned ground map to aid in the landing process. For example, FIG. 8A illustrates UAV 100 flying over a physical environment similar to that shown in FIGS. 5A-5B. As shown in FIG. 8A, environment 500 is populated with various objects, such as a person 802, a lake 804, a rock 806, and a tree 808. In some embodiments, semantic knowledge of these objects may be extracted, for example, by processing images captured by image capture device 114 / 115, and added to corresponding cells of ground map 510, for example, as labels or tags or by updating cost factors associated with the cells. For example, one or more cells overlapping lake 804 may be labeled "water." In some embodiments, landing system 150 may be configured to avoid landing in any footprint that includes a cell (or at least a predetermined number of cells) that includes a water label. Similarly, cells may be labeled with more specific information, such as walkable or traversable ground. The semantic smart landing techniques described below may be performed instead of or in addition to the geometric smart landing techniques described above. For example, characteristic data associated with height values ​​and / or semantic information may be added to the same ground map or to separate, overlapping ground maps. In either case, the landing system may perform certain steps of the geometric and semantic smart landing techniques in parallel to evaluate candidate footprints and identify footprints that meet prescribed landing criteria based on both height statistics and semantic information.

[0095] FIG. 9 shows a flowchart of an exemplary process for extracting semantic information about an object and inputting the semantic information into a ground map to thereby guide the landing of a UAV 100. One or more steps of this exemplary process 900 may be performed by any one or more of the components of the exemplary navigation system 120 shown in FIG. 2. For example, process 900 may be performed by the landing system 150 component of the navigation system 120. Furthermore, execution of process 900 may include any of the computing components of the exemplary computer systems of FIGS. 21 or 22. For example, process 900 may be represented in instructions stored in a memory, which are subsequently executed by a processing unit. The process 900 described with respect to FIG. 9 is an example provided for illustrative purposes and should not be construed as limiting. Other processes may include more or fewer steps than shown while remaining within the scope of the present disclosure. Furthermore, the steps shown in the exemplary process may be performed in an order different from that shown.

[0096] The example process 900 begins by receiving sensor data at step 902 from sensors onboard the UAV 100, for example, as described in step 402 of the example process 400. In the illustrated example, the sensors onboard the UAV 100 include a downward-facing stereo camera configured to capture images of the ground while the UAV 100 flies through the physical environment. The downward-facing stereo camera may be part of an array of stereoscopic navigation cameras that includes, for example, the image capture device 114 described above.

[0097] In step 904, the received sensor data is processed to detect one or more physical objects in the physical environment and extract semantic information associated with the detected one or more physical objects. In the illustrated example, landing system 150 may extract semantic information about a given object captured in an image based on an analysis of pixels in the image. The semantic information about the captured object may include information such as the object's category (i.e., class), location, shape, size, scale, pixel segmentation, orientation, inter-class appearance, activity, pose, etc. For example, techniques for detecting objects and extracting semantic information associated with the detected objects are described in more detail below with reference to FIGS. 15-20 under the section titled "Object Detection." The process of detecting objects in the physical environment is described in a later section as being performed by a separate tracking system 140 to facilitate tracking of the object. However, such a process may also be performed by landing system 150 as part of a smart landing process. In some embodiments, landing system 150 may communicate with tracking system 140 to receive semantic information about objects detected by tracking system 140. Alternatively or additionally, landing system 150 may execute certain processes independently and in parallel with tracking system 140 .

[0098] At step 906, the ground map is updated based on the extracted semantic information about the detected physical objects in the physical environment. As previously described, semantic information may include labels, tags, or any other data indicative of the detected objects. For example, the semantic information associated with rock 806 in FIG. 8 may include the label “rock.” The “rock” label may include additional information about the rock, such as its type, shape, size, orientation, and an instance identifier. For a mobile object, such as person 802, the associated label may include information about its current activity (e.g., “running,” “walking,” or “stopping”). Thus, in such an embodiment, step 906 may include adding data points in the form of labels to particular cells of the ground map corresponding to the locations of the detected objects. For example, FIG. 8B shows another diagram of physical environment 500 with various elements indicating semantic labels added to cells, such as element 826 indicating the semantic label “rock” added to the cell corresponding to the location of rock 806.

[0099] In some embodiments, the ground map may be updated based on the semantic information by calculating a cost value associated with the detected object and adding the cost value to the cell corresponding to the location of the detected object. In this context, a "cost value" may indicate a level of risk, for example, of damage to people or property, collision with an object, a failed landing, etc. Certain objects, such as trees 808 and / or lakes 804, may pose a greater risk than other objects, such as rocks 806, and therefore may be associated with higher cost values. In some embodiments, the cost values ​​of multiple detected objects in a given cell may be aggregated to generate cost statistics, such as an average cost value, a median cost value, a minimum cost value, a maximum cost value, or the like, similar to the height values ​​discussed above. Such cost values ​​may be added to cells in addition to or instead of semantic labels.

[0100] 9 , the exemplary process 900 continues, e.g., at step 908, similar to step 406 of exemplary process 400, by identifying footprints that meet one or more prescribed criteria. In some embodiments, the landing criteria are met if one or more cells of the candidate footprint do not contain semantic labels indicating known hazards, such as trees, bodies of water, rocks, etc. FIG. 8B illustrates a footprint 514 that may meet the prescribed landing criteria because the placement of cells within the footprint does not contain semantic labels for predetermined known hazards, such as trees, rocks, people, bodies of water, etc.

[0101] In some embodiments, predetermined characteristics of the detected object may affect the evaluation of the candidate footprint against the landing criteria. For example, even if a footprint contains a cell associated with a predetermined detected object, the candidate footprint may meet the landing criteria as long as the detected object meets a predetermined parameter threshold (e.g., based on size, type, movement, etc.). As an illustrative example, even if a footprint contains a cell labeled "rock," the candidate footprint may meet the indicated landing criteria as long as the detected rock is below a threshold size. Again, the predetermined parameter, such as size, may be indicated in the label.

[0102] In some embodiments, the landing criteria are met if the calculated cost statistic for the footprint is below a threshold cost level and / or is lower than other candidate footprints. Again, the cost statistic may be calculated based on detected objects at locations corresponding to cells in the footprint. For example, each of the semantic labels shown in FIG. 8B may be associated with a corresponding cost value based on the relative risk of landing within the area in which the corresponding object is located. The cost value is a number on a scale of 0 to 1, with 0 indicating no risk and 1 indicating maximum risk. As an illustrative example, the semantic label for "lake" may be associated with a cost value of 0.7 because landing in water poses a high risk of damage to the UAV 100. Conversely, the semantic label "rock" may be associated with a cost value of 0.3 (assuming the rock is relatively small) because landing near a rock poses a relatively low risk of damage to the UAV 100.

[0103] The one or more landing criteria utilized to evaluate footprint candidates based on cost statistics may remain static or may change dynamically based on, for example, user input, changing environmental conditions, the operating state of the UAV 100, etc. For example, a threshold cost level may gradually increase or an object cost value may decrease as the battery approaches zero charge to allow the UAV 100 to land in an area with a potentially dangerous object instead of crashing. However, certain overriding behavioral goals, such as avoiding collisions with people, may take priority over performing a safe landing.

[0104] Once the footprint is identified, the exemplary process 900 continues by, for example, designating a landing area on a surface within the physical environment based on the identified footprint in step 910, similar to steps 408 and 410 of process 400, and autonomously landing the UAV 100 in the designated landing area in step 912.

[0105] Learned Smart Landing In some embodiments, a user may be provided with options, for example, via an interface presented on the mobile device 104, to select a location for the UAV 100 to land. In one embodiment, the user may interact with a view of the physical environment captured by the UAV 100 and displayed via a touchscreen display on the mobile device 104. The user may select a location for landing by touching a portion of the displayed view that corresponds to the physical ground location where the UAV will land. This type of selection by the user, as well as other types of instructions or guidance regarding the landing (including monitoring a landing completely remotely controlled by the user), may be monitored by the landing system 150 for purposes of training its internal landing selection and motion planning process. The user may be provided with the option to opt in to such information collection. If the user opts in, information may be collected from any of the following: pixels in the view selected by the user for landing, images from the target camera 115 pointed downward with its gimbal pointing to the ground, the downward-pointing navigation camera 114 or any other navigation camera 114, the location of the UAV 100, the location of the most recent tracked target, the location of the mobile device 104, etc.

[0106] This collected information can be used to learn where users prefer to land their UAVs, thereby characterizing autonomous smart landing. In some embodiments, data is collected from multiple user UAVs to train and update the core algorithm of the smart landing algorithm. Landing algorithm updates may then be sent as software updates to the various UAVs, for example, over a network. Alternatively, data collection and training may be local to a particular UAV 100. The landing system 150 within a given UAV 100 may monitor landings to identify a particular user's preferences without sharing this information with a central data collection facility. The landing system 150 of the UAV 100 can then use the collected information to train its internal landing algorithm.

[0107] Handling edge cases Autonomous smart landing (and navigation in general) may present various edge cases that require special handling, including, for example, warnings issued to the user and / or specific actions to mitigate potential hazards.

[0108] Handling edge cases - caveats An alert generated in response to the detection of a predetermined condition may be displayed to the user via the connected mobile device 104 or otherwise presented to the user (e.g., through an audible alarm).

[0109] In an exemplary edge case, the UAV 100 may detect obstructions (e.g., from the ground or the user's hands) to the planned path of takeoff by checking the planned trajectory of takeoff against an obstacle map (e.g., a voxel-based occupancy map). In response to detecting a potential obstruction to the takeoff, the system may cause a warning about the obstruction to be presented to the user via the mobile device 104. The system may further be configured to restrict takeoff until the detected obstruction is no longer present.

[0110] Another example edge case involves the UAV 100 chasing a tracked object through a narrow area. At some point, the UAV 100 may determine that the limited space in the area will prevent the UAV 100 from achieving its goal of chasing the tracked object. In this situation, the system may cause a warning to be presented to the user via the mobile device 104 regarding the potential loss of the acquisition object. The warning may include instructions to navigate an alternate route and / or the option to manually control the UAV.

[0111] Another exemplary edge case involves the UAV 100 operating in low lighting conditions. As mentioned above, the UAV 100, in certain embodiments, may rely on images captured by one or more image capture devices to guide its autonomous flight. If the lighting conditions fall below a predetermined threshold level, the system may cause a warning to be presented to the user via the mobile device 104 regarding the insufficient lighting conditions.

[0112] Handling edge cases - mitigating actions As mentioned above, the UAV 100 may communicate with the user's mobile device 104 via a wireless communication link, such as Wi-Fi. Over distance, the wireless signal weakens and sometimes the connection is lost. In some embodiments, the navigation system 120 is configured to automatically and autonomously return the UAV 100 to the last known location (e.g., GPS location) of the mobile device 104 in response to detecting a weak wireless communication signal or a complete loss of the wireless communication link. By returning to (or at least moving toward) the last known location of the mobile device 104, the UAV 100 can more easily re-establish a wireless communication link with the mobile device 104. Alternatively or additionally, the system may allow the user to set a predefined location for the UAV to return to in the event of a loss of signal. This location could be a waypoint along a trajectory, a pin on a virtual map, the UAV's takeoff location, the user's home, etc. Information regarding remaining battery life may also help inform the UAV's response. For example, the UAV may select from one of a number of predetermined return locations (e.g., those listed above) based on remaining battery life and transmit a designation of its intended return location to the user (e.g., via the mobile device 104). The above features may also be used if, for some reason, the UAV 100 loses visual track of the target.

[0113] In some situations, the navigation system 120 may determine that it is no longer safe to continue following the tracked object. In particular, situations may arise where the operation of one or more components of the UAV 100 is impaired. For example, if one of the image capture devices 114 stops capturing images, the navigation system 120 can no longer rely on sensory input for obstacle avoidance. As another example, if a processing component (e.g., a CPU) becomes too hot, throttling by the processing component may occur, resulting in reduced processing performance. This may significantly impact the execution of the software stack and lead to autonomous navigation errors, particularly collisions. In such situations, the navigation system 120 may cause the UAV 100 to stop pursuing the tracked object and autonomously land the UAV 100 to avoid any damage from a collision.

[0114] FIG. 11 shows a flowchart of an example process 1000 for landing the UAV 100 based on the operating state of the UAV 100. One or more steps of this example process 1000 may be performed by any one or more of the components of the example navigation system 120 shown in FIG. 2. For example, the process 1000 may be performed by the tracking system 140 and landing system 150 components of the navigation system 120. Furthermore, execution of the process 1000 may include any of the computing components of the example computer systems of FIGS. 21 or 22. For example, the process 1000 may be represented in instructions stored in a memory, which are subsequently executed by a processing unit. The process 1000 described with respect to FIG. 10 is an example provided for illustrative purposes and should not be construed as limiting. Other processes may include more or fewer steps than shown while remaining within the scope of the present disclosure. Furthermore, the steps shown in the example process may be performed in an order different from that shown.

[0115] The example process 1000 begins at step 1002 by detecting and tracking a physical object in a physical environment and then causing the UAV 100 to autonomously follow the tracked object at step 1004. Steps 1002 and 1004 may be performed by the tracking system 140, for example, based on images captured from one or more image capture devices 114 / 115. The tracking system 140 may cause the UAV 100 to autonomously follow the tracked object by generating a tracking target, which is input to or otherwise communicated to the motion planner 130, which then processes the tracking target along with one or more other behavioral goals to generate a planned trajectory for the UAV 100.

[0116] In step 1006, the operating conditions of the UAV are monitored. In some embodiments, the operating conditions are continuously monitored while the UAV 100 is flying and autonomously pursuing the tracked object. In some embodiments, the landing system 150 performs step 1006 to monitor the operating conditions of the UAV 100, although this step may be performed by any other on-board system. Monitoring the operating conditions may include, for example, actively querying the various on-board systems for status updates, passively receiving notifications from the various on-board systems with status updates, or any combination thereof.

[0117] As part of monitoring the operating conditions in step 1006, landing system 150 may determine whether the operating conditions meet one or more operating criteria. If the operating conditions of UAV 100 do not meet one or more operating criteria, landing system 150 causes UAV 100 to stop pursuing the tracked object and perform an automatic smart landing in step 1008, for example, according to the presented techniques.

[0118] Examples of situations in which operating conditions do not meet the operating criteria include when the power source (e.g., a battery) onboard the UAV 100 falls below a threshold power level, when the storage device (e.g., for storing captured video) onboard the UAV 100 falls below a threshold available storage level, when any of the systems onboard the UAV 100 is malfunctioning or overheating, when the wireless communication link with the mobile device 104 is lost or of poor quality, or when visual contact with the tracked target is lost (the tracked target is lost). Other operating criteria may be implemented to guide the decision on when to initiate the smart landing process as well.

[0119] In some embodiments, causing UAV 100 to stop pursuing the tracked object may include landing system 150 communicating with the tracking system to cancel or otherwise modify a previously generated tracked target. Alternatively or additionally, landing system 150 may communicate a command to motion planner 130 to cancel or otherwise modify a previously input tracked target. In either case, landing system 150 may cause UAV 100 to perform a smart landing process by generating landing targets and communicating the landing targets to motion planner 130, as described above.

[0120] In some rare cases, the UAV 100 may need to land without time, and the ability to perform a smart landing using any of the aforementioned techniques may be necessary. For example, if state estimation by the navigation system 120 fails, the system may not be able to control the position of the UAV 100 to the extent necessary to maintain hover. In such a situation, the best course of action may be to immediately begin a controlled descent, while checking to see if the UAV has landed until the UAV 100 has safely landed on the ground.

[0121] 11A-11B show flowcharts of example processes 1100a-b for initiating a controlled descent in response to determining that a smart landing is not possible. One or more steps of the example processes 1100a-b may be performed by any one or more of the components of the example navigation system 120 shown in FIG. 2. For example, the process 1100 may be performed by the landing system 150 component of the navigation system 120. Furthermore, execution of the process 1100 may include any of the computing components of the example computer systems of FIGS. 21 or 22. For example, the process 1100 may be represented in instructions stored in a memory, which are subsequently executed by a processing unit. The process 1100 described with respect to FIGS. 11A-11B is an example provided for illustrative purposes and should not be construed as limiting. Other processes may include more or fewer steps than shown while remaining within the scope of the present disclosure. Furthermore, the steps shown in the example processes may be performed in an order different from that shown.

[0122] The example process 1100a begins by initiating a smart landing process at step 1102, for example, according to the example process 400 described with respect to Figure 4. During the smart landing process, the landing system 150 continuously monitors the operating conditions of the UAV 100 to determine whether a smart landing is possible. If the landing system 150 determines that a smart landing is possible, the smart landing process continues at step 1104 until the UAV 100 safely lands.

[0123] On the other hand, if during the smart landing process, the system determines that the smart landing process is no longer possible, landing system 150 will cause UAV 100 to perform a controlled descent until UAV 100 contacts the ground in step 1106. Performing a controlled descent of UAV 100 may include generating a controlled descent target and communicating the target to motion planner 130. Alternatively, if trajectory generation by motion planner 130 fails or is otherwise compromised, the landing system may send instructions to flight controller 160 causing the flight controller to generate control commands configured to cause UAV 100 to descend in a controlled manner. This may include gradually reducing the power output to each of one or more rotors of UAV 100 in a balanced manner to keep UAV 100 substantially level as it descends. When available, certain sensor data (e.g., from on-board gyroscopes and / or accelerometers) may be used by flight controller 160 to adjust power to the rotors to counter any rotational and / or lateral acceleration during controlled descent, even if the actual position and / or orientation of UAV 100 is unknown.

[0124] In some embodiments, landing system 150 may attempt to find an alternate landing area if a smart landing is not possible at the initial landing area. For example, process 1100b shown in FIG. 11B begins by first initiating a smart landing process to land on the initial landing area, e.g., similar to step 1102 of process 1100a. Also similar to process 1100a, example process 1100b proceeds to step 1124 to continue the smart landing process if landing system 150 determines that a smart landing is possible at the initial landing area.

[0125] However, if the landing system determines that a smart landing at the initial landing area is not possible, process 1100b proceeds to step 1126, where landing system 150 causes UAV 100 to initiate a controlled descent. During the controlled descent, the landing system attempts to identify an alternative landing area. This identification may be performed, for example, by identifying an alternative footprint in step 1128, designating the alternative landing area based on the alternative footprint in step 1130, and autonomously landing UAV 100 at the alternative landing area in step 1132. This process of searching for alternative landing areas may be performed continuously during the controlled descent until UAV 100 lands.

[0126] 11A or 11B, in some embodiments, landing system 150 may cause UAV 100 to perform a controlled descent to a particular height above the ground if such a maneuver is possible given the current operating state of UAV 100. Landing system 150 may then again attempt to initiate the smart landing process if such a landing is possible. If a smart landing is still not possible, landing system 150 will cause UAV 100 to continue the controlled descent until UAV 100 has landed.

[0127] localization The navigation system 120 of the UAV 100 may employ any number of systems and techniques for localization. FIG. 12 shows a diagram of an example localization system 1200 that may be utilized to guide the autonomous navigation of a transport vehicle, such as the UAV 100. In some embodiments, the position and / or orientation of the UAV 100 and various other physical objects within a physical environment may be estimated using any one or more of the subsystems illustrated in FIG. 12. By tracking changes in position and / or orientation over time (continuously or at regular or irregular time intervals (i.e., repeatedly)), the motion (e.g., velocity, acceleration, etc.) of the UAV 100 and other objects may also be estimated. Thus, any of the systems described herein for determining position and / or orientation may similarly be employed to estimate motion.

[0128] As shown in FIG. 12 , an exemplary positioning system 1200 may include a UAV 100, a global positioning system (GPS) including multiple GPS satellites 1202, a cellular system (with access to a source of positioning data 1206) including multiple cellular antennas 1204, a Wi-Fi system (with access to a source of positioning data 1206) including multiple Wi-Fi access points 1208, and / or a mobile device 104 operated by a user 106.

[0129] Satellite-based positioning systems, such as GPS, can provide an effective global position estimate (within a few meters) of any device equipped with a receiver. For example, as shown in FIG. 12, signals received from satellites of a GPS system 1202 at a UAV 100 can be used to estimate the global position of the UAV 100. Similarly, the position of other devices (e.g., mobile device 104) can be determined by communicating (e.g., via wireless communication link 116) and comparing the global position of the other devices.

[0130] Positioning techniques may also be applied in the context of various communication systems configured to transmit communication signals wirelessly. For example, various positioning techniques may be applied to estimate the position of the UAV 100 based on signals transmitted between the UAV 100 and either the cellular antenna 1204 of a cellular system or the Wi-Fi access points 1208, 1210 of a Wi-Fi system. Known positioning techniques that may be implemented include, for example, time of arrival (ToA), time difference of arrival (TDoA), round trip time (RTT), angle of arrival (AoA), and received signal strength (RSS). Furthermore, hybrid positioning systems that implement multiple techniques, such as TDoA and AoA, ToA and RSS, or TDoA and RSS, may be used to improve accuracy.

[0131] Some Wi-Fi standards, such as 802.11ac, allow for RF signal beamforming (i.e., directional signal transmission using a phase-shifted antenna array) from a transmitting Wi-Fi router. Beamforming may be achieved by transmitting RF signals at different phases from spatially distributed antennas ("phased antenna arrays"), which may cause constructive interference at certain angles and destructive interference at other angles, resulting in a target directional RF signal field. Such a target field is conceptually illustrated in FIG. 12 by the dotted line 1212 emanating from the Wi-Fi router 1210.

[0132] An inertial measurement unit (IMU) may be used to estimate the device's position and / or orientation. An IMU is a device that measures the angular velocity and linear acceleration of a transport vehicle. These measurements may be fused with other sources of information (e.g., those described above) to accurately estimate velocity, orientation, and sensor calibration. As described herein, the UAV 100 may include one or more IMUs. Using a method commonly referred to as "dead reckoning," the IMU (or associated system) may estimate a current position based on a previously measured position and the elapsed time since the previously measured position using the measured acceleration. While somewhat effective, the accuracy achieved by dead reckoning based on measurements from an IMU degrades rapidly due to the cumulative effect of errors in each predicted current position. The error is further exacerbated by the fact that each predicted position is based on a calculated integral of the measured velocity. To counteract such effects, embodiments utilizing IMU-based localization may include localization data from other sources (e.g., the GPS, Wi-Fi, and cellular systems mentioned above) to continuously update the object's last known position and / or orientation. Additionally, a nonlinear estimation algorithm (one embodiment is an "extended Kalman filter") may be applied to the series of measured positions and / or orientations to generate a real-time optimized prediction of the current position and / or orientation based on the assumed uncertainty in the observation data. Kalman filters are commonly applied in the fields of aircraft navigation, guidance, and control.

[0133] Computer vision may be used to estimate the position and / or orientation of the capture camera (and, by extension, the device to which the camera is coupled) and other objects in the physical environment. In this context, the term “computer vision” may generally refer to any method of acquiring, processing, analyzing, and “understanding” captured images. Computer vision may be used to estimate position and / or orientation using several different methods. For example, in some embodiments, raw image data received from one or more image capture devices (onboard or remote from UAV 100) may be received and processed to correct for certain variables (e.g., differences in camera orientation and / or intrinsic parameters (e.g., lens variations)). As discussed above with respect to FIG. 1 , UAV 100 may include two or more image capture devices 114 / 115. By comparing images captured from two or more viewpoints (e.g., at different time steps from a moving image capture device), a system employing computer vision may compute an estimate of the position and / or orientation of a vehicle (e.g., UAV 100) to which the image capture device is attached and / or of a captured object (e.g., a tree, a building, etc.) within the physical environment.

[0134] Computer vision can be applied to estimate position and / or orientation using a process called "visual odometry." FIG. 13 illustrates at a high level the working concept behind visual odometry. As an image capture device moves through space, multiple images are captured sequentially. Due to the motion of the image capture device, the captured images of the surrounding physical environment change from frame to frame. In FIG. 13, this is illustrated by an initial image capture FOV 1352 and a subsequent image capture FOV 1354 captured as the image capture device moves from a first position to a second position over a period of time. In both images, the image capture device may capture real-world physical objects, such as a house 1380 and / or a person 1302. Computer vision techniques are applied to the series of images to detect and match features of the physical objects captured in the image capture device's FOV. For example, a system employing computer vision may search for correspondences among pixels of digital images with overlapping FOVs. This correspondence may be identified using several different methods, including correlation-based and feature-based methods. As shown in FIG. 13 , features such as the head of a human subject 1302 or the corner of a chimney on a house 1380 may be identified, matched, and thereby tracked. By incorporating sensor data from an IMU (or accelerometer(s) or gyroscope(s)) associated with the image capture device into the tracked features of the image capture, estimates may be made about the position and / or orientation relative to the object 1380, 1302 captured in the image. Furthermore, these estimates may be used to calibrate various other systems, for example, by estimating differences in camera orientation and / or intrinsic parameters (e.g., lens changes) or IMU bias and / or orientation. Visual odometry may be applied to estimate the position and / or orientation of the UAV 100 and / or other objects, both in the UAV 100 and in any other computing device, such as a mobile device 104.Furthermore, by communicating these estimates between systems (e.g., via wireless communication link 116), estimates may be computed for their respective positions and / or orientations relative to each other. Position and / or orientation estimates based in part on sensor data from onboard IMUs may suffer from error propagation problems. As previously mentioned, optimization techniques may be applied to such estimates to counteract uncertainties. In some embodiments, a nonlinear estimation algorithm (one embodiment is an “extended Kalman filter”) may be applied to a series of measured positions and / or orientations to generate a real-time optimized prediction of the current position and / or orientation based on assumed uncertainties in the observation data. Such estimation algorithms may similarly be applied to generate smooth motion estimates.

[0135] In some embodiments, data received from sensors onboard the UAV 100 may be processed to generate a 3D map of the surrounding physical environment while estimating the relative position and / or orientation of the UAV 100 and / or other objects within the physical environment. This process is sometimes referred to as simultaneous localization and mapping (SLAM). In such embodiments, a system according to the present teachings may use computer vision processing to search for close correspondences between images with overlapping FOVs (e.g., images acquired during successive time steps and / or stereo images acquired within the same time step). The system may then use these close correspondences to estimate the depth or distance to each pixel represented in the respective images. These depth estimates may then be used to continuously update the generated 3D model of the physical environment, taking into account estimated motion of the image capture device (i.e., the UAV 100) within the physical environment.

[0136] In some embodiments, the 3D model of the surrounding physical environment may be generated as a 3D occupancy map including a plurality of voxels, each corresponding to a 3D volume of space in the physical environment that is at least partially occupied by a physical object. For example, FIG. 14 shows an example diagram of a 3D occupancy map 1402 of a physical environment including a plurality of cubic voxels. Each of the voxels in the 3D occupancy map 1402 corresponds to a space in the physical environment that is at least partially occupied by a physical object. The navigation system 120 of the UAV 100 may be configured to navigate the physical environment by planning a 3D trajectory 1420 through the 3D occupancy map 1402 that avoids the voxels. In some embodiments, the 3D trajectory 1420 planning using the 3D occupancy map 1402 may be optimized by applying an image-space motion planning process. In such an embodiment, the planned 3D trajectory 1420 of the UAV 100 is projected into the image space of a captured image for analysis regarding predetermined identified high-cost regions (e.g., regions with invalid depth estimates).

[0137] Computer vision may also be applied using sensing technologies other than cameras, such as light detection and ranging (LIDAR) technology. For example, a LIDAR-equipped UAV 100 may emit one or more laser beams up to 360 degrees around the UAV 100 in a scan. The light received by the UAV 100 as the laser beams reflect off physical objects in the surrounding physical world may be analyzed to construct a real-time 3D computer model of the surrounding physical world. Depth sensing through the use of LIDAR may, in some embodiments, enhance depth sensing through pixel correspondence, as described above. Additionally, images captured by a camera (e.g., as described above) may be combined with a laser-constructed 3D model to form a textured 3D model, which may be analyzed in real time or near real time (e.g., using computer vision algorithms) for physical object recognition.

[0138] The computer vision-assisted localization techniques described above may calculate the position and / or orientation of objects in the physical world in addition to the position and / or orientation of UAV 100. The estimated positions and / or orientations of these objects may then be fed to motion planner 130 of navigation system 120 to plan a path that avoids obstacles while meeting predetermined objectives (e.g., as described above). Additionally, in some embodiments, navigation system 120 may incorporate data from proximity sensors (e.g., electromagnetic, acoustic, and / or optical-based) to more accurately estimate the location of obstacles. Further improvement is possible through the use of stereoscopic computer vision with multiple cameras, as described above.

[0139] The localization system 1200 of Figure 12 (including all of the associated subsystems described above) is merely one example of a system configured to estimate the position and / or orientation of the UAV 100 and other objects within a physical environment. The localization system 1200 may include more or fewer components than shown, may combine two or more components, or may have a different configuration or arrangement of components. Some of the various components shown in Figure 12 may be implemented in hardware, software, or a combination of both hardware and software, and may include one or more signal processing and / or application specific integrated circuits.

[0140] Object Tracking The UAV 100 may be configured to track one or more objects, for example, to enable intelligent autonomous flight. In this context, the term "object" may include any type of physical object occurring in the physical world. Objects may include dynamic objects such as people, animals, and other vehicles. Objects may also include static objects such as terrain, buildings, and furniture. Additionally, certain descriptions herein may refer to a "subject" (e.g., a human object 102). As used in this disclosure, the term "subject" may simply refer to an object being tracked using any of the disclosed techniques. Thus, the terms "object" and "subject" may be used interchangeably.

[0141] 2, tracking system 140 associated with UAV 100 can be configured to track one or more physical objects based on images of the objects captured by image capture devices (e.g., image capture devices 114 and / or 115) onboard UAV 100. While tracking system 140 can be configured to operate solely based on input from the image capture devices, tracking system 140 can also be configured to incorporate other types of information to aid in tracking. For example, various other techniques for measuring, estimating, and / or predicting the relative position and / or orientation of UAV 100 and / or other objects are described with respect to FIGS. 12-20.

[0142] In some embodiments, tracking system 140 may be configured to fuse information related to two major categories: semantics and 3D geometry. When an image is received, tracking system 140 may extract semantic information about a given object captured in the image based on an analysis of pixels in the image. Semantic information about the captured object may include information such as the object's category (i.e., class), location, shape, size, scale, pixel segmentation, orientation, inter-class appearance, activity, and pose. In an exemplary embodiment, tracking system 140 may identify the object's general location and category based on the captured image and then determine or infer additional details about individual instances of the object based on further processing. Such a process may be performed as a series of separate operations, a series of parallel operations, or a single operation. For example, FIG. 15 illustrates an exemplary image 1520 captured by a UAV flying through a physical environment. As shown in FIG. 15, the exemplary image 1520 includes a capture of two physical objects, specifically, two people present in the physical environment. Exemplary image 1520 may represent a single frame in a sequence of frames of video captured by a UAV. Tracking system 140 may first identify a rough location of a captured object in image 1520. For example, pixel map 1530 shows two dots corresponding to the rough location of the captured object in the image. These rough locations may be expressed as image coordinates. Tracking system 140 may further process captured image 1520 to determine information about individual instances of the captured object. For example, pixel map 1540 shows the results of additional processing of image 1520 that identifies pixels corresponding to individual object instances (i.e., people in this case). Objects may be located and identified in the captured images, and semantic cues may be used to associate identified objects occurring in multiple images.For example, as mentioned above, the captured image 1520 shown in Figure 15 may represent a single frame in a sequence of frames of a captured video. Using semantic cues, tracking system 140 may associate regions of pixels captured in multiple images as corresponding to the same physical object occurring in the physical environment.

[0143] In some embodiments, tracking system 140 may be configured to utilize the 3D geometry of an identified object to associate semantic information about the object based on images captured from multiple views of the physical environment. Images captured from multiple views may include images captured by multiple image capture devices at different positions and / or orientations at a single time. For example, each image capture device 114 shown attached to UAV 100 in FIG. 1 may include multiple cameras at slightly offset positions (to achieve stereoscopic capture). Furthermore, even if not individually configured for stereoscopic image capture, multiple image capture devices 114 may be positioned at different positions relative to UAV 100, for example, as shown in FIG. 1. Images captured from multiple views may also include images captured by the image capture devices at multiple times as the image capture devices move through the physical environment. For example, as UAV 100 moves through the physical environment, either image capture device 114 and / or 115 attached to UAV 100 individually captures images from multiple views.

[0144] Using an online visual-inertial state estimation system, the tracking system 140 can determine or estimate the trajectory of the UAV 100 as it moves through the physical environment. Thus, the tracking system 140 can use the known or estimated 3D trajectory of the UAV 100 to associate semantic information in captured images, such as the location of detected objects, with information about the object's 3D trajectory. For example, FIG. 16 illustrates a trajectory 1610 of the UAV 100 moving through a physical environment. As the UAV 100 moves along the trajectory 1610, one or more image capture devices (e.g., devices 114 and / or 115) capture images of the physical environment at multiple views 1612a-c. Included within the images of the multiple views 1612a-c are captures of objects, such as a human subject 102. By processing the captured images from the multiple views 1612a-c, a trajectory 1620 of the object can also be resolved.

[0145] Object detection in a captured image generates rays along which the object lies, with some uncertainty, from the center position of the capturing camera to the object. The tracking system 140 may calculate depth measurements for these detections to create planes along which the object lies, parallel to the focal plane of the camera, with some uncertainty. These depth measurements may be calculated by a stereo vision algorithm operating on pixels corresponding to the object between two or more camera images of different views. The depth calculation may take into account, in particular, pixels labeled as being part of the object of interest (e.g., object 102). The combination of these rays and planes over time

[0146] While the tracking system 140 may be configured to rely solely on visual data from an image capture device mounted on the UAV 100, data from other sensors (e.g., sensors on the object, sensors on the UAV 100, or sensors in the environment) may be incorporated within this framework, if available. Additional sensors may include a GPS, an IMU, a barometer, a magnetometer, and a camera or other device, such as the mobile device 104. For example, a GPS signal from a mobile device 104 carried by a person may provide a rough location measurement of the person fused with visual information from an image capture device mounted on the UAV 100. IMU sensors on the UAV 100 and / or mobile device 104 may provide acceleration and angular velocity information, a barometer may provide relative altitude, and a magnetometer may provide orientation information. Images captured by a camera on the mobile device 104 carried by the person may be fused with images from the camera mounted on the UAV 100 to estimate the relative pose between the UAV 100 and the person by identifying common features captured in the images. Various other techniques for measuring, estimating, and / or predicting the relative position and / or orientation of UAV 100 and / or other objects are described with respect to FIGS.

[0147] In some embodiments, data from various sensors are input into a spatiotemporal factor graph to probabilistically minimize total measurement error using nonlinear optimization. FIG. 17 shows a diagrammatic representation of an exemplary spatiotemporal factor graph 1700 that can be used to estimate an object's 3D trajectory (e.g., including its pose and velocity over time). In this example, the spatiotemporal factor graph 1700 shown in FIG. 17 shows variable values ​​such as pose and velocity (represented as nodes (represented as 1702 and 1704, respectively)) connected by one or more motion model processes (represented as node 1706 along the connecting edge). For example, an estimate or prediction of the pose of UAV 100 and / or other objects at time step 1 (i.e., variable X(1)) may be computed by inputting the pose and velocity estimated at the previous time step (i.e., variables X(0) and V(0)), and may also be computed by inputting various sensory inputs, such as stereo depth measurements and camera image measurements, via one or more motion models. Spatiotemporal factor models may be combined with outlier rejection mechanisms, in which measurements that deviate too much from the estimated distribution are discarded. To estimate a 3D trajectory from measurements at multiple time instants, one or more motion models (or process models) are used to connect the estimation variables between each time step of the factor graph. Such motion models may include constant velocity, zero velocity, decaying velocity, or decaying acceleration. The applied motion model may be based on a classification of the type of object being tracked and / or learned using machine learning techniques. For example, a bicyclist may make large turns at high speeds but is not expected to move laterally. Conversely, small animals such as dogs may exhibit unpredictable movement patterns.

[0148] In some embodiments, tracking system 140 may generate an intelligent initial estimate of where a tracked object will appear in subsequently captured images based on the object's predicted 3D trajectory. FIG. 18 shows a diagram illustrating this concept. As shown in FIG. 18, UAV 100 is moving along trajectory 1810 while capturing images of the surrounding physical environment, including human subject 102. As UAV 100 moves along trajectory 1810, multiple images (e.g., frames of video) are captured from one or more onboard image capture devices 114 / 115. FIG. 18 shows a first FOV of the image capture device at a first pose 1840 and a second FOV of the image capture device at a second pose 1842. In this example, first pose 1840 may represent a previous pose of the image capture device at time t(0), and second pose 1842 may represent a current pose of the image capture device at time t(1). At time t(0), an image capture device captures an image of human subject 102 at a first 3D position 1860 in a physical environment. This first position 1860 may be the last known position of human subject 102. Given a first pose 1840 of the image capture device, human subject 102 while in first 3D position 1860 appears at a first image position 1850 in the captured image. Thus, an initial estimate of a second (or current) image position 1852 may be made based on projecting the last known 3D trajectory 1820a of human subject 102 into the future using one or more motion models associated with the object. For example, the predicted trajectory 1820b shown in FIG. 18 represents a projection of this future 3D trajectory 1820a. A second 3D position 1862 (at time t(1)) of the human subject 102 along this predicted trajectory 1820b may then be calculated based on the elapsed time from t(0) to t(1). This second 3D position 1862 may then be projected into the image plane of the image capture device at the second pose 1842, thereby estimating a second image position 1852 that would correspond to the human subject 102.Generating an initial estimate of the position of the tracked object in such newly captured images narrows the search space for tracking, enabling a more robust tracking system, particularly in the case of UAV 100 and / or tracked objects that exhibit rapid changes in position and / or orientation.

[0149] In some embodiments, tracking system 140 may utilize two or more types of image capture devices onboard UAV 100. For example, as described above with respect to FIG. 1 , UAV 100 may include image capture device 114 configured for visual navigation and image capture device 115 for capturing images viewed. Image capture device 114 may be configured for low latency, low resolution, and a high FOV, while image capture device 115 may be configured for high resolution. An array of image capture devices 114 around UAV 100 may provide low-latency information about objects up to 360 degrees around UAV 100 and may be used to calculate depth using stereo vision algorithms. Conversely, other image capture devices 115 may provide more detailed images (e.g., high resolution, color, etc.) in a limited FOV.

[0150] Combining information from both types of image capture devices 114 and 115 can be beneficial for object tracking purposes in several ways. First, high-resolution color information from image capture device 115 can be fused with depth information from image capture device 114 to create a 3D representation of the tracked object. Second, the low latency of image capture device 114 can enable more accurate detection of objects and estimation of object trajectories. Estimates such as these can be further improved and / or revised based on images received from high-latency, high-resolution image capture device 115. Image data from image capture device 114 can be fused with image data from image capture device 115 or can be used purely as an initial estimate.

[0151] By using image capture device 114, tracking system 140 can achieve up to 360-degree tracking of an object around UAV 100. Tracking system 140 can fuse measurements from either image capture device 114 or 115 when estimating the relative position and / or orientation of the tracked object as the positions and orientations of image capture devices 114 and 115 change over time. Tracking system 140 can also fluidly incorporate information from both image capture modalities, orienting image capture device 115 to obtain more accurate tracking of a particular object of interest. Using knowledge of the positions of all objects in a scene, UAV 100 can perform more intelligent autonomous flight.

[0152] As previously mentioned, the high-resolution image capture device 115 may be mounted on an adjustable mechanism, such as a gimbal, that provides one or more degrees of freedom of movement relative to the body of the UAV 100. Such a configuration may be useful for stabilizing image capture or tracking objects of particular interest. An active gimbal mechanism configured to adjust the orientation of the high-resolution image capture device 115 relative to the UAV 100 to track the position of an object in a physical environment may enable visual tracking at longer distances than would be possible using a lower-resolution image capture device 114 alone. An implementation of an active gimbal mechanism may involve estimating the orientation of one or more components of the gimbal mechanism at any given time. Such estimation may be based on a fusion of any hardware sensors coupled to the gimbal mechanism (e.g., accelerometers, rotary encoders, etc.), visual information from the image capture device 114 / 115, or any combination thereof.

[0153] The tracking system 140 may include an object detection system for detecting and tracking various objects. Given one or more classes of objects (e.g., people, buildings, cars, animals, etc.), the object detection system may identify instances of the various classes of objects occurring in captured images of the physical environment. The output by the object detection system may be parameterized in several different ways. In some embodiments, the object detection system processes the received images to output a dense pixel-by-pixel segmentation. Each pixel is associated with either a value corresponding to an object class label (e.g., people, buildings, cars, animals, etc.) and / or a likelihood of belonging to that object class. For example, FIG. 19 shows a visualization 1904 of a dense pixel-by-pixel segmentation of a captured image 1902 in which pixels corresponding to detected objects 1910a-b classified as people are set apart from all other pixels in the image 1902. Another parameterization may include resolving the image position of a detected object to specific image coordinates (e.g., as shown in map 1530 of FIG. 15), for example, based on the geometric center of the object's representation in the received image.

[0154] In some embodiments, the object detection system may utilize a deep convolutional neural network for object detection. For example, the input may be a digital image (e.g., image 1902), and the output may be a tensor with the same spatial dimensions. Each slice of the output tensor may represent a dense segmentation prediction, where the value of each pixel is proportional to the likelihood that the pixel belongs to the class of object corresponding to the slice. For example, visualization 1904 shown in FIG. 19 may represent a particular slice of the aforementioned tensor, where the value of each pixel is proportional to the likelihood that the pixel corresponds to a person. In addition, the deep convolutional neural network may also predict the geometric center location of each of the detected instances, as described in the next section.

[0155] Tracking system 140 may also include an instance segmentation system for distinguishing individual instances of objects detected by the object detection system. In some embodiments, the process of distinguishing individual instances of detected objects may include processing a digital image captured by UAV 100 to identify pixels that belong to one of multiple instances of a class of physical objects present in the physical environment and captured in the digital image. As described above with respect to FIG. 19 , a dense pixel-by-pixel segmentation algorithm may classify a given pixel in the image as corresponding to one or more classes of objects. The output of this segmentation process may allow tracking system 140 to distinguish objects represented in the image from the remainder of the image (i.e., the background). For example, visualization 1904 distinguishes pixels corresponding to people (e.g., contained within region 1912) from pixels that do not correspond to people (e.g., contained within region 1930). However, this segmentation process does not necessarily distinguish individual instances of detected objects. A person viewing visualization 1904 may conclude that the pixels corresponding to people in the detected image actually correspond to two separate people, but without further analysis, tracking system 140 may not be able to make this distinction.

[0156] Effective object tracking may involve distinguishing pixels corresponding to distinct instances of a detected object. This process is known as “instance segmentation.” FIG. 20 shows an example visualization 2004 of instance segmentation output based on a captured image 2002. Similar to the dense pixel-by-pixel segmentation process described with respect to FIG. 19, the output represented by visualization 2004 distinguishes pixels corresponding to detected objects 2010a-c of a particular class of object (in this case, people) (e.g., pixels included in regions 2012a-c) from pixels that do not correspond to such objects (e.g., pixels included in region 2030). Note that the instance segmentation process goes a step further to distinguish pixels corresponding to individual instances of a detected object from one another. For example, the pixels in region 2012a correspond to a detected instance of person 2010a, the pixels in region 2012b correspond to a detected instance of person 2010b, and the pixels in region 2012c correspond to a detected instance of person 2010c.

[0157] Distinguishing between detected object instances may be based on an analysis of pixels corresponding to the detected objects. For example, grouping methods may be applied by tracking system 140 to associate pixels corresponding to a particular class of object with a particular instance of that class. Such association may be made by selecting pixels that are substantially similar to certain other pixels corresponding to the instance, spatially clustered pixels, pixel clusters that fit an appearance-based model for the object class, etc. Again, this process may involve applying a deep convolutional neural network to distinguish between individual instances of the detected object.

[0158] While instance segmentation may associate pixels corresponding to particular instances of an object, such associations may not be consistent over time. Consider again the example described with respect to FIG. 20. As shown in FIG. 20, tracking system 140 has identified three instances of a given class of object (i.e., people) by applying an instance segmentation process to a captured image 2002 of a physical environment. This exemplary captured image 2002 may represent only one frame in a sequence of frames of a captured video. When the second frame is received, tracking system 140 may not recognize the newly identified object instance as corresponding to the same three people 2010a-c captured in image 2002.

[0159] To address this issue, tracking system 140 may include an identity recognition system. The identity recognition system may process received input (e.g., captured images) to learn the appearance of instances of predetermined objects (e.g., particular people). Specifically, the identity recognition system may apply machine learning appearance-based models to digital images captured by one or more image capture devices 114 / 115 associated with UAV 100. Instance segmentations identified based on processing of the captured images may then be compared to such appearance-based models to resolve one or more unique identities of the detected objects.

[0160] Identity recognition can be useful for a variety of different tasks related to object tracking. As mentioned above, recognizing the unique identity of a detected object enables temporal consistency. Furthermore, identity recognition can enable tracking of multiple different objects (discussed in more detail below). Identity recognition may also facilitate object persistence, which allows for the reacquisition of a previously tracked object that has fallen out of view due to limited FOV of the image capture device, object motion, and / or occlusion by another object. Identity recognition can also be applied to perform predetermined identity-specific behaviors or actions, such as recording video when a particular person is in view.

[0161] In some embodiments, the identity recognition process may employ a deep convolutional neural network to learn one or more effective appearance-based models of a given object. In some embodiments, the neural network may be trained to learn a distance metric that returns low distance values ​​for image crops that belong to the same instance of an object (e.g., a person) and high distance values ​​otherwise.

[0162] In some embodiments, the identity recognition process may also include learning the appearance of individual instances of objects, such as people. When tracking people, tracking system 140 may be configured to associate the person's identity either through user-entered data or external data sources, such as images associated with the individual available on social media. Such data may be combined with a detailed facial recognition process based on images received from any of one or more image capture devices 114 / 115 onboard UAV 100. In some embodiments, the identity recognition process may focus on one or more key individuals. For example, tracking system 140 associated with UAV 100 may specifically focus on learning the identity of the designated owner of UAV 100 to retain and / or improve that knowledge between flights for tracking, navigation, and / or other purposes (e.g., access control).

[0163] In some embodiments, tracking system 140 may be configured to focus tracking on a specific object detected in a captured image. In such a single-object tracking approach, an identified object (e.g., a person) is designated as the object to be tracked, while all other objects (e.g., other people, trees, buildings, landscape, etc.) are treated as distractors and ignored. While useful in some contexts, a single-object tracking approach may have some disadvantages. For example, from the perspective of the image capture device, overlapping trajectories of the tracked object and the distracting object may lead to an inadvertent switch of the tracked object, causing tracking system 140 to begin tracking the distractor instead. Similarly, spatially nearby false positives by an object detector may also lead to an inadvertent switch of tracking.

[0164] The multiple-object tracking approach addresses these shortcomings and offers several additional advantages. In some embodiments, a unique track is associated with each object detected in an image captured by one or more image capture devices 114 / 115. In some cases, it is computationally impractical to associate a unique track with every single object captured in an image. For example, a given image may contain hundreds of objects, including small features such as rocks, leaves, and trees. Instead, from a tracking perspective, a unique track may be associated with a given class of objects of interest. For example, tracking system 140 may be configured to associate a unique track with all detected objects belonging to a class that is generally mobile (e.g., people, animals, vehicles, etc.).

[0165] Each unique track may include an estimate of the spatial position and motion of the tracked object (e.g., using the spatiotemporal factor graph described above) and an estimate of the object's appearance (e.g., using an identity recognition function). Instead of pooling all other distractors together (i.e., as might be performed with a single object tracking approach), tracking system 140 may learn to distinguish between multiple individual tracked objects. By doing so, tracking system 140 may be less susceptible to inadvertent identity switching. Similarly, false positives by object detectors may be more robustly rejected, as they tend not to match any unique track.

[0166] Aspects to consider when performing multiple object tracking include the association problem. In other words, given a set of object detections based on captured images (including parameterization by 3D locations and regions within the images corresponding to segmentations), a problem arises as to how to associate each of the set of object detections with a corresponding track. To address this association problem, tracking system 140 may be configured to associate one of the multiple detected objects with one of the multiple estimated object tracks based on the relationship between the detected objects and the estimated object tracks. Specifically, this process may include calculating a "cost" value for one or more pairs of object detections and estimating the object tracks. The calculated cost value may take into account, for example, the spatial distance between the current position (e.g., in 3D space and / or image space) of a given object detection and the current estimate (e.g., in 3D space and / or image space) of a given track, the uncertainty of the current estimate of the given track, the difference between the appearance of the given detected object and the appearance estimate for the given track, and / or any other factors that may tend to suggest an association between the given detected object and the given track. In some embodiments, multiple cost values ​​are calculated based on various different factors and fused into a single scalar value, which is treated as a measure of how well the given detected object matches the given track. The foregoing cost formulation may then be used to determine an optimal association between the detected object and the corresponding track by treating the cost formulation as an instance of a minimum-cost complete bipartite matching problem. This complete bipartite matching problem may be solved using, for example, the Hungarian algorithm.

[0167] In some embodiments, effective object tracking by tracking system 140 may be improved by incorporating information about the object's state. For example, a detected object, such as a person, may be associated with one or more defined states. In this context, states may include activities performed by the object, such as sitting, standing, walking, running, jumping, etc. In some embodiments, one or more sensory inputs (e.g., visual input from image capture devices 114 / 115) may be used to estimate one or more parameters associated with the detected object. The estimated parameters may include activity type, motor ability, trajectory heading, contextual location (e.g., indoors vs. outdoors), interaction with other detected objects (e.g., two people walking together, a dog on a leash held by a person, a trailer pulled by a car, etc.), and any other semantic attributes.

[0168] Generally, object state estimation may be applied to estimate one or more parameters associated with the state of a detected object based on sensory input (e.g., images of the detected object captured by one or more image capture devices 114 / 115 onboard the UAV 100 or sensor data from other sensors onboard the UAV 100). The estimated parameters may then be applied to aid in predicting the motion of the detected object and thereby assisting in tracking the detected object. For example, a future trajectory estimate for a detected human may differ depending on whether the detected person is walking, running, jumping, biking, driving a car, etc. In some embodiments, a deep convolutional neural network may be applied to generate parameter estimates based on multiple data sources (e.g., sensory input) to aid in generating a future trajectory estimate and thereby assisting in tracking.

[0169] As previously mentioned, tracking system 140 may be configured to estimate (i.e., predict) the future trajectory of a detected object based on past trajectory measurements and / or estimates, current sensory input, a motion model, and any other information (e.g., an estimate of the object's state). Predicting the future trajectory of a detected object may be particularly useful for autonomous navigation by UAV 100. Effective autonomous navigation by UAV 100 may depend just as much on predicting future states in the physical environment as it does on predicting current states. Through a motion planning process, the navigation system of UAV 100 may generate control commands configured to steer UAV 100, for example, to avoid collisions, to maintain a distance from a moving tracked object, and / or to meet any other navigation goals.

[0170] Predicting the future trajectory of a detected object is generally a relatively difficult problem to solve. This problem can be simplified for objects moving according to a known, predictable motion model. For example, an object in free fall is expected to continue along its previous trajectory while accelerating at a rate based on a known gravitational constant and other known factors (e.g., wind resistance). In such cases, the problem of generating a prediction of the future trajectory can be simplified to simply propagating the past and present motion according to a known or predictable motion model associated with the object. Of course, the object may deviate from the predicted trajectory generated based on such assumptions for several reasons (e.g., due to a collision with another object). However, the predicted trajectory may still be useful for motion planning and tracking purposes.

[0171] Dynamic objects, such as people and animals, pose a more difficult challenge in predicting future trajectories because the motion of such objects is generally based on the environment and their own free will. To address such challenges, tracking system 140 may be configured to use precise measurements of the object's current position and motion, along with differentiated velocity and / or acceleration, to predict a trajectory for a short period of time (e.g., a few seconds) into the future and continuously update such predictions as new measurements are acquired. Furthermore, tracking system 140 may also use semantic information gleaned from analysis of captured images as clues to help generate a predicted trajectory. For example, tracking system 140 may determine that a detected object is a person riding a bicycle moving along a road. Using this semantic information, tracking system 140 may form a hypothesis that the tracked object is likely to follow a trajectory that roughly matches the path of the road. As another related example, tracking system 140 may determine that a person has begun to turn the handlebars of their bicycle to the left. Using this semantic information, tracking system 140 may form an assumption that the tracked object is likely to turn left before receiving any position measurements that reveal this movement. Another example, particularly relevant to autonomous objects such as people or animals, is assuming that the object tends to avoid collisions with other objects. For example, tracking system 140 may determine that the tracked object is a person on a trajectory that will result in a collision with another object, such as a utility pole. Using this semantic information, tracking system 140 may form an assumption that the tracked object is likely to change its current trajectory at some point before a collision occurs. Those skilled in the art will recognize that these are merely examples of how semantic information may be used as clues to guide predictions of a given object's future trajectory.

[0172] In addition to performing an object detection process in one or more captured images for each time frame, tracking system 140 may also be configured to perform an inter-frame tracking process, for example, to detect a particular set of motion or pixel regions within an image in a subsequent time frame (e.g., video frame). Such a process may involve the application of a mean-shift algorithm, a correlation filter, and / or a deep network. In some embodiments, inter-frame tracking may be applied by a system separate from the object detection system, where results from inter-frame tracking are fused into a spatiotemporal factor graph. Alternatively or additionally, the object detection system may perform inter-frame tracking, for example, if the system has sufficient computing resources (e.g., memory) available. For example, the object detection system may apply inter-frame tracking through iterations with a deep network and / or through passing multiple images at a time. The inter-frame tracking process and the object detection process may also be configured to complement each other, with one restoring the other when a failure occurs.

[0173] As mentioned above, the tracking system 140 may be configured to process images (e.g., raw pixel data) received from one or more image capture devices 114 / 115 onboard the UAV 100. Alternatively or additionally, the tracking system 140 may also be configured to operate by processing disparity images. Such disparity images tend to highlight regions of the image corresponding to an object in the physical environment because pixels corresponding to the object have similar disparity due to the object's 3D position in space. Thus, disparity images, which may be generated by processing two or more images according to separate stereo algorithms, may provide useful clues to guide the tracking system 140 in detecting objects in the physical environment. In many situations, particularly those in the presence of harsh lighting, disparity images may actually provide stronger clues regarding the object's location than images captured from the image capture devices 114 / 115. As mentioned above, disparity images may be calculated using a separate stereo algorithm. Alternatively or additionally, disparity images may be output as part of the same deep network applied by the tracking system 140. The disparity images may be used for object detection separately from the images received from the image capture devices 114 / 115, and the images may be combined into a single network for joint estimation.

[0174] In general, tracking systems 140 (e.g., including object detection systems and / or associated instance segmentation systems) may be primarily concerned with determining which pixels in a given image correspond to each object instance. However, these systems may not consider portions of a given object that are not actually captured in a given image. For example, pixels that would otherwise correspond to an occluded portion of an object (e.g., a person partially occluded by a tree) may not be labeled as corresponding to the object. This can be detrimental to object detection, instance segmentation, and identity recognition because the size and shape of the object may appear distorted in the captured image due to the occlusion. To address this issue, tracking system 140 may be configured to implicitly segment object instances in a captured image, even when the object instance is occluded by other object instances. Furthermore, object tracking system 140 may be configured to determine which pixels associated with an object instance correspond to the occluded portions of the object instance. This process is commonly referred to as "amodal segmentation," and the segmentation process considers the entire physical object, even if portions of the physical object are not necessarily recognized (e.g., in the received image captured by image capture device 114 / 115). Amodal segmentation can be particularly advantageous in performing identity recognition and in tracking system 140 configured for multiple object tracking.

[0175] Loss of visual contact is expected when tracking an object moving through a physical environment. A tracking system 140 based primarily on visual input (e.g., images captured by the image capture device 114 / 115) can lose track of an object when visual contact is lost (e.g., due to being occluded by another object or the object leaving the FOV of the image capture device 114 / 115). In such cases, the tracking system 140 may become uncertain about the object's location and may therefore declare the object lost. Human pilots typically do not have this problem, especially in the case of temporary occlusions, because they understand object permanence. Object permanence assumes that, given physical constraints, an object cannot suddenly disappear or instantly teleport to another location. Based on this assumption, if it is clear that all escape routes are clearly visible, the object is likely to remain within the occluded volume. This situation is most evident when a single occluding object (e.g., a boulder) is located on flat ground with open space all around. If a moving tracked object suddenly disappears at the location of another object (e.g., a boulder) in a captured image, leaving the object in a position occluded by the other object, the tracked object will emerge along one of one or more possible contrasting paths. In some embodiments, tracking system 140 may be configured to implement an algorithm that reduces the uncertainty in the tracked object's position given this concept. In other words, if visual contact with a tracked subject is lost at a particular location, tracking system 140 can reduce the uncertainty in the subject's position to its last observed location and one or more possible escape paths given its last observed trajectory. A possible implementation of this concept may include tracking system 140 generating a stereo-segmented occupancy map and particle filter-based segmentation on the possible escape paths.

[0176] Unmanned Aerial Vehicles - System Examples The UAV 100 according to the present teachings may be implemented as any type of UAV. A UAV, sometimes referred to as a drone, is generally defined as an aircraft capable of being piloted without a human pilot on board. A UAV may be controlled autonomously by an onboard computer processor or via a remote control by a human pilot located in a remote location. Similar to an airplane, a UAV may utilize a fixed aerodynamic surface in conjunction with a propulsion system (e.g., propellers, jets, etc.) to generate lift. Alternatively, similar to a helicopter, a UAV may directly use a propulsion system (e.g., propellers, jets, etc.) to counteract gravity and generate lift. Propulsion-driven lift (as in a helicopter) offers significant advantages in certain implementations, such as as a mobile filming platform, because it allows for controlled movement along all axes.

[0177] Multi-rotor helicopters, particularly quadcopters, have emerged as a popular UAV configuration. A quadcopter (also known as a quadrotor helicopter or quadrotor) is a multi-rotor helicopter powered by four rotors for lift and propulsion. Unlike most helicopters, quadcopters use two sets of fixed-pitch propellers. The first set of rotors rotates clockwise, while the second set rotates counterclockwise. By rotating in opposite directions, the first set of rotors can counteract the angular torque caused by the rotation of the other set, thereby stabilizing flight. Flight control is achieved by varying the angular velocity of each of the four fixed-pitch rotors. By varying the angular velocity of each rotor, the quadcopter can precisely adjust its position (e.g., altitude, and horizontal flight direction) and orientation (including pitch (rotation around the first horizontal axis), roll (rotation around the second horizontal axis), and yaw (rotation around the vertical axis)). For example, if all four rotors are rotating at the same angular velocity (two clockwise, two counterclockwise), the net aerodynamic torque about the vertical yaw axis is zero. If the four rotors rotate with sufficient angular velocity to provide vertical thrust equal to gravity, the quadcopter can maintain a hover. Yaw adjustments may be induced by varying the angular velocity of a subset of the four rotors, thereby unbalancing the cumulative aerodynamic torque of the four rotors. Similarly, pitch and / or roll adjustments are induced by varying the angular velocity of a subset of the four rotors, but in a balanced manner to increase lift on one side of the vehicle and decrease lift on the other side of the vehicle. Altitude adjustments from a hover may be induced by applying balanced changes to all four rotors, which may increase or decrease vertical thrust. Forward / backward and left / right positioning adjustments may be induced by a combination of pitch / roll maneuvers with the application of balanced vertical thrust. For example, to move forward on a horizontal plane, a quadcopter would vary the angular velocity of a subset of the four rotors to perform a pitch forward maneuver.During pitch forward, total vertical thrust can be increased by increasing the angular velocity of all rotors. Because of the pitch forward orientation, the acceleration caused by the vertical thrust maneuver will have a horizontal component, thereby accelerating the vehicle forward in the horizontal plane.

[0178] 21 shows a diagram of an exemplary UAV system 2100 including various functional system components that may be part of a UAV 100, according to some embodiments. The UAV system 2100 includes one or more propulsion systems (e.g., a rotor 2102 and one or more motors 2104), one or more electronic speed controllers 2106, a flight controller 2108, a peripherals interface 2110, one or more processors 2112, a memory controller 2114, a memory 2116 (which may include one or more computer-readable storage media), a power module 2118, a GPS module 2120, a communication interface 2122, audio circuitry 2124, and an accelerometer 2126. 21 may include a sensor 2126 (including subcomponents such as a gyroscope), an IMU 2128, a proximity sensor 2130, an optical sensor controller 2132 and associated optical sensor(s) 2134, a mobile device interface controller 2136 and associated interface device(s) 2138, and any other input controller 2140 and input device(s) 2142 (e.g., a display controller with associated display device(s)). These components may communicate via one or more communication buses or signal lines, as represented by the arrows in FIG. 21.

[0179] UAV system 2100 is only one example of a system that may be part of UAV 100. UAV 100 may include more or fewer components than those shown in system 2100, may combine two or more components as a functional unit, or may have a different configuration or arrangement of components. Some of the various components of system 2100 shown in FIG. 21 may be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application-specific integrated circuits. UAV 100 may also include off-the-shelf UAVs (e.g., currently available remote-controlled quadcopters) combined with modular add-on devices (e.g., those including the components within outline 2190) to perform the innovative functions described in this disclosure.

[0180] The propulsion system (e.g., including components 2102-2104) may include fixed-pitch rotors. The propulsion system may also include variable-pitch rotors (e.g., using a gimbal mechanism), variable-pitch jet engines, or any other propulsion mode that provides a force. The propulsion system may vary the applied thrust to vary the speed of each of the fixed-pitch rotors, for example, by using electronic speed controller 2106.

[0181] The flight controller 2108 may include a combination of hardware and / or software configured to receive input data (e.g., sensor data from the image capture device 2134, a generated trajectory from the autonomous navigation system 120, or any other input) and interpret the data to output control commands to the propulsion systems 2102-2106 and / or aerodynamic surfaces (e.g., fixed-wing control surfaces) of the UAV 100. Alternatively or additionally, the flight controller 2108 may be configured to receive control commands generated by another component or device (e.g., the processor 2112 and / or a separate computing device), interpret those control commands, and generate control signals for the propulsion systems 2102-2106 and / or aerodynamic surfaces (e.g., fixed-wing control surfaces) of the UAV 100. In some embodiments, the aforementioned navigation system 120 of the UAV 100 may include the flight controller 2108 and / or any one or more of the other components of the system 2100. Alternatively, the flight controller 2108 shown in FIG. 21 may exist as a separate component from the navigation system 120, similar to, for example, the flight controller 160 shown in FIG.

[0182] Memory 2116 may include high-speed random-access memory, and may also include non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Access to memory 2116 by other components of system 2100 (e.g., processor 2112 and peripherals interface 2110) may be controlled by memory controller 2114.

[0183] The peripherals interface 2110 may couple input and output peripherals of the system 2100 to the processor 2112 and memory 2116. The one or more processors 2112 execute or carry out various software programs and / or instruction sets stored in the memory 2116 to perform various functions and process data for the UAV 100. In some embodiments, the processor 2112 may include a general central processing unit (CPU), a specialized processing unit such as a graphics processing unit (GPU) particularly suited for parallel processing applications, or any combination thereof. In some embodiments, the peripherals interface 2110, the processor 2112, and the memory controller 2114 may be implemented on a single integrated chip. In other embodiments, they may be implemented on separate chips.

[0184] The network communication interface 2122 may facilitate the transmission and reception of communication signals, often in the form of electromagnetic signals. The transmission and reception of electromagnetic communication signals may be performed via a physical medium, such as a copper wire cable or a fiber optic cable, or wirelessly, for example, via a radio frequency (RF) transceiver. In some embodiments, the network communication interface may include RF circuitry. In such embodiments, the RF circuitry may convert electrical signals to and from electromagnetic signals to communicate with communication networks and other communication devices via electromagnetic signals. The RF circuitry may include well-known circuits for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, memory, etc. The RF circuitry may facilitate the transmission and reception of data over communication networks (including public, private, local, and wide area). For example, communication may occur over a wide area network (WAN), a local area network (LAN), or a network of networks, such as the Internet. Communication may be facilitated over a wired transmission medium (e.g., over Ethernet) or wirelessly, including via wireless cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and other modes of wireless communication.Wireless communication may use any of a number of communication standards, protocols, and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed ​​Downlink Packet Access (HSDPA), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0185] The audio circuitry 2124, including a speaker and microphone 2150, may provide an audio interface between the surrounding environment and the UAV 100. The audio circuitry 2124 may receive audio data from the peripherals interface 2110, convert the audio data into electrical signals, and transmit the electrical signals to the speaker 2150. The speaker 2150 may convert the electrical signals into sound waves that can be heard by humans. The audio circuitry 2124 may also receive electrical signals converted from sound waves by the microwaves 2150. The audio circuitry 2124 may convert the electrical signals into audio data and transmit the audio data to the peripherals interface 2110 for processing. The audio data may be retrieved from and / or transmitted by the peripherals interface 2110 from the memory 2116 and / or the network communication interface 2122.

[0186] I / O subsystem 2160 may couple input / output peripherals of UAV 100, such as optical sensor system 2134, mobile device interface 2138, and other input / control devices 2142, to peripheral interface 2110. I / O subsystem 2160 may include other input controller(s) 2140 for optical sensor controller 2132, mobile device interface controller 2136, and other input or control devices. One or more input controllers 2140 receive / send electrical signals from / to other input or control devices 2142.

[0187] The other input / control devices 2142 may include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, touchscreen displays, slider switches, joysticks, click wheels, etc. A touchscreen display may be used to implement virtual or soft buttons and one or more soft keyboards. A touch-sensitive touchscreen display may provide an input and output interface between the UAV 100 and a user. A display controller may receive and / or send electrical signals to and from the touchscreen. The touchscreen may display visual output to the user. The visual output may include graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output may correspond to user interface objects, which are described in more detail below.

[0188] A touch-sensitive display system may have a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. The touch-sensitive display system and display controller (together with any associated modules and / or instruction sets in memory 2116) may detect contacts (and any movement or disruption of the contacts) on the touchscreen and translate the detected contacts into interactions with user interface objects (such as one or more softkeys or images) displayed on the touchscreen. In an exemplary embodiment, the point of contact between the touchscreen and the user corresponds to the user's finger.

[0189] The touchscreen may use liquid crystal display (LCD) or light emitting polymer display (LPD) technology, although other display technologies may be used in other embodiments. The touchscreen and display controller may detect contact and its movement or interruption using any of a number of now known or later developed touch sensing technologies, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touchscreen.

[0190] The mobile device interface device 2138, in conjunction with the mobile device interface controller 2136, may facilitate the transmission of data between the UAV 100 and other computing devices, such as the mobile device 104. According to some embodiments, the communication interface 2122 may facilitate the transmission of data between the UAV 100 and the mobile device 104 (e.g., when the data is transferred over a Wi-Fi network).

[0191] The UAV system 2100 also includes a power system 2118 for providing power to the various components. The power system 2118 may include a power management system, one or more power sources (e.g., batteries, alternating current (AC), etc.), a recharging system, power loss detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in a computerized device.

[0192] The UAV system 2100 may also include one or more image capture devices 2134. The image capture device 2134 may be the same as the image capture device 114 / 115 of the UAV 100 described with respect to FIG. 1. FIG. 21 shows the image capture device 2134 coupled to the image capture controller 2132 in the I / O subsystem 2160. The image capture device 2134 may include one or more optical sensors. For example, the image capture device 2134 may include a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor of the image capture device 2134 receives light from the environment projected through one or more lenses (the combination of the optical sensor and lens may be referred to as a "camera") and converts the light into data representing an image. In cooperation with an imaging module located in the memory 2116, the image capture device 2134 may capture images (including still images and / or video). In some embodiments, the image capture device 2134 may include a single fixed camera. In other embodiments, image capture device 2140 may include a single adjustable camera (adjustable using a gimbal mechanism with one or more axes of motion). In some embodiments, image capture device 2134 may include a camera with a wide-angle lens to provide a wider FOV. In some embodiments, image capture device 2134 may include an array of multiple cameras to provide a field of view of up to 360 degrees in all directions. In some embodiments, image capture device 2134 may include two or more cameras (of any type described herein) positioned next to each other to provide stereoscopic vision. In some embodiments, image capture device 2134 may include multiple cameras in any combination as described above.In some embodiments, the cameras of image capture device 2134 may be positioned such that at least two cameras are provided with overlapping FOVs at multiple angles around UAV 100, thereby enabling stereoscopic (i.e., 3D) image / video capture and depth recovery (e.g., using computer vision algorithms) at multiple angles around UAV 100. For example, UAV 100 may include four sets of two cameras each positioned to provide stereoscopic views at multiple angles around UAV 100. In some embodiments, UAV 100 may include some cameras dedicated to image capture of objects and other cameras dedicated to image capture for visual navigation (e.g., via visual inertial odometry).

[0193] The UAV system 2100 may also include one or more proximity sensors 2130. Figure 21 shows a proximity sensor 2130 coupled to the peripherals interface 2110. Alternatively, the proximity sensor 2130 may be coupled to an input controller 2140 in the I / O subsystem 2160. The proximity sensor 2130 may generally include remote sensing technology for proximity detection, distance measurement, target identification, etc. For example, the proximity sensor 2130 may include radar, sonar, and LIDAR.

[0194] The UAV system 2100 may also include one or more accelerometers 2126. Figure 21 shows the accelerometer 2126 coupled to the peripherals interface 2110. Alternatively, the accelerometer 2126 may be coupled to an input controller 2140 in the I / O subsystem 2160.

[0195] The UAV system 2100 may include one or more IMUs 2128. The IMUs 2128 use a combination of gyroscopes and accelerometers (e.g., accelerometer 2126) to measure and report the speed, acceleration, orientation, and gravity of the UAV.

[0196] The UAV system 2100 may include a global positioning system (GPS) receiver 2120. Figure 21 shows the GPS receiver 2120 coupled to the peripherals interface 2110. Alternatively, the GPS receiver 2120 may be coupled to an input controller 2140 in the I / O subsystem 2160. The GPS receiver 2120 may receive signals from GPS satellites in orbit around the Earth and calculate (through the use of GPS software) the distance to each of the GPS satellites, thereby pinpointing the current global position of the UAV 100.

[0197] In some embodiments, software components stored in memory 2116 may include an operating system, a communications module (or set of instructions), a flight control module (or set of instructions), a localization module (or set of instructions), a computer vision module (or set of instructions), a graphics module (or set of instructions), and other applications (or sets of instructions). For clarity, one or more modules and / or applications may not be shown in FIG. 21.

[0198] An operating system (e.g., Darwin®, RTXC, Linux®, Unix®, Apple® OS X, Microsoft Windows®, or an embedded operating system such as VxWorks®) includes various software components and / or drivers for controlling and managing common system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware and software components.

[0199] The communications module may facilitate communication with other devices via one or more external ports 2144 and may include various software components for handling data transmission via network communications interface 2122. External port 2144 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) may be adapted to couple directly to other devices or indirectly via a network (e.g., the Internet, a wireless LAN, etc.).

[0200] The graphics module may include various software components for processing, rendering, and displaying graphics data. As used herein, the term "graphics" may include any object that can be displayed to a user, including text, still images, video, animations, icons (e.g., user interface objects including soft keys), and the like. The graphics module, in conjunction with the graphics processing unit (GPU) 2112, may process graphics data captured by the optical sensor 2134 and / or the proximity sensor 2130 in real time or near real time.

[0201] The computer vision module may be a component of the graphics module and provides analysis and recognition of graphics data. For example, while UAV 100 is flying, the computer vision module, in conjunction with the graphics module (if separate), GPU 2112, image capture device 2134, and / or proximity sensor 2130, may recognize and track captured images of objects located on the ground. The computer vision module may further communicate with the localization / navigation module and flight control module to update the position and / or orientation of UAV 100 and make course corrections to thereby fly through the physical environment along a planned trajectory.

[0202] The localization / navigation module may determine the position and / or orientation of the UAV 100 and provide this information for use by various modules and applications (e.g., to the flight control module to generate commands used by the flight controller 2108).

[0203] The image capture device 2134 , in conjunction with the image capture device controller 2132 and the graphics module, may be used to capture images (including still images and video) and store them in the memory 2116 .

[0204] The above-identified modules and applications each correspond to a set of instructions for performing one or more functions described above. These modules (i.e., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise reconfigured in various embodiments. In some embodiments, memory 2116 may store a subset of the above-identified modules and data structures. Additionally, memory 2116 may store additional modules and data structures not described above.

[0205] Exemplary Computer Processing System 22 is a block diagram illustrating an example computer processing system 2200 in which at least some operations described in this disclosure may be implemented. The example computer processing system 2200 may be part of any of the aforementioned devices, including, but not limited to, the UAV 100 and the mobile device 104. The processing system 2200 may include one or more central processing units (“processors”) 2202, a main memory 2206, a non-volatile memory 2210, a network adapter 2212 (e.g., a network interface), a display 2218, input / output devices 2220, control devices 2222 (e.g., a keyboard and pointing device), a drive unit 2224 including a storage medium 2226, and a signal generating device 2230 communicatively coupled to a bus 2216. The bus 2216 is illustrated as an abstraction representing any one or more separate physical buses, point-to-point connections, or both, connected by appropriate bridges, adapters, or controllers. Thus, bus 2216 may include, for example, a system bus, a Peripheral Component Interconnect (PCI) bus or PCI-Express bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), an IIC (I2C) bus, or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus (also known as "Firewire"). A bus may also relay data packets (over full-duplex or half-duplex wires) between components of a network appliance, such as a switching fabric, network port(s), and tool port(s).

[0206] Although main memory 2206, non-volatile memory 2210, and storage medium 2226 (also referred to as "machine-readable medium") are shown as a single medium, the terms "machine-readable medium" and "storage medium" should be interpreted to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions 2228. The terms "machine-readable medium" and "storage medium" should also be interpreted to include any medium capable of storing, encoding, or transmitting a set of instructions for execution by a computing system, causing the computing system to perform any one or more of the methodologies of the embodiments disclosed herein.

[0207] In general, the routines executed to implement embodiments of the present disclosure may be implemented as part of an operating system, or as a specific application, component, program, object, module, or sequence of instructions referred to as a “computer program.” A computer program typically includes one or more instructions (e.g., instructions 2204, 2208, 2228) stored at various times in various memory and storage devices of a computer that, when read and executed by one or more processing units or processors 2202, cause the processing system 2200 to perform operations to implement elements including various aspects of the present disclosure.

[0208] Furthermore, while the embodiments are described in the context of fully functional computers and computer systems, those skilled in the art will understand that various embodiments may be distributed in various forms as program products, and that the present disclosure applies equally regardless of the particular type of machine or computer-readable medium used to actually accomplish the distribution.

[0209] Further examples of machine-readable storage media, machine-readable media, or computer-readable (storage) media include recordable-type media such as volatile and non-volatile memory devices 2210, floppy and other removable disks, hard disk drives, optical disks (e.g., compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs)), and transmission-type media such as digital and analog communications links.

[0210] Network adapter 2212 enables computer processing system 2200 to broker data within network 2214 with entities external to computer processing system 2200, such as a network appliance, via any known and / or conventional communication protocol supported by computer processing system 2200 and the external entity. Network adapter 2212 may include one or more of a network adapter card, a wireless network interface card, a router, an access point, a wireless router, a switch, a multi-layer switch, a protocol converter, a gateway, a bridge, a bridge router, a hub, a digital media receiver, and / or a repeater.

[0211] Network adapter 2212 can include a firewall, which in some embodiments can govern and / or manage permissions to access / proxy data within a computer network and track changes in trust levels between different machines and / or applications. A firewall can be any number of modules comprising any combination of hardware and / or software components capable of enforcing a set of predefined access rights between specific sets of machines and applications, between sets of machines, and / or between sets of applications (e.g., regulating the flow of traffic and resource sharing between these various entities). A firewall may also manage and / or access access control lists that detail permissions, including, for example, the rights of individuals, machines, and / or applications to access and manipulate objects, and the circumstances under which the permissions are valid.

[0212] As noted above, the techniques presented herein may be implemented, for example, by programmable circuitry (e.g., one or more microprocessors) programmed by software and / or firmware, by entirely dedicated hardwired (i.e., non-programmable) circuitry, or any combination thereof. The dedicated circuitry may be in the form of, for example, one or more application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), etc.

[0213] It should be noted that any of the above-described embodiments may be combined with other embodiments unless otherwise stated above or unless any such embodiments are considered to be mutually exclusive in function and / or structure.

[0214] While the invention has been described with reference to certain exemplary embodiments, it will be recognized that the invention is not limited to the described embodiments, but can be practiced with modification and alteration within the spirit and scope of the appended claims. The specification and drawings are, therefore, to be regarded in an illustrative rather than a restrictive sense.

Claims

1. 1. A method of operating an aircraft for autonomous landing, the method comprising: generating a ground map of a portion of a surface within a physical environment based on images captured by an image capture device mounted on the aircraft, the ground map including a plurality of cells, each of the plurality of cells including characteristic data; identifying a landing footprint based on the ground map, the landing footprint including a subset of the plurality of cells having characteristic data that meets prescribed landing criteria; designating a landing area on a surface within the physical environment corresponding to the identified landing footprint; autonomously landing the aircraft in the designated landing area; A method comprising:

2. processing the image captured by the image capture device to determine a plurality of data points indicative of height values ​​at a plurality of points; setting or updating the characteristic data for a cell in the ground map based on one or more of the plurality of data points corresponding to the cell; The method of claim 1 further comprising:

3. Processing the image captured by the image capture device to determine a plurality of data points includes: processing the images to generate a disparity image; mapping pixels in the disparity image to three-dimensional (3D) points in space corresponding to points along the surface in the physical environment, each 3D point having a height value based on the 3D point's respective position in space; Including, The method of claim 2 , wherein one or more of the plurality of data points is based on the height value of the 3D point.

4. The method of claim 3 , wherein the characteristic data for a particular cell comprises statistical height information for a plurality of 3D points along a portion of the surface of the physical environment corresponding to the particular cell.

5. The step of setting or updating characteristic data for the specific cell includes: processing the data points corresponding to one or more of said particular cells to calculate an average height value for said particular cells; and processing the data points corresponding to one or more of said particular cells to calculate a sum of squared differences in height values ​​for said particular cells; The method of claim 4, comprising:

6. monitoring operating conditions of the aircraft while the aircraft is flying through the physical environment; determining, based on the monitoring, that the operating conditions of the aircraft do not meet operating criteria; responsively generating, by the aircraft, a command to land; further comprising The method of claim 1 , wherein the aircraft generates the ground map in response to the command to land.

7. 1. An apparatus including one or more non-transitory computer-readable storage media having stored thereon program instructions that, when executed by a processor, cause the processor to: processing images to determine data points representing height values ​​at a plurality of points along a surface within a physical environment, the images being captured by an image capture device mounted on the aircraft while the aircraft is flying through the physical environment; generating a ground map of a portion of the surface, the ground map comprising a grid of cells, each of the cells comprising height statistics based on the data points associated with the cell; identifying a landing footprint based on the ground map, the landing footprint including a subset of the plurality of cells having characteristic data that meets designated landing criteria; designating a landing area on the surface within the physical environment corresponding to the identified landing footprint; generating control commands to autonomously land the aircraft at the designated landing area; A device that instructs.

8. To generate the control instructions to autonomously land the aircraft at the designated landing area, the program instructions, when executed by the processor, cause the processor to: generating a behavioral goal for landing on the designated landing area; inputting the generated behavioral goals into a motion planner, the motion planner being configured to process multiple behavioral goals to generate a planned trajectory; Instruct the The apparatus of claim 7 , wherein the control commands are generated based on the planned trajectory.

9. To generate the ground map of the portion of the surface in the physical environment, the program instructions, when executed by the processor, cause the processor to: processing the images captured by the image capture device to generate a parallax image; mapping pixels in the generated disparity image to three-dimensional (3D) points in space corresponding to points along the surface in the physical environment, each 3D point having a height value based on the 3D point's respective position in space; Instruct the The apparatus of claim 7 , wherein the data points are based on the height values ​​of the 3D points.

10. To generate the ground map of the portion of the surface in the physical environment, the program instructions, when executed by the processor, cause the processor to: continually adding data points to one or more cells in said ground map as new images are processed; setting or updating the height statistics for the one or more cells when a data point is added; The device of claim 9 , wherein the device indicates:

11. To set or update the height statistics for the one or more cells, the program instructions, when executed by the processor, cause the processor to: The apparatus of claim 10 , wherein when a newer data point is added to a particular one of the one or more cells, the apparatus indicates that a weight of a height value associated with an older data point is decreased.

12. The apparatus of claim 7 , wherein the height statistics include any of the average height, median height, minimum height, or maximum height at points along the portion of the surface corresponding to the particular cell.

13. The apparatus of claim 7 , wherein the prescribed landing criteria is met if the height statistics associated with the subset of the plurality of cells have a variance below a threshold level.

14. The program instructions, when executed by the processor, cause the processor to: processing the images to detect physical objects within the physical environment; extracting semantic information associated with the detected physical objects; adding said semantic information to one or more cells in said ground map corresponding to the location of said physical object in said physical environment; The device of claim 7, wherein the device indicates:

15. 8. The apparatus of claim 7, wherein the identified footprint size and / or shape is based on any of the size or shape of the aircraft, a user preference, or a characteristic of a portion of the physical environment proximate to the aircraft.

16. a propulsion system; an image capture device configured to capture an image of a physical environment; a navigation system communicatively coupled to the image capture device and the propulsion system, the navigation system comprising: generating a ground map of a portion of a surface within the physical environment based on the image, the ground map including a plurality of cells, each of the plurality of cells including characteristic data; identifying a landing footprint based on the ground map, the landing footprint including a subset of the plurality of cells having characteristic data that meets designated landing criteria; designating a landing area on a surface within the physical environment corresponding to the identified landing footprint; commanding the propulsion system to autonomously land the aircraft at the designated landing area; The navigation system is configured as follows: Including, aircraft.

17. The navigation system processing the image captured by the image capture device to determine a plurality of data points indicative of height values ​​at a plurality of points; setting or updating the characteristic data for a cell in the ground map based on one or more of the plurality of data points corresponding to the cell.

17. An aircraft according to claim 16, configured as follows:

18. to process the image captured by the image capture device to determine a plurality of data points, the navigation system processing said images to generate a disparity image; Mapping pixels in the disparity image to three-dimensional (3D) points in space corresponding to points along the surface in the physical environment, each 3D point having a height value based on the 3D point's respective position in space. It is structured as follows: The aircraft of claim 17 , wherein one or more of the plurality of data points is based on the height value of the 3D point.