Multi-uav visual content capture system
By working together with the flight controller and camera controller, and combining SLAM and multi-view triangulation technology, the challenges of multi-UAV image integration and control were solved, achieving efficient and automated 3D scene reconstruction and improving the system's reliability and output quality.
Patent Information
- Application Number
- CN202180006219.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2021-06-25
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-06-25
AI Technical Summary
Existing technologies struggle to effectively control and integrate images captured by multiple drone cameras, especially in complex scene reconstruction and formation changes, and require too many human operators, impacting system reliability and output quality.
By employing the collaborative operation of the flight controller and camera controller, and utilizing SLAM and multi-view triangulation techniques, the system automatically estimates and adjusts the UAV's flight path and camera pose to generate high-precision 3D scene reconstructions that can be completed by a single operator.
It enables efficient and automated image capture and 3D scene reconstruction for multi-UAV systems, improving system reliability and output quality while reducing the need for manual operation.
Smart Images

Figure CN114651280B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 16 / 917,013 (020699-116500US) entitled “Multi-UAV Visual Content Capture System”, filed June 30, 2020, which is incorporated herein by reference as if it were set forth in its entirety for all purposes. Background Technology
[0003] The increasing prevalence of camera-equipped drones has inspired a new approach to cinematography, based on capturing images of previously inaccessible scenes. While professionals traditionally capture high-quality images using precise camera trajectories with well-controlled extrinsic parameters, cameras on drones are constantly in motion, even when the drone is hovering. This is due to the aerodynamic characteristics of drones, making continuous motion fluctuations unavoidable. If only a single drone is involved, camera pose (a 6D combination of position and orientation) can still be estimated using Simultaneous Localization and Mapping (SLAM), a technique well-known in robotics. However, it is often desirable to use multiple cameras simultaneously from different viewpoints to enable complex editing and complete 3D scene reconstruction. Traditional SLAM methods are suitable for single-drone, single-camera situations but not for estimating all poses involved in multi-drone or multi-camera scenarios.
[0004] Other challenges in multi-drone photography include the complexity of integrating image and video streams captured by multiple drones, the need to control the flight paths of all drones to achieve the desired formation (or swarm pattern), and any expected changes in that formation over time. In current professional photography practices involving drones, a human operator must operate two separate controllers on each drone: one to control flight parameters and one to control camera pose. This has numerous negative impacts on the drones (in terms of size, weight, and cost), the reliability of the entire system, and the quality of the output scene reconstruction.
[0005] Therefore, improved systems and methods are needed to integrate images captured by cameras on multiple mobile drones and to accurately control these drones (and potentially the cameras independently of the drones) to efficiently capture and process the visual content necessary to reconstruct scenes of interest. Ideally, visual content integration will be performed automatically at a location outside the drones, and control will also be executed at a location outside the drones, but not necessarily in the same location. Control will involve automatic feedback control mechanisms to achieve high accuracy in drone positioning and to accommodate aerodynamic noise caused by factors such as wind. Sometimes, minimizing the number of human operators required to operate the system can also be beneficial. SUMMARY
[0006] Embodiments generally relate to methods and systems for imaging a scene in 3D based on images captured by multiple drones.
[0007] In one embodiment, a system includes multiple drones, a flight controller, and a camera controller, where the system is fully operational with only one human operator. Each drone moves over the scene along a corresponding flight path, and each drone has a drone camera that captures a corresponding first image of the scene at a corresponding first pose and at a corresponding first time. The flight controller controls the flight path of each drone, in part, by using an estimate of the first pose of each drone camera provided by the camera controller, to create and maintain a desired drone pattern over the scene with desired camera poses. The camera controller receives the corresponding plurality of captured images of the scene from the multiple drones, processes the received images to generate a 3D representation of the scene as a system output, and provides the estimate of the first pose of each drone camera to the flight controller.
[0008] In another embodiment, a method of imaging a scene includes deploying multiple drones, each drone moving over the scene along a corresponding flight path, each drone having a camera that captures a corresponding first image of the scene at a corresponding first pose and at a corresponding first time; controlling the flight path of each drone, in part, by using an estimate of the pose of each camera provided by a camera controller, using a flight controller, to create and maintain a desired drone pattern over the scene with desired camera poses; and receiving the corresponding plurality of captured images of the scene from the multiple drones, and processing the received images to generate a 3D representation of the scene as a system output, using the camera controller, and providing the estimate of the pose of each camera to the flight controller. All operations of the method are performed with only one human operator.
[0009] In another embodiment, an apparatus includes one or more processors; and logic encoded in one or more non-transitory media for execution by the one or more processors. When executed, the logic is operable to image a scene by deploying a plurality of drones, each drone moving over the scene along a corresponding flight path, and each drone having a camera that captures a corresponding first image of the scene at a corresponding first pose and at a corresponding first time; using a flight controller portion to control the flight path of each drone using an estimate of the pose of each camera provided by a camera controller to create and maintain a desired drone pattern over the scene having desired camera poses; and using the camera controller to receive a corresponding plurality of captured images of the scene from the plurality of drones, and to process the received images to generate a 3D representation of the scene as a system output, and to provide the estimate of the pose of each camera to the flight controller. No more than one human operator is needed to fully operate the apparatus to image the scene.
[0010] Further understanding of the nature and advantages of certain embodiments disclosed herein can be realized by reference to the remaining portions of the specification and the attached drawings. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 Imaging a scene is shown in accordance with some embodiments.
[0012] Figure 2 Imaging a scene is shown in accordance with Figure 1 embodiments.
[0013] Figure 3 An example of how a drone agent works in accordance with some embodiments is shown.
[0014] Figure 4 An overview of the transformation computation between a pair of drone cameras is shown in accordance with some embodiments.
[0015] Figure 5 Mathematical details of a least squares method applied to estimate the intersection of a plurality of vectors between two camera positions is given in accordance with some embodiments.
[0016] Figure 6 An initial solution for how scaling is achieved for two cameras is shown in accordance with some embodiments.
[0017] Figure 7 An initial rotation between the coordinates of two cameras can be computed as shown in accordance with some embodiments.
[0018] Figure 8 The last step of the computation to fully align the coordinates (position, rotation, and scaling) of two cameras is summarized in accordance with some embodiments.
[0019] Figure 9 How a drone agent generates a depth map is shown, according to some embodiments.
[0020] Figure 10 Interaction between the flight controller and the camera controller is shown, in some embodiments.
[0021] Figure 11 How flight and pose control of a swarm of drones is implemented, according to some embodiments, is illustrated.
[0022] Figure 12 High-level data flow between components of the system is illustrated, in some embodiments. DETAILED DESCRIPTION
[0023] Figure 1 A system 100 for imaging a scene 120 is shown, according to some embodiments of the invention. Figure 2 Components of the system 100 are shown at different levels of detail. Multiple drones are shown, each drone 105 moving along a corresponding path 110. Figure 1 A flight controller 130 operated by a human 160 is shown, in wireless communication with each drone. The drones are also in wireless communication with a camera controller 140, to which they transmit captured images. Data is sent from the camera controller 140 to the flight controller 130 to facilitate flight control. Other data can optionally be sent from the flight controller 130 to the camera controller 140 to facilitate image processing therein. The system outputs in the form of a 3D reconstruction 150 of the scene 120.
[0024] Figure 2 Some internal organization of the camera controller 140 is shown, including multiple drone agents 142 and a global optimizer 144, as well as data flow between system components, including feedback loops. For simplicity, the scene 120 and the scene reconstruction 150 are represented in a more abstract manner than Figure 1
[0025] Each drone agent 142 is "matched" to one and only one drone, receiving images from within or connected to the drone 105 camera 115. For simplicity, Figure 2 The UAV cameras are shown as being in the same relative position and orientation on each UAV, but this is not necessarily the case. Each UAV agent processes each image received from the corresponding UAV camera (or frame from a video stream) (in some cases, in conjunction with flight command information received from the flight controller 130) along with data characterizing the UAV, the UAV camera, and the captured image, to generate (e.g., using SLAM techniques) an estimate of the UAV camera pose in the local coordinate frame of the UAV, which for the purposes of this disclosure is defined as the combination of a 3D position and a 3D orientation. The aforementioned feature data typically includes the UAV ID, intrinsic camera parameters, and image capture parameters such as image timestamp, size, encoding, and capture rate (fps).
[0026] Each UAV agent then cooperates with at least one other UAV agent to compute a coordinate transformation specific to its own UAV camera so that the estimated camera pose can be represented in a global coordinate frame shared by each UAV. This computation can be performed using a novel robust coordinate alignment algorithm, discussed in more detail below with reference to Figure 3 and Figure 4 .
[0027] Each UAV agent also generates a dense depth map of the scene 120 (the word "dense" in this context refers to the resolution of the depth map being equal to or very close to the resolution of the RGB image from which the depth map is obtained. Generally, modalities like LiDAR or RGB-D generate much lower resolution (less than VGA) than RGB. Vision keypoint based methods generate even sparser points with depth) viewed by the corresponding UAV camera for each pose from which a corresponding image is captured. The depth map is computed and expressed in the global coordinate frame. In some cases, the map is generated by a UAV agent processing a pair of images received from the same UAV camera at slightly different times and poses, the fields of view of which overlap enough to act as a stereo pair of images. As Figure 9 shown, the UAV agent can use well-known techniques to process such a pair of images to generate a corresponding depth map, as described below. In other cases, the UAV can include some type of depth sensor such that depth measurements are sent along with the RGB image pixels, forming an RGBD image (rather than a simple RGB image), which the UAV agent processes to generate the depth map. In yet other cases, both options can be present, with information from the depth sensor used as an aid to refine the depth map previously generated from processing a stereo pair of images. Examples of built-in depth sensors include LiDAR systems, time-of-flight sensors, and sensors provided by stereo cameras.
[0028] Each drone agent sends its own drone camera pose estimate and corresponding depth map (both in global coordinates) as well as data intrinsic to the corresponding drone to the global optimizer 144. Upon receiving all this data and RGB images from each drone agent, the global optimizer 144 collectively processes this data, generating a 3D point cloud representation that can be extended, corrected and refined over time as more images and data are received. A keypoint of an image is said to be "registered" if it already exists in the 3D point cloud and a match is confirmed. The main purpose of the processing is to validate the 3D point cloud image data across multiple images and adjust the estimated pose and depth map of each drone camera accordingly. In this way, a joint optimization of the "structure" of the imaged scene reconstruction and the "motion" or positioning of the drone cameras in space and time can be achieved.
[0029] The global optimization part depends on the use of any of the various state-of-the-art SLAM or Structure from Motion (SfM) optimizers now available, such as the graph-based optimizer BundleFusion that generates a 3D point cloud reconstruction from multiple images captured at different poses.
[0030] In the present invention, such an optimizer is embedded in a process-level iterative optimizer that sends the updated (improved) camera pose estimates and depth maps after each cycle to the flight controller, which can use it to adjust the flight path and pose if necessary. As mentioned above, the subsequent images sent by the drones to the drone agents are then processed by the drone agents, involving each drone agent cooperating with at least one other drone agent to produce further improved depth maps and drone camera pose estimates, which are sent to the global optimizer for use in the next iteration cycle, and so on. Thus, the accuracy of the camera pose estimates and depth maps is improved cycle by cycle, in turn improving the control of the drone flight path and the quality of the 3D point cloud reconstruction. When this reconstruction is deemed to satisfy a predetermined quality threshold, the iterative cycle can stop, and the reconstruction at that point is provided as the final system output. Many applications of this output can be readily envisaged, including, for example, 3D scene reconstruction for photography or view change experiences.
[0031] More details on how the drone agents 142 shown in the system 100 operate in various embodiments will now be discussed.
[0032] The problem of how to control the positioning and motion of multiple drone cameras is solved in the present invention by a combination of SLAM and multi-view triangulation (MVT). Figure Three Figure 3 The advantages and disadvantages of the two techniques employed respectively are shown, as well as the details of one embodiment of the proposed combination, which assumes that the image sequences (or videos) have been temporarily synchronized, comprising first running a SLAM process (e.g.: ORBSLAM2) on each drone to generate a local drone camera pose on each image (local SLAM pose from now on), then loading some (e.g.: 5) RGB image frames and their corresponding local SLAM poses for each drone. This determines a consistent "local" coordinates and "local" scale for that drone camera. Next, a robust MT algorithm is run against the multiple drones - the result is a transformation (rotation, scale and translation) that aligns the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras. Figure 4 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras.
[0033] Figures 5-8 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras.
[0034] Figure 5 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras. Figure 6 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras. Figure 7 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras. Figure 8 Mathematical details of the steps involved in determining the transformation necessary to align the local SLAM poses of the second drone with the SLAM defined coordinates of the first drone are shown schematically. This is then extended to each other drone in the multiple drones. Then the transformation that fits each local SLAM pose is applied. The result is that spatial and temporal consistency is achieved from the images captured by all the multiple drone cameras.
[0035] For simplicity, one of the drone agents can be considered the "master" drone agent, representing the "master" drone camera, whose coordinates can be considered global coordinates, against which all other drone camera images are aligned using the techniques described above.
[0036] Figure 9The internal functional steps that the drone agent can perform after aligning the images of the corresponding camera with the main drone camera using techniques such as the ones described above and roughly estimating the corresponding camera pose in the process are shown in the form of a schematic. According to some embodiments, the pose estimation post-steps represented in the four blocks below the picture generate a depth map based on a pseudo stereo pair of consecutively captured images (e.g. first and second images). The order of operations that are then performed is image modification (compare images taken by the drone camera at slightly different times), depth estimation using any of various known tools (such as PSMnet, SGM, etc.), and finally demodification to assign the computed depth to the pixels of the first image of the image pair.
[0037] Figure 10 The high-level aspects of the interaction between the flight controller and the camera controller in some embodiments of the system 100 are summarized. These interactions take the form of a feedback loop between the flight controller and the camera controller, where the flight controller uses the latest measured visual pose by the camera controller to update its control model, and the camera controller takes into account the commands sent by the flight controller in the SLAM computation of the camera pose.
[0038] Figure 11 More details are provided of the typical procedure that implements control of the flight paths or poses of multiple drones - called feedback swarm control, as it relies on continuous feedback between the two controllers. Key aspects of the resulting system of the invention can be laid out as follows.
[0039] (1) Control stems from a 3D map that is globally optimized, which serves as the latest, most accurate visual reference for camera localization. (2) The flight controller uses the 3D map information to generate commands for each drone to compensate for apparent localization errors in the map. (3) Upon arrival of images from the drones, the drone agent starts computing a "measured" position "around" the expected position, which can avoid unlikely solutions. (4) For a swarm queue of drones, the feedback mechanism always adjusts the pose of each drone by visual means, and queue deformation due to drift is limited.
[0040] Figure 12The information flow is labeled, showing the "external" control feedback loop between the flight controller 130 and the camera controller 140 for the two main components of the integrated system 100, and the "internal" feedback loop between the global optimizer 144 and each drone agent 142. The global optimizer in the camera controller provides fully optimized pose data (rotation + position) to the flight controller as observation channels, and the flight controller considers these observations in its control parameter estimation, so the drone commands sent by the flight controller will respond to the latest pose uncertainty. Continuing the external feedback loop, the flight controller shares its motion commands with the drone agents 142 in the camera controller. These commands are prior information that constrains and accelerates the next camera pose computation optimization within the camera controller. The internal feedback loop between the global optimizer 144 and each drone agent 142 is indicated by the double-headed arrow between these components in the diagram.
[0041] The embodiments described herein provide various benefits in systems and methods for capturing and integrating visual content using multiple camera-equipped drones. In particular, the embodiments enable automatic spatial alignment or coordination of drone trajectories based on camera-captured images and these camera poses, and computation of consistent 3D point clouds, depth maps, and camera poses across all drones, as facilitated by the proposed iterative global optimizer. Successful operation does not depend on the presence of depth sensors (although they can be useful additions), as the SLAM-MT mechanism in the proposed camera controller can simply use visual content from multiple (even much greater than 2) drones continuously capturing images to generate scale-consistent RGB-D image data. These data are extremely useful in modern high-quality 3D scene reconstruction.
[0042] The novel local-to-global coordinate transformation method described above is based on matching multiple pairs of images, thus performing a many-to-one global matching, which provides robustness. In contrast to prior art systems, the image processing performed by the drone agents to compute their corresponding camera poses and depth maps does not depend on the availability of a global 3D map. Each drone agent can generate a dense depth map on its own given a pair of RGB images and their corresponding camera poses, then convert the depth map and camera poses to global coordinates, and then pass the results to the global optimizer. Thus, the operation of the global optimizer of the present invention is simpler, as it processes camera poses and depth maps in a unified coordinate system.
[0043] It is noted that this involves two data transmission loops. The external loop operates between the flight controller and the camera controller to provide global positioning accuracy, while the internal loop (composed of multiple sub-loops) operates between the drone agents and the global optimizer within the camera controller to provide structure and motion accuracy.
[0044] While the specification has been described in terms of specific embodiments, these specific embodiments are illustrative only and not restrictive. Applications include professional 3D scene capture, digital content asset generation, real-time review tools for studio capture, and drone swarm formation and control. In addition, since the present invention can handle multiple drones performing complex 3D motion trajectories, it can also be applied to cases that handle low-dimensional trajectories of scans, for example, of a team of robots.
[0045] Routines of particular embodiments can be implemented using any suitable programming technology including C, C++, Java, assembly language, etc. Different programming technologies can be used such as procedural or object oriented. These routines can be executed on a single processing device or multiple processors. Although steps, operations, or calculations can be presented in a particular order, that order can be changed in different particular embodiments. In some particular embodiments, multiple steps shown as sequential in this specification can be performed at the same time.
[0046] Particular embodiments can be implemented in a computer readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or device. Particular embodiments can be implemented in the form of control logic in software or hardware or a combination of both. The control logic, when executed by one or more processors, is operable to perform the methods described in the particular embodiments.
[0047] Particular embodiments can be implemented through the use of programming or general purpose digital computers, through the use of application specific integrated circuits, programmable logic devices, field programmable gate arrays, using optical, chemical, biological, quantum or nanoengineered systems, components and mechanisms. In general, the functions of particular embodiments can be implemented using any hardware, software, systems, or combinations thereof. Distributed, networked systems, components and / or circuitry can be used. Communication or transfer between data can be wired, wireless, or by any other means.
[0048] It should also be understood that one or more elements depicted in the figures / illustrations can also be implemented in a more separated or integrated manner, or even removed or presented in a different manner as is useful in a certain application. Implementations can be stored in a machine-readable medium allowing a computer to execute any of the methods described above.
[0049] A "processor" includes any suitable hardware and / or software system, mechanism or component that processes data, signals or other information. A processor can include a system with a general- purpose central processing unit, multiple processing units, dedicated circuitry for implementing specific functions, or other systems. Processing need not be limited to a geographic location, or have temporal limitations. For example, a processor can perform its functions in "real time," "offline," in a "batch mode," etc. Portions of processing can be performed at different times and at different locations, by different (or the same) processing systems. Examples of processing systems can include servers, clients, end user devices, routers, switches, networked storage devices, and the like. A computer can be any processor in communication with a memory. A memory can be any suitable processor-readable storage medium, such as random access memory (RAM), read only memory (ROM), a magnetic or optical disc, or other non-transitory medium suitable for storing instructions for execution by a processor.
[0050] As used in the description of the application herein and the appended claims, the articles "a", "an" and "the" include plural references unless the context clearly dictates otherwise. Furthermore, "said" is used as a pronoun wherein "said" takes the place of one or more antecedents and is used in the manner that is most close in meaning to the antecedent(s) it replaces.
[0051] Accordingly, although specific embodiments have been described herein, many modifications, variations and alternatives are possible therein, and it is intended that the foregoing disclosure be interpreted as merely illustrative of the breadth of the contemplated aspects and in a manner that is most abiding with the principles and concepts of the various embodiments. Accordingly, numerous modifications can be made to provide specific conditions or materials to adapt the basic range and concepts to particular situations.
Claims
1. A system for imaging a scene, the system comprising: Multiple drones, each moving on the scene along a corresponding flight path, and each drone having a drone camera, which captures a corresponding first image of the scene at a corresponding first pose and a corresponding first time. The flight controller, in part, controls the flight path of each drone by using an estimate of the first pose of each drone camera provided by the camera controller, thereby creating and maintaining the desired pattern of the drone with the desired camera pose on the scene. and A camera controller receives multiple captured images of a scene from the plurality of drones, processes the received multiple captured images to generate a 3D representation of the scene as the system output, and provides an estimate of the first pose of each drone camera to the flight controller. The camera controller includes: Multiple drone agents, each drone agent communicatively coupled to one and only one corresponding drone to receive the corresponding captured first image; and A global optimizer that communicates with each drone agent and flight controller; In this process, the drone agent and global optimizer in the camera controller work together to iteratively improve the estimation of the first pose and the depth map representing the scene imaged by the corresponding drone camera for each drone, and use the estimation and depth map from all drones to create a 3D representation of the scene. The flight controller receives the estimate of the first pose of each UAV camera from the camera controller, and adjusts the corresponding flight path and UAV camera pose accordingly. Wherein, the local coordinates of the drone camera corresponding to one of the plurality of drone agents are used as global coordinates, and the local coordinates of the other drone cameras of the other drones are aligned with the global coordinates; and Each drone agent: collaborates with another drone agent to process the first image captured by the corresponding drone using data and image capture parameters characterizing the corresponding drone to generate an estimate of the first pose of the corresponding drone; and collaborates with a global optimizer to iteratively improve the first pose estimate of the drone camera of the drone coupled with the drone agent, and iteratively improve the corresponding depth map.
2. The system according to claim 1, wherein, The depth map corresponding to each drone is generated by the corresponding drone agent based on the first and second images of the processing scene. The second image is captured by the corresponding drone camera in the corresponding second pose and the corresponding second time, and is received by the corresponding drone agent.
3. The system according to claim 1, wherein, The depth map corresponding to each drone is generated by the corresponding drone agent based on processing the first image and the depth data generated by the depth sensor in the corresponding drone.
4. The system according to claim 1, wherein, The estimation of the first pose for each UAV camera involves transforming the pose-related data represented in the local coordinate system of each UAV into a global coordinate system shared by multiple UAVs. The transformation includes a combination of simultaneous localization and mapping (SLAM) and multi-view triangulation (MT).
5. The system as claimed in claim 1, wherein, The global optimizer: A 3D representation of the scene is generated and iteratively improved based on inputs from each of the plurality of drone agents, the inputs including data characterizing the corresponding drone, and the corresponding processed first image, first pose estimate and depth map; and The pose estimation of the drone cameras of the plurality of drones is provided to the flight controller.
6. The system according to claim 5, wherein, The iterative improvements performed by the global optimizer include a loop process in which the UAV camera pose estimation and depth map are continuously and iteratively improved until the 3D representation of the scene meets a predetermined quality threshold.
7. A method for imaging a scene, the method comprising: Multiple drones are deployed, each drone moves on the scene along a corresponding flight path, and each drone is equipped with a camera, which captures a corresponding first image of the scene at a corresponding first pose and a corresponding first time. Using a flight controller, the flight path of each drone is controlled in part by using an estimate of the first pose of each camera provided by the camera controller, thereby creating and maintaining the desired pattern of the drone with the desired camera pose on the scene. and The camera controller receives multiple captured images of the scene from the multiple drones and processes the received multiple captured images to generate a 3D representation of the scene as the system output, and provides the first pose estimate of each camera to the flight controller. The camera controller includes: Multiple drone agents, each drone agent communicatively coupled to one and only one corresponding drone to receive the corresponding captured first image; and A global optimizer that communicates with each drone agent and flight controller; and In this process, the drone agent and global optimizer in the camera controller work together to iteratively improve the estimation of the first pose and the depth map representing the scene imaged by the corresponding drone camera for each drone, and use the estimation and depth map from all drones to create a 3D representation of the scene. The flight controller receives an estimate of the improved first pose of each UAV camera from the camera controller, and adjusts the corresponding flight path and UAV camera pose accordingly. Wherein, the local coordinates of the drone camera corresponding to one of the plurality of drone agents are used as global coordinates, and the local coordinates of the other drone cameras of the other drones are aligned with the global coordinates; and Each drone agent: collaborates with another drone agent to process the first image captured by the corresponding drone using data and image capture parameters characterizing the corresponding drone to generate an estimate of the first pose of the corresponding drone; and collaborates with a global optimizer to iteratively improve the first pose estimate of the drone camera of the drone coupled with the drone agent, and iteratively improve the corresponding depth map.
8. The method of claim 7, wherein, The depth map corresponding to each drone is generated by the corresponding drone agent based on the first and second images of the processing scene. The second image is captured by the corresponding drone camera in the corresponding second pose and the corresponding second time, and is received by the corresponding drone agent.
9. The method of claim 7, wherein, The depth map corresponding to each drone is generated by the corresponding drone agent based on processing the first image and the depth data generated by the depth sensor in the corresponding drone.
10. The method of claim 7, wherein generating an estimate of the first pose of each UAV camera comprises transforming pose-related data represented in the local coordinate system of each UAV to a global coordinate system shared by multiple UAVs, said transformation comprising a combination of simultaneous localization and mapping (SLAM) and multi-view triangulation (MT).
11. The method according to claim 8, wherein, The global optimizer: A 3D representation of the scene is generated and iteratively improved based on inputs from each of the plurality of drone agents, the inputs including data characterizing the corresponding drone, and the corresponding processed first image, first pose estimate and depth map; and The first pose estimate of the plurality of UAV cameras is provided to the flight controller.
12. The method according to claim 11, wherein, The iterative improvements performed by the global optimizer include a loop process in which the UAV camera pose estimation and depth map are continuously and iteratively improved until the 3D representation of the scene meets a predetermined quality threshold.
13. The method of claim 7, further comprising: Prior to cooperation, temporal and spatial relationships were established among the multiple drones through the following operations: Electrical or visual signals from each of the plurality of drone cameras are compared to achieve time synchronization; Run a SLAM process for each drone to establish a local coordinate system for each drone; and Run a multi-view triangulation process to define a global coordinate frame shared by the multiple UAVs.
14. An apparatus for imaging a scene, comprising: One or more processors; and Logic encoded in one or more non-transient media, the logic being executed by the one or more processors, and operable at execution to image a scene in such a way as: Multiple drones are deployed, each drone moves on the scene along a corresponding flight path, and each drone is equipped with a camera, which captures a corresponding first image of the scene at a corresponding first pose and a corresponding first time. Using a flight controller, the flight path of each drone is controlled in part by using an estimate of the first pose of each camera provided by the camera controller, thereby creating and maintaining the desired pattern of the drone with the desired camera pose on the scene. and The camera controller receives multiple captured images of the scene from the multiple drones and processes the received multiple captured images to generate a 3D representation of the scene as the system output, and provides the first pose estimate of each camera to the flight controller. The camera controller includes: Multiple drone agents, each drone agent communicatively coupled to one and only one corresponding drone to receive the corresponding captured first image; and A global optimizer that communicates with each drone agent and flight controller; and In this process, the drone agent and global optimizer in the camera controller work together to iteratively improve the estimation of the first pose and the depth map representing the scene imaged by the corresponding drone camera for each drone, and use the estimation and depth map from all drones to create a 3D representation of the scene. The flight controller receives an estimate of the improved first pose of each UAV camera from the camera controller, and adjusts the corresponding flight path and UAV camera pose accordingly. Wherein, the local coordinates of the drone camera corresponding to one of the plurality of drone agents are used as global coordinates, and the local coordinates of the other drone cameras of the other drones are aligned with the global coordinates; and Each drone agent: collaborates with another drone agent to process the first image captured by the corresponding drone using data and image capture parameters characterizing the corresponding drone to generate an estimate of the first pose of the corresponding drone; and collaborates with a global optimizer to iteratively improve the first pose estimate of the drone camera of the drone coupled with the drone agent, and iteratively improve the corresponding depth map.
15. The apparatus according to claim 14, wherein, The depth map corresponding to each drone is generated by the corresponding drone agent based on the following operations: Processing the first and second images of the scene, the second image is captured by the corresponding drone camera in the corresponding second pose and at the corresponding second time, and received by the corresponding drone agent; or Process the first image and the depth data generated by the depth sensor in the corresponding UAV.