Removal of Irrelevant Content from Images of Scenes Captured by a Multi-Drone Group

By predicting and masking irrelevant drone content in multi-drone image capture, the method improves image reconstruction quality and efficiency, addressing the disruption caused by other drones or objects in the field of view.

JP7711220B2Active Publication Date: 2025-07-22SONY GROUP CORP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023571173
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-10
Filing Date
2022-05-27
Publication Date
2025-07-22
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The presence of irrelevant content, such as other drones or objects, in images captured by a multi-drone system disrupts accurate image alignment and reconstruction, leading to low-quality and incomplete scene reconstruction, and current methods are laborious and computationally intensive.

Method used

A method for each drone to predict the 3D position of other drones, define a region of interest (ROI), generate a drone mask, and apply it to remove irrelevant content from captured images, utilizing real-time 3D position data and learning-based detection for improved accuracy.

Benefits of technology

Enables efficient and automated removal of irrelevant content, reducing computational requirements and enhancing the quality of image reconstruction for applications like AR/VR and video capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711220000001
    Figure 0007711220000001
  • Figure 0007711220000002
    Figure 0007711220000002
  • Figure 0007711220000003
    Figure 0007711220000003
Patent Text Reader

Abstract

A method for removing extraneous content in a first plurality of images of a scene in which a second drone is present, captured by a first drone at corresponding plurality of poses and corresponding first plurality of time instants, includes the following steps for each of the first plurality of captured images: The first drone predicts a 3D position of the second drone at the time of capture of the image; The first drone defines a region of interest (ROI) in an image plane corresponding to the captured image, the region of interest including a projection of the predicted 3D position of the second drone at the time of capture of the image; A drone mask of the second drone is generated and applied to the defined ROI to generate an output image that is free of extraneous content attributable to the second drone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Patent Application No. 17 / 344,659 (Client Reference No.: SYP339160US01) entitled "Extraneous Content Removal From Images of a Scene Captured by a Multi - Drone Swarm", filed on June 10, 2021, and this document is incorporated herein by reference as if the entire text thereof were set forth herein for all purposes.

[0002] This application is related to U.S. Patent Application Serial No. 16 / 917,013 (020699 - 116500US) entitled "System of Multi - Drone Visual Content Capturing" filed on June 30, 2020, and U.S. Patent Application Serial No. 16 / 917,671 (020699 - 117000US) entitled "Method of Multi - Drone Camera Control" filed on June 30, 2020, and these documents are incorporated herein by reference as if the entire text thereof were set forth herein for all purposes.

Background Art

[0003] When using a group of drones instead of just one drone to capture multiple image streams over the same period, the reconstruction of a large-scale 3D scene from the captured images can be much more efficient. However, due to the presence of multiple drones each following its own trajectory through the scene, there is a possibility that one or more images captured by any one drone may accidentally contain visual content corresponding to one or more other drones that were in the field of view at the moment of image capture. This irrelevant content, which provides no useful information about the scene of interest being captured, may show part or all of one or more other drones. Even if all drone trajectories and drone poses are very strictly controlled, some of the images within the captured image stream will almost certainly contain such content.

[0004] The presence of such content causes problems when subsequently processing the image stream to reconstruct the scene, whether it is 2D or 3D. One problem is that visual features indicating moving drones, extracted from the image pixels, disrupt the assumption of fixed points that is the basis of accurate image alignment schemes. Another problem is that during the 3D reprojection of rays captured by different drones after alignment, the drones and the scene texture become mixed, resulting in inconsistent color data for the rays projected from different drones during image reconstruction, leading to a low-quality and incomplete reconstruction of the actual scene. Note that although the present disclosure mainly focuses on the specific case of drones for simplicity, similar problems occur when there are irrelevant objects other than drones.

[0005] Typically, current approaches to address these issues rely on manually editing each frame of the captured image. This approach is clearly a time-consuming and laborious process, and it is also costly. Additionally, more automated methods that include the detection and removal of visual content are based solely on visual information and are limited to the boundaries of individual images. When only a portion of the drone that is visually "interfering" is within the field of view, and thus when "cropping" this, typical detection algorithms may not be useful.

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] Therefore, there is a need for an improved method for removing irrelevant content, particularly content related to "other drones," from images captured by a given drone. These methods should operate automatically and preferably do not impose high requirements on the computer memory or processing power, either within the drone itself, at a ground control station involved in trajectory control, or in a post-processing stage of image processing.

[0007] Similar methods are useful in situations where irrelevant objects present in the captured image are not other drones, but are nevertheless desirable to be "erased" from the image before performing alignment and scene reconstruction processes. An example of such an object is a drone pilot or other observer who, for flight safety reasons, is present or necessary but has been captured within the scene's image. In the context of the present invention, if such objects are equipped and measured as necessary for the drone, they can also be automatically removed from the captured visual image (based on their own ROI and mask). Similarly, there may be cases where a person holding a camera rather than a second capturing drone is an irrelevant object. The present invention can also be applied to situations where the flexibility of a "multi-crane shot" is required to conveniently capture multiple viewpoints, but there is a possibility that at least a portion of one or more cranes may be visible within a shot captured by another crane.

Means for Solving the Problem

[0008] Embodiments generally relate to a system and method for removing irrelevant content in an image of a scene in which there are other drones or other unrelated objects captured by one drone. In one embodiment, a method for removing irrelevant content in a first plurality of images of a scene in which a second drone exists, the first plurality of images captured by a first drone at a first plurality of corresponding times in corresponding postures, includes, for each of the first plurality of captured images, the step of the first drone predicting the 3D position of the second drone at the time of image capture; the step of the first drone defining a region of interest (ROI) in the image plane corresponding to the captured image, the ROI including the projection of the predicted 3D position of the second drone at the time of image capture; the step of generating a drone mask for the second drone; and the step of applying the generated drone mask to the defined ROI to generate an output image that does not include irrelevant content caused by the second drone.

[0009] In another embodiment, a method for removing irrelevant content in a first plurality of images of a scene in which a plurality of other drones exist, the first plurality of images captured by a first drone (scene drone) at a first plurality of corresponding times in corresponding postures, includes, for each of the first plurality of captured images, the step of the first drone predicting the 3D position of each of the other drones at the time of image capture; the step of the first drone defining a region of interest (ROI) for each of the other drones in the image plane corresponding to the captured image, the ROI including the projection of the predicted 3D position of each of the other drones; the step of generating a drone mask for each of the other drones; and the step of applying these drone masks to the corresponding defined ROIs to generate an output image of the scene that does not include irrelevant content caused by these other drones.

[0010] In yet another embodiment, an apparatus for removing irrelevant content in a first plurality of images of a scene in which a second drone exists, captured by a first drone at a first plurality of corresponding time points in corresponding poses, includes one or more processors and logic encoded on one or more non-transitory media for execution by the one or more processors. The logic, when executed, for each of the first plurality of captured images, causes the first drone to predict the 3D position of the second drone at the time of image capture, define a region of interest (ROI) in the image plane corresponding to the captured image that includes the projection of the predicted 3D position of the second drone at the time of image capture, generate a drone mask for the second drone, and apply the generated drone mask to the defined ROI to generate an output image that does not include irrelevant content caused by the second drone. The predicting step and the defining step are executed by one or more of the one or more processors located within the first drone, and the generating step and the applying step of the drone mask are at least partially executed by one or more processors located remotely from the first drone.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

BEST MODE FOR CARRYING OUT THE INVENTION

[0012] By referring to the remainder of this specification and the accompanying drawings, the characteristics and advantages of the specific embodiments disclosed herein can be further understood.

[0013] FIG. 1 is a diagram showing how the problems addressed by the present invention can occur in the manner described below. This figure is a schematic depiction showing a three-dimensional scene of interest represented by block 100. The first drone 105 is capturing images of 100 while following the trajectory indicated by arrow A. On the other hand, the second image-capturing drone 110 approaches the target scene 100 while following the trajectory indicated by arrow B and views 100 from a different set of viewpoints. At the instant shown, drone 110 is clearly about to enter the field of view of drone 105 and thus may appear in one or more images of scene 100 captured by 105.

[0014] FIG. 2 shows an exemplary two - dimensional image 210 captured by a drone such as drone 105 of FIG. 1. Image 210 includes pixels having content representing the scene of interest, shown here as a mottled grey pattern, and also includes a region 220 of pixels having content representing another drone such as drone 110 of FIG. 1. As will be described in more detail below, embodiments of the present invention can be used to operate on image 210 to define a region of interest (ROI) 230 around suspect pixels and form a mask for selectively removing this content. In the case shown in the upper right, as a simple approach, a mask 240A that exactly matches ROI 230 is applied to blank out all pixels corresponding to 220, and also blank out the surrounding pixels that reach the boundary of ROI 230. As a result, a masked image 250A is obtained. In the case shown in the lower right, a highly precise mask 240B using a learning - based method described later is used. All pixels corresponding to 220 are blanked out as before, but this time the number of surrounding pixels to be blanked out is reduced, and more of the scene itself can be seen in the masked image 250B.

[0015] FIG. 3 is a flowchart of a method 300 for removing irrelevant content in each image captured at a specific drone pose at a specific point in time in a stream of drone - captured images according to some embodiments of the present invention. Note that in most of this disclosure, for simplicity, it is assumed that the pose (position and orientation) of the drone completely determines the pose of the camera on that drone according to a known calibration factor. Of course, in the most general case, these two poses (especially the orientation) do not need to be exactly the same. In fact, the pose of the drone camera is the important pose for the matters of interest in the present invention, but for convenience, this is simply referred to as the "drone pose" in this specification.

[0016] In step 310, the first drone predicts the 3D position of the second drone at the time of capturing the i-th image captured by the camera on the first drone. In step 320, a projection of this predicted position is performed onto the image plane corresponding to the captured image to define a region of interest (ROI) within this image plane. Usually, these two steps are performed by the first drone while it is flying because they require few computational resources and can be easily provided by one or more drone-mounted processors, enabling rapid processing for each frame.

[0017] In step 330, a mask of the second drone is generated. In step 340, the generated drone mask is applied to the defined ROI to generate an output image that does not contain irrelevant content caused by the second drone. In step 350, the index "i" is incremented and the method returns to step 310. For applications that require very high visual quality, such as video capture or aerial inspection of a construction site, it is best to perform mask generation and application offline, possibly at a studio or "post-processing" location that can provide high computational power. However, for less demanding applications in augmented reality or virtual reality (AR / VR), mask generation on the drone may be sufficient in some cases. In such cases, generally the mask shape is encoded and sent as an image to the post-processing location.

[0018] For simplicity, the initialization of "i" to 1 and the termination of the method when there are no more images to process are omitted from the flowchart of FIG. 1.

[0019] Next, the embodiment of step 310 will be considered in more detail. An important feature of many embodiments of the present invention is to use sensors other than cameras to provide real-time 3D position information about the second drone to the first drone. This is a significant advantage over prior art methods that relied on the second drone accurately following a pre-planned trajectory relative to or in an absolute sense with respect to the first drone. In embodiments of the present invention, regardless of whether the initial position of the second drone is established in advance or by real-time measurement, the subsequent estimated position of the second drone is continuously updated according to the real-time measurement data received thereafter.

[0020] Multiple 3D global positioning methods that can be used to provide real-time positioning data have been established, including GNSS, RTK-GNSS, and RTK-GNSS-IMU. Data from other sensors such as LIDAR and RADAR can provide additional accuracy. In some cases, optionally, the 5G communication protocol can be used to directly transmit data from the second drone to the first drone, and in other cases, the data can be transmitted indirectly via a ground control station or a "master" drone. Other data generation and transmission options will be apparent to those skilled in the art. Of course, the positioning data must be timestamped so that the position of the second drone at the time of image capture by the first drone can be estimated.

[0021] Position estimation based on timestamped data for a second drone received sequentially can involve the use of models such as simple linear or quadratic interpolation / extrapolation, spline trajectory fitting, filter-based predictors such as Kalman filters and their variants, or sequential regression models such as RNNs and their variants trained on the actual trajectories of drones similar to the second drone. Prior knowledge of the planned drone trajectory can serve as an additional constraint. The end result for each case is a 3D position estimate in the global coordinate system.

[0022] Next, the predicted 3D position is projected onto the 2D image plane, and the surrounding area is set to define the ROI in step 320, which will be described in more detail.

[0023] The first drone needs to know its own pose (3D position and 3D orientation) at the time of image capture for each captured image. This pose is provided by real-time measurement, preferably by an RTK-GNSS-IMU system. Usually, a measurement frequency suitable for such measurements is 10 Hz or higher. The first drone can calculate the orientation of the captured image in the global 3D coordinate system based on this data. Figure 4 shows the relationships between various coordinate systems of interest in an exemplary case. The axes on the left side of the figure are those of the global coordinate system, which in this case is the "North-East-Down" or "NED" coordinate system. On the right side of the figure, the directions of the axes of the drone camera are shown as the same as those of the central drone itself. As described above, it is not necessary to necessarily apply this equivalence, but it is assumed in this way here for simplicity. At the bottom of the figure, the 2D axes characterizing the captured image are shown.

[0024] Regarding the intrinsic parameters (such as focal length, sensor size, etc.) of the camera on the first drone, it is assumed to be known because they determine the relationship between the position in the real 3D world and the position in the 2D image captured by this camera. Assume a projective camera model.

[0025] The first drone also requires information regarding the physical dimensions of the second drone, which is usually determined offline before the deployment of the drone. This information must at least include the maximum span of the second drone when viewed from the orientation in which the second drone appears largest, which is generally when viewed directly from above or below during flight. Figure 5 shows two examples where the approximate estimated values of the perceivable maximum diameters of drones 510 and 520 (shown exaggerated for clarity in the figure) are taken as parameters D1 and D2 respectively.

[0026] Figure 6 is an overview of how the first drone can then use data regarding its own attitude, its own intrinsic characteristics, the estimated global 3D position (P os ) of the second drone, and the known dimensions of the second drone to first calculate four maximum span points that characterize the maximum span of the second drone as seen by the first drone as P1, P2, P3, and P4 in the global coordinate system as shown in the upper left of the figure first, and then project them onto the 2D image plane as p1, p2, p3, and p4 (around the corresponding projection from P os to p0) shown in the central part of the figure. The diagonal dashed arrows indicate the central ray projection along the camera optical axis. Note that most drones are far from being spherically symmetric and the projection position p0 does not usually exist at their midpoint, so usually the separation of span points p3 and p4 is not equal to the separation of span points p 1及 and p2.

[0027] After the 2D image projections of the five points are established, a rectangular contour (not shown to scale with respect to the 2D image on its left in the figure) surrounding these points as shown by the thick dashed boundary line in the right part of the figure can be defined.

[0028] Typically, the actual size of the rectangle defining the ROI 600 is enlarged from the minimum size including the points, considering timing, positioning, and other uncertainties. Note that the ROI can be completely included within the boundaries of the captured image as suggested by the thick outline 620, but in some cases (not shown), although it is natural to be in the same plane, it can also extend beyond this boundary. This is because the second drone is very close to the image boundary and, in some cases, crosses the boundary. In fact, the possibility that the ROI extends beyond the image boundary reduces the possibility of "losing sight of" the second drone when the second drone is close to the boundary to such an extent that it is difficult to recognize the cut-out part visible in the image, thus providing an additional advantage to the present invention compared to prior art techniques.

[0029] Returning now to step 330 of method 300, a mask representing the second drone must be generated in some meaningful way. The mask has a width and height that match the width and height of the captured image. In the simplest case, one value, such as zero, is set or labeled for any mask pixel at a position in the mask that corresponds to the position of an image pixel within the defined ROI, and another value, such as 1 (unity), is labeled for all other mask pixels. FIG. 7 shows how such a mask 700 can be applied to the corresponding captured image 710 with the ROI 720 defined to "blank out" or remove a set of image pixels that may contain irrelevant content caused by the second drone. This method can be acceptable depending on the available time and computational power resources for the application. However, applying such a simple rectangular mask may also result in discarding a significant amount of image content in nearby pixels representing the underlying scene itself.

[0030] In some embodiments, a detection system or detector that identifies a subset of pixels within the ROI as having a high probability of containing content due to the presence of the second drone (compared to other pixels within the ROI) can be used to generate a mask with more complex features.

[0031] One such detection method implements a rule-based method for performing this detection using, for example, some combination of the size, shape, color, dynamics of the second drone, and / or additional information received from other sensors. Another method uses a learning-based drone detector to classify the pixels within the ROI according to the likelihood of belonging to the internal drone. This likelihood can be determined based on the image itself and the position data of the second drone, specifically the projection center position P0 within the ROI of that drone. See ROI600 in FIG. 6.

[0032] One such learning-based detection method relies on training the detector to recognize the shape of the drone using these combinations after preparing heatmap inputs and visual capture image inputs. Preparation of the heatmap depends on collecting a set of drone capture images, defining the ROI as described above, and then manually annotating these with the ground truth drone center positions. This enables the calculation of error vectors and standard deviations of the positioning errors along the x and y axes, allowing the generation of heatmap images such as 810 in the upper left of FIG. 8. Visual images used in heatmap preparation, such as 820 in the center of FIG. 8, can be augmented in a manner well-known in the art to include variations such as cropped or partial drone images like 830 in FIG. 8 to improve the training set. Subsequently, the resulting "4-channel" input image, in which each visual image is represented by three channels (R - G - B) and the heatmap data constitutes a fourth channel, can be used to train the detector.

[0033] Subsequently, the detector thus trained is used to generate a more detailed mask, which is applied to the newly captured (untrained) image to identify the pixels related to the drone and remove the corresponding visual content with improved accuracy and efficiency.

[0034] Of course, in applications such as those described in the background section of the present invention, where the irrelevant object to be removed, for example, is another person or a crane with another camera attached, the training data of the detector should be appropriately modified so that a mask of an appropriate shape can be generated and applied.

[0035] After estimating the second drone position in the 3D space, projecting the ROI defined around that position into the 2D image space, generating and applying the mask, for each image captured by the first drone, the resulting masked image can be aligned and reconstructed to reproduce the scene in 2D or 3D form. Of course, the essence of the above-described process can also be implemented for images captured by a second drone that may potentially include the first drone within its field of view. In the most common case, each drone within a group of drones operating over an overlapping period to image a specific 3D scene can basically use the same process steps to remove content corresponding to any of the other drones in the group from its set of captured images.

[0036] In the art, methods for reconstructing the original scene from the processed (masked) images provided by a group of drones are well known and have been developed for use in applications such as movies, TV, video games, AR / VR, or visual content editing software, which include image alignment, point cloud generation, and mesh, texture, or other similar reconstruction models. Therefore, the present invention can have great value in providing images that do not contain irrelevant visual content for all of these applications. As another application area, it is conceivable to monitor the positioning of an actual drone or group of drones using method 300 as a means of providing real-time visual feedback regarding a planned or expected trajectory.

[0037] Embodiments of the present invention offer many advantages. Specifically, embodiments of the present invention enable multiple drones to capture images of a scene in a relatively short time without the need for extremely accurate drone trajectory control because each drone can quickly and efficiently perform at least the initial stage of the process of removing irrelevant content from the images captured by each drone. The present invention also includes an improved method of generating and applying a mask to perform subsequent stages of the process, enhancing the quality of the results of subsequent reconstruction and reorganization efforts.

[0038] Although specific embodiments have been described, these specific embodiments are merely illustrative and not limiting.

[0039] For the implementation of the routines of specific embodiments, any suitable programming language can be used, including C, C++, Java, assembly language, etc. Different programming techniques such as procedural or object-oriented can be used. These routines can be executed on a single processing device or multiple processors. Although steps, operations, or calculations may be shown in a specific order, this order can be changed in different specific embodiments. In some specific embodiments, multiple steps shown sequentially herein can also be executed simultaneously.

[0040] Specific embodiments can be implemented in a computer-readable storage medium used by, or connected to, an instruction execution system, apparatus, system, or device. Specific embodiments can also be implemented in the form of control logic in software or hardware or a combination thereof. The control logic can perform what has been described in specific embodiments when executed by one or more processors.

[0041] Certain embodiments can be implemented by using a programmed general-purpose digital computer, or by using application-specific integrated circuits, programmable logic devices, field programmable gate arrays, optical, chemical, biological, quantum, or nano engineering systems, components, and mechanisms. In general, the functions of certain embodiments can be realized by any means well known in the art. Distributed, networked systems, components, and / or circuits can also be used. The communication or transfer of data can be by wire, wireless, or any other means.

[0042] Also, when useful for a particular application, it will be understood that one or more of the elements shown in the drawings / figures can be implemented in a more separated or integrated form, or in some cases removed or made inoperable. Implementing a program or code storable in a machine-readable medium that enables a computer to execute any of the above-described methods is also within the spirit and scope of the present invention.

[0043] "Processor" includes any suitable hardware and / or software system, mechanism, or component that processes data, signals, or other information. The processor can include a general-purpose central processing unit, multiple processors, a dedicated circuit for realizing functions, or other systems having such. The processing need not be limited by geographical location or have a time limit. For example, the processor can execute its functions in "real time", "offline", "batch mode", etc. Some parts of the processing can also be executed by different (or the same) processing systems at different times and in different locations. Examples of processing systems can include servers, clients, end-user devices, routers, switches, networked storage, etc. A computer can be any processor that communicates with a memory. The memory can be any suitable processor-readable storage medium such as random access memory (RAM), read-only memory (ROM), magnetic or optical disks, or other non-transitory media suitable for storing instructions executed by the processor.

[0044] As used throughout this specification and the following claims, the indefinite article "a" and the definite article "the" include references in the plural unless the context clearly dictates otherwise. Also, as used throughout this specification and the following claims, the meaning of "in" includes the meanings of "in" and "on" unless the context clearly dictates otherwise.

[0045] Above, specific embodiments have been described in this specification. However, in the above disclosure, modifications, various changes, and substitutions are intended. In some examples, it should be understood that some features of a specific embodiment are used without being accompanied by the use of corresponding other features without departing from the described scope and spirit. Therefore, many modifications can be made to adapt to a particular situation or material to the basic scope and spirit.

Description of Reference Numerals

[0046] 100 scenes 105 First drone 110 Second drone

Claims

1. A method for removing irrelevant content in a first plurality of captured images of a scene in which a second drone exists, captured at a first plurality of corresponding points in time in a plurality of corresponding poses by a first drone, comprising: For each of the first plurality of captured images, the first drone predicting a 3D position of the second drone at the time of capturing the image; the first drone defining a region of interest (ROI) including a projection of the predicted 3D position of the second drone at the time of capturing the image in an image plane corresponding to the captured image; generating a drone mask of the second drone; applying the generated drone mask to the defined ROI to generate an output image that does not include irrelevant content caused by the second drone; including, defining a region of interest (ROI) including a projection of the predicted 3D position of the second drone in an image plane corresponding to the captured image comprises: calculating the orientation of the captured image in a global 3D coordinate system using the pose of the first drone determined at the time of capturing the image; calculating four maximum span points characterizing the second drone in a global coordinate system using the determined pose of the first drone and a predetermined dimension of the second drone; projecting the four maximum span points and the predicted 3D position of the second drone onto the image plane corresponding to the captured image; defining a rectangular box including the projected four maximum span points and the predicted 3D position of the second drone in the image plane; enlarging the rectangular box by a scaling factor considering a timing coefficient and a measurement uncertainty to define the ROI; characterized by including.

2. Predicting the position of the second drone at the time of capturing the captured image comprises: receiving time-stamped position data regarding the second drone, and utilizing prior knowledge regarding the planned trajectory of the second drone, The method according to claim 1, comprising at least one of the above.

3. Receiving the timestamped position data regarding the second drone includes receiving a stream of timestamped position data items transmitted directly or indirectly to the first drone by the second drone, The method according to claim 2.

4. Predicting the position of the second drone at the time of capturing the image includes updating an estimated value of the position of the second drone at the time of capturing based on one or more items of the received stream of timestamped position data items, The method according to claim 2.

5. The timestamped position data regarding the second drone is at least partially generated by 3D global positioning method, The method according to claim 2.

6. The first and second drones use a 5G communication protocol, The method according to claim 1.

7. Generating a drone mask includes defining a shape within the ROI that includes at least some pixels that are likely to contain visual content indicating the second drone within the field of view of the first drone in the pose of the first drone at the time of capturing the image, The method according to claim 1.

8. The shape is defined to be equal to the ROI, The method according to claim 7.

9. The shape is defined as a subset of the pixels of the ROI, and the subset is determined by one or more predetermined rules, The method according to claim 7.

10. The shape is defined as a subset of the pixels of the ROI, and the subset is determined by a learning-based drone detection model, The method according to claim 7.

11. The learning-based drone detection model is trained using a combination of visual training data and heatmap training data, and the heatmap training data is generated using position measurement values of the second drone or another drone identical to the second drone, and the position measurement values are acquired by a non-visual sensor system, The method according to claim 10.

12. The non-visual sensor system is an RTK-GNSS-IMU system, The method according to claim 11.

13. A method for removing irrelevant content in a first plurality of captured images of a scene in which there are a plurality of other drones, captured by a first drone of a scene drone at a first plurality of corresponding points in time corresponding to a plurality of corresponding poses, For each of the first plurality of captured images, the first drone predicting the 3D position of each of the other drones at the time of capturing the image; the first drone defining, in the image plane corresponding to the captured image, a region of interest (ROI) for each of the other drones, including a projection of the predicted 3D position of each of the other drones; generating a drone mask for each of the other drones; applying the generated drone mask to the corresponding defined ROI to generate an output image of the scene that does not include irrelevant content caused by the other drones; including, Defining a region of interest (ROI) including a projection of the predicted 3D position of each of the other drones in the image plane corresponding to the captured image, calculating the orientation of the captured image in a global 3D coordinate system using the pose of the first drone determined at the time of capturing the image; calculating, in the global coordinate system, four maximum span points characterizing each of the other drones using the determined pose of the first drone and the predetermined dimensions of each of the other drones; projecting the four maximum span points and the predicted 3D position of each of the other drones onto the image plane corresponding to the captured image; defining a rectangular box in the image plane including the projected four maximum span points and the predicted 3D position of each of the other drones; enlarging the rectangular box by a scaling factor considering a timing coefficient and a measurement uncertainty to define the ROI; A method characterized by including.

14. An apparatus for removing irrelevant content in a first plurality of captured images of a scene in which a second drone exists, captured by a first drone at a first plurality of corresponding points in time corresponding to a plurality of corresponding poses, one or more processors; logic encoded on one or more non-transitory media for execution by the one or more processors, comprising, the logic, when executed, for each of the first plurality of captured images, The first drone predicts the 3D position of the second drone at the time of capturing the image, the first drone determines a region of interest (ROI) including a projection of the predicted 3D position of the second drone at the time of capturing the image within an image plane corresponding to the captured image, generates a drone mask, applies the generated drone mask to the determined ROI to generate an output image that does not include irrelevant content caused by the second drone, is operable to, the predicting and the determining are executed by one or more of the one or two or more processors located within the first drone, the generating and applying the drone mask are at least partially executed by one or two or more processors located away from the first drone, determining a region of interest (ROI) including a projection of the predicted 3D position of the second drone within an image plane corresponding to the captured image, calculating the orientation of the captured image in a global 3D coordinate system using the attitude of the first drone determined at the time of capturing the image, calculating four maximum span points characterizing the second drone in a global coordinate system using the determined attitude of the first drone and a predetermined dimension of the second drone, projecting the four maximum span points and the predicted 3D position of the second drone onto the image plane corresponding to the captured image, defining a rectangular box including the projected four maximum span points and the predicted 3D position of the second drone within the image plane, enlarging the rectangular box by a scaling factor considering a timing coefficient and a measurement uncertainty to define the ROI, characterized by including.

15. The first drone receives time-stamped position data regarding the second drone, and the time-stamped position data enables at least partially the one or two or more processors located within the first drone to predict the position of the second drone at the time of capturing the captured image. The apparatus according to claim 14.

16. The one or more processors located within the first drone can access prior knowledge of the planned trajectory of the second drone, which may improve the prediction of the position of the second drone at the time of capturing the captured image. The apparatus according to claim 15.

17. The timestamped position data regarding the second drone is at least partially generated by 3D global positioning. The apparatus according to claim 15.

18. The timestamped position data regarding the second drone includes a stream of timestamped position data items transmitted directly or indirectly by the second drone to the first drone. The apparatus according to claim 15.

19. The first drone includes a global navigation satellite system (GNSS) receiver and an inertial measurement unit (IMU) operable to determine the attitude of the first drone when capturing the plurality of first captured images. The apparatus according to claim 14.

Citation Information

Patent Citations

  • System and method for dynamic image masking

    CN105526916A

  • Dynamic image masking system and method

    JP2017027571A

  • Route selecting unit, unmanned aircraft, data processing unit, route selection processing method, and program for route selection processing

    JP2019067252A

  • Method, system, server, mobile device, unmanned aircraft, and program for controlling movement of mobile device

    JP2020091851A

  • Dynamic image masking system and method

    US20170018058A1