Point of interest tracking and estimation

By combining image processing systems and machine learning models with multiple images and GPS data, the problem of 3D object location recognition on mobile devices has been solved, achieving accurate 3D location calculation and reducing reliance on bulky equipment.

CN122249837APending Publication Date: 2026-06-19TOMAHAWK ROBOTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TOMAHAWK ROBOTICS CO LTD
Filing Date
2024-07-10
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine the 3D position of objects within an image on moving vehicles, especially on devices such as drones. This is particularly true when the device vibrates or the distance changes, making it difficult to precisely identify and measure the 3D position of objects.

Method used

By employing an image processing system combined with a machine learning model, the system receives target location markers, identifies objects, determines their orientation and real-world dimensions, records objects from different angles and distances using multiple images, generates image metadata, and calculates the object's accurate location in three-dimensional space using methods such as GPS.

Benefits of technology

It enables accurate 3D positioning of objects within images on mobile devices, improving the accuracy and reliability of position recognition and reducing reliance on bulky equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122249837A_ABST
    Figure CN122249837A_ABST
Patent Text Reader

Abstract

This paper describes a method and system for determining the 3D location of objects within an identified portion of an image. The image processing system receives an image and identifiers of locations within the image. The image can be input into a machine learning model to detect one or more objects within the identified locations. Multiple images can then be used to generate location estimates for those objects. Based on the location estimates, accurate 3D locations can be calculated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 350,722, filed July 11, 2023. The contents of the foregoing application are incorporated herein by reference in their entirety. Background Technology

[0003] Stable and reliable robotic systems are becoming increasingly common, contributing to the recent advancements and widespread adoption of unmanned systems technology. In many cases, these systems are equipped with recording devices (e.g., video, infrared, thermal, audio, point cloud, and / or other recording devices). For example, drones equipped with cameras (e.g., video cameras, night vision cameras, infrared cameras, or other suitable cameras) can provide an operator or another observer with a good tactical view of what is happening in the operational area. In many cases, identifying objects (e.g., other operators, vehicles, etc.) in camera footage and indicating their position and orientation in three-dimensional space (e.g., whether the object is facing a particular direction) can be useful. However, because cameras are mounted on moving (e.g., flying) vehicles (e.g., drones) and the distance to the object may change as the vehicle moves, determining the accurate three-dimensional position of the object within the image can be difficult. Furthermore, determining the distance from the camera to the object, which is fundamental to determining the object's three-dimensional position, can also be challenging.

[0004] For example, an operator might be controlling a drone with a mounted camera hovering above an operational area. The drone can transmit video feeds of the operational area to the operator and / or to command and control positions. Multiple operators and multiple vehicles may be present within the operational area. Determining the three-dimensional position of each object within the operational area can be useful. Some methods of determining three-dimensional position may include using bulky and expensive laser systems. These systems can reduce the operational time of vehicles (e.g., battery consumption when using and carrying equipment). Furthermore, equipping each vehicle with such a system may be impractical, especially when the vehicle may never return. Therefore, without specialized, bulky, and expensive equipment, it may be difficult to accurately determine the three-dimensional position of objects within an image in a video feed. This determination is even more difficult because the vehicle is moving and susceptible to vibration. Summary of the Invention

[0005] Therefore, this document describes methods and systems for determining the three-dimensional position of an object within an image received from a camera mounted on a vehicle (e.g., an unmanned vehicle such as a flying / hovering drone). For example, an image processing system can be used to perform the operations described herein. The image processing system can be hosted by the vehicle (e.g., an aerial drone), hosted at a central location (e.g., a command and control point), or hosted on a computing device used by an operator. For example, the command and control point can be located in a vehicle equipped with one or more computing devices, in a data center housing the computing devices, or in another suitable environment.

[0006] An image processing system can (e.g., at an unmanned vehicle) receive identifiers of target locations within images recorded by the unmanned vehicle. The target location can be an object, an area, a group of objects, or another suitable location. The image processing system can receive the identifiers from a command center, from an operator, or from another suitable source. For example, an operator can use an input system on their user device (e.g., a touchscreen) to circle an area, touch the screen where the object is located, or perform another suitable selection. This selection can indicate to the image processing system which location within the image is of interest to the user and can be stored as an identifier.

[0007] The image processing system can then identify objects within the target location. Specifically, the image processing system can input images into a machine learning model to obtain object identifiers and the orientation of the objects within the target location. The machine learning model can be trained to detect objects within the received images. In some embodiments, the image processing system can provide the target location to the machine learning model, and the machine learning model can output object identifiers and the orientation of the corresponding objects. For example, the machine learning model can output object identifiers (e.g., tanks and / or tank models) and orientations (e.g., turrets / tanks facing south).

[0008] The image processing system can then determine the real-world dimensions of each identified object. Specifically, the image processing system can determine the set of real-world dimensions associated with an object based on its object identifier. For example, the image processing system can use the object's identifier to query a database and retrieve the width, length, and height of an object (e.g., a tank) from the database. The real-world dimensions of the object can be used to help determine the distance from the vehicle recording the image to the object.

[0009] The image processing system can then collect more images of the target location from slightly different distances or positions. Therefore, the image processing system can receive an image stream comprising multiple images from a camera mounted on the unmanned vehicle, each image indicating the target location. As described above, the image stream can be recorded by the camera as the unmanned vehicle moves relative to the object and / or target area.

[0010] When recording images, an image processing system can generate the data needed to estimate the distance between the camera and the object. Therefore, the image processing system can generate multiple sets of image metadata for multiple images. Each set of image metadata can include the object's orientation, the object's image size, camera data associated with the camera, the camera's orientation, and the camera's position in three-dimensional space when each image is recorded. That is, the image processing system can estimate the range from the camera to the object in its field of view by recognizing the object's scale and using view data from the camera. For example, a camera can record ten images of a tank at different distances and angles to it, and can generate metadata associated with each image (e.g., the camera's orientation when the image is recorded, the camera's position in three-dimensional space, camera data, and / or other metadata). The image processing system can also determine the image size of the object (e.g., based on the object's real-world size and its orientation). For example, if the tank is facing a certain direction, the image processing system can determine that a particular side of the object is visible and can measure the size of that side from the image. However, given the orientation, the other side of the object may not be visible, and the image processing system can calculate the size of that side within the image.

[0011] The image processing system can then use metadata to determine an estimate of the object's location for each image. Specifically, the image processing system can determine multiple location estimates of the object in three-dimensional space. For example, the image processing system can generate an estimated distance from the camera to the object in three-dimensional space for each image based on camera settings such as field of view and focal length, real-world object size, image object size, etc. The image processing system can then combine this distance from the camera to the object with the camera's location (e.g., determined based on GPS coordinates and / or other methods).

[0012] When determining the estimates, the image processing system can use a combination of estimates to calculate a more accurate position of the object in three-dimensional space. Specifically, the image processing system can determine the object's position in three-dimensional space based on multiple position estimates. That is, all individual estimates may not be precise, but combining estimates can result in a much more accurate position in three-dimensional space. In some embodiments, the image processing system can calculate the average latitude and longitude of each three-dimensional estimate to arrive at the object's position in three-dimensional space. Additionally, for example, for a hovering object, the image processing system can also determine the object's height.

[0013] In some cases, as the vehicle moves around and records images, the orientation of an object may change (e.g., due to object movement or vehicle movement). Therefore, the image processing system takes each image as input to a machine learning model to calculate the object's updated orientation. Additionally, if the object is moving, the image processing system can instruct the unmanned vehicle to maneuver to obtain a more accurate estimate of the object's 3D position.

[0014] Furthermore, the image processing system may be unable to determine the size of an object because the object is not in the database available to the system. In this case, the image processing system can identify another object with known dimensions near the unidentified object and calculate the size of the unidentified object based on the dimensions of the known object.

[0015] Various other aspects, features, and advantages of the system will become apparent from the detailed description and accompanying drawings. It should also be understood that the foregoing general description and the following detailed description are exemplary and do not limit the scope of this disclosure. As used in the specification and claims, the singular forms “a,” “an,” and “the” include plural indicators unless the context clearly specifies otherwise. Additionally, as used in the specification and claims, the term “or” means “and / or” unless the context clearly specifies otherwise. Furthermore, as used in the specification, “part” means a portion or all (i.e., the entire part) of a given item (e.g., data) unless the context clearly specifies otherwise. Attached Figure Description

[0016] Figure 1 An exemplary system for determining the three-dimensional position of an object in an image is shown according to one or more embodiments of the present disclosure.

[0017] Figure 2 An example fragment (excerpt) of a data structure that can store a target location and the shape of the target location according to one or more embodiments of the present disclosure is shown.

[0018] Figure 3An exemplary machine learning model according to one or more embodiments of this disclosure is shown.

[0019] Figure 4 Example fragments of data structures that can be imaged to generate three-dimensional position estimates according to one or more embodiments of the present disclosure are shown.

[0020] Figure 5 A computing device according to one or more embodiments of the present disclosure is shown.

[0021] Figure 6 This is a flowchart of operations for identifying the three-dimensional positions of objects within a video stream and linking these objects to known objects, according to one or more embodiments of this disclosure. Detailed Implementation

[0022] In the following description, numerous specific details are set forth for illustrative purposes in order to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details or using equivalent arrangements. In other instances, well-known models and devices are shown in block diagram form to avoid unnecessarily obscuring the disclosed embodiments. It should also be noted that the methods and systems disclosed herein are also applicable to applications unrelated to source code programming.

[0023] Figure 1 This is an example of an environment 100 used to identify the three-dimensional position of one or more objects within a video stream. Environment 100 includes an image processing system 102, a data node 104, and recording devices 108a-108n. The image processing system 102 can execute instructions for identifying the three-dimensional position of one or more objects within the video stream. The image processing system 102 can include software, hardware, or a combination of both. In some embodiments, although shown separately, the image processing system 102 and the data node 104 may reside on the same computing device.

[0024] Data node 104 can store various types of data. For example, data node 104 can store a repository of machine learning models that can be accessed by image processing system 102. In some embodiments, data node 104 can also be used to train machine learning models and / or tune parameters (e.g., hyperparameters) associated with those machine learning models. Data node 104 can include software, hardware, or a combination of both. For example, data node 104 can be a physical server or a virtual server running on a physical computer system. In some embodiments, data node 104 can reside in a data center for use by a commander for situational awareness. Network 150 can be a local area network, a wide area network (e.g., the Internet), or a combination of both. Recording devices 108a-108n can be devices attached to unmanned vehicles and can include cameras, infrared cameras, microphones, thermal imaging devices, and / or other suitable devices.

[0025] Image processing system 102 can receive an identifier of a target location within an image. For example, image processing system 102 may be hosted on an unmanned vehicle (e.g., an aerial drone). Image processing system 102 can transmit images or video streams of images to an operator, command and control center, or another suitable target. For example, one or more images transmitted may be part of an image stream captured by a camera (e.g., a recording device in recording devices 108a-108n) mounted on a drone or other suitable vehicle. In some embodiments, the drone may be wirelessly connected to a network (e.g., network 150) and may transmit image data (e.g., video footage) to the image processing system and / or data node 104. When one or more images are received by a device (e.g., a user device, a device at a command and control center, or another suitable device), the operator of the device can select a target location based on the image. For example, the operator may use a finger, stylus, or other suitable selection tool to select a location (e.g., a circle, square, or other suitable shape). In some embodiments, the target location may be automatically selected by a computer system. Furthermore, multiple target locations may be selected, and each target location may be processed based on the disclosure below.

[0026] When a target location is selected, it can be sent to an image processing system 102 (e.g., hosted at an unmanned vehicle). In some embodiments, the image processing system may be hosted on a device in a command and control center and / or on an operator's device. The image processing system 102 may receive an identifier of the target location using a communication subsystem 112. The communication subsystem 112 may include software components, hardware components, or a combination of both. For example, the communication subsystem 112 may include a network interface card (e.g., a wired / wireless network interface card / processor) coupled to software to drive the card / processor. The network interface card may be built into a server or another suitable computing device.

[0027] Figure 2 A sample fragment of a data structure 200 is shown that can store the target location and the shape of the target location. Field 203 can store an image identifier that identifies the image in which the target location is selected. Field 206 can store the target location as indicated by a set of coordinates. Field 209 can store a shape identifier. For example, Figure 2 The target location is shown with four coordinates indicating a rectangular shape. Therefore, the coordinates can be pixel locations within the image. Another shape that can be used is a circle or an ellipse. For a circle, there can be a single coordinate indicating the center of the circle (e.g., XY coordinates within the image), where a measure of diameter or radius indicates the size of the circle. These measures can be in pixels or other suitable units. In some embodiments, the image processing system 102 can use other methods to receive and store the target location. For example, if an operator uses a stylus or finger to draw the target location, the image processing system 102 can simply store the coordinates of each pixel touched by the user when using the stylus or finger. The communication subsystem 112 can pass each image and image metadata or a pointer to an address in memory to the object detection subsystem 114.

[0028] The object detection subsystem 114 may include software components, hardware components, or a combination of both. The object detection subsystem 114 may encompass a machine learning model or may be enabled to access a machine learning model. The object detection subsystem 114 may input an image into the machine learning model to obtain object identifiers and the orientation of the object within a target location. The machine learning model may be trained to detect objects within a received image. In some embodiments, in addition to inputting an image into the machine learning model, the object detection subsystem 114 may also input an indication of the target location into the machine learning model. In some embodiments, the object detection subsystem 114 may use the target location in conjunction with the output of the machine learning model, as described later in this disclosure.

[0029] The machine learning models used in this disclosure can take many forms. Figure 3An exemplary machine learning model is illustrated. Machine learning model 302 can take input 304 (e.g., an image and / or target location) and can output 306 one or more object identifiers of objects within the image. In some embodiments, the machine learning model can output the probability that an object has been detected and the object's location within the image. In some embodiments, the machine learning model can output an identifier of an object found within a target location input to the machine learning model. Output parameters can be fed back to the machine learning model as input to train it (e.g., user indications of the accuracy of the output alone or in combination with labels or other reference feedback information associated with the input). The machine learning model can update its configuration (e.g., weights, biases, or other parameters) based on evaluations of its (e.g., those of an information source) predictions and reference feedback information (e.g., user indications of accuracy, reference labels, or other information). For example, if the machine learning model is a neural network, connection weights can be adjusted to reconcile the difference between the neural network's predictions and the reference feedback. One or more neurons in the neural network may need to backpropagate their respective errors through the neural network to facilitate the update process (e.g., error backpropagation). For example, updates to connection weights can reflect the magnitude of the error backpropagated after the forward pass is complete. In this way, for example, machine learning models can be trained to generate better predictions from information sources in response to queries.

[0030] In some embodiments, a machine learning model may include an artificial neural network. In such embodiments, the machine learning model may include an input layer and one or more hidden layers. Each neuron in the machine learning model may be connected to one or more other neurons in the machine learning model. Such connections may have an enhancing or inhibitory effect on the activation state of the connected neurons. Each individual neuron may have a summation function that combines the values ​​of all its inputs. Each connection (or the neuron itself) may have a threshold function that a signal must exceed to propagate to other neurons. The machine learning model may be self-learning and / or trained, rather than explicitly programmed, and may perform significantly better in a specific problem-solving domain compared to a computer program that does not use machine learning. During training, the output layer of the machine learning model may correspond to a classification of the machine learning model, and inputs known to correspond to that classification may be fed into the input layer of the machine learning model during training. During testing, inputs of unknown classifications may be fed into the input layer, and a determined classification may be output.

[0031] Machine learning models can include embedding layers, where each feature of a vector is transformed into a dense vector representation. These dense vector representations of each feature can be pooled at one or more subsequent layers to transform the set of embedding vectors into a single vector.

[0032] Machine learning models can be constructed as factorization machine models. Machine learning models can be nonlinear models capable of performing classification and / or regression, and / or supervised learning models. For example, a machine learning model can be a general supervised learning algorithm used by the system for both classification and regression tasks. Alternatively, machine learning models can include Bayesian models configured to perform variational inference on graphs and / or vectors.

[0033] When the object detection subsystem 114 receives output from the machine learning model, it can determine a set of real-world dimensions associated with objects based on object identifiers. In some embodiments, the machine learning model can receive images as input for detecting objects within the images. The machine learning model can output object identifiers for the objects detected in the images. Additionally, the machine learning model can output coordinates within each image associated with each object. The coordinates can be two-dimensional coordinates (e.g., XY coordinates) based on pixel counts within the image. In one example, only one XY coordinate (e.g., 230 x 340) indicating the center or center location of the object can be output. In one example, multiple XY coordinates can be output. For example, XY coordinates can indicate the outline of the object. The object detection subsystem 114 can then determine which objects(s) are located within the target location based on the coordinates and the target location.

[0034] In some embodiments, the machine learning model may receive an image and a target location as input and may output only one or more objects within the target location. In both cases, the machine learning model may output the probability that the object has been correctly identified. The image processing system 102 may use the probability to determine whether to process the object or not. Furthermore, the machine learning model may output the orientation of the object. The orientation of the object may indicate which direction the object is facing. For example, if the object identifier is a tank, the machine learning model may output which direction the turret is facing or the front of the tank is facing within the image. This indication may be in degrees, where the top of the image indicates north, the bottom of the image may indicate south, the left side of the image may indicate west, and the right side of the image may indicate east. For example, if the tank is facing the lower left corner of the image, the indication may be 45 degrees southwest. However, other schemes for indicating orientation may be implemented.

[0035] Therefore, the object detection subsystem 114 can determine the set of real-world dimensions associated with an object based on the object identifier. For example, the object identifier could be a specific type of object (e.g., a tank). In another example, the object identifier could be a specific model of the object (e.g., an M1A2 Abrams main battle tank). Thus, the object detection subsystem 114 can send a request for the real-world dimensions of the object to a database server (e.g., data node 104). The database server can perform a lookup of the object identifier and respond with dimensions (e.g., length, width, and / or height). The object detection subsystem 114 can store the set of real-world dimensions, for example, in memory.

[0036] The image processing system 102 can then continue to determine the object's position in three-dimensional space by recording images from different locations and estimating the three-dimensional position from those locations to reach the position within the three-dimensional space as the vehicle (e.g., an aerial drone) moves around. Specifically, the object detection subsystem 114 can (e.g., via the communication subsystem 112) receive an image stream containing multiple images, each of which indicates the target location. The image stream can be recorded by a camera as the unmanned vehicle moves relative to the object. For example, if the image processing system 102 is hosted on the unmanned vehicle (e.g., an aerial drone), the image processing system can receive the image stream directly from the camera (e.g., via an electronic connection). If the image processing system is not hosted on the unmanned vehicle (e.g., hosted on an operator's device or hosted at a command and control center), the image processing system can receive images wirelessly from the unmanned vehicle. The images may include camera settings for recording the images (e.g., focal length and / or other suitable settings).

[0037] The object detection subsystem 114 can pass image streams and camera settings to the position estimation subsystem 116. The position estimation subsystem 116 can include software components, hardware components, or a combination of both. For example, the position estimation subsystem 116 can include software components that access data in memory and / or storage devices, and its operation can be performed using one or more processors. The position estimation subsystem 116 can generate multiple sets of image metadata for multiple images. Each set of image metadata can include one or more of the following: the orientation of the object, the image size of the object, camera data associated with the camera (e.g., focal length, field of view, and / or other data), the orientation of the camera mounted on the unmanned vehicle, and the position of the camera mounted on the unmanned vehicle in three-dimensional space at the time each image is recorded. The position estimation subsystem 116 can determine the orientation of the camera by querying payload data (e.g., data associated with the gimbal used to mount the camera). Furthermore, the position estimation subsystem 116 can determine the position of the camera based on the position of the unmanned vehicle. For example, the position estimation subsystem 116 can query the position from the navigation system of the unmanned vehicle.

[0038] In some embodiments, the orientation of a camera may be referred to as its angular position. The angular position of a camera can be expressed using roll, pitch, and yaw. The position of a camera can be expressed using a three-dimensional position. The three-dimensional position may include latitude, longitude, and altitude.

[0039] In some embodiments, the position estimation subsystem 116 may perform the following operations when generating multiple sets of image metadata. This process may be performed for each image within an image stream or for some images within an image stream. In some embodiments, the position estimation subsystem 116 may select a first image and infer known dimensions associated with the object from the first image based on the object's orientation. For example, an object (such as a tank) may be oriented within an image with its front facing the top left corner of the image. Thus, depending on the angle at which the image was captured, the system may determine only the length of one or both sides of the object, and may be able to determine the object's height. Therefore, the position estimation subsystem 116 may not be able to determine all dimensions of the object within the image.

[0040] To infer the known dimensions associated with an object, the position estimation subsystem 116 can match the correct dimensions based on the object's orientation. Specifically, the position estimation subsystem 116 can determine a first object dimension from the known dimensions. For example, the position estimation subsystem 116 can perform image analysis to determine the size of the first object dimension. Image analysis may involve color comparison to determine where a specific dimension of the object begins and ends. For example, the position estimation subsystem 116 can determine that the first object dimension is 3.12 inches (e.g., the length of a tank as shown in the image).

[0041] Then, the position estimation subsystem 116 can determine a first real-world size that matches the first object size based on the object's orientation. For example, if the object's orientation indicates that the object is facing the top-left corner of the image, the position estimation subsystem 116 can determine that the first object size corresponds to the object's real-world length. Therefore, the position estimation subsystem 116 can assign a size label to the first object size. For example, the position estimation subsystem 116 can assign the label "length" to the first object size.

[0042] However, based on determining at least one size of the object and its real-world size, the position estimation subsystem 116 can determine other sizes. Specifically, the position estimation subsystem 116 can determine a size modification factor based on one or more known sizes and real-world sizes. For example, the position estimation subsystem 116 can determine a first size of the object within the image and match that size to the same size of the real-world object. Based on the ratio of sizes, the position estimation subsystem 116 can determine the size modification factor of the object (e.g., the ratio of the object within the image to the real-world object). For example, a tank may have an actual length of 26 feet. Furthermore, in the image, the tank may have a length of 3.12 inches. Therefore, the ratio or size modification factor can be 100 times (100x), which can be the ratio of the 3.12-inch image length to the 26-foot (312-inch) real-world length.

[0043] Therefore, the position estimation subsystem 116 can generate size values ​​for unknown dimensions associated with the object based on the size modification factor to generate a set of image dimensions for the object in the first image. For example, if the width of a real-world object is 12 feet (144 inches), the image width of the object could be 1.44 inches. The position estimation subsystem 116 can perform the same calculation for each dimension of the object to determine the complete set of object dimensions.

[0044] Figure 4An example fragment of a data structure is shown that can image metadata to generate a 3D position estimate. Field 403 can store an image identifier for each image used to determine the real-world location of an object. Field 406 can include object metadata. Object metadata can include attributes such as object type (e.g., person, vehicle, etc.). For a vehicle, object metadata can include the type of vehicle (e.g., air, ground, sea (underwater, above water, etc.)). Field 406 can store camera data, including, for example, the 3D (e.g., real-world) location of the camera (e.g., the location of an unmanned vehicle on which the camera is mounted), camera settings, and / or other suitable camera data. Camera settings can include things such as focal length, shutter speed, and / or other suitable camera settings. Field 409 can store image metadata, including, for example, the orientation of the object. Additionally, field 409 can store the real-world object size (e.g., as retrieved from a database) and the image object size, as determined, for example, above.

[0045] In some embodiments, the location estimation subsystem 116 can determine the orientation of an object for one or more images in a data stream. In some cases, the location estimation subsystem 116 can determine the orientation of an object for each image in the data stream. This can make the estimation more accurate because the object may move in ways that could affect the measurement (e.g., rotate or otherwise manipulate) with different orientations. Furthermore, movement of a vehicle (e.g., an unmanned vehicle) may cause the vehicle's position to change sufficiently that the object's orientation may change relative to the vehicle, resulting in a less accurate estimation. Therefore, the location estimation subsystem 116 can input a first image from a plurality of images into a machine learning model (e.g., the machine learning model described above). In some embodiments, a single machine learning model may exist that outputs the object's orientation and object identifier along with the image. However, in some embodiments, multiple machine learning models may exist that perform these tasks. The location estimation subsystem 116 can receive object identifiers and updated orientations of objects from the machine learning models. For example, the machine learning model may output a set of identifiers for one or more objects detected within an image, along with their orientations. The location estimation subsystem 116 can determine the object being estimated (e.g., as described above) based on the object identifiers and store the object's orientation. Then, the location estimation subsystem 116 can add the updated orientation to the corresponding set in multiple image metadata sets. For example, if the updated orientation is for, for example, Figure 4 If the image 121 is shown in field 403, then the location estimation subsystem 116 can add a new orientation to the corresponding field 409.

[0046] When generating image metadata, the position estimation subsystem 116 can determine multiple estimates of the three-dimensional position of an object, such that each estimate is based on a corresponding image. In some embodiments, estimates can be generated in real time, for example, when each image is received and a set of image metadata for that image is generated. However, in some embodiments, estimates can be generated only after a certain number of images have generated associated image metadata (e.g., 3 images, 5 images, 10 images, etc.). Therefore, the image processing subsystem 116 can determine multiple position estimates of the object in three-dimensional space, such that each of the multiple position estimates is generated for a corresponding image among the multiple images based on a corresponding set of image metadata. In some embodiments, the position estimation subsystem 116 can use the camera's focal length, the object's real-world size, the image size, the object size, and the sensor size to estimate the distance from the camera to the object in each image. Once each distance to the object is generated, the image processing subsystem 116 can use the camera's position (e.g., generated via GPS or another suitable method) to determine a three-dimensional estimate of the object within each image.

[0047] In some embodiments, the location estimation subsystem 116 can add estimates to... Figure 4 The data structure can be modified as follows. For example, the position estimation subsystem 116 can add another field (not shown) and store the position estimate of a specific image in that field. When generating position estimates, the position estimation subsystem 116 can determine the position of an object in three-dimensional space based on multiple position estimates. For example, the position estimation subsystem 116 can execute a function on one or more position estimates to determine the position of the object in the three-dimensional system. This function can be an averaging function, a mean function, a pattern function, or other suitable function.

[0048] In some embodiments, each location estimate may be a combination of the longitude and latitude coordinates of an object's location. Additionally, each location estimate may include altitude (e.g., for an aerial object). Therefore, the location estimation subsystem 116 can retrieve multiple latitude coordinates and multiple longitude coordinates from multiple location estimates. For example, the location estimation subsystem 116 may access... Figure 4 The data structure retrieves multiple latitude and longitude coordinates. In some embodiments, the position estimation subsystem 116 can also use elevation or altitude to identify objects in three-dimensional space. For example, the object could be a hovering aircraft or other aircraft. Therefore, the position estimation subsystem 116 can also determine the elevation of the object (e.g., its position above the ground). Elevation or elevation estimates can be determined in the same manner as the estimation of longitude and latitude coordinates.

[0049] Then, the location estimation subsystem 116 can generate the object's position in three-dimensional space based on average latitude and average longitude coordinates. For example, the location estimation subsystem 116 can calculate the average longitude and average latitude of coordinates within an image and use the average coordinates as the object's three-dimensional position. In some embodiments, the location estimation subsystem 116 can perform an averaging operation on the estimated elevation coordinates within the image. Therefore, the object's three-dimensional position can be the object's latitude, the object's longitude, and the object's altitude (e.g., height).

[0050] In some embodiments, the location estimation subsystem 116 can determine whether the location estimation has converged in time, and if so, generate a 3D location based on the convergence. Specifically, the location estimation subsystem 116 can sort multiple location estimates based on corresponding timestamps. For example, each image can have a corresponding timestamp. Therefore, each location estimate can be associated with the timestamp of the corresponding image. The location estimation subsystem 116 can sort the location estimates based on those timestamps, where earlier timestamps are ranked earlier.

[0051] Then, the location estimation subsystem 116 can determine whether multiple location estimates converge to a given value over time. For example, the location estimation subsystem 116 can determine that the location estimates get closer and closer to a specific location (e.g., latitude, longitude, and / or altitude) over time. If this is the case, the location estimation subsystem 116 can record that location as the three-dimensional location of the object.

[0052] However, if the position estimates do not converge over time, the position estimation subsystem 116 can enable the camera to record more images. Specifically, based on the determination that multiple position estimates do not converge at a given time value, the position estimation subsystem 116 can generate a first command for the camera to record more images and a second command for the unmanned vehicle to perform more maneuvers. The position estimation subsystem 116 can repeat this process until convergence is determined.

[0053] In some cases, the image processing system 102 can detect objects it cannot recognize. For example, the object could be a specific building, a vehicle, or another suitable object that the machine learning model cannot identify. The image processing system 102 can also perform localization operations on those objects. Specifically, the image processing system 102 can identify a second object within a target location. For example, the machine learning model can identify that another object has already been detected at that location. Although the object could be a building, the machine learning model might not be able to recognize the type of object. Therefore, the image processing system 102 can determine that the second object does not have a corresponding known size.

[0054] In some embodiments, based on the determination that a second object does not have a corresponding known size, image processing system 102 can (e.g., via object detection subsystem 114) compare the first image size of an object and the second image size of a second object. For example, if the first object in the location has been identified as a tank and the second object in the location has not yet been identified, image processing system 102 can use the ratio of the known size of the identified object to the size of the object in the image to determine the size of the unidentified object. Therefore, image processing system 102 can (e.g., via object detection subsystem 114) compare the sizes of the two objects. These comparisons can be based on the size of the object in the image. For example, the width of the known object could be one inch, while the width of the unknown object could be ten inches.

[0055] Based on a comparison of the first image size and the second image size, the image processing system 102 can (e.g., via the object detection subsystem 114) determine a second size modification factor for the second object. For example, if the width of the first object is one inch and the width of the second object is ten inches, the second size modification factor could be multiplied by ten. Therefore, the image processing system 102 can determine the second three-dimensional position of the second object based on the second size modification factor. When determining the second object modification factor, the image processing system 102 can combine the first and second size modification factors to achieve the object's real-world size. For example, if the real-world width of the original object (e.g., a tank) is seven feet and the image width is one inch, the image processing system 102 can determine that a second object (e.g., an unidentified building) that is ten inches in size on the image is ten times larger in the real world. Therefore, the width of the unidentified building could be seventy feet.

[0056] In some embodiments, the image processing system 102 can determine that the object of interest is moving. Because it is difficult to calculate the distance to the moving object while the observation point is also moving, the image processing system 102 can instruct the vehicle hosting the camera to stop moving. Specifically, the image processing system 102 can determine that the object is moving. The image processing system can use any available means to make this determination. In some embodiments, the image processing system can use images recorded by the camera to determine whether the object appears larger (moving closer) or smaller (moving farther away) in images taken over time to determine whether the object is moving.

[0057] Based on the determination that the object is moving, the image processing system 102 can generate a first command for the unmanned vehicle to stop maneuvering. For example, if the vehicle is moving away from or towards the object, the image processing system 102 can instruct the vehicle to stop. The image processing system 102 can then generate a second command for the camera to record more images. Once new images are generated, the image processing system 102 can repeat the process discussed above to generate a position estimate, and then generate the object's position. Because the object is moving, the image processing system 102 can adjust its calculations based on the object's movement. Therefore, the image processing system 102 can update the object's position over time based on the object's movement.

[0058] In some embodiments, other vehicles (e.g., unmanned vehicles) may be within communication range of the vehicle hosting the camera. Image processing system 102 may use recording devices (e.g., cameras) and processors on those vehicles to aid in calculating the three-dimensional position of the object. For example, recording devices 108a-108n may be used in these embodiments. Specifically, image processing system 102 may detect a set of unmanned vehicles capable of recording images of the target location. This set of unmanned vehicles may include one or more vehicles. In some embodiments, image processing system 102 may detect manned or unmanned vehicles that support communication with image processing system 102 and are capable of calculating estimates.

[0059] The image processing system 102 can send commands to the ensemble of unmanned vehicles to establish point-to-point communication. For example, there may be no available network connectivity (e.g., cellular or satellite connectivity) near the ensemble of unmanned vehicles. Therefore, an unmanned vehicle can establish one or more point-to-point communications with another unmanned vehicle. In some embodiments, any of these vehicles can be a manned or unmanned vehicle. Point-to-point communication can be conducted via known protocols, such as Bluetooth, Wi-Fi, Wi-Max, and / or other suitable protocols.

[0060] Image processing system 102 can receive additional images and additional image metadata of the target location from the set of unmanned vehicles, and can use the additional images and additional image metadata to determine an estimate of the location of additional objects. For example, each vehicle in the set of unmanned (or manned) vehicles can send an image to image processing system 102. The image may include image metadata (e.g., such as...). Figure 4 (As shown). In some embodiments, each vehicle in this set can... Figure 4The data structure is sent to the image processing system 102. In some embodiments, the image processing system 102 may instruct each unmanned vehicle to perform position estimation using corresponding images recorded by cameras on those vehicles and send those estimates to the image processing system 102. When the image processing system 102 receives those estimates, it can calculate the three-dimensional position of the object.

[0061] When the three-dimensional position of an object is calculated, the image processing system 102 can use the output subsystem 118 to send the three-dimensional position to one or more other devices. The output subsystem 118 may include software components, hardware components, or a combination of both. For example, the output subsystem 118 may include software components that access data in memory and / or storage devices, and may use one or more processors to generate an overlay on the image. The output subsystem 118 may generate an indicator at the three-dimensional position of each object within the image. The indicator may include an identifier of the object and / or other information associated with the object. In some embodiments, the indicator may be an augmented reality indicator. For example, an operator may wear an augmented reality device (e.g., augmented reality glasses). The augmented reality device may receive an image (e.g., drone video footage) and may display the video footage to the operator's augmented reality device. Along with the drone video footage, the output subsystem 118 may display an augmented reality indicator overlaid on the drone video footage.

[0062] In some embodiments, the output subsystem 118 may select different indicators based on object type. For example, the output subsystem 118 may determine the type associated with an object based on metadata associated with a known object. This type may be an operator, a land vehicle, a water vehicle, an air vehicle, or other suitable type. The output subsystem 118 may retrieve an augmented reality identifier associated with that type. For example, each type of object may have a different associated indicator. For an operator, the indicator may include a human silhouette, while each vehicle may include a unique silhouette associated with that particular vehicle. The output subsystem 118 may then generate an augmented reality identifier associated with that type at the object's location within the image for display.

[0063] Computing environment

[0064] Figure 5 An example computing system is illustrated that can be used according to some embodiments of this disclosure. In some cases, computing system 500 is referred to as a computer system. The computing system may be hosted on a device that can be controlled by an operator (e.g., a smartphone, tablet, or other suitable device). In some embodiments, the computing system may be hosted on a server at a data center. Those skilled in the art will understand that these terms are used interchangeably. Figure 5 The components can be used to perform actions related to... Figures 1-4 Some or all of the operations discussed herein. Furthermore, various parts of the systems and methods described herein may include or be executed on one or more computer systems similar to computing system 500. Additionally, the processes and modules described herein may be executed by one or more processing systems similar to computing system 500.

[0065] Computing system 500 may include one or more processors (e.g., processors 610a-610n) coupled to system memory 520, input / output (I / O) device interface 530, and network interface 540 via I / O interface 550. The processor may include a single processor or multiple processors (e.g., a distributed processor). The processor may be any suitable processor capable of executing or otherwise performing instructions. The processor may include a central processing unit (CPU) that executes program instructions to perform arithmetic, logical, and input / output operations of computing system 500. The processor may execute code that creates an execution environment for the program instructions (e.g., processor firmware, protocol stack, database management system, operating system, or a combination thereof). The processor may include a programmable processor. The processor may include a general-purpose or special-purpose microprocessor. The processor may receive instructions and data from memory (e.g., system memory 520). Computing system 500 may be a single-processor system including one processor (e.g., processor 610a), or a multiprocessor system including any number of suitable processors (e.g., 610a-610n). Multiple processors may be employed to provide parallel or sequential execution of one or more portions of the techniques described herein. The processes described herein (such as logical flows) can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating corresponding outputs. The processes described herein can be executed by special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)), and the devices can also be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs). The computing system 500 may include multiple computing devices (e.g., a distributed computer system) to implement various processing functions.

[0066] I / O device interface 530 can provide an interface for connecting one or more I / O devices 560 to computer system 500. I / O devices may include, for example, devices that receive input from a user or output information to a user. I / O devices 560 may include, for example, graphical user interfaces displayed on a display (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor), pointing devices (e.g., a computer mouse or trackball), keyboards, keypads, touchpads, scanning devices, voice recognition devices, gesture recognition devices, printers, audio speakers, microphones, cameras, etc. I / O devices 560 can be connected to computer system 500 via wired or wireless connections. I / O devices 560 can be connected to computer system 500 from remote locations. For example, I / O devices 560 located on a remote computer system can be connected to computer system 500 via a network and network interface 540.

[0067] Network interface 540 may include a network adapter that provides a connection between computer system 500 and a network. Network interface 540 can facilitate data exchange between computer system 500 and other devices connected to the network. Network interface 540 may support wired or wireless communication. The network may include electronic communication networks such as the Internet, local area network (LAN), wide area network (WAN), cellular communication network, etc.

[0068] System memory 520 may be configured to store program instructions 570 or data 580. Program instructions 570 may be executed by a processor (e.g., one or more of processors 610a to 610n) to implement one or more embodiments of the present invention. Program instructions 570 may include modules for implementing one or more technologies described herein with respect to various processing modules. Program instructions may include computer programs (which are referred to in some forms as programs, software, software applications, scripts, or code). Computer programs may be written in programming languages, including compiled or interpreted languages, or declarative or procedural languages. Computer programs may include units suitable for use in a computing environment, including as standalone programs, modules, components, or subroutines. Computer programs may or may not correspond to files in a file system. Programs may be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). Computer programs may be deployed to execute on one or more computer processors located locally at a single site or distributed across multiple remote sites interconnected by a communication network.

[0069] System memory 520 may include a tangible program carrier on which program instructions are stored. The tangible program carrier may include a non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium may include a machine-readable storage device, a machine-readable storage substrate, a memory device, or any combination thereof. A non-transitory computer-readable storage medium may include non-volatile memory (e.g., flash memory, ROM, PROM, EPROM, EEPROM memory), volatile memory (e.g., random access memory (RAM), static random access memory (SRAM), synchronous dynamic RAM (SDRAM)), mass storage memory (e.g., CD-ROM and / or DVD-ROM, hard disk drive), etc. System memory 520 may include a non-transitory computer-readable storage medium that may have program instructions stored thereon, which can be executed by a computer processor (e.g., one or more processors 610a-610n) to implement the subject matter and functional operations described herein. Memory (e.g., system memory 520) may include a single memory device and / or multiple memory devices (e.g., distributed memory devices).

[0070] I / O interface 550 can be configured to coordinate I / O traffic between processors 610a-610n, system memory 520, network interface 540, I / O devices 560, and / or other peripheral devices. I / O interface 550 can perform protocol, timing, or other data transformations to convert data signals from one component (e.g., system memory 520) into a format suitable for use by another component (e.g., processors 610a-610n). I / O interface 550 may include support for devices attached via various types of peripheral buses, such as the Peripheral Component Interconnect (PCI) bus standard or variants of the Universal Serial Bus (USB) standard.

[0071] Embodiments of the techniques described herein can be implemented using a single instance of computer system 500 or multiple computer systems 500 configured to host different portions or instances of the embodiments. Multiple computer systems 500 can provide parallel or sequential processing / execution of one or more portions of the techniques described herein.

[0072] Those skilled in the art will understand that computer system 500 is merely illustrative and not intended to limit the scope of the techniques described herein. Computer system 500 may include any combination of devices or software capable of performing or otherwise providing the performance of the techniques described herein. For example, computer system 500 may include or be a combination of cloud computing systems, data centers, server racks, servers, virtual servers, desktop computers, laptop computers, tablet computers, server equipment, client devices, mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, in-vehicle computers, global positioning systems (GPS), etc. Computer system 500 may also be connected to other devices not shown, or may operate as a standalone system. Furthermore, in some embodiments, the functionality provided by the illustrated components may be combined in fewer components or distributed across other components. Similarly, in some embodiments, some of the functionality of the illustrated components may not be provided, or additional functionality may be provided.

[0073] Operating procedures

[0074] Figure 6 This is a flowchart 600 of the operation used to generate a composite frame of objects detected in multiple different types of data streams. Figure 6 The operation can be used with regard to Figure 5 The components described. In some embodiments, the image processing system 102 may include one or more components of the computing system 500. At 602, the image processing system 102 receives an identifier of a target location within an image. For example, the image processing system may receive the identifier from an operator or command and control center. The image processing system 102 may receive the identifier via network 150 using network interface 540. In some embodiments, the image processing system 102 may receive the identifier from data node 104.

[0075] At 604, image processing system 102 inputs an image into a machine learning model to obtain the orientation of an object within the target. For example, image processing system 102 may use one or more processors 510a, 510b, and / or 510n to perform the input. At 606, image processing system 102 generates a set of dimensions associated with the object based on its orientation. For example, image processing system 102 may use one or more processors 510a-510n to perform this operation, and the set of dimensions may be stored in system memory 520.

[0076] At 608, the image processing system 102 receives multiple images indicating the target location. The image processing system 102 can receive the images via a wired or wireless connection between a recording device (e.g., a camera) and the hardware hosting the image processing system. At 610, the image processing system 102 generates multiple sets of image metadata for the multiple images. Each image may include the corresponding position of the camera in three-dimensional space at the time of recording each image. The image processing system 102 can perform this operation using one or more processors 510a, 510b and / or 510n and / or system memory 520.

[0077] At 612, image processing system 102 determines multiple position estimates of an object in three-dimensional space based on multiple images. Image processing system 102 may perform this operation using one or more processors 510a, 510b, and / or 510n, and store the position estimates in system memory 520. At 614, image processing system 102 determines the position of an object in three-dimensional space based on the multiple position estimates. Image processing system 102 may perform this operation using one or more processors 510a, 510b, and / or 510n, and store the position estimates in system memory 520.

[0078] Although the invention has been described in detail for illustrative purposes based on embodiments currently considered to be the most practical and preferred, it should be understood that such details are for that purpose only, and the invention is not limited to the disclosed embodiments, but rather is intended to cover modifications and equivalent arrangements within the scope of the appended claims. For example, it should be understood that the invention contemplates that, to the extent possible, one or more features of any embodiment may be combined with one or more features of any other embodiment.

[0079] The embodiments described above are presented for illustrative purposes and not for limitation, and this disclosure is limited only by the appended claims. Furthermore, it should be noted that the features and limitations described in any embodiment can be applied to any other embodiment herein, and flowcharts or examples associated with one embodiment can be combined with any other embodiment in a suitable manner, performed in a different order, or performed in parallel. Additionally, the systems and methods described herein can be performed in real time. It should also be noted that the above systems and / or methods can be applied to other systems and / or methods, or used according to other systems and / or methods.

[0080] The present invention will be better understood by referring to the following examples:

[0081] 1. A method comprising: receiving an identifier of a target location within an image; inputting the image into a machine learning model to obtain an orientation of an object within the target location, wherein the machine learning model is trained to detect objects within the received image; generating a set of real-world dimensions associated with the object based on the object's orientation; receiving a plurality of images from a recording device mounted on an unmanned vehicle, wherein each of the plurality of images indicates the target location; generating a plurality of image metadata sets for the plurality of images, wherein each set of image metadata includes a corresponding position of a camera in a three-dimensional space at the time of recording each image; determining a plurality of position estimates of the object in the three-dimensional space, wherein each of the plurality of position estimates is generated for a corresponding image among the plurality of images based on the set of real-world dimensions and the corresponding set of metadata; and determining the position of the object in the three-dimensional space based on the plurality of position estimates.

[0082] 2. According to any of the foregoing embodiments, determining the position of an object in three-dimensional space based on multiple position estimates further includes: retrieving multiple latitude coordinates, multiple altitudes and multiple longitude coordinates from the multiple position estimates; and generating the position of the object in three-dimensional space based on the average latitude coordinates, average altitudes and average longitude coordinates.

[0083] 3. According to any of the foregoing embodiments, generating a plurality of image metadata sets further includes: inferring a known size associated with an object from a first image based on the object's orientation; determining a dimension modification factor based on one or more known sizes and the set of real-world sizes; and generating a size value of an unknown size associated with the object based on the size modification factor to generate the image size of the object in the first image.

[0084] 4. According to any of the foregoing embodiments, the method further includes: determining that the orientation of the object has changed; updating the size modification factor based on the determination that the orientation of the object has changed; and updating the known dimensions associated with the object to generate an updated set of dimensions.

[0085] 5. According to any of the foregoing embodiments, determining that the orientation of the object has changed further includes: inputting a first image from a plurality of images into a machine learning model; receiving an updated orientation of the object from the machine learning model; and determining that the orientation does not match the updated orientation.

[0086] 6. According to any of the foregoing embodiments, the method further includes: identifying a second object within the target location; determining that the second object does not have a corresponding known size; comparing a first image size of the object and a second image size of the second object based on the determination that the second object does not have a corresponding known size; determining a second size modification factor of the second object based on the comparison of the first image size and the second image size; and determining a second three-dimensional position of the second object based on the second size modification factor.

[0087] 7. According to any of the foregoing embodiments, wherein determining the position of an object in three-dimensional space based on multiple position estimates further includes: sorting the multiple position estimates based on corresponding timestamps; determining whether the multiple position estimates converge to a given value over time; and generating a first command for the recording device to record more images and a second command for the unmanned vehicle to perform more operations based on the determination that the multiple position estimates do not converge to a given time value.

[0088] 8. According to any of the foregoing embodiments, the method further includes: determining that the object is moving; generating a first command for stopping the operation of the unmanned vehicle based on determining that the object is moving; generating a second command for the recording device to record more images; and adjusting the position of the object based on the movement of the object.

[0089] 9. According to any of the foregoing embodiments, further comprising: detecting a set of unmanned vehicles capable of recording images of a target location; sending a command to the set of unmanned vehicles to establish point-to-point communication; receiving additional images and additional image metadata of the target location from the set of unmanned vehicles; and using the additional images and additional image metadata to determine an additional object location estimate.

[0090] 10. A tangible, non-transitory, machine-readable medium storing instructions that, when executed by a data processing apparatus, cause the data processing apparatus to perform operations including any one of embodiments 1-9.

[0091] 11. A system comprising: one or more processors; and a memory storing instructions that, when executed by the processor, cause the processor to perform operations including any of embodiments 1-9.

[0092] 12. A system comprising components for performing any one of embodiments 1-9.

[0093] 13. A system comprising cloud-based circuitry for performing any one of embodiments 1-9.

Claims

1. A system for estimating the position of an object in three-dimensional space and its orientation within an image, the system comprising: One or more processors; and A non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: Receive the identifier of the target location within the image recorded by the unmanned vehicle at the unmanned vehicle; The image is input into a machine learning model to obtain object identifiers and the orientation of objects within the target location, wherein the machine learning model is trained to detect objects within the received image; The set of real-world dimensions associated with the object is determined based on the object identifier; The image stream, comprising multiple images, is received from a camera mounted on the unmanned vehicle, each of the multiple images showing the target location, wherein the image stream is recorded by the camera as the unmanned vehicle moves relative to the object; Multiple image metadata sets are generated for the multiple images, wherein each image metadata set includes the orientation of the object, the image size of the object, camera data associated with the camera, the orientation of the camera, and the position of the camera in three-dimensional space when recording each image; Determine multiple position estimates of the object in the three-dimensional space, wherein each of the multiple position estimates is generated for a corresponding image among the multiple images based on a corresponding set of image metadata; and The position of the object in the three-dimensional space is determined based on the multiple position estimates.

2. The system of claim 1, wherein, Instructions for determining the position of the object in the three-dimensional space based on the plurality of position estimates also cause the one or more processors to perform operations including: Retrieve multiple latitude coordinates, multiple altitudes, and multiple longitude coordinates from the multiple location estimates; and The position of the object in the three-dimensional space is generated based on the average latitude coordinates, average altitude, and average longitude coordinates.

3. The system of claim 1, wherein, The instructions for generating the plurality of image metadata sets also cause the one or more processors to perform operations including the following: The known dimensions associated with the object are inferred from the first image based on the object's orientation; The size modification factor is determined based on one or more known dimensions and the set of real-world dimensions; as well as Generate a size value of an unknown size associated with the object based on the size modification factor, so as to generate the image size of the object in the first image.

4. The system of claim 3, wherein, Instructions for inferring known dimensions associated with the object from the first image also cause the one or more processors to perform operations including: Determine the size of the first object with the known dimensions; A first real-world size matching the size of the first object is determined based on the orientation of the object. as well as Assign a size label to the first object size.

5. The system of claim 1, wherein, The instructions for generating the plurality of image metadata sets also cause the one or more processors to perform operations including the following: The first image from the plurality of images is input into the machine learning model; Receive the object identifier and the direction of the object's update from the machine learning model; as well as The updated orientation is added to the corresponding set in the plurality of image metadata sets.

6. The system of claim 1, wherein, The instructions also cause the one or more processors to perform operations including the following: Identify a second object within the target location; It is determined that the second object does not have a corresponding known size; Based on the determination that the second object does not have the corresponding known size, the first image size of the object and the second image size of the second object are compared; A second size modification factor for the second object is determined by comparing the size of the first image and the size of the second image. as well as The second three-dimensional position of the second object is determined based on the second size modification factor.

7. The system of claim 1, wherein, Instructions for determining the position of the object in the three-dimensional space based on the plurality of position estimates also cause the one or more processors to perform operations including: The multiple location estimates are sorted based on their corresponding timestamps; Determine whether the plurality of location estimates converge to a given value over time; as well as Based on the determination that the multiple position estimates do not converge to a given time value, a first command is generated for the camera to record more images and a second command is generated for the unmanned vehicle to perform more maneuvers.

8. The system of claim 1, wherein, The instructions also cause the one or more processors to perform operations including the following: It is determined that the object is moving; Based on the determination that the object is moving, a first command is generated for the unmanned vehicle to stop operation; Generate a second command for the camera to record more images; as well as The position of the object is adjusted based on the object's movement.

9. The system of claim 1, wherein, The instructions also cause the one or more processors to perform operations including the following: Detect a collection of unmanned vehicles capable of recording one or more images of the target location; Commands are sent to the set of unmanned vehicles to establish point-to-point communication; Receive additional images and additional image metadata of the target location from the set of unmanned vehicles; as well as The additional image and its metadata are used to determine the location estimate of the additional object.

10. A method comprising: Receive the identifier of the target location within the image recorded by the unmanned vehicle at the unmanned vehicle; The image is input into a machine learning model to obtain object identifiers and the orientation of objects within the target location, wherein the machine learning model is trained to detect objects within the received image; The set of real-world dimensions associated with the object is determined based on the object identifier; Multiple images are received from a recording device mounted on the unmanned vehicle, each of the multiple images showing the target location; Multiple image metadata sets are generated for the multiple images, wherein each image metadata set includes the orientation of the object, the image size of the object, the orientation of the camera, and the position of the camera in three-dimensional space when recording each image; Determine multiple position estimates of the object in the three-dimensional space, wherein each of the multiple position estimates is generated for a corresponding image in the multiple images based on a corresponding set of image metadata; as well as The position of the object in the three-dimensional space is determined based on the multiple position estimates.

11. The method of claim 10, wherein, Determining the position of the object in the three-dimensional space based on the multiple position estimates also includes: Retrieve multiple latitude coordinates and multiple longitude coordinates from the multiple location estimates; and The position of the object in the three-dimensional space is generated based on the average latitude coordinates and the average longitude coordinates.

12. The method of claim 10, wherein, Generating the multiple image metadata sets also includes: The known dimensions associated with the object are inferred from the first image based on the object's orientation; The size modification factor is determined based on one or more known dimensions and the set of real-world dimensions; and Generate a size value of an unknown size associated with the object based on the size modification factor, so as to generate the image size of the object of the first image.

13. The method of claim 12, further comprising: It has been determined that the orientation of the object has changed; Based on the determination that the orientation of the object has changed, the size modification factor is updated; as well as Update the known dimensions associated with the object to generate an updated set of dimensions.

14. The method of claim 10, wherein, Determining that the orientation of the object has changed also includes: The first image from the plurality of images is input into the machine learning model; Receive the updated orientation of the object from the machine learning model; and It is determined that the orientation and the updated orientation do not match.

15. The method of claim 10, further comprising: Identify a second object within the target location; It is determined that the second object does not have a corresponding known size; Based on the determination that the second object does not have the corresponding known size, the first image size of the object and the second image size of the second object are compared; A second size modification factor for the second object is determined by comparing the size of the first image and the size of the second image. as well as The second three-dimensional position of the second object is determined based on the second size modification factor.

16. The method of claim 10, wherein, Determining the position of the object in the three-dimensional space based on the multiple position estimates also includes: The multiple location estimates are sorted based on their corresponding timestamps; Determine whether the plurality of location estimates converge to a given value over time; and Based on the determination that the plurality of position estimates do not converge to a given time value, a first command is generated for the recording device to record more images and a second command is generated for the unmanned vehicle to perform more maneuvers.

17. The method of claim 10, further comprising: It is determined that the object is moving; Based on the determination that the object is moving, a first command is generated for the unmanned vehicle to stop operation; Generate a second command for the recording device to record more images; as well as The position of the object is adjusted based on the movement of the object.

18. The method of claim 10, wherein, Also includes: Detect a set of unmanned vehicles capable of recording images of the target location; Commands are sent to the set of unmanned vehicles to establish point-to-point communication; Receive additional images and additional image metadata of the target location from the set of unmanned vehicles; as well as The additional image and its metadata are used to determine the location estimate of the additional object.

19. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: Receive the identifier of the target location within the image; The image is input into a machine learning model to obtain object identifiers and the orientation of objects within the target location, wherein the machine learning model is trained to detect objects within the received image; Based on the object identifier and the object's orientation, generate a set of real-world dimensions associated with the object; Receive multiple images from a recording device, each of which shows the target location; Multiple image metadata sets are generated for the multiple images, wherein each image metadata set includes the position of the recording device in three-dimensional space when recording each image; Determine multiple position estimates of the object in the three-dimensional space, wherein, for a corresponding image among the multiple images, each position estimate among the multiple position estimates is generated based on the set of real-world dimensions, the corresponding position of the object in the corresponding image, and the corresponding position of the recording device; and The position of the object in the three-dimensional space is determined based on the multiple position estimates.

20. The non-transitory computer-readable medium of claim 19, wherein the instructions further cause the one or more processors to perform operations including: Detect a set of unmanned vehicles capable of recording images of the target location; Commands are sent to the set of unmanned vehicles to establish point-to-point communication; Receive additional images and additional image metadata of the target location from the set of unmanned vehicles; as well as The additional image and its metadata are used to determine the location estimate of the additional object.