Devices, systems and / or methods for passive optical radar sensing
Patent Information
- Application Number
- PCT/AU2026/050154
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-26
- Filing Date
- 2026-02-25
- Publication Date
- 2026-09-03
Smart Images

Figure AU2026050154_03092026_PF_FP_ABST
Abstract
Description
DEVICES, SYSTEMS AND / OR METHODS FOR PASSIVE OPTICAL RADAR SENSINGFIELD
[0001] The present disclosure relates generally to devices, systems and / or methods that may be used for detecting and / or tracking one or more targets.BACKGROUND
[0002] The increased use of targets, (for example, but not limited to, drones and / or unmanned vehicles) has led to an increased need for accurate and / or reliable systems for detecting and / or tracking targets in a three-dimensional (3D) space. The targets, (for example, but not limited to, drones and / or unmanned vehicles) which may be aerial, marine, and / or ground-based, present challenges for detection and / or tracking due to at least one or more of the following: varying sizes, speeds, and operational environments.
[0003] Active radar systems are widely used for detecting and tracking objects, including drones. Many of these systems emit radio waves and analyse the reflected signals to determine the position and movement of the target. While active radar may be effective in certain scenarios, it has several limitations. The emitted signals may be detected by the target, potentially alerting it to the presence of the tracking system. Additionally, active radar systems may be affected by environmental factors, for example, weather conditions and / or terrain, which may degrade their accuracy and / or reliability.
[0004] Laser ranging, or Light Detection and Ranging (LIDAR), is another known method for detecting and tracking drones. LIDAR systems use laser beams to measure the distance to the target by analyzing the time it takes for the light to return after hitting the object. The use of LIDAR may provide high-resolution data and may be effective in clear weather conditions. However, LIDAR systems are often expensive and may struggle in adverse weather conditions such as fog, rain, and / or snow.
[0005] Passive radar systems, which rely on analysing signals from existing sources such as commercial broadcast stations, offer a stealthier alternative to active radar. These systems do not emit their own signals, making them less detectable by the target.However, passive radar systems may be less accurate and reliable than active radar, as they typically depend on the availability and quality of external signals. Furthermore, the complexity of processing these signals may lead to increased computational requirements and / or potential delays in detection and / or tracking.
[0006] Monocular camera-based estimation techniques may involve determining the range of a target based on its known size. For example, if a specific drone model is identified the system may estimate its distance based on the apparent size of the drone in the captured image. While this method may be useful in certain situations, it often has significant drawbacks. The accuracy of scale-based estimation is highly dependent on the correct identification of the target model, which may be challenging in real-world scenarios with diverse and rapidly evolving drone designs and is often difficult at longer ranges. Another variation of the monocular camera-based estimation is the use of Al processing to estimate range to a target based learnt model or the targets and / or background scene. Again, these approaches may be inaccurate where there is not enough context in the scene or around the target to enable a reliable estimate of the range.
[0007] In certain embodiments, the systems and / or methods disclosed herein may use a set of cameras that are configured to 3D detect and / or track one or more targets using triangulation to locate the one or more the targets. By capturing images from different angles and analyzing the data to determine the precise position of the one or more targets in three-dimensional space, the devices, systems and / or methods disclosed herein offers several advantages. For example, at least one of the exemplary disclosed embodiments provides acceptable levels of accuracy and / or reliability without the need for emitting detectable signals, making the exemplary embodiment difficult for an adversary to detect. Additionally, the use of multiple cameras allows for detection and tracking in various environmental conditions and may handle a wide range of target sizes and / or shapes.
[0008] The present disclosure is directed to overcome and / or ameliorate at least one or more of the disadvantages of the prior art, as will become apparent from the discussionherein. The present disclosure also provides other advantages and / or improvements as discussed herein.
[0009] It is desired to address or alleviate one or more disadvantages or limitations of the prior art, or to at least provide a useful alternative.SUMMARY
[0010] This summary is not intended to be limiting as to the embodiments disclosed herein and other embodiments are disclosed in this specification. In addition, limitations of one embodiment may be combined with limitations of other embodiments to form additional embodiments.
[0011] Certain embodiments are directed to a system fortracking in a real-life scene one or more targets in three dimensions comprising:at least a first sensor that is configured to generate at least a first two-dimensional (2D) image of the real-life scene and at least a second sensor that is configured to generate at least a second two-dimensional image of the real-life scene, wherein the two-dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene;a processor that is configured to:receive the at least first two-dimensional image and the at least second two-dimensional image;detect, from the images if one or more target candidates are present in the images; determine that a candidate from the at least first image matches with a candidate from the at least second image; and determine the location of the one or more targets in the real-life scene.
[0012] Certain embodiments are directed to one or more computer-readable non-transitory storage media embodying software that is operable when executed using any of the systems or methods disclosed herein.
[0013] Certain embodiments are directed to a method for tracking in a real-life scene one or more targets in three dimensions comprising:configuring at least a first sensor to generate at least a first two-dimensional image of the real-life scene and at least a second sensor to generate at least a second two-dimensional image of the real-life scene;generating the at least first two-dimensional image of the real-life scene from the at least first sensor and the at least second two-dimensional image of the real-life scene from the at least second sensor, wherein the two-dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene; sending the at least first two-dimensional image and the at least second two-dimensional image to the processor;receiving the at least first two-dimensional image and the at least second two-dimensional image at the processor;detecting, from the received images if one or more target candidates are present in the images;determining that a candidate from the at least first image matches with a candidate from the at least second image; anddetermining the location of the one or more targets in the real-life scene.BRIEF DESCRIPTION OF DRAWINGS
[0014] One or more embodiments of the present invention are hereinafter described, by way of example only, with reference to the accompanying drawings in which:
[0015] FIG. 1 is a top-level system diagram for a passive optical radar, according to at least one embodiment.
[0016] FIG. 2 is an illustration of an exemplary real-world scene observed by a sensor array, according to at least one embodiment.
[0017] FIG. 3 is a flow chart of the process to calibrate, deploy and operate the system, according to at least one embodiment.
[0018] FIG. 4 is a flow chart of a process for onsite calibration of the system, according to at least one embodiment.
[0019] FIG. 5 is a flow chart of a process for operation of the system, according to at least one embodiment.
[0020] FIG. 6 illustrates how a pixel-based background model is built, according to at least one embodiment.
[0021] FIG. 7 shows output displayed to assist evaluation of calibration, according to at least one embodiment.
[0022] FIG. 8 is an illustration of 2D tracking and redetection, according to at least one embodiment.
[0023] FIG. 9 is an illustration of stereo matching, according to at least one embodiment.
[0024] FIG. 10 is an illustration of lost track recovery, according to at least one embodiment.
[0025] FIG. 11 shows an exemplary user interface for the system including a representative view, a map view and a target list, according to at least one embodiment.DETAILED DESCRIPTION
[0026] The subject headings used in the detailed description are included only for the ease of reference of the reader and should not be used to limit the subject matter found throughout the disclosure or the claims. The subject headings should not be used in construing the scope of the claims or the claim limitations.
[0027] Certain embodiments may be useful for the detection, identification, ranging and / or tracking of targets, for example but not limited to, drones, aircraft, airplanes, rockets, helicopters, unmanned aerial vehicles, birds, other flying objects, people, animals, motorcycles, cars, trucks, armoured vehicles, tanks, trains, unmanned ground vehicles, other ground-based objects, boats, yachts, ferries, ships, surfaced submarines, unmanned surface vessels, other surface vessels, or combinations thereof.
[0028] In certain embodiments, these capabilities may be used in a system to provide alerts regarding the presence, approach and / or behaviour of targets of interest.
[0029] In certain embodiments, the system may be used provide real time target information to other systems including, but not limited to, systems for commissioning of vehicles orvessels, systems for managing a space (e.g., a port, an airport, an airspace), effector systems (e.g., projectile, directed energy, intercept systems).
[0030] Certain embodiments may be used to inform and / or or warn of potential threats. Certain embodiments may be used to guide an effector (e.g. projectile, directed energy, intercept systems) to engage with identified threats.
[0031] Certain embodiments may be used to identify one or more targets within a real-life scene.
[0032] Certain embodiments may be used to track over a period of time one or more targets within a real-life scene (for example they may track a target for a period of time between 30ms to 1000ms or between 1 to 10 seconds, or between 10 and 60 seconds, or between 1 minute and 10 minutes, or between 10 minutes and 60 minutes).
[0033] The term “scene” (also referred to as a “real-life scene” ) means a subset of the three dimensional real-world (i.e. , 3D physical reality) as perceived through the field of view of one or more sensors. In certain embodiments, there may be at least 2, 3, 4, 5, 10, 15, 20, 25, 30, 35 or 40, 100, 1000, or more sensors.
[0034] The term “target” means an element in a scene that is of interest. A target may be an element in the scene that is at least one of the following: stationary, substantially stationary, moving or may be an element in the scene that may move in the future or may have moved in the past or is capable of moving. A target may be a drone, aircraft, airplane, rockets, helicopter, unmanned aerial vehicle, bird, other flying object, person, animal, motorcycle, car, truck, armoured vehicle, tank, train, unmanned ground vehicle, other ground-based object, boat, yacht, ferry, ship, surfaced submarine, unmanned surface vessel, other surface vessel, or combinations thereof.
[0035] The term “3D point” or “3D coordinates” means a representation of the location of a point in the scene defined at least in part by at least three parameters that indicate distance in three dimensions from an origin reference to the point, for example, in three directions from the origin where the directions may be substantially perpendicular (at leastnot co-planar or co-linear), or as an alternative example using a spherical coordinate system consisting of a radial distance, a polar angle, and an azimuthal angle.
[0036] The term “each” as used herein means that at least 95%, 96%, 97%, 98%, 99% or 100% of the items or functions referred to perform as indicated.
[0037] Exemplary items or functions include, but are not limited to, one or more of the following: location(s), image pair(s), cell(s), pixel(s), pixel location(s), layer(s), element(s), point(s), neighbourhood(s), and 3D point(s).
[0038] The term “at least a substantial portion” as used herein means that at least 60%, 70%, 80%, 85%, 95%, 96%, 97%, 98%, 99%, or 100% of the items or functions referred to. Exemplary items or functions include, but are not limited to, one or more of the following: location(s), image pair(s), cell(s), pixels(s), pixel location(s), layer(s), element(s), point(s), neighbourhood(s), and 3D point(s).Certain Exemplary Advantages
[0039] In addition to other advantages disclosed herein, one or more of the following advantages may be present in certain exemplary embodiments:
[0040] One advantage may be that the 3D coordinates of one or more targets may be determined without use of active signals enabling the system to avoid (or substantially avoid) being detected and / or located.
[0041] One advantage may be that the 3D coordinates of one or more targets may be determined with insubstantial use of active signals enabling the system to avoid (or substantial avoid) being detected and / or located.
[0042] Another advantage may be that the 3D coordinates of one or more targets may be determined without use of active signals that may interfere with other ranging systems.
[0043] Another advantage may be that the detection and / or tracking of a number of targets may be performed at the same, or substantially the same, time.
[0044] Another advantage may be that the detection and tracking of several targets may be performed where those targets pass closely to each other.
[0045] Another advantage may be that the 3D coordinates of a target may be determined with increased accuracy (which may be in terms of range or angle from the passive optical sensor system or both).
[0046] Another advantage may be that the 3D coordinates of a target may be determined with increase speed, reduced latency and / or reduced processing.
[0047] Another advantage may be that calibration steps may be performed in the field and may be more effective and / or efficient in determining an accurate calibration for the system.
[0048] Another advantage may be that the system may be updated or refined in the field quickly and efficiently enabling reducing time and increasing operational availability of the system.
[0049] Another advantage may be that the system may be moved, or the system configuration may be changed, and in field calibration re-performed to determine a new calibration for the system.
[0050] Another advantage may be that accuracy of the system can be improved or refined by adding new observed data while in the field.
[0051] System Diagram
[0052] FIG. 1 shows a system diagram 100 of an exemplary embodiment. The system includes a sensor array 101 and a computer system 102.
[0053] The sensor array 101 includes sensors 110, 112, 115, 120, 122.
[0054] In certain embodiments, the sensor array 101 consists of a single type of sensor (e.g., visible light camera, RGB camera, thermal infrared camera, short-wave infrared camera, mediumwave infrared camera). In certain embodiments, the sensor array 101 may include related circuitry to ensure synchronised capture of data from a portion, a substantial portion, or all of the sensors of the sensor array 101. In certain embodiments, the sensor array 101 may include related circuitry to ensure synchronised capture of data from one or more of the sensors of the sensor array. In certain embodiments the sensor array 101 may comprise two sensors. In certain embodiments the sensor array 101 may comprise between2 and 10 sensors. In certain embodiments the sensor array 101 may comprise between 10 and 100 sensors. In certain embodiments, the sensor array 101 may be comprise of at least two or more different types of sensors. For example, the sensor array 101 may be made up of RGB cameras and thermal infrared cameras or RGB cameras, shortwave infrared cameras and longwave infrared cameras, or other combinations of visible light cameras, RGB cameras, shortwave infrared cameras, medium wave infrared cameras, long wave infrared cameras.
[0055] The computer system 102 includes a receiving unit 150 for communication with the sensors in the sensor array 101. The receiving unit 150 is connected via communication bus 151 with the processor unit 160, and via communication bus 152 with a memory unit 170. The processor unit 160 may be a general-purpose Central Processing Unit (CPU) or Graphics Processing Unit (GPU) or may be customised hardware such as a Field-Programmable Gate Array (FPGA) or Application-Specific Integrated Circuit (ASIC) designed to perform the required processing. In certain embodiments, the processing unit may comprise a number of processing elements including multiple CPUs, and / or GPUs and / or customised hardware. The memory unit 170 may include volatile and / or non-volatile memory. The memory unit 170 and the processor unit 160 may be communicatively connected via a communications bus 161. The memory unit 170 may store instructions for the processor unit 160 as well as image data received from the receiving unit 150. The processor unit 160 may also be connected to a data store 190 via communications bus 162. The processor unit 160 may also be connected to (or in communication with) an external communications unit 180 via communications bus 163. The memory unit 170 may also be connected to (or in communication with) the external communications unit 180 via communications bus 171. The external communications unit 180 may be used to output a data stream for the use of external systems via communication mechanism 182, 181 (which may be by wired communication such as ethernet or wireless communication such as Wi-Fi or radio link). External systems may include Command and Control (“C2”) systems, security management systems, alarm systems, effector systems (e.g., projectile, directed energy,intercept systems) and counter-UAS systems. The external communications unit 180 may also receive data from external sources including position data, map data and / or previously recorded 3D data regarding the scene.
[0056] Sensors in the sensor array 101 may be connected to the computer system 102. Sensors may have a communication channel indicated for example by 111, 121 to output data and / or to accept control and / or synchronisation signals. Synchronous capture of the sensor output may be useful and may be enabled by the communication channels 111, 121.The communication channels 111, 121 may be wired using connection such as ethernet, GiGE Vision, USB, IEEE1394 or GMSL. The communication channels 111, 121 may be wireless such as WiFi, cellular or radio-based communication.
[0057] The computer system 102 may receive and may process data from sensors in the sensor array 101 to enable the detection and tracking of targets in a scene.
[0058] FIG. 1 shows an exemplary system 100, according to at least one embodiment. FIG. 1 includes an exemplary configuration of a sensor array 101 and an exemplary computer system 102. In certain embodiments, one or more computer systems perform one or more steps of one or more methods described or disclosed herein. In certain embodiments, one or more computer systems provide functionality described or shown in this disclosure. In certain embodiments, software configured to be executable running on one or more computer systems performs one or more steps of one or more methods disclosed herein and / or provides functionality disclosed herein. Reference to a computer system may encompass a computing device, and vice versa, where appropriate.
[0059] This disclosure contemplates a suitable number of computer systems. As example and not by way of limitation, computer system 102 may be an embedded computer system, a system-on- chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a main-frame, amesh of computer systems, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of thereof. Where appropriate, computer system 102 may include one or more computer systems; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centres; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 102 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, one or more computer systems 102 may perform in real time or in batch mode one or more steps of one or more methods disclosed herein.
[0060] The computer system 102 may include a processor unit 160, memory unit 170, data store 190, a receiving unit 150, and an external communications unit 180.
[0061] The processor unit 160 may include hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor unit 160 may retrieve the instructions from an internal register, an internal cache, memory unit 170, or data store 190; decode and execute them; and then write one or more results to an internal register, an internal cache (not shown), memory unit 170, or data storage 190. The processor unit 160 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor units 160 including a suitable number of suitable internal caches, where appropriate. The processor unit 160 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory unit 170 or data store 190, and the instruction caches may speed up retrieval of those instructions by processor unit 160.
[0062] The memory unit 170 may include main memory for storing instructions for processor to execute or data for processor to operate on. The computer system 102 may load instructions from data store 190 or another source (such as, for example,another computer system) to memory unit 170. The processor unit 160 may then load the instructions from memory unit 170 to an internal register or internal cache. T o execute the instructions, the processor unit 160 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, the processor unit 160 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. The processor unit 160 may then write one or more of those results to the memory unit 170. The processor unit 160 may execute only instructions in one or more internal registers or internal caches or in the memory unit 170 (as opposed to data store 190 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory unit 170 (as opposed to data store 190 or elsewhere). One or more memory buses may couple processor unit 160 to memory unit 170. The Bus (not shown) may include one or more memory buses. The memory unit 170 may include random access memory (RAM). This RAM may be volatile memory, where appropriate Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM).Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. Memory unit 170 may include one or more memories, where appropriate.
[0063] The data store 190 may include mass storage for data or instructions. The data store 190 may include a hard disk drive (HDD), flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination therein. Data store 190 may include removable or non-removable (or fixed) media, where appropriate. Data store 190 may be internal or external to computer system 102, where appropriate. Data storage may include read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination thereof.
[0064] In certain embodiments, I / O interface (not shown) may include hardware, software, or both, providing one or more interfaces for communication betweencomputer system and one or more I / O devices. Computer system may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system. An I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination thereof. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces for them. Where appropriate, I / O interface may include one or more device or software drivers enabling the processor unit 160 to drive one or more of these I / O devices. I / O interface may include one or more I / O interfaces, where appropriate. In one embodiment, the computer system 102 may further comprise a display 185 and controls 186 to control information presented on the display.Exemplary Illustrative Scene
[0065] FIG. 2 shows an exemplary scene in the real world 200 according to at least one embodiment. In the scene may be some elements that are largely stationary for example the trees 270 and fence 250. In the scene there are targets including aerial drones 260, 220 and a ground vehicle 240 and a person 245. There may be some markers 230, 232 placed into the scene. A sensor array 101 is shown oriented to observe the scene. In this example, the sensor array 101 is illustrated with two sensors 110 and 120 and fields of view for these sensors illustrated with dashed lines projecting out from the sensors 110, 120 into the scene. The sensor array 101 may be located on a vehicle (not shown) and may itself be moving. A 3D co-ordinate system 210 (e.g., axes 210 indicating a reference position and orientation) may be used to describe the locations of targets in the scene. The sensor array 101 is placed into the field with sensors, for example sensors 120 and 110, separated with an approximate baseline 215.Overall Process
[0066] FIG. 3 shows an exemplary process 300 for a passive optical radar system, according to at least one embodiment.
[0067] Starting from 305, the first step in the process 300 is to perform the offsite calibration (step 310). Typically, the offsite calibration may include intrinsic calibration of each sensor e.g., 110, 120 in the sensor array 101. Intrinsic calibration may be the determining of the intrinsic parameters of the sensor such as focal length, principal point, and / or lens distortion coefficients. Intrinsic calibration is a known procedure in the art and there are number of methods that may be applied. In certain embodiments, the intrinsic parameters may include a lens distortion calibration to characterise distortion in the sensors. In certain embodiments, intrinsic calibration may be a manual process. In certain embodiments, intrinsic calibration may be an automated process and may be performed by the computer system 102 working in conjunction with the sensor array 101. In certain embodiments, intrinsic calibration may use test charts or specialised targets with known size and / or configuration. Typically, the offsite calibration (step 310) may be performed as a late stage of product manufacturing. In certain embodiments, the intrinsic calibration may be performed in the field. In certain embodiments, the intrinsic calibration may be performed initially offsite and then refined in the field. Following the offsite calibration (step 310), the exemplary embodiment proceeds to deploy (step 320).
[0068] At step 320, the passive optical radar system may be deployed in the field. The sensor array 101 is placed into the field with sensors, for example sensors 120 and 110, separated with an approximate baseline 215. Typically, the length of the baseline may depend on the application of one or more of the following: the passive optical radar, the required range, the required field of view, the number of sensors in the sensor array 101 , physical size and / or space constraints of the deployment, or other factors. In certain embodiments, the baseline maybe 10cm, 20cm, 50cm or 100cm. In certain embodiments the baseline may be 1m, 2m, 5m or 10m. In certain embodiments the baseline may be 10m, 20m, 50m or 100m. In certain embodiments the baseline may be 200m, 500m or 1000m. The sensors may be oriented so that the field of view of one or more of the sensors covers a portion of the scene 200 so that there may be overlap in the field of view of the one or more sensors. In certain embodiments, the one or more sensors of the sensor array 101may be separately mounted onto tripods or other stable supports and may be deployed by mounting the one or more sensors and tripods (or other supports) to the ground, buildings or other structures with approximately the baseline between sensors in the sensor array 101. In certain embodiments, the sensor array 101 may comprise a ridged structure to support the sensors and the sensor array 101 may be placed on the ground or mounted to a vehicle or a building or other suitable support.
[0069] Following deployment (step 320), the exemplary embodiment process 300 proceeds to the onsite calibration (step 330). At the onsite calibration step 330, an extrinsic calibration of the sensors in the sensor array 101 may be performed. Extrinsic calibration involves determining the relative position and orientation between the sensors in the sensor array 101. A process for the onsite calibration is described elsewhere in this specification.Following the onsite calibration (step 330), the exemplary embodiment process 300 proceeds to operate (step 340).
[0070] At the operate step 340, the system 100 may operate as a passive optical radar. The system 100 uses the data from the sensor array 101 to detect and locate targets in the scene. The detailed description of this step is provided elsewhere in this specification. In certain embodiments, the targets are displayed as annotations on a map view of the scene. In certain embodiments, target’ s 3D coordinates, geolocation, bearing, range, altitude, speed, direction of travel, size, classification and / or other information, combinations thereof may be displayed on the display 185. In certain embodiments, an operator may select which information may be displayed on the display 185 using controls 186, which may be buttons, touch panel screen, keyboard, mouse, voice controls or other mechanisms to input user control into the system 100. In certain embodiments, data describing the target’ s 3D coordinates, geolocation, bearing, range, altitude, speed, direction of travel, size, classification and / or other information may be transmitted to an external system via the external communications unit 180. A user interface presented on display 185 is described indetail elsewhere in this specification. The operate step 340 may continue indefinitely or until an operator terminates operation (step 350) using controls 186.Onsite Calibration Process
[0071] FIG. 4 shows an exemplary process 400 for performing the onsite calibration 330, according to at least one embodiment. Starting from 410, the first step is determine reference coordinates (step 420).
[0072] At the step 420 to determine reference coordinates, a 3D co-ordinate system 210 that may be used to describe the locations of targets in the scene may be determined. In certain embodiments, the co-ordinate system 210 may use the location of the sensor array 101 as a reference location. In certain embodiments, the co-ordinate system 210 selected may use the location of one of the sensors, for example sensor 110 as reference location. In certain embodiments, the co-ordinate system 210 selected may use the reference location as its origin. In certain embodiments, an operator may determine the world coordinates e.g., latitude and longitude, of the reference location. In certain embodiments, GPS / GNNS may be used to determine world coordinates e.g., latitude and longitude, of the reference location. In certain embodiments, the world coordinates may be one or more of the following: latitude & longitude, grid reference. If the world coordinates of the reference location are known, then targets in the scene may be geo-located in that world coordinate system. If the system 100 is deployed without determining the world coordinates of the reference location, then then targets in the scene may be located with reference to the location of the reference location. Following the step 420 of determining reference coordinates, the exemplary embodiment process 400 proceeds to step 425 for determining distance to calibration points.
[0073] At the step 425 for determining calibration point distances, measurements may be made to determine calibration point distances. A calibration point distance may be the distance from the reference point to a calibration point in the scene. Calibration points in the scene may be markers 230, 232 placed in the scene, natural elements e.g., tree 270, or man-made elements e.g., fence 250. Calibration points may be selected to be conspicuousand distributed in field of view of the sensors in the sensor array 101. Measurement of a calibration point distance may be made by laser range finder. In certain embodiments, the calibration point distance may be measured by determining the location of the calibration point using GPS or GPS / RTK and then calculating the distance to the location of the reference point. In certain embodiments, the calibration point distance may be determined by locating points on a map corresponding to the reference point and measuring the scaled distance between points on the map. In certain embodiments, 1, 2, 4, 5, 6, 7, 8, 9, or 10 or more calibration points may be selected, and calibration point distance may be determined. Following determine calibration point distances (step 425), the exemplary embodiment process 400 proceeds to the step 430 to determine correspondences.
[0074] At the step 430 to determine correspondences, images from the sensors of the sensor array 101 may be examined and the 2D location in two or more images of a calibration point may be determined. Calibration points may be given an identity (ID). The ID of the calibration point, and the 2D location(s) of the calibration point in the two or more images may be recorded. In certain embodiments, images from the sensors of the sensor array 101 may be shown to the operator on display 185 and using controls 186, the operator marks the corresponding locations of a calibration point in the images manually. In certain embodiments, marked calibration points may be given an annotation of the ID of the calibration point. In certain embodiments, calibration points may be distinctive and correspondences in the images may be determined by the processor unit 160 matching the calibration points in an image. In certain embodiments, calibration markers may be used that have a distinctive shape, marking or signal embedded that identifies a calibration marker uniquely and correspondences in the images may be determined by the processor unit 160 matching the calibration marker seen in each image. In certain embodiments, a moving target such as a drone may be used as a calibration marker. In certain embodiments, calibration markers may have a distinctive motion or distinctive patterns of motion that identifies a calibration marker uniquely and correspondences in the images may be determined by the processor unit 160 matching the calibration marker seen in eachimage at least in part through the distinctive motion of the calibration marker. Following the determine correspondences step 430, the exemplary embodiment process 400 proceeds to step 435 of performing initial calibration.
[0075] At the perform initial calibration step 435, the set of 2D locations for calibration points in a sensor image, the set of intrinsic calibration for the sensor and set of calibration distances for calibration points are processed by the processor 160 to determine a calibration of the sensor array 101. The calibration may include information, extrinsic calibration parameters, describing the location and orientation of a camera in the sensor array 101. In certain embodiments, the location and orientation of one or more cameras in the extrinsic calibration parameters may be with respect to a local reference point. Each calibration point may be projected from a sensor at the reference location by the distance measured at step 425 (and using the intrinsic calibration parameters) to determine an initial estimated location of the calibration point. In certain embodiments, a portion ora substantial portion of the calibration points may be projected from a sensor at the reference location by the distance measured at step 425 (and using the intrinsic calibration parameters) to determine an initial estimated 3D location of the calibration points. In certain embodiments, the extrinsic calibration parameters and initial 3D point positions are then regressed via a least-squares optimization to find the best fit extrinsic parameters. In certain embodiments, the processor unit 160 may apply a perspective N point algorithm to determine the extrinsic calibration. In certain embodiments, the calibration performed at step 435 may include performing an intrinsic calibration to determine the intrinsic parameters of the cameras in the sensor array 101 or may include refining intrinsic parameters determined during offsite calibration (step 310). Following the step 435 of performing initial calibration, the exemplary embodiment process 400 proceeds to step 440 of annotating additional points.
[0076] At the annotating additional points step 440, additional points in the scene may be selected. These points may be selected to be conspicuous, distinct and distributed in field of view of the sensors in the sensor array 101. Additional points may be moving or maybe points on a target in the scene that is moving, for example additional points could belocated on an aerial drone or bird flying in the sky or on a ground vehicle traveling on a road or on a ship or unmanned surface vessel travelling across the sea. In certain embodiments, the system 100 shows still images from the sensors of the sensor array 101 captured at the same time or substantially the same time. The images from the sensors of the sensor array 101 may be examined and the location in an image of an additional point may be determined. The additional point may be given an ID. The ID of the additional point and the location of the additional point in the image may be recorded. In certain embodiments, images from the sensors of the sensor array 101 may be shown to the operator on display 185 and using controls 186, the operator may mark the corresponding locations of an additional point in the images manually. In certain embodiments, additional points may be found by processor unit 160 using an image feature detection algorithm and a feature matching algorithm may be used to find the appearance of the additional point in the images and determine an annotated set of additional points without requiring manual intervention. Following the annotating additional points step 440, the exemplary embodiment process 400 proceeds to the step 445 of performing refined calibration.
[0077] At the step 445 of performing refined calibration, an extrinsic calibration may be performed using the set of 2D locations for a calibration point in a sensor image, the set of intrinsic parameters for the sensor, the initial calibration determined at perform initial calibration step 435, the set of calibration distances for the calibration point and set of 2D locations for additional points in the sensor image. These may be processed by the processor 160 to determine refined extrinsic parameters for the sensor array 101. The extrinsic parameters may include information about the location and orientation of a camera in the sensor array 101. In certain embodiments, the location and orientation of a camera in the extrinsic parameters may be with respect to a local reference point. The extrinsic parameters and 3D point positions may then be regressed via a least-squares optimization to find the best fit extrinsic parameters. In certain embodiments, the processor 160 may apply the perspective N point algorithm to determine the extrinsic parameters. At this step 445, the extrinsic calibration may be improved from that obtained at the perform initialcalibration step 435 because of the additional constraints provided by the additional points. Following the step 445 of performing refined calibration, the exemplary embodiment process 400 proceeds to the step 450 of reviewing reprojection and distance errors.
[0078] At the step 450 of reviewing reprojection and distance errors, the quality of the sensor array 101 extrinsic calibration may be checked. Typically, a poor extrinsic calibration may occur if the 2D locations of the calibration points or the additional points in the sensor images are inaccurate, for example if calibration points or additional points are incorrectly placed by more than a few (e.g., 2, 3, 4, 5 or more) pixels then the calibration may be inaccurate. At this step 450, the reprojection error may be calculated by averaging the distance between labelled image points and the projections of those points in the scene. In certain embodiments, the operator may be shown a value for the reprojection error. In certain embodiments, if the reprojection error exceeds a threshold, the system 100 will instruct the operator to repeat the calibration process. Display 185 may show a bullseye plot of the reprojection errors, an example is illustrated in FIG. 7 at 700. For a sensor in the sensor array, a bullseye plot 710 may be displayed where the magnitude of reprojection error is the distance from the center enabling visual evaluation against the concentric error rings. Calibration points 720 and additional points 730 may be plotted in location (du, dv) with reprojection error du, dv in u and v image coordinates respectively and distinguished via markers. Displayed calibration points and additional points may be labelled with the ID labels 740, 750 of the calibration points and additional points. Bad points such as 760, those which exceed a given error tolerance (for example where the reprojection error is greater than 0.5, 1, 2, 3, 4, 5 or more pixels), may be highlighted for the attention of the operator.
[0079] Reprojection errors may be evaluated to ensure there are no gross outliers and that reprojection error is bounded to a suitable level. Given an anticipated pixel accuracy from the system 100, an expected distance error profile may be generated and displayed to the operator. An example of an error profile is shown in FIG. 7 at 790.
[0080] If there are outliers or significant errors in reprojection error, the image annotations can be revisited to improve the extrinsic calibration best fit and re-evaluation can be performed.
[0081] Following the step 450 of reviewing reprojection and distance errors, the exemplary embodiment process 400 proceeds to the validate model step 455.
[0082] At the validate model step 455, the operator may be presented images from the sensors of the sensor array 101 and may mark a validation point in the scene on one of the images using controls 186 to mark the point. On the other images (being from the other sensors), an epi-polar line may be shown. By visual inspection, the operator may check that the displayed epi-polar line intersects or nearly intersects with the validation point. The quality of the calibration may be judged by the operator by observing a number of validation points and seeing how closely the epi-polar lines are to the validation points. If the operator judges that the calibration is insufficient, then they may choose to repeat the extrinsic calibration process entirely or in part. Following the completion of this step 455, the process 400 may complete at 490.Operating Process
[0083] FIG. 5 shows an exemplary process 500 for performing the operate step 340.Starting from 510, the first step is to receive sensor data (step 520).
[0084] At the receive sensor data step 520, sensor data may be captured by sensors (e.g., 110, 120) in the sensor array 101 and may be transmitted via communications bus (e.g., 111, 121) to the receiving unit 150 and then may be stored in the memory unit 170 and made accessible to the processor unit 160. Following the receive sensor data step 520, the exemplary embodiment process 500 proceeds to the step 525 of update background model.
[0085] At the step 525 of update background model, a pixelwise background model may be updated for one or more sensors in the sensor array 101. A process for updating the background model for a sensor is described with reference to FIG. 6. FIG. 6 illustrates how a pixel-based background model 600 is built, according to at least one embodiment. Sensor data for a sensor, e.g., 110, may be stored in memory unit 170 and may be represented in arectangular grid of pixels as shown at 610 where three pixels 611, 612, 613 are specifically illustrated for this example. For one or more sensor pixels, the system holds a circular buffer, as shown at 620, that may be used to hold historical samples of the pixel values. Additionally for the one or more pixels in the sensor data 610, the system holds a histogram of the held historical pixel values, shown at 630. In step 525, a new pixel value may be added to the circular buffer 620 when the circular buffer is at capacity then adding a sample may remove the oldest sample from the circular buffer. Also in step 525, the histogram 630 may be updated as follows: for the value of the sample being removed from the circular buffer, the histogram 630 has the corresponding bucket decremented (indicated at 631) and for a new sample value being added to the circular buffer, the histogram has the corresponding bucket incremented (indicated at 632), thus the histogram is updated without recalculation of the historical data stored in the circular buffer. In certain embodiments, the histogram may store sample counts for one or more pixel values, for example, 0, 1, 2, •••, 254, or 255.
[0086] Once the histogram (for example 630) for a pixel is determined then a background model for the background may be calculated. In certain embodiments, the background model is a centre + radius model and is calculated with the centre being the middle of the two quartiles of the histogram and radius being half the distance between the two quartiles of the histogram.
[0087] In certain embodiments, an alternative background model may be used such as a reference image, temporal average model, a mixture of gaussians model or other background modelling approach.
[0088] Following the updating background model step 525, the exemplary embodiment process 500 proceeds to step 530 of detecting targets.Detect T argets
[0089] At step 530 of detecting targets, for one or more sensors (e.g. , 110, 120) of the sensor array 101 the pixel data captured at step 520 may be compared to the background model to determine if the pixels are considered as foreground pixels.
[0090] In certain embodiments, update background model step 525 is performed using pixel data captured in previous iterations of the operation process 500. In certain embodiments, the detecting targets step 530 is performed using background model determined on a previous iteration of the operation of the process 500. In certain embodiments, in the process 500, the detecting targets step 530 may be performed prior to the update background model step 525.
[0091] In certain embodiments, a centre + radius model for the background is used and a foreground pixel may be determined as one where the value of the pixel is outside of the centre + radius by a margin (the margin may be set as a parameter for operation of the system). In certain embodiments, the margin may be a multiplier of x15 or may be between x1 and x20 with higher values reducing the number of false positive detections. In certain embodiments, a bin gap test may be used. In the bin gap test, a pixel may be considered foreground if there is an absolute gap of N or more histogram bins between the histogram bin corresponding to the value of the pixel and the nearest non-empty histogram bin in the background model forthat pixel location, where N may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The determined foreground pixels are represented in a binary foreground mask, where a non-zero pixel value indicates that the corresponding captured pixel is considered foreground, and a zero value indicates that the corresponding captured pixel is not considered foreground.
[0092] Once a binary foreground mask has been constructed, connected component analysis may be performed to label one or more connected components (e.g., each set of connected pixels). In certain embodiments, clustering is (or may be) performed to associate together neighbouring connected components. The centroid of the one or more connected components may be calculated from the binary mask and used for the nominal location of the detected target in the 2D image. In certain embodiments, the centroid calculation is (ormay be) performed using a score weighted average of the pixel location. In certain embodiments, a detected target is labelled with a track ID used to consistently label detected targets over iterations of the operation process 500. FIG. 8 is an illustration of 2D tracking and redetection, according to at least one embodiment. FIG. 8 at 810 illustrates an image with 5 detections marked by crosses. In certain embodiments, a Kalman filter uses a set of detections from previous images to predict a 2D location as illustrated in 820 with a dashed arrow at 821. Current detections that reasonably match a predicted location (such as 822) may be presumed to be of the same target and assigned a consistent 2D track ID. In certain embodiments, the appearance of detected target may be compared to the appearance of previously detected targets and be matched by additionally using patchbased similarity measure such as sum of absolute differences, local descriptors such as SIFT or SURF descriptors, learned patch similarity metrics, or other similarity measure. The current detection may be used to predict the location of the target for a subsequent iteration.
[0093] Following the detect targets step 530, the exemplary embodiment process 500 proceeds to the redetection (salience) step 535.Redetection with Salience
[0094] At the redetection (salience) (also referred to as “redetect targets”) step 535, previous 2D tracks that have not found matching detection at step 530 may be evaluated as follows with a saliency detection method. At times the detect targets step 530 may fail to detect a target. This may occur when the target moves too far away from the sensor array 101 or is too faint to enable a reliable detection distinct from the background. FIG. 8 at 830 illustrates a prediction with a dashed arrow and a dot-point but there is no target detection. A search area or region 831 is established about the predicted location and a saliency score calculated for pixels in the search region 831. As illustrated at 840, the saliency score is calculated with an N X N window 842 about each pixel in the search area or region 831. The saliency score for a pixel is calculated as the absolute difference between the value of the pixel, and the mean value of the pixels in the N x N window. In certain embodiments,the N x N window may be a 5 x 5, 7 x 7, 9 x 9, 11 x 11 or another size. The pixel with the highest salience score in the search area or region 831 is compared to a threshold and if greater than the threshold is taken to be a detection of the target. The 2D track of the target may then be updated with this detection which may also be used to update the tracker prediction for the next iteration. Through performing the step of 535, the range of the system 100 may be increased.
[0095] Following the redetect targets step 535, the exemplary embodiment process 500 proceeds to the stereo track matching step 560.Stereo Matching
[0096] At the stereo track matching step 560, targets detected and tracked in the 2D data from different sensors may be matched. When a target has been detected in images from two or more sensors and matched, then the 3D coordinates of the target may be calculated by geometric projection. The target detected in each image may be projected from the 2D images into the 3D scene using one or more of the following: the intrinsic calibration parameters and extrinsic calibration parameters. The intersection of the projection from the two or more sensor’ s images is the 3D location of the target in the scene. In certain embodiments, the appearance of targets in two or more 2D images may be compared and matched using patch-based similarity measure such as sum of absolute differences, local descriptors such as SIFT or SURF descriptors, learned patch similarity metrics, or other similarity measure. Where three or more sensors have an overlapping field of view, a cost matrix may be used to find the optimal set of target matches over the three or more sensors. In FIG. 9 at 910 are illustrated detections from three sensors and epi-polar lines such as 912 and 913. In certain embodiments, matching may be performed using epi-polar, i.e. , geometric, constraints without using the appearance characteristics of the detected targets. An epi-polar line may be considered as the ray projected by a detected target from the 2D image of a first sensor as viewed from a second sensor, and as determined using the intrinsic and extrinsic calibration parameters of the first and second sensors. In certainembodiments, 2D detections may be matched between sensors where the epi-polar line extended from one detection passes within a threshold distance on the image plane of a 2D detection from a different sensor. In certain embodiments, the matching of 2D detections may be optimised for minimal cost over a set of 2D detections. In FIG. 9 at 920 are illustrated a set of three cost matrix that may be used to determine matches over a set of detections in images from three cameras. In certain embodiments, stereo matching may be applied iteratively, removing matches and resolving matches that were previously ambiguous in one or more iterations. A target in one sensor may be an ambiguous match where the epi-polar line passes within a threshold distance on the image plane of two or more targets in a different sensor. At the conclusion of the Stereo Track Matching step 560, the targets that have been matched across two or more sensors may have a target ID assigned and their location may have been calculated through geometric projection.Following the stereo track matching step 560, the exemplary embodiment process 500 proceeds to the lost track recovery step 565.Lost T rack Recovery
[0097] At the lost track recovery step 565, the system 100 attempts to repair 2D tracks and 3D tracks where a 2D track has been lost. Sometimes 2D tracking of a target in one view may fail but 2D tracking for the same target in a different view may continue to track successfully. With reference to FIG. 10 at 1010, two camera views are shown with a target matched shown by a small circle 1012, 1013 on the epi-polar line 1011. At a later time, shown at 1020, the target is detected correctly at 1022, but not in the second camera where only a prediction is available at 1023. Then at a further later time, shown at 1030, the target is detected successfully in both camera views as shown at 1032, 1033. The lost recovery tracking associates the newly successful matching detections 1032 and 1033 with the previous successful 3D track and may interpolate intermediate 2D tracking, shown at 1040 with the 2D track at 1038 now repaired, to build a 3D track without the gaps where 2D tracking was lost. Thus, the 2D tracks and the 3D track may be made continuous and may retain their IDs even though tracking was lost for a time in one view. Following the lost trackrecovery step 565, the exemplary embodiment process 500 proceeds to classify targets step 570.Classify Targets
[0098] At the classify targets step 570, the system 100 may classify tracked targets. The system 100 may use the target’ s 2D track, 3D track, velocity, altitude, size, and / or other attributes to classify the target. In certain embodiments, a set of 2D or 3D locations for the target may be retained and a classification determined at least in part from the target’ s accumulated pattern of motion. In certain embodiments, the characteristics of the target’ s path over time, such as the local path deviation, may be used to classify the target. In certain embodiments, a set of 2D or 3D locations for the target may be retained and a local path deviation score calculated from the target’ s accumulated deviation from a straight-line path over a set of tracked locations. In certain embodiments, visual appearance may be used to classify the target. The target may be classified into one or more classes which may, for example, be aircraft, car, person, bird, aerial drone (quad-copter), aerial drone (winged), and / or helicopter. The classification may be performed on the processor unit 160. The classification may use heuristics for size, speed, and / or acceleration to classify a target. In certain embodiments, the classification may use the appearance of the target, using one or more of the images from the sensor array 101 or other source of image data. In certain embodiments, a neural network and / or other artificial intelligence (Al) may be used to perform the classification of the target. Where a classification is made then a class and level of confidence may be determined and may be associated with the 3D tracking data for the target.
[0099] In certain embodiments, a secondary camera system 115 with a Pan-Tilt-Zoom (PTZ) or gimble facility may capture images of targets that are being tracked. Such a secondary camera system may be able to use telephoto lenses to capture high resolution images of targets, these images may capture more pixels on the target than are available in the images from a sensor array 101 (for example the secondary camera system maycapture between 5 and 10 pixels across a target or between 10 and 50 pixels across a target or between 50 and 200 pixels across a target). In certain embodiments, the 3D coordinates of the target may be used to direct a secondary camera system to capture images of a tracked target. In certain embodiments, images captured by a secondary camera system may be input to a target classification system as described previously.Output
[0100] At the output step 575, the system 100 may output track information to external systems via the external communications unit 180 and communication mechanisms 182, 181. In certain embodiments, the output to external system may be information on one or more targets being tracked including one or more of the following: a timestamp, a target ID, classification, altitude, latitude, and / or longitude. In certain embodiments, the output may be in a standardised format such as CivTAK or ATAK. In certain embodiments, the output may include images of the target.
[0101] At the output step 575, the system 100 may also display a representation of the track information on the display 185. The operator may be able to control the information presented on the display using controls 186.
[0102] Following the Output step 575, the exemplary embodiment process returns to Receive Sensor Data 520. Thus, the process 500 loops indefinitely.
[0103] User Interface
[0104] FIG. 11 shows an exemplary user interface 1100 that may be presented on display 185 and operated using controls 186. The user interface 1100 includes a representative view 1120, a map view 1130, a target list 1160 and associated target properties panel 1180, target image panel 1140 and a system status panel 1190. The representative view 1120 may show an image from one of the cameras or may show a composite image generated from two or more cameras. In the representative view, detected targets such as 1124, 1122 are each marked with box and labelled with a target ID. The map view 1130 presents a top down, two-dimensional representation of area forward of the system which may be at 1136.The background image for the map view 1130 may be from satellite imagery or from a map or maybe blank. The targets being tracked are shown on the map view, for example 1132 and 1134 marked with squares and an ID. A region 1138 indicates the field of view of the system 100. Thus, the map view shows the targets, e.g. 1132, 1134, location and range with respect to the system location at 1136. The target list 1160 shows detected targets in rows with properties in columns that may include ID, range (from the system), age (time that the target has been known), affiliation (e.g. friendly, foe, neutral, unknown), classification (e.g. civilian drone, airplane, etc), stream to TAK (enable streaming the tracking data to an external system), and / or visible (show on the representative view and / or on the map view). The target attributes panel 1180 shows some attributes of a selected target. The attributes may include one or more of the ID, classification, range, location, altitude, and / or time stamp. The target attributes panel 1180 may show other attributes of the target including the age (time that the target has been known), affiliation (e.g., friendly, foe, neutral, and / or unknown), and / or classification (e.g., civilian drone, airplane, etc.). The target image panel shows a zoomed in view of the target. The image may be image from one of the cameras or may show a composite image generated from two or more cameras. The image may be generated from multiple images captured over a period of time which may be composited using super-resolution techniques. A target may be selected by using a movable pointer (controlled via a mouse, touch pad etc.) to choose a target in the representative view 1120, or in the target list 1160 or in the map view 1130. Where a target is selected in one of the views then it may be shown as selected in the other views, as shown at 1124, 1164 and 1134. Additionally, the properties for the selected target may be shown in the target attributes panel 1180 and a zoomed in image of the target may be shown in the target image panel 1140 at 1144. In certain embodiments, the user interface may show the velocity of the targets. In certain embodiments, the user interface may use different markings, e.g., rectangle, vs circle vs triangle or different colours to mark targets depending on their classification, affiliation or on an estimate of their criticality.
[0105] Further advantages and / or features of the claimed subject matter will become apparent from the following examples describing certain embodiments of the disclosed subject matter.Example 1. A system for tracking in a real-life scene one or more targets in three dimensions comprising:at least a first sensor that is configured to generate at least a first two-dimensional image of the real-life scene and at least a second sensor that is configured to generate at least a second two-dimensional image of the real-life scene, wherein the two-dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene;a processor that is configured to:receive the at least first two-dimensional image and the at least second two-dimensional image;detect, from the images if one or more target candidates are present in the images; determine that a candidate from the at least first image matches with a candidate from the at least second image; and determine the location of the one or more targets in the real-life scene.2. The system of example 1, wherein the determination is based at least in part on a geometric projection.3. The system of examples 1 or 2, wherein the determination is based at least in part on a plurality of images comprising: the at least first two-dimensional image and the at least second two-dimensional image.4. The system of any of examples 1 to 3, wherein the determination is based at least in part on multiple images from the at least first sensor and the at least second sensor.5. The system of any of examples 1 to 4, wherein the plurality of the multiple images from the at least first sensor and the at least second sensor are generated over a time period.. The system of any of examples 1 to 5, wherein a plurality of sensors is used in the system and the plurality of sensors comprises: the at least first sensor and the at least second sensor.. The system of any of examples 1 to 6, wherein the plurality of sensors comprises: at least 2, 3, 4, 10, or 50 sensors.. The system of any of examples 1 to 7, wherein the at least first sensor and the at least second sensor are both cameras.. The system of any of examples 1 to 8, wherein the cameras are one or more of the following: long-wave infrared cameras, medium wave infrared cameras, short wave infrared cameras, monochrome visible light cameras, visible light cameras, and multispectral cameras.0. The system of any of examples 1 to 9, wherein the system is configured to use extrinsic calibration.1. The system of any of examples 1 to 10, wherein the extrinsic calibration is configured to be refined with one or more of the following: one or more calibration points and additional calibration points.2. The system of example 11 , wherein the one or more calibration points and the additional calibration points are in a known range.3. The system of any of examples 1 to 12, wherein the at least first two-dimensional image and the at least second two-dimensional image of the system are communicated to an external system.4. The system of any of examples 1 to 13, wherein the processor is configured to generate a track of the one or more targets in the real-life scene, based at least in part on multiple images taken from the at least first sensor over a period of time, the at least second sensor over a period of time, or combinations thereof.5. A method using the system of any one of examples 1 to 14.16. One or more computer-readable non-transitory storage media embodying software that is operable when executed using the system of any one of examples 1 to 14 or the method in example 15.17. A method fortracking in a real-life scene one or more targets in three dimensions comprising:configuring at least a first sensor to generate at least a first two-dimensional image of the real-life scene and at least a second sensor to generate at least a second two-dimensional image of the real-life scene;generating the at least first two-dimensional image of the real-life scene from the at least first sensor and the at least second two-dimensional image of the real-life scene from the at least second sensor, wherein the two-dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene; sending the at least first two-dimensional image and the at least second two-dimensional image to a processor;receiving the at least first two-dimensional image and the at least second two-dimensional image at the processor;detecting, from the received images if one or more target candidates are present in the images;determining that a candidate from the at least first image matches with a candidate from the at least second image; anddetermining the location of the one or more targets in the real-life scene.18. The method of example 17, wherein the determination is based at least in part on a geometric projection.19. The method of examples 17 or 18, wherein the determination is based at least in part on a plurality of images comprising: the at least first two-dimensional image and the at least second two-dimensional image.20. The method of any of examples 17 to 19, wherein the determination is based at least in part on multiple images from the at least first sensor and the at least second sensor.1. The method of any of examples 17 to 20, wherein the plurality of the multiple images from the at least first sensor and the at least second sensor are generated over a time period.2. The method of any of examples 17 to 21 , wherein a plurality of sensors is used and the plurality of sensors comprises: the at least first sensor and the at least second sensor. 3. The method of example 22, wherein the plurality of sensors comprises: at least 2, 3, 4, 10, or 50 sensors.4. The method of any of examples 17 to 23, wherein the at least first sensor and the at least second sensor are both cameras.5. The method of example 24, wherein the cameras are one or more of the following: longwave infrared cameras, medium wave infrared cameras, shortwave infrared cameras, monochrome visible light cameras, visible light cameras, and multispectral cameras.6. The method of any of examples 17 to 25, further comprising using extrinsic calibration.7. The method of example 26, wherein the extrinsic calibration is configured to be refined with one or more of the following: one or more calibration points and additional calibration points.8. The method of example 27, wherein the one or more calibration points and the additional calibration points are in a known range.9. The method of any of examples 17 to 28, wherein the at least first two-dimensional image and the at least second two-dimensional image of the system are communicated to an external system.0. The method of any of examples 17 to 29, wherein the processor is configured to generate a track of the one or more targets in the real-life scene, based at least in part on multiple images taken from the at least first sensor over a period of time, the at least second sensor over a period of time, or combinations thereof.
[0106] Any description of prior art documents herein, or statements herein derived from or based on those documents, is not an admission that the documents or derived statements are part of the common general knowledge of the relevant art.
[0107] While certain embodiments have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only.
[0108] In the foregoing description of certain embodiments, specific terminology has been resorted to for the sake of clarity. However, the disclosure is not intended to be limited to the specific terms so selected, and it is to be understood that a specific term includes other technical equivalents which operate in a similar manner to accomplish a similar technical purpose. Terms such as “left” and right” , “front” and “rear” , “above” and “ below” and the like are used as words of convenience to provide reference points and are not to be construed as limiting terms.
[0109] In this specification, the word “comprising” is to be understood in its “open” sense, that is, in the sense of “including” , and thus not limited to its “closed” sense, that is the sense of “consisting only of” . A corresponding meaning is to be attributed to the corresponding words “comprise” , “comprised” and “comprises” where they appear.
[0110] It is to be understood that the present disclosure is not limited to the disclosed embodiments and is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the present disclosure. Also, the various embodiments described above may be implemented in conjunction with other embodiments, e.g., aspects of one embodiment may be combined with aspects of another embodiment to realize yet other embodiments. Further, independent features of a given embodiment may constitute an additional embodiment. In addition, a single feature or combination of features in certain of the embodiments may constitute additional embodiments. Specific structural and functional details disclosed herein are not to beinterpreted as limiting, but merely as a representative basis forteaching one skilled in the art to variously employ the disclosed embodiments and variations of those embodiments.
[0111] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0112] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
Claims
THE CLAIMS DEFINING THE INVENTION ARE AS FOLLOWS:
1. A system for tracking in a real-life scene one or more targets in three dimensions comprising:at least a first sensor that is configured to generate at least a first two-dimensional image of the real-life scene and at least a second sensor that is configured to generate at least a second two-dimensional image of the real-life scene, wherein the two- dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene;a processor that is configured to:receive the at least first two-dimensional image and the at least second two-dimensional image;detect, from the images if one or more target candidates are present in the images;determine that a candidate from the at least first image matches with a candidate from the at least second image; anddetermine the location of the one or more targets in the real-life scene.
2. The system of claim 1, wherein the determination is based at least in part on a geometric projection.
3. The system of claims 1 or 2, wherein the determination is based at least in part on a plurality of images comprising: the at least first two-dimensional image and the at least second two-dimensional image.
4. The system of any of claims 1 to 3, wherein the determination is based at least in part on multiple images from the at least first sensor and the at least second sensor.
5. The system of any of claims 1 to 4, wherein the plurality of the multiple images from the at least first sensor and the at least second sensor are generated over a time period.
6. The system of any of claims 1 to 5, wherein a plurality of sensors is used in the system and the plurality of sensors comprises: the at least first sensor and the at least second sensor.
7. The system of any of claims 1 to 6, wherein the plurality of sensors comprises: at least 2, 3, 4, 10, or 50 sensors.
8. The system of any of claims 1 to 7, wherein the at least first sensor and the at least second sensor are both cameras.
9. The system of any of claims 1 to 8, wherein the cameras are one or more of the following: long-wave infrared cameras, medium wave infrared cameras, short wave infrared cameras, monochrome visible light cameras, visible light cameras, and multispectral cameras.
10. The system of any of claims 1 to 9, wherein the system is configured to use extrinsic calibration.
11. The system of any of claims 1 to 10, wherein the extrinsic calibration is configured to be refined with one or more of the following: one or more calibration points and additional calibration points.
12. The system of claim 11 , wherein the one or more calibration points and the additional calibration points are in a known range.
13. The system of any of claims 1 to 12, wherein the at least first two-dimensional image and the at least second two-dimensional image of the system are communicated to an external system.
14. The system of any of claims 1 to 13, wherein the processor is configured to generate a track of the one or more targets in the real-life scene, based at least in part on multiple images taken from the at least first sensor over a period of time, the at least second sensor over a period of time, or combinations thereof.
15. A method using the system of any one of claims 1 to 14.
16. One or more computer-readable non-transitory storage media embodying software that is operable when executed using the system of any one of claims 1 to 14 or the method of claim 15.
17. A method fortracking in a real-life scene one or more targets in three dimensions comprising:configuring at least a first sensor to generate at least a first two-dimensional image of the real-life scene and at least a second sensor to generate at least a second two-dimensional image of the real-life scene;generating the at least first two-dimensional image of the real-life scene from the at least first sensor and the at least second two-dimensional image of the real-life scene from the at least second sensor, wherein the two-dimensional images of the at least first sensor and the at least second sensor are at least in part overlapping views of the real-life scene;sending the at least first two-dimensional image and the at least second two- dimensional image to the processor;receiving the at least first two-dimensional image and the at least second two-dimensional image at the processor;detecting, from the received images if one or more target candidates are present in the images;determining that a candidate from the at least first image matches with a candidate from the at least second image; anddetermining the location of the one or more targets in the real-life scene.