Detection and tracking of moving objects
The method addresses the challenges of UAV surveillance by employing region-based registration and hybrid tracking techniques to enhance object detection and tracking in UAV systems, achieving precise separation and tracking of moving objects despite rapid camera movements and low frame rates.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2011-12-15
- Publication Date
- 2026-05-07
AI Technical Summary
Existing surveillance systems using unmanned aerial vehicles (UAVs) face challenges in detecting and tracking moving objects due to rapid camera movements, non-uniform motion, low frame rates, and fluctuating lighting conditions, which impair object separation from the background and hinder precise registration.
A method involving region-based image registration, motion decomposition, and hybrid target tracking using Kanade-Lucas-Tomasi feature tracker and mean shift, combined with automatic dynamic threshold estimation and multi-target tracking algorithms, to enhance object detection and tracking in UAV surveillance.
The method effectively separates moving objects from the background, handles geometric distortions, and maintains precise tracking even with low frame rates and lighting fluctuations, enabling accurate detection and tracking of small targets.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of invention
[0001] Embodiments of the invention generally relate to information technology and in particular to image analysis of objects in a video. Background of the invention
[0002] In recent years, reconnaissance, surveillance, disaster relief, search and rescue, agricultural intelligence gathering, and rapid remote sensing and mapping for civilian and military purposes have gained increasing attention. For example, unmanned aerial vehicles (UAVs) can be an attractive platform for performing such operations due to their small size and cost-effective sensor platform. However, UAVs present significant challenges when used in surveillance systems. For instance, the background changes significantly when the camera moves quickly and rotates erratically, and the movement of a UAV is generally not uniform.Furthermore, the frame rate is very low (for example, 1 frame per second), which increases the difficulty of detecting and tracking targets moving on the ground, and a small object size presents a further challenge for object detection and tracking. Strong fluctuations in lighting conditions and camera noise can also severely impair the separation of actual moving objects from the background.
[0003] Existing approaches also have object initialization problems and are also unable to obtain highly precise registration results, handle rotation and size changes of a target, and cope with a similar distribution between target and background.
[0004] US patent 2007 / 0104383A1 discloses the use of multiple cameras to simultaneously or sequentially capture multiple views of an image to record the images of a scene. It also discloses the creation and use of a generation model to decompose image sequences into a group of layered 2-dimensional appearance images and masks of moving objects, allowing the deletion of sprite pixels determined to be fully transparent, while background pixels obscured by the sprite are deleted accordingly.
[0005] Furthermore, a layer extraction system based on dominant motion estimation and global registration is known from Kang et al., published in IEEE, International Conference on Multimedia and Expo (ICME), Vol. 1, 2004. Additionally, methods for motion estimation for video compression are known from Jasinschi et al., published in Journal of the Franklin Institute, 1998, Vol. 335, No. 8. Moreover, an image registration method is known from EP 2 006 805 A2. Brief description
[0006] The present invention is defined by the patent claims.
[0007] One or more embodiments or elements thereof can be realized in the form of a computer product comprising a physical, computer-readable storage medium with computer-usable program code for carrying out the specified process steps. Furthermore, one or more embodiments or elements thereof can be realized in the form of a device comprising a memory and at least one processor connected to and operable by the memory to carry out exemplary process steps.In another aspect, one or more embodiments or elements thereof may be realized in the form of means for carrying out one or more of the process steps described herein; the means may comprise (i) hardware module(s), (ii) software module(s), or (iii) a combination of hardware and software modules; each of (i) to (iii) implements the specific techniques set forth herein, and the software modules are stored on a physical, computer-readable storage medium (or several such media).
[0008] Following a first exemplary aspect, an exemplary method for carrying out visual monitoring of one or more moving object(s) is provided, wherein the method may include: the registration of one or more images taken by one or more camera(s), wherein the registration of the image(s) includes region-based registration of the image(s) in two or more adjacent frames; the motion decomposition of the image(s) to detect one or more moving object(s) and one or more background region(s) in the image(s); and tracking of the moving object(s) to facilitate visual monitoring of the moving object(s).
[0009] Preferably, an exemplary method is provided in which the registration of an image or multiple images features the recursive global and local geometric registration of the image or images.
[0010] Preferably, an exemplary method is provided in which the registration of one or more images involves the use of one or more subpixel image comparison techniques.
[0011] Preferably, an exemplary method is provided in which performing the motion decomposition of the image or images includes forward and backward single-image differentiation.
[0012] Preferably, an exemplary method is provided in which the forward and backward frame differences feature an automatic dynamic threshold estimation based on temporal and / or spatial filtering.
[0013] Preferably, an exemplary method is provided in which the forward and backward frame differences involve the removal of one or more incorrect moving pixels based on independent movement of one or more image features.
[0014] Preferably, an exemplary method is provided in which the forward and backward frame differences involve performing a morphological operation and generating one or more motion pixels.
[0015] Preferably, an exemplary method is provided in which tracking the moving object(s) involves performing hybrid target tracking, wherein the hybrid target tracking includes the use of a Kanade-Lucas-Tomasi (KLT) feature tracker and mean shift, the use of auto-kernel scale estimation and update, and the use of one or more feature trajectories.
[0016] Preferably, an exemplary method is provided in which the tracking of the moving object(s) involves the use of one or more multiple target tracking algorithms based on feature comparison and distance matrices for one or more target(s).
[0017] Preferably, an exemplary procedure is provided in which tracking the moving object(s) includes: generating a motion image; identifying one or more moving object(s); performing object initialization and object verification; identifying one or more object region(s) in the motion image; extracting one or more features; defining a search region in the motion image; identifying one or more candidate region(s) in the motion image; tracking the mean shift; identifying one or more moving object(s) in the candidate region(s); performing a Kanade-Lucas-Tomasi feature comparison; performing an affine transformation; performing an end-region determination using the Bhattacharyya coefficient; and updating a target model and the trajectory information.
[0018] Preferably, an exemplary method is provided in which the tracking of the moving object(s) features reference plane-based registration and tracking.
[0019] Preferably, an exemplary method is provided, further comprising relating each camera view to one or more other camera views.
[0020] Preferably, an exemplary method is provided, further comprising forming a panoramic view from the image or images taken by the camera(s).
[0021] Preferably, an exemplary method is provided, furthermore comprising estimating the movement of each camera based on video information about one or more static objects in the panoramic view.
[0022] Preferably, an exemplary method is provided, furthermore comprising the estimation of one or more background structures in the panorama view based on the detection of linear structures and the statistical analysis of the moving object(s) over a period of time.
[0023] Preferably, an exemplary method is provided, further comprising automatic feature extraction, wherein the automatic feature extraction includes: decomposing a video image into individual frames; performing a Gaussian smoothing operation; using a Canny detector to extract one or more feature edge(s); implementing a Hough transform for feature analysis; determining a maximum address result to reduce the influence of multiple peaks in a transformation space; determining whether a feature length is greater than a certain threshold, and performing feature extraction and pixel removal if the feature length is greater than the threshold.
[0024] Preferably, an exemplary procedure is provided in which the automatic feature extraction also includes performing single-image differentiation and testing using motion history images (MHI).
[0025] Preferably, an exemplary method is provided, further comprising performing outlier removal to eliminate one or more false matches of moving objects.
[0026] Preferably, an exemplary procedure is provided, further comprising the filtering of false BLOBs or large binary objects, wherein the filtering of false BLOBs comprises: generating a motion image; applying a connected component process to link all BLOB data; generating a motion BLOB table; extracting one or more features for each BLOB in a previously recorded frame; and applying a Kanade-Lucas-Tomasi procedure to estimate the motion of each BLOB, and deleting the BLOB from the BLOB table if no motion occurs for the BLOB.
[0027] Preferably, an exemplary procedure is provided, furthermore comprising the updating of a target model in a temporal domain and / or a spatial domain.
[0028] Preferably, an exemplary method is provided, also featuring the generation of an index of object appearances and object tracking in a panoramic view.
[0029] Preferably, an exemplary procedure is provided, also showing the determination of a similarity metric between a query and an entry in the index.
[0030] Preferably, an exemplary method is provided, further comprising the provision of a system, wherein the system comprises one or more separate software modules, each of the one or more separate software modules being executed on a physical, computer-readable writable storage medium, and wherein the separate software module(s) comprise a geometric registration module, a motion extraction module, and an object tracking module, which are executed on a hardware processor.
[0031] According to another exemplary aspect, an exemplary computer program product is provided, comprising a physical, computer-readable, writable storage medium with computer-usable program code for performing visual monitoring of one or more moving objects, wherein the computer program product includes: computer-usable program code for registering an image or images captured by one or more cameras, wherein the registration of the image or images includes region-based registration of the image or images in two or more adjacent frames; computer-usable program code for performing motion decomposition of the image or images in order to detect one or more moving objects and one or more background regions in the image or images;and computer-usable program code for tracking the moving object(s) to facilitate visual monitoring of the moving object(s).
[0032] According to a further exemplary aspect, an exemplary system for performing visual monitoring of one or more moving objects is provided, comprising: a memory; and at least one processor connected to the memory and operable to: register one or more images captured by one or more camera(s), wherein the registration of the image(s) includes region-based registration of the image(s) in two or more adjacent frames; perform motion decomposition of the image(s) to detect one or more moving object(s) and one or more background region(s) in the image(s); and track the moving object(s) to facilitate visual monitoring of the moving object(s). Brief description of the drawings
[0033] A preferred embodiment of the present invention will now be described by way of example only with reference to the accompanying drawings, wherein: Fig. 1 is a diagram illustrating the subpixel position estimation according to a preferred embodiment of the present invention; Fig. 2 is a diagram illustrating the subregion selection according to a preferred embodiment of the present invention; Fig. 3 is a diagram illustrating the geometric forward and backward registration according to a preferred embodiment of the present invention; Fig. 4 is a flowchart illustrating forward and backward single-frame differentiation according to a preferred embodiment of the present invention; Fig. 5 is a flowchart illustrating the filtering of false BLOBs according to a preferred embodiment of the present invention; Fig. 6 is a flowchart illustrating the multi-object tracking according to a preferred embodiment of the present invention; Fig. 7 is a diagram illustrating the reference plane-based registration and tracking according to a preferred embodiment of the present invention; Fig. 8 is a flowchart illustrating the automatic extraction of traffic routes according to a preferred embodiment of the present invention; Fig. 9 is a block diagram illustrating the architecture of an object detection and tracking system according to a preferred aspect of the invention; Fig. 10 is a flowchart illustrating techniques for performing visual monitoring of one or more moving objects according to a preferred embodiment of the invention; and Fig. 11 is a system diagram of an exemplary computer system on which a preferred embodiment of the invention can be implemented. Detailed description of preferred embodiments
[0034] The principles of the invention include the detection, tracking, and searching of moving objects during visual surveillance. In an example configuration with moving objects and one or more moving camera(s), one or more embodiments of the invention include motion decomposition (motion BLOBs versus background region), multi-object tracking (for example, consistent tracking over time), and reference-plane-based registration and tracking.As explained in detail herein, one or more embodiments of the invention include the use of multiple (for example, adjacent) cameras mounted, for example, on mobile platforms (for example, videos from unmanned aerial vehicles (UAVs)) to detect, track, and search for moving objects by generating a panoramic view based on global / local geometric registration, motion decomposition, tracking of moving objects, reference plane-based registration and tracking, and automatic extraction of urban traffic routes from the images received by the cameras.
[0035] The techniques described herein include recursive geometric registration, which involves region-based image registration for adjacent frames rather than an entire frame, subpixel image comparison techniques, and region-based geometric transformation to handle geometric lens distortion. One or more embodiments of the invention also include bidirectional motion detection and hybrid target tracking based on colors and features. Bidirectional motion detection includes forward and backward frame differentiation, automatic dynamic threshold estimation based on temporal and / or spatial filtering, and the elimination of false moving pixels based on independent feature movements.The hybrid target tracking includes a Kanade-Lucas-Tomasi feature tracker (KLT) and mean shift, auto-kernel scale estimation and updating, and consistent tracking over time based on the coherent movement of feature trajectories.
[0036] Furthermore, the techniques described herein include multi-target tracking algorithms based on feature comparison and distance matrices for small targets, as well as, for example, a low frame rate (1 frame / second) UAV surveillance system implementation for detecting and tracking small targets (for example, without a known shape model).
[0037] As mentioned herein, one or more embodiments of the invention include the local / global geometric registration of videos (for example, UAV videos). To reduce the camera motion effect, a frame-by-frame video registration process is implemented. A precise way to register two images could involve comparing every pixel in each image. However, the high computational cost is not feasible. An efficient way is to find a relatively small set of feature points in the image that are easily retrievable and to use only these points to estimate the frame-to-frame homography. For example, for an image of 1280 × 1280 pixels, only 500 to 600 feature points can be extracted.
[0038] The Harris corner detector, due to its invariance with respect to variations in size, rotation, and illumination, can be used for image registration and motion detection. In one or more embodiments of the invention, the Harris corner detector can be used as a feature point detector. Its algorithm can be described as follows: 1. For a pixel in an image I, calculate its x and y direction derivatives I x and I y, and I xy = I x I y. 2. Apply a window function A, that is, hx = Al x, hy = Al y, hxy = Al xy. 3. Calculate H=hxhy−hxy2−κ(hx+hy)2 (κ is a constant) to measure changes in both directions. 4. Determine the threshold H and find local maxima to obtain a corner.
[0039] To compare the windows, one or more embodiments of the invention include the use of a normalized correlation coefficient, which is an efficient statistical method. The actual feature comparison is achieved by maximizing the correlation coefficient across small windows surrounding the points. The correlation coefficient is given by: ρ=∑r=1R∑c=1C[g1(r,c)−u1]•[g2(r,c)−u2]∑r=1R∑c=1C[g1(r,c)−u1]2∑r=1R∑c=1C[g2(r,c)−u2]2;−1≤ρ≤1 where: g1(r,c) represent individual grey values of the template matrix; u1 represents the average gray value of the template matrix; g2(r,c) represents individual grey values of the corresponding part of the search matrix; u2 represents the average gray value of the corresponding part of the search matrix; and R, C represents the number of rows and columns of the template matrix.
[0040] The block matching process can therefore be achieved as follows. For each point in a reference frame, all points in the selected frame are checked, and the most similar point is chosen. Next, it is tested whether the correlation achieved is high enough. The point with the maximum correlation coefficient is taken as a candidate point.
[0041] Video recording must be performed in real time. In one or more embodiments of the invention, the block-matching algorithm is implemented only for the features. This significantly reduces the computational effort.
[0042] One or more embodiments of the invention also include checking relevant features and removing outliers. Feature-based block matching can sometimes lead to a mismatch. To avoid a mismatch problem, one or more embodiments of the invention include using a forward search to process the mismatch data that occurs one too many times, retaining the corresponding candidate feature with the maximum gradient value and removing the others. A backward search is also used to solve the remaining mismatch problem using the same approach.
[0043] In many cases, a pair of features with similar attributes is accepted as a match. However, false matches can occur. Therefore, in one or more embodiments of the invention, a random sample consensus (RANSAC) procedure is performed to eliminate outliers, thereby removing false matches and increasing registration accuracy.
[0044] The techniques described herein can also include coarse to fine feature comparison. Multi-resolution feature comparison can reduce the search space and false matches. In the coarsest resolution layer, the feature comparison is performed and the search scope is determined. In the current resolution layer, the comparison results from the previous layer can be used as initial results, and the comparison process can be carried out using equation (1) given above. In one or more embodiments of the invention, the search scope is limited to 1 to 3 pixels. Furthermore, the same operation can be repeated until the highest resolution layer is reached.
[0045] As further described herein, one or more embodiments of the invention may include precise position determination. For video recording and motion detection, pixel-level accuracy may not be sufficient. In such cases, a subpixel position approach is considered, and a distance-dependent weighted interpolation to the tip is determined. The horizontal and vertical positions of the tip for the feature can be estimated separately. The one-dimensional horizontal and vertical correlation curves can also be obtained. Furthermore, the correlation value in the x,y directions is interpolated separately, and the precise position of the tip is calculated. Fig. Figure 1, for example, is a diagram illustrating the subpixel position estimation according to an embodiment of the present invention.
[0046] The techniques described herein can also include local geometric registration. For example, geometric registration of a subregion can be chosen, and the entire single image can be divided into 2 x 2 subregions. Fig. Figure 2 illustrates two selection models.
[0047] Fig. Figure 2 is a diagram illustrating the subregion selection according to one embodiment of the present invention. For illustration, it shows Fig. 2 a subregion selection model 202 and a subregion selection model 204.
[0048] One or more embodiments of the invention also include an affine local transformation, such as the following: [xy]=[a0+a1u+a2vb0+b1u+b2v]
[0049] Where (x, y) is the newly transformed coordinate of (u, v) and (a j , b k) (j, k = 1, 2, 3) is the set of transformation parameters. To determine the local transformation parameters for each subregion, one or more embodiments of the invention also include the use of a least squares technique for calculating the transformation parameters.
[0050] One or more embodiments of the invention also include single-frame reverse / forward registration. For example, single-frame reverse / forward registration is performed in cases with rapid camera movement, strong fluctuations in lighting conditions, and severe banding noise for multi-frame differentiation in order to avoid residual error propagation. Fig. Figure 3 illustrates one approach.
[0051] Fig. Figure 3 is a diagram illustrating the geometric forward and backward registration according to one embodiment of the present invention. For illustration, it shows Fig. 3 a single image 302 (F i-1 ), a single image 304 (F i ) and a single image 306 (F i+1 ). To show the object movement in frame 304 (F i ) to estimate, which is taken as the reference frame, the previous frame 302 (F i-1 ) and the next single image 306 (F i+1 The motion is geometrically registered to the reference frame. Motion estimation for each frame is performed in this way.
[0052] Forward / backward frame differentiation can also be implemented for motion detection. A diagram of the approach used in one or more embodiments of the invention is shown in Fig. 4 shown. Fig. Figure 4 is a flowchart illustrating forward / backward frame differentiation according to one embodiment of the present invention. After the forward / backward frames (for example, frame 402, frame 404, and frame 406) have been geometrically registered and aligned in steps 408 and 410, difference images are calculated. Instead of a simple subtraction between the aligned frames, one or more embodiments of the invention use forward / backward frame differentiation in steps 412 and 414 to reduce motion noise and, for example, to compensate for fluctuations in lighting conditions by means of automatic gain control.
[0053] Additionally, step 416 includes the performance of image arithmetic by I new = ΔI i-1,i AND Δ i,i+1Step 418 includes median filtering, which can reduce random motion noise. To extract moving pixels of moving objects, step 420 performs automatic dynamic threshold estimation based on spatial filtering. Step 422 also includes performing a morphological operation to remove small isolated areas and fill holes in the foreground image, and step 424 includes generating motion pixels (for example, a motion image).
[0054] To further reduce random noise and the effect of changing lighting conditions, a logical AND operation is performed on forward / backward difference images to obtain a final difference image. {Di−1,i(x,y)=|Fi−1(x,y)−Fi(x,y)|;Di,i+1(x,y)=|Fi(x,y)−Fi+1(x,y)|; Di(x,y)=Di−1,i(x,y)∩Di,i+1(x,y);i=1,2,⋯,N
[0055] A threshold value for each pixel, based on statistical properties and spatial high-frequency data of the difference image, is automatically calculated statistically. Furthermore, a morphological step can be applied to remove small isolated areas and fill holes in the foreground image.
[0056] As described herein, one or more embodiments of the invention also include a motion test. Fig. Figure 5 is a flowchart illustrating the filtering of erroneous BLOBs according to an embodiment of the present invention; Step 502 includes generating a motion image. Step 504 includes applying a connected component process to link all BLOB data. Step 506 includes generating a motion BLOB table. Step 508 includes performing an optical flow estimation. Step 510 includes performing a displacement determination. If displacement is present, the process proceeds to Step 512, which includes performing post-processing such as data mapping, object tracking, trajectory maintenance, and tracking data management. If no displacement is present, the process proceeds to Step 514, which includes filtering erroneous BLOBs.
[0057] After generating a BLOB table, all BLOB data is checked to remove incorrect motion BLOBs from the BLOB table. One or more embodiments of the invention apply a KLT process to estimate the motion of each BLOB after frame-by-frame forward / backward recording. An incorrect BLOB is deleted from the BLOB table. The process steps can include, for example: applying a connected component process to link all BLOB data, generating a BLOB table, extracting features for each BLOB from a previously recorded frame, applying the KLT method to estimate the motion of each BLOB, and deleting the BLOB from the BLOB table if no motion is detected. The above steps can also be repeated for all BLOBs.
[0058] As also explained herein, one or more embodiments of the invention include multi-object tracking. Fig. Figure 6 is a flowchart illustrating multi-object tracking according to an embodiment of the present invention. Step 602 includes generating a motion image. Step 604 includes identifying moving BLOBs. Step 606 includes object initialization, and step 608 includes object verification. Step 610 includes identifying object regions. Step 612 includes identifying candidate regions. Step 614 also includes tracking mean shift, and step 616 includes identifying new positions.
[0059] After identifying object regions in step 610, features can be extracted in step 618. Once a search region has been defined in step 620, moving BLOBs can be identified as potential object candidates in step 622. In step 624, KLT matching is performed, and in step 626, outliers are removed using an affine transformation with RANSAC. In step 628, a new candidate region is identified. Meanwhile, the mean shift in step 614 is applied to calculate the inter-frame translation. This yields a candidate region location in step 616. From steps 628 and 616, the process can proceed to step 630, where the position of the final region is determined based on the Bhattacharyya coefficient.Step 632 also includes an update of the target model to solve drift problems, and step 634 includes an update of the trajectory. To track moving objects, a hybrid tracking model based on a combination of the KLT and mean-shift methods is also applied from steps 618 to 630.
[0060] As mentioned, the techniques described herein include object initialization. Motion detection results from forward / backward frame differentiation and may contain some correct, genuine moving objects, some incorrect objects, and miss some true objects. For example, in a UAV video with a low frame rate (e.g., 1 frame / second), a moving object has no overlapping regions between two consecutive frames, which is why conventional object initialization methods do not work. To efficiently isolate promising moving objects from all detection results for a given frame, one or more embodiments of the invention combine a distance matrix with a similarity measurement for initializing moving objects. The processing steps may include, for example, the following: A search radius, a threshold for the degree of match, and a minimum length for the tracking history are defined. The distance matrix between the objects (including object candidates) and all the BLOBs in the table is calculated. If the object trajectory length is less than the preset value, a kernel-based algorithm is applied to find the match between the object candidate and BLOBs with respect to a preset degree of match. If the object candidate appears in several consecutive frames, this candidate is initialized and stored in the object table. Otherwise, the object candidate is considered a false object.
[0061] One or more embodiments of the invention include projecting the previous BLOB from the previous frame into a current frame after its geometric registration. The movement of each object according to its previous position can be estimated by a KLT tracking process. In a KLT tracking process, a motion model is represented approximately by an affine transformation such that I curr (A· x + T) = I prev (x), where A is a two-dimensional (2D) transformation matrix and T is the translation vector.
[0062] In one or more embodiments of the invention, the affine transformation parameters can be calculated starting from only four feature points. To determine these parameters, a least squares technique can be used for their calculation.
[0063] An accuracy estimate can be performed, for example, when a number of non-matching pairs occur. One measure of tracking accuracy is the mean squared error (RMSE) between the matching points before and after the affine transformation formula. This measure is used as a criterion for eliminating matches deemed inaccurate.
[0064] To eliminate outliers, one or more embodiments of the invention additionally include performing the RANSAC algorithm to remove mismatches successively in an iterative manner until the RMSE value is less than the desired threshold.
[0065] The techniques described herein also include mean shift tracking and object representation. For example, in a UAV tracking system, traditional intensity-based target representation is no longer suitable for multi-object tracking due to the large size variations and perspective geometric distortion. To efficiently identify the object, a histogram-based feature space can be chosen. In one or more embodiments of the invention, a metric based on the Bhattacharyya coefficient is used to define a similarity measure between a reference object and a candidate for multi-object tracking. For a given object region histogram q in the reference single image, the objective function based on the Bhattacharyya coefficient is given by: ρ(p,q)=∑u=1Mpu(x)qu(x0) where M is the histogram dimension and x0 is the 2D center.
[0066] The candidate region histogram p u (x) in the 2D center x of the current frame is defined as: pu(x)=∑k(‖x−xih‖2)δ(b(xi),u)∑k(‖x−xih‖2)
[0067] Here, u = 1, 2, ..., M. k(x) denotes a non-negative, non-increasing, and piecewise differentiable kernel profile for weighting the pixel position, h is a 2D bandwidth vector of k(x), δ is the Kronecker delta function, and each pixel value is represented by b(x). i ) designated.
[0068] Additionally, in one or more embodiments of the invention, the Bhattacharyya distance can be used to determine a similarity measure between distributions. B(Ix,Iy)=1−ρ(px,py) include, whereby ρ(px,py)=∫p^x(u)p^y(u) du and ρx and py to represent the target and candidate distributions.
[0069] The techniques described herein can also include object localization. To search for the position corresponding to the object from one frame to the next, one or more embodiments of the invention include the application of a mean-shift tracking algorithm based on gradient optimization rather than a comprehensive search. The strengths of the mean-shift method include its computational efficiency and suitability for real-time application. However, a target may be lost, for example, due to an inherent limitation of the investigated local maxima, especially if the tracked object is moving rapidly. The candidate region histogram p u (x) can be obtained from the equation above.
[0070] The new position of the tracked object can be estimated as: y^1=∑i=1nXiωig(‖y^0−Xih‖2)∑i=1nωig(‖y^0Xih‖2) where: ω1=∑u=1mδ[b(Xi)−u]q^up^u(y^0) g(x) = -k(x), the derivative of k(x).
[0071] One or more embodiments of the invention may also include updating the target model in a temporal domain. In some cases, a mean-shift approach without updating the target model can suffer from sudden changes in the target model. On the other hand, updating the model for each individual frame can lead to a decrease in the reliability of the tracking results due to a cluttered environment, occlusion, random noise, etc. One way to modify the target model is to periodically update the target distributions.
[0072] To obtain a precise tracking result, the target model can be dynamically updated. Accordingly, one or more embodiments of the invention can include a model update that uses both recent tracking results and an older target model to influence a current target model for object tracking. The update procedure is formulated as follows: qunew=(1−α)quold+α•pus
[0073] Here, the upper indices "new" and "old" denote the newly obtained target model and the old model, respectively. s represents the more recent tracking result. α weights the contribution of the more recent tracking result (usually < 0.1). q and p represent the target model and the candidate model, respectively.
[0074] Furthermore, one or more embodiments of the invention include updating the target model in a spatial domain. Normally, mean-shift tracking provides little to no precise boundary position of the tracked object due to the lack of spatial data. Fortunately, the detection results derived from the KLT tracker and the motion detection results can provide much more precise information, such as the exact position and object size, compared to mean-shift tracking.
[0075] No single algorithm, on its own, can perfectly perform the task of multi-object tracking. Therefore, a merging of their data can be used in a multi-object tracking procedure. Depending on the strengths of each method, one or more embodiments of the invention use the following hybrid method: Output={result by motion det ector;if Overlapping≥TKLT result;if Outlier for MS occursresult by meanshift;otherwise where "overlapping" represents the degree of overlap of the region.
[0076] Fig. Figure 7 is a diagram illustrating the reference plane-based registration and tracking according to an embodiment of the present invention. For illustration, it shows Fig. 7 a georeference plane 702. The first single image 704 is registered on the georeference plane 702, and the second single image 706 is created from the first registered single image and with corresponding single-image transformation parameters TC, (equation 712 in Fig. 7) registered on georeference plane 702. In this way, individual images 708 and 710 are registered on georeference plane 702. In addition, each object is projected onto georeference plane 702 using navigation data.
[0077] Fig. Figure 8 is a flowchart illustrating the automatic extraction of traffic routes according to an embodiment of the present invention. Step 802 includes splitting an image into individual frames. Step 804 includes performing a Gaussian smoothing operation. Step 806 also includes the use of a Canny detector, and Step 808 includes performing a Hough transform. Step 810 includes determining a maximum address result. Step 812 includes determining whether the strip length is greater than a predefined threshold. If the strip length is not greater than the predefined threshold, the process terminates at Step 814. If the strip length is greater than the predefined threshold, the process continues with Step 816, which includes performing the extraction of a straight line.Furthermore, step 818 includes performing the removal of stripe pixels (which may, for example, lead to a return to step 808).
[0078] As in Fig. As shown in Figure 8, step 820 includes performing single-frame differentiation, and step 822 includes examination using motion history images (MHIs) (which may, for example, lead back to step 816). Additionally, one or more embodiments of the invention may also include the extraction of road stripes by iterative Hough transform.
[0079] As explained herein, one or more embodiments of the invention include recursive geometric registration with subpixel matching accuracy, which can handle various residual geometric errors of an uncalibrated camera. Additionally, the techniques described herein include motion detection based on forward / backward frame differentiation, which can efficiently separate moving objects from the background. Furthermore, a hybrid object tracker can be implemented that uses colors, features, and intensity statistics over time to detect and track multiple small objects.
[0080] Fig. Figure 9 is a block diagram illustrating the architecture of an object detection and tracking system according to one aspect of the invention. An exemplary software architecture setup for a detection and tracking system (for example, a UAV system) can consist of several services to provide a tracking database for object search and intelligent analysis. As shown in Fig. As illustrated in Figure 9, the software architecture can include several sensor modules 904, video streaming service modules 906, tracking suite service modules 908, a tracking database (DB) server module 910, a user interface module 902, and a visual console 912. A video streaming module 906 is used to capture and make available images from multiple sensors. The captured images are used by a tracking suite module 908 as the basis for multi-object detection and tracking. The tracking suite modules 908 include a geometric registration submodule 914, a motion extraction submodule 916, an object tracking submodule 918, a tracking data submodule 920, and a geocoordinate mapping submodule 922.
[0081] By processing real-time images from multiple sensors, a technically sophisticated conversion of data into tracking information is achieved. The 910 tracking database server manages tracking metadata. The 912 display console generates graphical overlays, indexes them on the display, and presents them to the user. These overlays can contain graphical information of any type, supporting higher-level components such as class types, directions of movement, trajectories, and object sizes. The 902 user interface allows the user to access and operate the data.
[0082] Fig. Figure 10 is a flowchart illustrating techniques for performing visual monitoring of one or more moving objects according to an embodiment of the present invention. Step 1002 includes registering one or more images acquired by one or more cameras, wherein the registration of the image or images comprises region-based registration of the image or images in two or more adjacent frames. This step can be performed, for example, using a geometric registration submodule 914 in the tracking suite service module 908. The image registration can include recursive global and local geometric registration of the image or images (for example, region-based geometric transformation to handle geometric lens distortion). The image registration can also include the use of subpixel image comparison techniques.
[0083] Step 1004 includes performing motion decomposition of the image(s) to detect one or more moving objects and one or more background regions in the image(s). This step can be performed, for example, using a motion extraction submodule 916 in the Tracking Suite service module 908. Performing motion decomposition of the images can include forward and backward frame differentiation. Forward and backward frame differentiation can include, for example, automatic dynamic threshold estimation based on temporal and / or spatial filtering, removal of false moving pixels based on independent motions of image features, and performing a morphological operation and generating motion pixels.
[0084] Step 1006 includes tracking the moving object(s) to facilitate visual monitoring of the moving object(s). This step can be performed, for example, using an object tracking submodule 918 in the tracking suite service module 908. Tracking moving objects can include performing hybrid target tracking, which may involve using a Kanade-Lucas-Tomasi feature tracker and mean shift, auto-kernel scale estimation and updates, and feature trajectories. One or more embodiments of the invention may also include using colors for tracking. Tracking moving objects may additionally include using multi-target tracking algorithms based on feature comparison and distance matrices for one or more (small) targets.
[0085] Tracking moving objects can also include generating a motion map, identifying one or more moving object(s) (BLOBs), performing object initialization and verification, identifying object regions in the motion map, extracting features, defining a search region in the motion map, identifying candidate regions in the motion map, tracking mean shifts, identifying moving objects in the candidate regions, performing a Kanade-Lucas-Tomasi feature comparison, performing an affine transformation (using RANSAC), determining end regions using the Bhattacharyya coefficient, and updating a target model and trajectory information. Tracking moving objects can additionally include reference plane-based registration and tracking.
[0086] The in Fig. The techniques shown in Figure 10 can also include relating each camera view to one or more other camera views and creating a panoramic view from the images captured by one or more cameras. One or more embodiments of the invention additionally include estimating the movement of each camera based on the video information about static objects in the panoramic view, as well as estimating one or more background structures (for example, streets) in the panoramic view based on the recognition of linear structures and the statistical analysis of moving objects over a period of time.
[0087] Furthermore, the in Fig. The technique shown in Figure 10 introduces automatic feature extraction (for example, of a road), where the automatic feature extraction includes decomposing an image into individual frames, performing a Gaussian smoothing operation, using a Canny detector to extract one or more feature edges (for example, road), implementing a Hough transform for feature analysis (for example, road strips), determining a maximum address result to reduce the influence of multiple peaks in a transformation space, determining whether the length of a feature (for example, road strips) is greater than a certain threshold, and if the length of the feature is greater than the threshold, performing feature extraction and pixel removal. The automatic feature extraction can additionally include performing frame-by-frame differentiation and verification using motion history images.
[0088] One or more embodiments of the invention also include performing outlier removal to eliminate false matches of moving objects (and thus increase registration accuracy). The in Fig. The 10 techniques shown can additionally include filtering for false BLOBs. Filtering for false BLOBs involves generating a motion image, applying a connected component process to join all BLOB data, generating a motion BLOB table, extracting features for each BLOB from a previously recorded frame, applying a Kanade-Lucas-Tomasi method to estimate the motion of each BLOB, and deleting a BLOB from the BLOB table if no motion occurs for that BLOB.
[0089] Additionally, one or more embodiments of the invention can include updating a target model in a temporal domain and / or a spatial domain, as well as generating an index (for example, a searchable index) of object appearances and object tracking in a panoramic view. A template index of the object appearances and tracking can also be stored in a template data store with a pointer to the corresponding video segments for easy retrieval. Furthermore, one or more embodiments of the invention can include determining a similarity metric between a query and an entry in the index, which can facilitate searching for the object appearance and tracking in a template data store / index based on the similarity metric, and outputting / listing the search results for an operator based on the similarity of the query.
[0090] The in Fig. The techniques shown in Figure 10 may also include, as described herein, the provision of a system comprising separate software modules, each of which is executed on a physical, computer-readable, writable storage medium. All modules (or some of them) may, for example, reside on the same medium, or each may reside on a different medium. The modules may include some or all of the components shown in the figures. In one or more embodiments, the modules include sensor modules, video streaming service modules, tracking suite service modules (including the submodules described herein), a tracking database (DB) server module, a user interface module, and a visual console module, which may, for example, run on one or more hardware processors.The process steps can then be performed as described above using the system's separate software modules, which run on the one or more hardware processor(s). Furthermore, a computer program product can include a physical, computer-readable, writable storage medium containing code suitable for executing one or more of the process steps described herein, including providing the system with the separate software modules.
[0091] Additionally, the in Fig. The techniques shown in Figure 10 are implemented via a computer program product that may comprise computer-usable program code stored in a computer-readable storage medium in a data processing system, wherein the computer-usable program code has been downloaded from a remote data processing system via a network. In one or more embodiments of the invention, the computer program product may comprise computer-usable program code stored in a computer-readable storage medium in a server data processing system, wherein the computer-usable program code is downloaded via a network to a remote data processing system for use in a computer-readable storage medium with the remote system.
[0092] As those skilled in the art will recognize, aspects of the present invention can be implemented as a system, a method, or a computer program product. Therefore, aspects of the present invention can take the form of a complete hardware embodiment, a complete software embodiment (including firmware, memory-resident software, microcode, etc.), or an embodiment that combines software and hardware aspects, all of which may herein be generally referred to as a "circuit," "module," or "system." Furthermore, aspects of the present invention can take the form of a computer program product implemented on one or more computer-readable media with computer-readable program code executed thereon.
[0093] One or more embodiments of the invention or elements thereof can be realized in the form of a device comprising a memory and at least one processor which is connected to the memory and is operable to perform exemplary process steps.
[0094] One or more embodiments may use software that runs on a general-purpose computer or workstation. Referring to Fig. 11. Such an implementation can, for example, use a processor 1102, a memory 1104, and an input / output interface consisting, for example, of a screen 1106 and a keyboard 1108. The term "processor," as used herein, is intended to include any processing unit, such as one with a CPU (central processing unit) and / or other forms of processing circuitry. Furthermore, the term "processor" can refer to more than one individual processor. The term "memory" is intended to include memory associated with a processor or CPU, such as RAM (random access memory), ROM (read-only memory), a solid-state storage device (e.g., a hard disk), a removable storage device (e.g., a floppy disk), flash memory, and the like.Additionally, the term "input / output interface," as used herein, shall include, for example, one or more mechanisms for inputting data into the processing unit (for example, a mouse) and one or more mechanisms for outputting the results of that processing unit (for example, a printer). The processor 1102, the memory 1104, and the input / output interface, such as the screen 1106 and the keyboard 1108, may, for example, be interconnected as part of a data processing unit 1112 via a bus 1110. Suitable intermediate connections, for example via the bus 1110, may also exist to a network interface 1114, such as a network card, intended for connection to a computer network, and to a data storage interface 1116, such as a floppy disk or CD-ROM drive, intended for accessing data storage media 1118.
[0095] Accordingly, computer software containing instructions or code for carrying out the methods of the invention, as described herein, can be stored in one or more of the associated memory units (for example, ROM, fixed or removable memory) and, when available, partially or completely loaded (for example, into RAM) and executed by a CPU. Such software can include, but is not limited to, firmware, resident software, microcode, and the like.
[0096] A data processing system suitable for storing and / or executing program code will have at least one processor 1102, which is directly or indirectly connected to memory elements 1104 via a system bus 1110. The memory elements may include local memory, which is used during the actual execution of the program code, mass storage, and cache memory, which provides temporary storage of at least certain parts of the program code in order to reduce the frequency with which the code has to be retrieved from mass storage during implementation.
[0097] Input / output or I / O devices (including, but not limited to, keyboards 1108, screens 1106, pointing devices and the like) may be connected to the system either directly (e.g. via the bus 1110) or through intermediate I / O controllers (not shown for clarity).
[0098] Network adapters, such as the 1114 network interface, can also be connected to the system to allow the data processing system to connect to other data processing systems or remote printers or storage devices via intermediary private or public networks. Modems, cable modems, and Ethernet cards are just some of the network adapter types currently available.
[0099] As used herein, including in the claims, a “server” comprises a physical data processing system (for example, the System 1112, as in Fig. (shown in 11) on which a server program is running. It is understood that such a physical server may or may not have a screen and a keyboard.
[0100] As mentioned, aspects of the present invention can take the form of a computer program product implemented on one or more computer-readable media with computer-readable program code executed thereon. Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signal carrier or a computer-readable storage medium. A computer-readable storage medium can be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, a corresponding device or unit, or any suitable combination of the foregoing. The storage media block 1118 is a non-limiting example.More specific examples (a non-exhaustive list) of computer-readable storage medium include: an electrical connection with one or more conductors, a portable computer disk, a hard disk, working memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, computer-readable storage medium can be any physical medium capable of containing or storing a program for use by or in conjunction with a command-execution system or similar device or unit.
[0101] A computer-readable signal carrier can be a disseminated data signal containing computer-readable program code, for example, in the baseband or as part of a carrier wave. Such a disseminated signal can take various forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal carrier can be any computer-readable medium that is not a computer-readable storage medium and that can transmit, disseminate, or transport a program for use by or in conjunction with a command-execution system or a corresponding device or unit.
[0102] Program code executed on a computer-readable medium may be transmitted by any suitable medium, including, but not limited to, wireless, wire, fiber optic cable, radio frequency (RF), etc., or any suitable combination of the foregoing.
[0103] The computer program code for performing operations for aspects of the present invention can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages such as the programming language "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.In the latter scenario, the remote computer can be connected to a user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, via the Internet through an Internet service provider).
[0104] Aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the invention. It is understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a specialized computer, or any other programmable data processing device for the manufacture of a machine such that the instructions executed by the processor of the computer or other programmable data processing device provide means for carrying out the functions / operations specified in the block or blocks of the flowcharts and / or block diagrams.
[0105] These computer program instructions may also be stored in a computer-readable medium that can instruct a computer, other programmable data processing device, or other unit to function in a particular manner, such that the instructions stored in the computer-readable medium produce a product of instructions that implement the functions / operations specified in the block or blocks of the flowcharts and / or block diagrams.
[0106] The computer program instructions can also be loaded into a computer, other programmable data processing device, or other units to effect the execution of a series of operations on the computer, other programmable device, or other units to produce a computerized process, such that the instructions executed on the computer or other programmable device produce processes for performing the functions / operations specified in the block or blocks of the flowcharts and / or block diagrams.
[0107] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, processes, and computer program products according to various embodiments of the present invention. In this context, each block in the flowcharts or block diagrams can represent a code module, code segment, or code part comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that the functions mentioned in the blocks may occur in a different order than that shown in the figure in some alternative implementations. For example, two blocks shown sequentially may actually be executed essentially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved.It should also be noted that each block of the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can be executed by systems based on special hardware that perform the specified functions of the operations, or by combinations of special hardware and computer commands.
[0108] It should be noted that each of the methods described herein may include an additional step of providing a system comprising separate software modules running on a computer-readable storage medium; the modules may, for example, include some or all of the Fig.The 9 components shown are included. The process steps can then be carried out as described above using separate software modules and / or submodules of the system, which are executed on one or more hardware processors 1102. Furthermore, a computer program product can include a computer-readable storage medium containing code suitable for execution to perform one or more of the process steps described herein, including the provision of the system with the individual software modules.
[0109] In any case, it is understood that the components illustrated herein can be implemented in various forms of hardware, software, or combinations thereof; for example, application-specific integrated circuits (ASICs), functional circuits, one or more suitably programmed general-purpose digital computers with associated memory, and the like. Based on the teachings of the invention given herein, a person skilled in the art will be able to consider other implementations of the components of the invention.
[0110] The terminology used herein serves only to describe certain embodiments and is not intended to limit the invention in any way. The singular forms "a" and "the" as used herein are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, it is understood that the expressions "has" and / or "having" when used in this patent specification indicate the presence of the aforementioned features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0111] The corresponding structures, materials, processes, and correspondences of all means or steps and functional elements in the following claims are intended to include all structures, materials, and processes for carrying out the function in combination with other claimed elements, as specifically claimed. The description of the present invention is intended for illustration and description purposes, but is not exhaustive or limited to the invention as disclosed. Many modifications and variations will occur to those skilled in the art without departing from the scope and spirit of the invention. The embodiment has been chosen and described to best explain the principles of the invention and its practical application, and to enable other skilled persons to understand the invention for various embodiments with different modifications suitable for the intended specific application.
[0112] At least one embodiment can provide one or more advantageous effects, such as automatic dynamic threshold determination based on the temporal and / or spatial domain.
[0113] It is understood that the exemplary embodiments of the invention described above can be implemented in numerous different ways. Based on the teachings of the invention presented herein, a person skilled in the art will be able to consider other implementations of the components of the invention. Although illustrative embodiments of the present invention have been described herein with reference to the accompanying drawings, it is understood that the invention is not limited to these specific embodiments and that a person skilled in the art may make various other changes and modifications.
Claims
[1] Method for carrying out visual monitoring of one or more moving objects, wherein the method comprises: Registering one or more images taken by multiple cameras of an unmanned aerial vehicle, wherein the registration of the image or images comprises a recursive global and local geometric registration of the one or more images in two or more adjacent frames (302, 304, 306; 402, 404, 406), wherein the recursive global and local geometric registration includes: (i) Dividing each of the two or more adjacent frames (302, 304, 306; 402, 404, 406) into several sub-areas, each comprising one or more sub-areas associated with a candidate frame and one or more sub-areas associated with a reference frame; (ii) Determining a corner for each of the sub-areas associated with a candidate image and each of the sub-areas associated with a reference image by implementing a multi-resolution technique; (iii) Establishing a correspondence between the individual sub-areas of a candidate image and the sub-areas of a reference image with subpixel accuracy; (iv) Estimating local transformation parameters (TC) i ) for the individual sub-areas using recursive outlier removal and least squares method; (v) Registering all pixels of the individual sub-areas of a candidate image with a reference image; and (vi) Implementing forward and backward frame-to-frame registration by repeating steps (i) to (v) for each of the two or more adjacent frames (302, 304, 306; 402, 404, 406); Performing motion decomposition of the image or images to detect one or more moving objects and one or more background region(s) in the image or images, wherein the performance includes an automatic estimation of a dynamic motion threshold based on spatial filtering; Combining a distance matrix with a similarity measure to (i) initialize a moving object from one or more detected moving objects that satisfies one or more parameters, and (ii) ignore an object from the detected moving object(s) as a false moving object that does not satisfy the one or more parameters; and Tracking the initialized moving object(s) to facilitate visual monitoring of the initialized moving object(s). [2] Method according to claim 1, wherein the registration of one or more images comprises the use of one or more subpixel image comparison techniques. [3] Method according to claim 1, wherein performing the motion decomposition of the image or images includes forward and backward single-image differentiation. [4] Method according to claim 3, wherein the forward and backward frame differences have an automatic dynamic threshold estimation based on temporal filtering and / or spatial filtering. [5] Method according to claim 3, wherein the forward and backward frame differences comprise performing a morphological operation and generating one or more motion pixels. [6] Method according to claim 1, wherein the tracking of the moving object(s) comprises performing hybrid target tracking, wherein the hybrid target tracking comprises the use of a Kanade-Lucas-Tomasi feature tracker and mean shift, the use of auto-kernel scale estimation and update, and the use of one or more feature trajectories. [7] Method according to claim 1, wherein the tracking of the initialized moving object(s) comprises the use of one or more multi-target tracking algorithms based on feature comparison and distance matrices for one or more targets. [8] Method according to claim 1, wherein the tracking of the moving object(s) comprises: Generating a motion image; Identifying one or more moving objects; Performing object initialization and object verification; Identifying one or more object regions in the motion image; Extracting one or more features; Defining a search region in the motion image; Identifying one or more candidate regions in the motion image; Tracking the mean shift; Identifying one or more moving objects in the candidate region(s); Performing the Kanade-Lucas-Tomasi trait comparison; Performing an affine transformation; Performing a determination of end regions using the Bhattacharyya coefficient; and Updating a target model and trajectory information. [9] Method according to claim 1, wherein the tracking of the initialized moving object(s) comprises reference plane-based registration and tracking. [10] Method according to claim 1, further comprising relating each camera view to one or more other camera views. [11] Method according to claim 1, further comprising forming a panoramic view from the image or images taken by the cameras. [12] Method according to claim 11, further comprising estimating the movement of each camera based on the video information about one or more static objects in the panoramic view. [13] Method according to claim 11, further comprising estimating one or more background structures in the panorama view based on the detection of linear structures and the statistical analysis of the moving object(s) over a period of time. [14] Method according to claim 1, further comprising an automatic feature extraction, wherein the automatic feature extraction comprises: Breaking down an image into individual images; Performing a Gaussian smoothing operation; Using a Canny detector to extract one or more feature edges; Implementing a Hough transformation for feature analysis; Determining a maximum response result to reduce the influence of multiple peaks in a transformation space; Determine if the length of a feature is greater than a certain threshold, and perform feature extraction and pixel removal if the length of the feature is greater than the threshold. [15] Method according to claim 14, wherein the automatic feature extraction further comprises performing single-image differentiation and testing on motion history images. [16] Method according to claim 1, further comprising filtering false BLOBs, wherein the filtering of false BLOBs comprises: Generating a motion image; Applying a related component process to link all BLOB data; Creating a motion BLOB table; Extracting one or more features for each BLOB in a previously registered frame; and Applying a Kanade-Lucas-Tomasi procedure to estimate the movement of each BLOB, and deleting the BLOB from the BLOB table if no movement occurs for the BLOB. [17] Method according to claim 1, further comprising updating a target model in a temporal domain and / or a spatial domain. [18] Method according to claim 1, further comprising generating an index of object appearances and object tracking in a panoramic view. [19] Method according to claim 10, further comprising determining a similarity metric between a query and an entry in the index. [20] The method of claim 1, further comprising providing a system, wherein the system comprises one or more separate software modules, each of the one or more separate software modules being executed on a material, computer-readable writable storage medium, and wherein the separate software module(s) comprise a geometric registration module, a motion extraction module and an object tracking module, which are executed on a hardware processor. [21] System for carrying out visual monitoring of one or more moving objects, comprising: a storage facility; and at least one processor that is connected to the memory and capable of being used to: to register one or more images taken by multiple cameras of an unmanned aerial vehicle, wherein the registration of the image or images comprises a recursive global and local geometric registration of the one or more images in two or more adjacent frames (302, 304, 306; 402, 404, 406), wherein the recursive global and local geometric registration includes: (i) Dividing each of the two or more adjacent frames (302, 304, 306; 402, 404, 406) into several sub-areas, each comprising one or more sub-areas associated with a candidate frame and one or more sub-areas associated with a reference frame; (ii) Determining a corner for each of the sub-areas associated with a candidate image and each of the sub-areas associated with a reference image by implementing a multi-resolution technique; (iii) Establishing a correspondence between the individual sub-areas of a candidate image and the sub-areas of a reference image with subpixel accuracy; (iv) Estimating local transformation parameters (TC) i ) for the individual sub-areas using recursive outlier removal and least squares method; (v) Registering all pixels of the individual sub-areas of a candidate image with a reference image; and (vi) Implementing forward and backward frame-to-frame registration by repeating steps (i) to (v) for each of the two or more adjacent frames (302, 304, 306; 402, 404, 406); to perform the motion decomposition of the image or images in order to detect one or more moving objects and one or more background region(s) in the image or images, wherein the performance includes an automatic estimation of a dynamic motion threshold based on spatial filtering; to combine a distance matrix with a similarity measure to (i) initialize a moving object from one or more detected moving objects that satisfies one or more parameters, and (ii) ignore an object from the detected moving object(s) as a false moving object that does not satisfy the one or more parameters; and to track the initialized moving object(s) in order to facilitate visual monitoring of the initialized moving object(s). [22] Computer program comprising computer program code to perform all steps of the method according to claims 1 to 20 when loaded into a computer system and executed.
Citation Information
Patent Citations
Stabilization of objects within a video sequence
US20070104383A1
Image registration method
EP2006805A2