Stable tracking in long inspection videos

The method addresses labor-intensive and unstable tracking issues in video analytics by using local descriptors and dynamic time warping for stable tracking in long inspection videos, enhancing reliability and reducing manual effort.

WO2025207083A1PCT designated stage Publication Date: 2025-10-02HITACHI AMERICA LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/021515
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Current inspection methods for infrastructure and industrial spaces are labor-intensive, time-consuming, and prone to human error, with existing video analytics systems facing challenges in stable tracking due to annotation costs, tracking instability, and the need for extensive manual labor in generating training data.

Method used

A method for stable tracking in video sequences using local descriptors, keyframe estimation, and dynamic time warping to enhance object tracking stability, reduce fragmented tracks, and improve reliability by aggregating tracks and rejecting outliers.

Benefits of technology

Enhances object tracking stability and reduces labor and time by improving anchor point identification, reducing fragmented tracks, and ensuring accurate, reliable tracking in long-form videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000023_0000
    Figure 00000023_0000
  • Figure 00000024_0000
    Figure 00000024_0000
Patent Text Reader

Abstract

Systems and methods provide for stable tracks and reduced fragmented tracks with visually consistent and transformation invariant anchors. Longfomi video leverages a hash table for feature lookup and track assignment using the geometric and perceptual match over a memory window. A video quality profile assessment module generates video profiles with good-quality frames related to observations that are used to enable fracking. A scene track adjustment module estimates an event window for low and high-frequency tracks associated with observations and reduces scene-level outliers to prevent noise over a long memory video window.
Need to check novelty before this filing date? Find Prior Art

Description

STABLE TRACKING IN LONG INSPECTION VIDEOSBACKGROUNDField

[0001] The present disclosure is generally directed to inspection systems and methods for identifying events in image data, and more specifically, to stable tracking in inspection videos.Related Art

[0002] Regular inspection of infrastructure, such as sewer lines, tunnels, and bridges, is vital for ensuring their safety and reliability. In North America alone, there are over 1.2 million miles of sewage pipelines, but only a fraction of this vast network is inspected annually. The current pipeline assessment processes are labor-intensive and time-consuming, which limits the extent of inspections that can be carried out. During inspections, inspectors assess the condition of these structures and document any observed damages, their severity, and precise locations. This information aids in forecasting potential future collapses and strategizing repair priorities for issues that pose significant risks. Such inspections are critical in ensuring uninterrupted utility services and reducing the risk of severe injuries. This is not limited to sewage systems. For instance, amusement parks have motion ride experiences that need to be monitored for safe operations. This process is manual and prone to human error. Similarly, industrial-grade inspections of large spaces using video where a single frame does not provide the needed coverage also present challenges.

[0003] Visual inspection is a widely used method for infrastructure inspections. It involves documentation of visually observed dama ges . The process typically involves the analysis of video recordings of the target infrastructure by certified inspectors. This method allows inspectors to visually assess the surfaces of the infrastructure, which improves the accuracy of damage evaluations and enables infrastructure owners to take prompt action to prevent severe damage. However, the detailed video analysis required by inspectors is both time-consuming and labor- intensive. Further, it carries the risk of overlooking existing damages, thus potentially compromising the accuracy of the evaluations.

[0004] To address these challenges, video analytics-based systems have been developed for automation in industrial settings. These systems use machine learning algorithms to analyze videos and provide inspection information such as the locations, sizes, and other details of defects. This information can be used to identify frames with defects and other issues, thereby helping inspectorsquickly find damages and bypass non-problematic frames. As a result, machine learning-based video analysis can significantly reduce inspection time and labor costs. One existing approach involves a vision-based defect detection method that feeds surface images into machine-learning models to detect defects in input images. If the model detects any defects, their locations are displayed inside images using visualization methods such as bounding boxes around defect locations. Yet, to accurately evaluate infrashuetme conditions, it is essential to report each defect, necessitating a method for linking detection results across video frames. One way to achieve this is by using tracking-by-detection, which is a common method to associate detected defects between frames. In this technique, defect detectors ar e first applied to each frame of an input video, and output detection results are obtained for each frame. Subsequently, detection results such as bounding boxes are assoc iated between frames based on the loca t ion of the box and its appearance. Through this tracking, one track result is obtained for each defect.

[0005] The process of generating training data for these models requires extensive manual labor. For example, recent deep-neural network-based methods require tens of thousands of annotated images to achieve accurate prediction results, meaning that tens of thousands of images need to be annotated with bounding boxes and their corresponding defect class labels. This can result in substantial annotation costs. The problem is further exacerbated when defect locations are not well-defined by bounding boxes, which can cause confusion and prolong the annotation process. Additionally, if trackers need to be trained to obtain highly accurate hacking results, ground truth bounding boxes must be annotated with frame association information so that trackers can extract crucial information to associate bounding boxes during training. This annotation requirement further increases the labor involved.

[0006] Therefore, it is desirable to have systems and methods that overcome the limitations of existing methods.SUMMARY

[0007] In some aspects of the disclosure, a method for tracking an object in video sequences comprises: for a region of interest within a video frame that comprises an object, defining a set of local descriptors for neighboring cells within a grid of cells of the video frame to improve identification of an anchor point on the object in a video sequence, thereby increasing object tracking stability and reducing fragmented tracks: using at least one of a full reference-based metric or a no-reference metric to perform a keyframe estimation for representative frames to assess avideo quality; and adjusting an object tack by performing steps comprising: tacking signal metrics to aggregate tracks; estimating an event window for high and low frequency tracks to detect and reject outliers in the aggregated tacks: using dynamic time warping to characterize and compare tracks across a time series: and applying track matching to the event window to sample and identify tacks that have detailed features to enhance object tracking reliability.

[0008] The keyframe estimation may be based on a quality attribute range, and the hack assignment module may use a recognition measure or a similarity measure of track matching to merge or assign an existing track identifier to enhance stability. The signal metrics may include a frequency, a size, or a distribution statistic,

[0009] The frack matching may use geometric hashing, perceptual hash, or nearest neighbor to identify, independent of transforma tions or occlusion, a vector in a subsequent frame to establish a stable track of the object. The geometric hashing may be used to recognize discrete points representing the object in subsequent frames. The discrete points are recognized where detections were available in prior frames but are now unavailable or detected with low confidence.

[0010] In a pre-processing step, previous detection and neighboring local keypoints may be encoded by basis in a hash table to be used later for next fr ame recognition by geometric matching. The perceptual hashing may be implemented by' a deterministic hash table-based search or a deep neural network that assesses feature similarity- of the region of interest in the neighborhood of prior detections. Perceptual hashing may further serve as a feature similarity test by computing localitysensitive hash of the next frame and comparing it with the previous features available in the hash table for a minimum hamming distance. Motion vectors associated with local features or keypoints may be estimated to update the track, the motion vectors representing a movement of the object across video frames.

[0011] In some aspects, the techniques described herein relate to a method, further including using a track assignment module that uses matching information of the local features to connect the fragmented tracks of the object, thereby’ maintaining fr ack continuity. Further, a tracking tool may- be used to treat the object as an anchor point to estimate a frack of the object in the event window. In some aspects, for each frame in the video, the no-reference metric may’ provide a profile of the video quality.

[0012] Aspects of the present disclosure can involve a system, which can involve means for performing steps comprising: defining a set of local descriptors for neighboring cells within a gridof cells of the video frame for a region of interest within a video frame that comprises an object to improve identification of an anchor point on the object in a video sequence to increase object tracking stability and reduce fragmented tracks; means for performing steps comprising using a full reference-based metric or a no-reference metric to perform a keyframe estimation for representative frames to assess a video quality; and means for performing steps comprising adjusting an object track by performing steps comprising: tracking signal metrics to aggregate tracks; estimating an event window for high and low frequency hacks to detect and reject outliers in the aggregated tracks; using dynamic time warping to characterize and compare tracks across a time series; and applying track matching to the event window to sample and identify tracks that have detailed features to enhance object tracking reliability.BRIEF DESCRIPTION OF DRAWINGS

[0013] FIG. 1 illustrates a conventional flow for exemplary automated detection and reporting use cases.

[0014] FIG. 2A - FIG. 2C depict sample defect tracks a sequence of frames in a CCTV inspection video.

[0015] FIG. 3 provides an overview of a stable tracking system for fracking an object in a video sequence, according to various embodiments of the present disclosure.

[0016] FIG. 4 illustrates an exemplary system for stable tracking of an object in a video sequence, according to various embodiments of the present disclosure.

[0017] FIG. 5 illustrates an exemplary process for stable anchor track estimation, according to various embodiments of the present disclosure.

[0018] FIG. 6 illustrates an exemplary workflow for track assignment, according to various embodiments of the present disclosure.

[0019] FIG. 7 is a flowchart illustrating an exemplary process for visual quality profile crea tion, according to various embodiments of the present disclosure.

[0020] FIG. 8 is a flowchart illustrating an exemplary process for scene track adjustment and track selection, according to various embodiments of the present disclosure.

[0021] FIG. 9 illustrates an exemplary user interface concept for human augmentation of periodic or continuous tracks in inspection videos, according to various embodiments of the present disclosure.

[0022] FIG. 10 is a flowchart illustrating an exemplary process for tracking an object in video sequences, according to various embodiments of the present disclosure.

[0023] FIG. 11 illustrates an example computing environment with an example computer device suitable for use in various embodiments of the present disclosure.DETAILED DESCRIPTION

[0024] The following detailed description provides details of the figures and example implementations of the present application. Reference numerals and descriptions of redundant elements between figures are omitted for clarity. Terms used throughout the description are provided as examples and are not intended to be limiting. For example, the use of the term “automatic” may involve fully automatic or semi-automatic implementations involving user or administrator control over certain aspects of the implementation, depending on the desired implementation of one of ordinary skill in the art practicing implementations of the present application. Selection can be conducted by a user through a user interface or other input means, or can be implemented through a desired algorithm. Example implementations as described herein can be utilized either singularly or hi combination and the functionality of the example implementations can be implemented through any means according to the desired implementations .

[0025] FIG. 1 illustrates a conventional flow for exemplary automated detection and reporting use cases. The use cases depicted in FIG. 1 comprise long video applications 102, such as CCTV deployment and video collection; quality assurance / quality control 104 that employ experts to detect defects offline; standardized reporting 106 that uses CCTV video; and planning operations 108 for activities such as maintenance planning based on reports. Automated detection 112 and reporting 114 necessitate identification, localization, and tracking for observations of interest. Handling scale and translation in defect detections for inspection scenarios is a complex problem. Existing approaches focus on defect detection but lack the ability to tack periodic or continuous observations in long-form video. Furthermore, the tracking-by-detection approach has room for improvements in observing conditions over extensive video information. Other known methods focus primarily on a signal-processing approach for video compression.

[0026] Existing technical problems comprise tracking instability or interruptions when either the camera or the object of interest is in motion. Each time an event is observed, the object can be spread out on the surface or as multiple standalone spots with inconsistent observations or even occlusions. Scale and tanslation-related transformations can also disrupt tracks, hi the special case of a smooth inner surface of the inspection environment, there exists no consistent feature to track. Additionally, the lack of video quality information reduces confidence in track processing results. Oftentimes, minor variations can result in no detections being observed. Conventional tackingdoes not take into account environmental noise, spatial structures, or objects in the field of view, or data collection events, leading to unstable report tracks. Moreover, in long- form video, tracking is unbounded and exploratory rattier than standardized event reporting 114. There is no understanding of temporal windows or periodic features.

[0027] FIG. 2A - FIG. 2C depict sample defect tracks a sequence of frames in a CCTV inspection video. FIG. 2 A - FIG. 2C illustrate the issue of changes in detected defects as encountered in conventional designs and, in particular, how changes in count, size, and accuracy associated with an initial detection can create inconsistencies in localized results across frames. As shown, detection boxes 201 comprising the detection result changes over time within frame FIG. 2A and from FIG. 2A to detection box 202 in FIG. 2B. This is especially true if the system is not well-trained. Further, the detection can be misaligned in subsequent frame FIG. 2B, or missing altogether, thereby creating disconnected tracks in an inspection sequence. Moreover, detections 203 in subsequent flames may be transformed and scaled with pan, tilt, zoom, and unwanted perspective changes within a video sequence, as illustrated in FIG. 2C. Therefore, it is desirable to have systems and methods that overcome the limitations of existing methods.

[0028] FIG. 3 provides an overview of a stable tracking system for tracking an object in a video sequence, according to various embodiments of the present disclosure. System 300 comprises visual quality profile assessment module 302, which represents the first component of system 300; feature stabilizing module 304, feature compression module 306, feature memory module 308, feature matching module 310, and track assignment module 312, which represent the second component of system 300; and outlier detection module 314 and scene track module 316, which represent the third component of system 300. It is understood that outlier detection module 314 may serve as a preprocessing step for scene track module 316.

[0029] As depicted, system 300 receives, at visual quality profile assessment module 302, inspection video data in the form of raw and unprocessed data generated by a long-form CCTV pipeline and outputs stable tracks associated with the inspection video data.

[0030] As discussed in greater detail with reference to FIG. 4, visual quality profile assessment module 302 processes or pre-processes the received video data to create a visual quality profile by performing a frame-level video quality assessment for occlusion, motion blur, and the like. Advantageously, this enables the selection of high-quality video frames to improve overall system accuracy.

[0003] J The processed video data is provided to the second component (denoted as the stable anchor track computation module in FIG. 4). which enhances the stability of features by improving anchor point identification using local descriptors where detections are unavailable or noisy. This is achieved by leveraging transformed structure recognition, such as geometric matching, and feature similarity technology, such as perceptual hashing, as feature compressiondiashing, storage and matching for searching a vector in a next frame to drive tracking. This component also provides memory for track features and has the capability to compare within or across multiple long-form videos. It is noted that geometric matching recognizes a structure based on a scale and translation-like transformation, and that perceptual hashing preserves the feature similarity in the encoded / compressed bytecode .

[0032] The output of the stable anchor track component is provided to the third component of system 300 (denoted as the scene track adjustment module in FIG. 4), which offers scene-level temporal object tracking, dynamic time warping to characterize the tracks over time series, and estimates periodic or similar tracks for event windows estimation.

[0033] FIG. 4 illustrates an exemplary system for stable tracking of an object in a video sequence, according to various embodiments of the present disclosure. System 400 comprises interface 402, external system 406, closed-circuit Television (CCTV) inspection system 408, video database 410, and defect database 420. In operation, system 400 receives, at interface 402, raw CCTV7inspection data 404, which may have been obtained from the CCTV inspection system 408. This system may be communicatively coupled to video database 410 or any other external system 406. Raw video data 404 is input to video processing batch pipeline 430 and may be queued as job 432 and executed as job 436 in a distributed task queue 434. In response to receiving job 436, inspection video decoder 440 decodes raw video data 404 to output decoded video data 441 to video quality profile assessment module 442. As discussed in gr eater detail with respect to FIG. 7, video quality profile assessment module 442 processes the received video data to create a visual quality profile for decoded video data 441.

[0034] Stable anchor track computation module 444, which corresponds to feature stabilizing module 304, feature compression module 306, feature memory module 308, feature matching module 310, and track assignment module 312 in FIG. 3, enhances the stability of features associated with the visual quality profile to generate an input for scene track adjustment module446. This module performs scene-level temporal object tacking to generate stable inspection tacks 460.

[0033] In addition, system 400 may store information in defect database 420 for various data analysis tasks 450, such as assisted review of defects 452, inspection auto Pipeline Assessment Certification Program (PACP) coding 454, audit and quality assessment 456, or defect feature analysis 458.

[0036] In detail, stable anchor track computation module 444 stabilizes tracks by using features or anchor compression and memory and hash matching over a temporal sequence. It may leverage detections where available or, alternatively, rely on local descriptors and optical flow to guide tracking. Stable anchor track computation module 444 may use deterministic approaches like geometric and / or perceptual hashing for feature or anchor encoding. A hash table may guide tracking over a memory window.

[0037] Video quality profile assessment module 442, which corresponds to visual quality profile assessment module 302 in FIG. 3, processes decoded video data 441 by performing a framelevel video quality assessment and representative keyframe estimation for a track profile. This may comprise performing an environment-based representative frame selection process and an outlier rejection process, to create a visual quality profile for decoded video data 441.

[0038] Scene track adjustment module 446, which corresponds to outlier detection module 314 and scene track module 316 in FIG. 3, tracks and extracts metrics, such as frequency, size, or distribution statistics to aggregate tracks for observation reporting. The tracks are characterized, e.g., using dynamic time warping or other statistical methods for a video time series profile.

[0039] Advantageously, stable tracks provided by stable anchor tack computation module 444 reduce fragmented tracks with visually consistent arid transformation invariant anchors. Longform video leverages hash tables for feature lookup and track assignment using the geometric and perceptual match over a memory window. Further, video quality profile assessment module 442 generates video profiles with good-quality frames related to the observation that are used to enable tracking. Environmental factors like occlusion, motion, exposure, noise, and blur that affect video quality are considered in the video profile. Furthermore, scene track adjustment module 446 estimates an event window for low and high-frequency tacks associated with observations and reduces scene-level outliers that otherwise would introduce noise into the vision system over a long memory window in the video.

[0040] It is noted that examples herein are presented in the context of signal, image, or s tatistical processing systems, embodiments of the present disclosure may equally be implemented using machine learning, e.g., deep neural network-based techniques. For example, in accurate automated video coding, an Al-based review or coding system may use tracking to automatically code points and periodic or continuous defects in the inspection videos.

[0041] Further, embodiments of the present disclosure may be used in related applications to understand the behavior of inspection personnel by analyzing track choices, camera velocity and stopping frequency, and the like. Query or retrieval of similar anchors and tracks may be used to search for or retrieve similar frames or tracks from historical data, e.g., by utilizing the matching approach and a hash table.

[0042] In embodiments, tracked features may be used to compare across inspections, e.g., to identify and flag defect patterns across inspections within a geometrical area (e.g., a city) or regional pipeline infrastructure. In embodiments, information a quality7assessment and audit system may be used to score data and evaluate high and low-quality data, e.g., to ensure service quality7and facilitate audits. In embodiments, blockchain techniques may be used to generate trusted sensing metadata. For example, blockchain extensions for a single source of truth for always-on inspection edge devices may exchange spatio-temporal defect condition metadata with an loT sensor network.

[0043] A person of skill in the art will appreciate that, although the examples herein use only one camera, the teachings of the present disclosure may equally be extended to perform object tracking by using two or more cameras.

[0044] FIG. 5 illustrates an exemplary7process for stable anchor track estimation, according to various embodiments of the present disclosure . Process 500 comprises spatial component 520 and temporal component 522. Spatial component 520 comprises object detection module 510 and local descriptor module 530, which each receives an input from inspection video decoder 502. Spatial component 520 further comprises anchor selection module 540. Temporal component 522 comprises object track module 542, local flow module 544, summation module 546, and track matching module 580.

[0045] In operation, object detection module 510 is configured to detect spatial defects or regions of interest (Rol) 516, e.g., in a grid of cells (represented as bounding boxes in start frame 514). Local descriptor module 530 is configured to identify local descriptors / features 530, e.g., forcells in a neighborhood or separated across the grids of ftame 514 in a video of length “n” and local key points 534. Anchor selection module 540 uses the detected Rol 516 and local key points 534 to select, as anchor points, those local keypoints 534 that are located in the proximity of defects 516 in start frame 514. Advantageously, using detected Rol 516 and local descriptors 530 in this manner improves anchor points and, thus, increases stability.

[0046] Object track module 542 uses a tracker, such as SORT or DeepSORT, to estimate tracks for the detected objects as anchors in the video sequence. Local flow module 544 uses a tool such as Optical Flow to estimate motion vectors on local features or keypoints 534 for estimating tracks in the video sequence. In response to receiving current 560 and next frame 564, in which the lack of defect 570 indicates a broken track, track matching module 580 uses geometric hashing, perceptual hash, or nearest neighbor for searching a vector in the next frame regardless of transformations or occlusions to enable stable tracking for an event window or the entire video sequence and output a track similarity. Finally, process 500 outputs stable tracks 590-594. Track matching thus serves as a function for streamlining broken tracks.

[0047] A workflow for track matching and hack assignment in broken tracks leveraging a hash table is shown in FIG. 6. Track matching 580 comprises motion trajectory 602, which is utilized in geometric matching module 604 that performs a geometric hashing process to recognize discrete points. These discrete points represent an object in a next (i+i) frame of a video where detections were available in prior frames but ar e now unavailable or detected with low confidence.

[0048] In a pre-processing step, previous detection and neighboring local keypoints are transformed by basis in a hash table for subsequent use in next frame recognition by geometric matching. Perceptual hashing module 606 assesses feature similarity of the Rol in the neighborhood of prior detections. This serves as a feature similarity test by computing locality7sensitive hash of the next frame and comparing it with the previous features available in the hash table for a minimum hamming distance. Structure matching (e.g.. geometric matching) 620 and similarity matching (e.g., perceptual matching) 622 may be implemented by a deterministic hash table-based search technique or modeled as a deep neural network for recognition and similarity. Track assignment module 608 uses the recognition and similarity7measure of track matching to merge or assign an existing track identifier for stability.

[0049] FIG. 7 is a flowchart illustrating an exemplary process for visual quality profile creation, according to various embodiments of the present disclosure. Process 700 begins at step702 when a reference measure is received. Video quality assessment (VQA) for inspection may be a full reference (FR), e.g., in a well-controlled data collection environment to obtain profiles having several measures of video quality for each frame in a video. Conversely, in inspection video applications that are more likely prone to noise, a more general method such as a No-Reference (NR) metric, e.g., VQI or NIQE may be used to obtain profiles having a single measure of video quality, such as a scalar score for each video frame. After a profile is selected at step 704, a keyframe estimation may be performed, at step 706, e.g., based on a selected quality attribute range to obtain representative video frames. The resulting representative video frames are then output at step 70S.

[0050] FIG. 8 is a flowchart illustrating an exemplary process for scene track adjustment and track selection, according to various embodiments of the present disclosure. Process 800 may start at step 802, when tracks are aggregated for observation reporting, e.g., by using signal parameters, such as frequency, size, or distribution statistics extracted from tracks in the video sequence.

[0051] At step 804, based on, e.g., a distribution obtained at step 802, signal characteristics between video sequences, such as position values, may be tracked, e.g., by using dynamic time warping or any other statistical methods for video time series analysis that allows comparing tracks. To accomplish this, dynamic time warping may utilize any existing libraries.

[0052] At step 806, based on steps 802 and 804, high and low-frequency features or signals for a given estimate period maybe used, e.g., in combination with thresholding or any other known outlier detection technique known in the art, to detect and / or reject unwanted scene-level outliers.

[0053] At step 808, hi response to inspection videos being sampled and processed across an extended-time video sequence, feature-rich tracks are selected by using track matching. While this allows for offline processing for long event window tracks, it is noted that present disclosure is not so limited. In fact, depending on the inspection use case, systems for short event windows may also be implemented to allow for near real-time applications.

[0054] Finally, at step 810, representative video tracks are output and may be used for visualization (as shown in FIG. 9) and / or coding and reporting.

[0055] FIG. 9 illustrates an exemplary user interface concept for human augmentation of periodic or continuous tracks in inspection videos, according to various embodiments of the present disclosure. Depicted are bounding boxes and key points for video frame key AI / ML detections 902: short event signal visualizations 904, where the amplitude of the dark linerepresents camera motion. FIG. 9 further comprises stable tracking inspection visualizations 906: and tirneiine 908, where stable tracking inspection visualizations 906 correspond to different tracks indicating different events, such as cracks and taps in video frame detections 902. It is noted that manual intervention is enabled to allow for scene track adjustments or corrections of the output, e.g., in scenarios where shape transformations are not properly depicted.

[0056] FIG. 10 is a flowchart illustrating an exemplary process for tracking an object in video sequences, according to various embodiments of the present disclosure. Process 1000 may start at step 1002, when, for a region of interest within a video frame that comprises an object, a set of local descriptors are defined for neighboring cells, e.g., within a grid of cells of the video frame, to improve identification of an anchor point on the object in a video sequence to increase object tracking stability and reduce fragmented tracks.

[0057] At step 1004, at least one of a full reference-based metric or a no-reference metric is used to perform a keyframe estimation for representative frames to assess a video quality.

[0058] At step 1006, to adjust an object track, signal metrics are tracked to aggregate tracks;

[0059] At step 1008, an event window is estimated for high and low-frequency tracks to detect and reject outliers in the aggregated tracks;

[0060] At step 1010, dynamic time warping is employed to characterize and compare tracks across a time series.

[0061] Finally, at step 1012, track matching is applied to the event window to sample and identify tracks that have detailed features to enhance object tracking reliability.

[0062] One skilled hi the art shall recognize that: (1) certain steps herein may optionally be performed; (2) steps may not be limited to the specific order set forth herein: (3) certain steps may be performed in different orders: and (4) certain steps may be performed concurrently.

[0063] FIG. 11 illustrates an example computing environment with an example computer device suitable for use in various embodiments of the present disclosure. Computer device 1105 hi computing environment 1100 can include one or more processing units, cores, or processors 1117, memory 1115 (e.g., RAM, ROM, and / or the like), internal storage 1120 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interface 1125, any of which can be c oupled on a communication mechanism or bus 1130 for communicating hiformation or embedded in the computer device 1105. I / O interface 1125 is also configured to receive images from cameras or provide images to projectors or displays, depending on the desired implementation.

[0064] Computer device 1105 can be communicatively coupled to input / user interface 1135 and output device / interface 1140. Either one or both of input riser interface 1135 and output device / interface 1140 can be a wired or wireless interface and can be detachable. Input / user interface 1135 may include any device, component, sensor, or interface, physical or virtual, that can be used to provide input (e.g., buttons, touch-screen interface, keyboard, a pointing / ciii'sor control, microphone, camera, braille, motion sensor, optical reader, and / or the like). Output device / interface 1140 may include a display, television, monitor, printer, speaker, braille, or the like. In some example implementations, input / user interface 1135 and output device / interface 1140 can be embedded with or physically coupled to the computer device 1105. In other example implementations, other computer devices may function as or provide the functions of input / user interface 1135 and output device / interface 1140 for a computer device 1105.

[0065] Examples of computer device 1105 may include highly mobile devices (e.g., smartphones, devices in vehicles and other machines, devices carried by humans and animals, and the like), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more processors embedded therein and / or coupled thereto, radios, and the like).

[0066] Computer device 1105 can be communicatively coupled (e.g., via I / O interface 1125) to external storage 1145 and network 1150 for communicating with any number of networked components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1105 or any connected computer device can be functioning as. providing sendees of, or referred to as a server, client, thin server, general machine, special-purpose machine, or another label.

[0067] I / O interface 1125 can include wired and / or wireless interfaces using any communication or L-'O protocols or standards (e.g., Ethernet, 802.1 lx. Universal System Bus, WiMax, modem, a cellular network protocol, and the like) for communicating information to and / or from at least all the connected components, devices , and network in computing environment 1100. Network 1150 can be any network or combmation of networks (e.g., the Internet, local area network, wide area network, a telephonic network, a cellular network, a satellite network, and the like).

[0068] Computer device 1105 can use and / or communicate using computer-usable or computer-readable media, including transitory media and non-transitory media. Transitory media include transmission media (e.g., metal cables, fiber optics), signals, carrier waves, and the like. Non-transitory media include magnetic media (e.g., disks and tapes), optical media (e.g., CD ROM, digital video disks, Blu-ray disks), solid-state media (e.g., RAM, ROM. flash memory, solid-state storage), and other non-volatile storage or memory.

[0069] Computer device 1105 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some example computing environments. Computer-executable instructions can be retrieved from transitory media, and stored on and retrieved from non-transitory media. The executable instructions can originate horn one or more of any programming, scripting, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).

[0070] Processors) 1110 can execute under any operating system (OS) (not shown), in a native or virtual environment. One or more applications can be deployed that include logic unit 1160, application programming interface (API) unit 1165, input unit 1170, output unit 1175, and inter-unit communication mechanism 1195 for the different units to communicate with each other, with the OS, and with other applications (not shown). The described units and elements can be varied in design, function, configuration, or implementation and are not limited to the descriptions provided. Processor(s) 1110 can be in the form of hardware processors such as central processing units (CPUs) or a combination of hardware and software units.

[0071] In some example implementations, when information or an execution instruction is received by API unit 1165, it may be communicated to one or more other units (e.g., logic unit 1160, input unit 1170, output unit 1175). In some instances, logic unit 1160 may be configured to control the information flow among the units and direct the services provided by API unit 1165, input unit 1170, and output unit 1175, in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled by logic unit 1160 alone or in conjunction with API unit 1165. The input unit 1170 may be configured to obtain input for the calculations described in the example implementations, and the output unit 1175 may be configured to provide output based on the calculations described in example implementations.

[0072] Processors) 1110 can be configured to execute a method or computer instructions which can involve, defining a set of local descriptors for neighboring cells within a grid of cells ofthe video frame for a region of interest within a video frame that comprises an object such as to improve identification of an anchor point on the object in a video sequence, increase object tracking stability, and reduce fragmented tracks, as illustrated in FIG. 5 and FIG. 9.

[0073] Processors) 1110 can be further be configured to execute a method or computer instructions which can involve, using at least one of a full reference-based metric or a no-reference metric to perform a keyframe estimation for representative frames to assess a video quality, as illustrated in FIG. 5 and FIG. 7.

[0074] Processors) 1110 can be further be configured to execute a method or computer instructions which can involve, adjusting an object track by performing steps comprising fracking signal metrics to aggregate tracks; estimating an event window for high and low-frequency tracks to detect and reject outliers in the aggregated tracks; using dynamic time warping to characterize and compare tracks across a time series; and applying track matching to the event window to sample and identify tracks that have detailed features to enhance object tracking reliability, as illustrated in FIG. 5, FIG. 8, and FIG. 9.

[0075] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a series of defined steps leading to a desired end state or result. In example implementations, the steps earned out require physical manipulations of tangible quantities to achieve a tangible result.

[0076] Unless specifically stated otherwise, as apparent from the discussion, it is appreciated that throughout the description, discussions utilizing terms such as “processing,” “computing,” “calculating,” “determining,” “displaying,” or the like, can include the actions and processes of a computer system or other information processing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system’s registers and memories into other data similarly represented as physical quantities within the computer system’s memories or registers or other information storage, transmission or display devices.

[0077] Example implementations may7also relate to an apparatus for performing the operations herein. This apparatus may7be specially7constructed for the required purposes, or it may include one or more general-purpose computers selectively7activated or reconfigured by one or more computer programs. Such computer programs may be stored in a computer-readable medium, suchas a computer -readable storage medium or a computer -readable signal medium. A computer- readable storage medium may involve tangible mediums such as optical disks, magnetic disks, read-only memories, random access memories, solid-state devices, drives, or any other types of tangible or non-transitory media suitable for storing electronic information. A computer-readable signal medium may include mediums such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Computer programs can involve pure software implementations that involve instructions that perform the operations of the desired implementation.

[0078] Various general-purpose systems may be used with programs and modules in accordance with the examples herein, or it may prove convenient to construct a more specialized apparatus to perform desired method steps. In addition, the example implementations are not described with reference to any particular programming language, ft will be appreciated that a variety of programming languages may be used to implement the techniques of the example implementations as described herein. The instructions of the programming langtiage(s) may be executed by one or more processing devices, e.g., central processing units (CPUs), processors, or controllers.

[0079] As is known in the art, the operations described above can be performed by hardware, software, or some combination of software and hardware. Various aspects of the example implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software), which if executed by a processor, would cause the processor to perform a method to carry out implementations of the present application. Further, some example implementations of the present application may be performed solely in hardware, whereas other example implementations may be performed solely in software. Moreover, the various functions described can be performed in a single unit, or can be spread across a number of components in any number of ways. When performed by software, the methods may be executed by a processor, such as a general -purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions can be stored on the medium in a compressed and / or encrypted format.

[0080] Moreover, other implementations of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the techniques of the present application. Various aspects and / or components of the described example implementationsmay be used singly or in any combination. It is intended that the specification and example implementations be considered as examples only, with the true scope and spirit of the present application being indicated by the following claims

Claims

CLAIMSWhat is claimed is:

1. A method for tracking an object in video sequences, the method comprising: for a region of interest within a video frame that comprises an object, defining a set of local descriptors for neighboring cells within a grid of cells of the video frame to improve identification of an anchor point on the object in a video sequence, thereby increasing object tracking stability and reducing fragmented tracks; using at least one of a frill reference-based metric or a no-reference metric to perform a keyframe estimation for representative frames to assess a video quality; and adjusting an object track by performing steps comprising: tracking signal metrics to aggregate tracks; estimating an event window for high and low-frequency tracks to detect and reject outliers in the aggregated fracks; using dynamic time warping to characterize and compare tracks across a time series; and applying track matching to the event window to sample and identify tracks that have detailed features to enhance object tracking reliability.

2. The method according to claim 1, wherein, for each frame in the video, the no-reference metric provides a profile of the video quality.

3. The method according to claim 1, wherein the signal metrics comprise a t least one of a frequency, a size, or a distribution statistic.

4. The method according to claim 1 , further comprising using a tracking tool that treats the object as an anchor point to estimate a track of the object in the event window.

5. The method according to claim 1 , further comprising estimating motion vectors associated with local features or keypoints to update the track, the motion vectors representing a movement of the object across video frames.

6. The method according to claim 1 , further comprising using a track assignment module that uses matching information of the local features to connect the fragmented tracks of the object, thereby maintaining track continuity.

7. The method according to claim 1 , wherein track matching uses at least one of geometric hashing, perceptual hash, or nearest neighbor to identify, independent of transformations or occlusion, a vector in a subsequent frame to establish a stable track of the object.

8. The method according to claim 7, wherein geometric hashing is used to recognize discrete points representing the object in subsequent frames.

9. The method according to claim 8, wherein the at least some of the discrete points are recognized where detections were available in prior frames but are now unavailable or detected with low confidence.

10. The method according to claim 8, wherein, in a pre-processing step, previous detection and neighboring local keypoints are encoded by basis in a hash table to be used later for next fiame recognition by geometric matching.

11. The method according to claim 7, wherein the perceptual hashing is implemented by at least one of a deterministic hash table-based search or a deep neural network that assesses feature similarity of the region of interest in the neighborhood of prior detections.

12. The method according to claim 8, wherein the perceptual hashing serves as a feature similarity test by computing locality sensitive hash of the next fiame and comparing it with the previous features available in the hash table for a minimum hamming distance.

13. The method according to claim 1, wherein the keyframe estimation is based on a quality attribute range.

14. The method according to claim 1, wherein the track assignment module uses at least one of a recognition measure or a similarity measure of track matching to merge or assign an existing track identifier to enhance stability.

15. A non -transitory computer-readable medium for storing instructions for executing a process, the instructions comprising: for a region of interest within a video frame that comprises an object, defining a set of local descriptors for neighboring cells within a grid of cells of the video frame to improveidentification of an anchor point on the object in a video sequence, thereby increasing object tracking stability and reducing fragmented tracks; using at least one of a frill reference-based metric or a no-reference metric to perform a keyframe estimation for representative frames to assess a video quality; and adjusting an object track by performing steps comprising; tracking signal metrics to aggregate tracks; estimating an event window for high and low-frequency tracks to detect and reject outliers in the aggregated tracks; using dynamic time warping to characterize and compare fracks across a time series; and applying track matching to the event window to sample and identify fracks that have detailed features to enhance object tracking reliability7.

Citation Information

Patent Citations

  • Video Analysis Based on Sparse Registration and Multiple Domain Tracking

    US20130335635A1

  • Tracking biological objects over time and space

    US20200364857A1

  • Obstacle recognition method for autonomous robots

    US20220066456A1

  • Vehicle speed estimation systems and methods

    US20230019731A1

  • System for and method of real-time nonrigid mosaicking of laparoscopy images

    WO2022192540A1