FORENSIC VIDEO ANALYSIS AND UTILIZATION TOOLS

MX434000BActive Publication Date: 2026-05-19MASSACHUSETTS INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2021014250
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-20
Filing Date
2021-11-19
Publication Date
2026-05-19
Estimated Expiration
2040-05-20

AI Technical Summary

Technical Problem

Existing video surveillance systems lack effective tools for locating and tracking objects across multiple camera views, especially when camera fields of view do not overlap, and require manual effort to switch between camera feeds, leading to inefficiencies in video review.

Method used

The development of forensic video analysis tools that utilize transition zones and anchor points to automatically track objects across multiple camera views, allowing seamless switching between camera feeds and generating composite videos for efficient review.

Benefits of technology

Enables rapid and accurate tracking of objects across non-overlapping camera views, reducing the time and effort required to locate and review video footage, and enhancing the efficiency of video surveillance operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX434000B0
    Figure MX434000B0
Patent Text Reader

Abstract

This document describes systems and methods for locating an object detected in a video; the system detects a bounding box at least partially around an object in a first frame at a first time in the video and a second frame in the video corresponding to a second time; the system determines if there is no movement within the bounding box of the second frame; the system compares edge information, or color information, or intensity information associated with one or more pixels in the first frame with edge information, or color information, or intensity information associated with one or more pixels within each bounding box; the system generates a score according to the comparison; the system further determines based on the score whether the object is present in the second frame; the system also determines an estimated time period of a first appearance of the object.
Need to check novelty before this filing date? Find Prior Art

Description

FORENSIC VIDEO ANALYSIS AND UTILIZATION TOOLS CROSS-REFERENCE TO RELATED PATENT APPLICATIONS This application claims priority over U.S. provisional application no. 62 / 850,384, filed on May 20, 2019, the contents of which are incorporated herein by reference. DECLARATION OF GOVERNMENTAL INTEREST The present invention was made with government support pursuant to contract number FA8702-15-D-0001 awarded by the United States Air Force. The government holds certain rights with respect to the invention. TECHNICAL FIELD This description refers to techniques for video surveillance. More specifically, this description refers to methodologies, systems, and devices for locating an image at a specific point in a video transmission, for example, the first occurrence of an object, and, if necessary, tracking the object within the field of view of two or more cameras. BRIEF DESCRIPTION OF THE FIGURES The foregoing and other functions and advantages provided by this description will be more fully understood from the following description of example modalities when read in conjunction with the accompanying figures, in which: Figure 1 is a flowchart that illustrates an example method for creating a summary video, according to the modalities of the present description. Figure 2 is an example video timeline that illustrates the formation of a summary video frame according to the example modalities of the present description. Figure 3 represents a screenshot of a graphical user interface for displaying a summary video, according to examples of modalities in this description. Figures 4A to 4C represent screenshots of a graphical user interface for displaying a summary video, according to examples of modalities in this description. Figure 5A represents a screenshot of a graphical user interface for displaying a summary video, according to examples of modalities in this description. Figure 5B represents a screenshot of a graphical user interface for displaying a source video, according to examples of modalities in this description. Figure 6 is a diagram of an example of a backtracking process according to examples of modalities in the present description. Figure 7 represents screenshots of a graphical user interface to display a back function according to examples of modalities in the present description. Figure 8 represents a screenshot of multiple fields of view corresponding to different cameras, and transition zone icons that can be activated by a user to move through the multiple fields of view according to modality examples in this description. Figure 9 is a block diagram of an example computer system that can perform example processes according to example modalities in the present description. Figure 10 is a diagram of an example network environment suitable for a distributed implementation of example modalities of the present description. Figure 11 is a flowchart that illustrates an example method for checking when an object was placed in a camera's field of view, according to the modalities of the present description. Figure 12 is a flowchart that illustrates an example method for tracking an object or person across multiple cameras, according to the modalities of the present description. Figure 13 illustrates a screenshot of a graphical user interface for navigating a camera network, according to embodiments of the present invention. Figure 14 illustrates a screenshot of a graphical user interface for selecting an object when it is in the field of view of multiple cameras and generating anchor points, according to embodiments of the present invention. Figure 15 is a flowchart that illustrates an example method for navigating a camera network using a graphical user interface, according to the modalities of the present description. Figure 16 illustrates a screenshot of a graphical user interface for reviewing composite video recording of an object in the fields of view of multiple cameras, according to the modalities of the present description. Figure 17 illustrates a timeline corresponding to segments of a composite video comprising where each segment is a part of a video transmission from a different camera, according to the modalities of the present description. Figure 18 illustrates a screenshot of a reconstruction tool IVIA / t / ZUZZ / UII Doo trajectory that detects an object moving in the field of view of two cameras and a logical diagram that indicates the location of the transition zone icons that link the two cameras, according to the modalities of the present description. Figure 19 includes a first graph representing the amount of video recording reviewed by a reviewer, using a lite version of a video forensics tool, over a period of time, and a second graph representing the amount of video recording reviewed by a user, when using a full version of the video forensics tool according to the modalities of the present description. Figure 20 includes a first graph representing the amount of video recording reviewed by a reviewer, using a lite version of a video forensics tool, over a period of time, and a second graph representing the amount of video recording reviewed by a user, when using a full version of the video forensics tool according to the modalities of the present description. Figure 21 represents the amount of time a reviewer spends reviewing video recording, produced by one or more cameras, of an object, person, or animal, as the object, person, or animal moves in the field of view of one or more cameras according to the modalities of the present description. Figure 22 represents the amount of time a reviewer spends reviewing video recording, produced by one or more cameras, of an object, person, or animal, as the object, person, or animal moves in the field of view of one or more cameras according to the modalities of the present description. DETAILED DESCRIPTION OF THE INVENTION This document describes tools for the real-time, on-demand forensic analysis of multiple video streams. These tools enable forensic video analysis to, for example, identify when an object was placed within the field of view of one or more cameras. The moment the object entered the field of view of one or more cameras can be approximated by comparing multiple frames of video recording from the cameras to determine any differences in information between frames. For example, information about edge, color, and intensity in the frames is compared across frames over a period of time that begins at a reference point when the object is first identified in the field of view of one or more cameras, working backward to the moment when the object was no longer within the field of view of one or more cameras.Information about the edge, color, and intensity in the frames is also compared between frames within the same time period, starting from when the object was no longer in the field of view of one or more cameras and continuing until a time prior to the corresponding reference point. The tool can continue this process of moving backward and forward in time and comparing information about the edge, color, and intensity in smaller time increments to determine when the object came into view of one or more cameras. As another example of the video forensics tools described herein, a tool for reconstructing video from one or more cameras is also provided. This tool allows a user to track an object as it moves from one camera's field of view to another's using transition zones that connect the two cameras and anchor points to record the object's presence at a particular moment in the field of view of one or more cameras. The transition zones enable the tool to switch between displaying video footage from one camera and another. As another example of video forensics tools described herein, a video summarization tool is described. Video summarization begins with an activity detection stage. The goal of this stage is to process the source video, represented as a three-dimensional spacetime function / (x, t), and extract an activity function A(x, t) that indicates the degree of apparent activity at each pixel. Activity levels can be measured using, for example, a pixel-adaptive background subtraction model followed by local environment morphological operations such as dilation and erosion to remove noise and fill in gaps. The adaptive background model is a characterization of the background within a specified number of frames before and after each frame of the source video.As such, the adaptive background model is not a static image, but rather an adaptive model of the background that updates over time and can account for changes in ambient light. In addition to the adaptive background model, a background image can be generated, which can be formed by taking the median value of each pixel in the source video of interest. This background can be specified as a background image / B(x), which forms the background of the summary video, onto which the active foreground pixels are copied. In general, a video summary can be defined by the following parameters: the time interval of the source video sequence that encompasses Noframes, the length of the summary video frame, which can be determined based on a time compression ratio, and a motion sensitivity threshold value ω. In example modes, the video summary can be displayed to a user via a graphical user interface (GUI) that includes parameter controls, allowing a user to dynamically adjust the video summary parameters, including motion sensitivity and the time compression ratio. In example modes, such controls allow the user to move from sparse visual representations to a dense single-frame representation, which is a static map of all activity on the site, and anywhere in between.Adjusting motion sensitivity settings allows the user to find a balance between activity detection and interference suppression, enabling them to capture the most valuable content in the summary view. These parameter controls, along with the viewing interface, foster a remarkably effective interactive style of video review and dynamic content exploration. The following are example modalities described with reference to the figures. A person skilled in the art will recognize that the example modalities are not limited to the illustrative modalities and that the component of example systems, devices, and methods is not limited to the illustrative modalities described below. As used herein, the term "object" refers to a physical object such as a box, a bag, a piece of luggage, a vehicle, a human being, an animal, etc. Figure 1 is a flowchart illustrating an example method 100 for creating a summary video, according to the modalities of this description. Example method 100 is described with reference to block diagram 900, described in more detail later. In step 102, a source video is received. In example modalities, the source video can be received from a video input device 924, such as one or more surveillance cameras, or from a database or storage that archived or stored video data. In example modalities, the source video includes a number of source frames, No. A background image IB is generated in step 104. The background image can be generated, for example, using the background pixel detection module 930. As described above, the background image IB can be generated as the set of median pixel values ​​of the No frames of the source video of interest.To facilitate the description, examples are provided herein with RGB pixel values. Once the background image 1B has been generated, the method can proceed to step 105 and generate an adaptive background model. This adaptive background model can be generated using a background pixel detection module 930 and is a characterization of the background over a local time period of the source video. In example modes, the adaptive background model includes average pixel values ​​from a period or subset of source video frames surrounding a specific source video frame, such that the adaptive background model reflects the changing lighting conditions of a field. Once the adaptive background model has been generated, the activity level for the pixels in the source video is determined in step 106. This step can be performed, for example, using the active pixel detection module 928. In example modes, the segment of source video to be reviewed is inspected for motion at the pixel level.In one mode, each pixel in the source video is assigned a specific activity score of either zero (indicating a static background pixel) or 1 to 255 (indicating the degree of apparent motion), using an adaptive background subtraction model. The background subtraction model compares the value of each pixel in a frame of the source video with the corresponding pixel in that space of an adaptive background model. The activity level of each pixel can be saved, in some modes, as an activity map A(x, t) for each frame. The pixel activity level can be stored, for example, in the active pixel storage module 938. In many surveillance scenarios, non-zero values ​​in this activity map are sparsely distributed because most pixels represent components of the static scope. Therefore, only the non-zero activity map values ​​are stored.For each pixel where A(x, t > 0), the location value x, the activity level A(x, t), and the Red Green Blue (RGB) pixel value l(x, t) are stored. In each frame of the source video, the list of active pixels can be sorted in ascending order by activity score to facilitate the efficient retrieval of active pixels that exceed the user-controlled motion sensitivity threshold. As described above, the video summary is typically defined by the following parameters: the time interval of the source video sequence spanned by Nframes, the length of the summary video frame N2, and a motion sensitivity threshold value (represented by “ω” in the equations below). Once the pixel activity level has been calculated in stage 108, it is computationally determined whether the activity level of each pixel in the source video is greater than the motion sensitivity threshold value. In example modalities, this can be achieved by extracting the relevant subset of active pixels (without activity scores exceeding the threshold value ω) for each source frame.In modes where the pixel activity level A(x, t) is pre-organized in ascending order, this is equivalent to finding the first pixel that exceeds the motion sensitivity threshold value and extracting all subsequent pixels in the set. If the activity level of a pixel is greater than the motion sensitivity threshold, then in step 110 the selected pixel is added to a binary activity mask. In example modes, the binary activity mask function M(x, t) can be defined according to equation (1) below: ... . (1 if A(x, t) > ω) if A(x, t) < ω) This binary activity mask can be stored using an efficient data structure, for example, sparse sets of motion pixels, sorted into lookup tables by frame number and motion score. The data structure is designed to minimize the time required to access active pixels during the synthesis of the video summary frames. In example modes, the goal of the video summary is to map all pixels with a binary activity mask value of 1 to the video summary. Since this mapping is done at the pixel level and not at the activity tube level, no tracing is required at this stage. Instead, a cumulative activity count function, C(x, t), can be generated with the same time period as the video summary (0 < t < N^), by summing the periodic frames of the mask function according to equation (2) below: (2) íVh C(x, ü) = M(x, kN¡ + t) k=oOnce the selected pixel is added to the binary activity mask in stage 110, a time compression ratio is determined in stage 112, which determines the final length of the summary video. For example, a compression ratio of two cuts the source video in half, resulting in a summary video with half the frame count of the source video. Similarly, a compression ratio of eight results in a summary video one-eighth the length of the source video. If, however, it is determined in stage 108 that the pixel activity level does not exceed the motion sensitivity threshold, the time compression ratio is determined in stage 112 without adding the pixel to the binary mask in stage 110.Once the two parameters of the time compression ratio and the motion sensitivity threshold are known, the generation of the summary video frames can be carried out by remapping the active pixels onto the background image in step 114. In example modes, the summary frames and the summary video can be created using summary frame creation module 946 and summary video creation module 934. In example modes, a summary sequence is a sequence of individual summary video frames, each a composite of the background image and the remapped foreground components. The time it takes to compute a single summary frame depends on the length and amount of activity in the original surveillance video. For example, creating a summary video from a one-hour surveillance video source can range from approximately five to fifteen milliseconds in some modes, which is fast enough to support real-time frame generation. In example modes, the summary sequence can be generated according to equation (3) below: (3) ls^. O = ( / « Wlf 0=0) l / F(x, t) otherwise) where IF is calculated by collecting any activity at that pixel location in all frames that are evenly spaced in the timeline of the source video, according to equation (4) below: (4)IρΆ t) = This is a cyclic mapping procedure where a foreground pixel that appears at a location (x, t) in the source video appears at a location (x, mod(t, WJ)) in the summary video, mixed with any other foreground pixels mapped to the same location. This pixel-based remapping technique preserves the following important temporal continuity property of the source video: If pixel p1 is located at (xpt) and pixel p2 is located at (x2, t + At), with (0 < At - N^), then pixel p2 appears At frames after px in the resulting summary video (assuming the video is looped). Therefore, although the remapping is done at the pixel level, rather than at the object or scan level, foreground objects and local activity sequences remain intact in the summary video. Once the active pixels are remapped into a summary video, the summary video can be displayed to a user via a GUI in step 116. As described earlier, in example modes, the GUI allows a user to dynamically adjust the time compression ratio, the motion sensitivity threshold value, or both. While the summary video is being displayed via the GUI, the time compression ratio can be adjusted in real time by the user in the GUI in step 118. If the time compression ratio is adjusted in step 118, the method can return to step 114 and remap the active pixels based on the new time compression ratio and display the new summary video to the user via the GUI in step 116. If the time compression ratio is not adjusted, then in step 120, it is computationally determined whether the motion sensitivity threshold value is adjusted by the user in real time via the GUI. If the motion sensitivity threshold value is not adjusted, the method continues displaying the summary video to the user via the GUI in step 116. If, however, the motion sensitivity threshold value is adjusted, the method can return to step 108 and computationally determine whether the activity levels of each pixel are greater than the new threshold value. The method then continues with the following steps 110–116, displaying the new summary video to the user via the GUI based on the new motion sensitivity threshold value. In some modalities, the GUI can be generated using GUI 932 of a 900 example computing device. Figure 2 is a diagram illustrating the formation of a summary video frame 214 according to the example modalities in this description. In example modalities, IVIA / I 1300 takes a source video 200, which has a number of frames N, and divides it into smaller segments, each with a number of frames equal to the number of frames in the summary video. In this mode, the source video is tested simultaneously within each of the smaller segments, and summary video frames 202, 204, and 206, in which activity is detected, are displayed. Specifically, activity 208 is detected in frame 202, activity 210 is detected in frame 204, and activity 212 is detected in frame 206. As described above, the active pixels associated with activities 208, 210, and 212 are combined and remapped onto a background image to produce summary frame 216, which includes a composite of all the detected movements or activities 208, 210, and 212. This process is carried out for all frames of the video summary 214 to form the final video summary.Note that, as described above, activities that happen at the same time in source video sequence 200, such as activities 212, also happen at the same time in summary video 214. Figure 3 shows a screenshot of a sample GUI 300 for displaying a video summary, according to the example modes described herein. In these example modes, the GUI displays the video summary to a user and gives the user immediate control over key parameters of the summary video's formatting. For example, the GUI might display a first slider to control the duration of the summary clip and, therefore, the time compression ratio, and a second slider to control the motion sensitivity threshold that determines which pixels are considered part of the foreground and are mapped onto the video summary. Additionally, the GUI might allow the viewer to click on a specific pixel in the summary clip and jump to the corresponding frame in the original source video that contains that activity.In some modes, if the camera or video input device moves, a new background can be computed along with a new summary video for that particular viewpoint. The GUI can be generated using GUI 932 of a sample computing device 900, as described in more detail below. In sample modes, the GUI includes a window 301 for displaying a summary video to a user. The GUI also includes a playback speed control bar 302, a time compression control bar 304, a motion sensitivity control bar 306, and a summary video duration indicator 308. The playback speed control bar determines the speed at which new summary frames of the summary video sequence are displayed, which the user can speed up or slow down. In sample modes, a time compression slider or control bar 304 is associated with the time compression ratio (or summary video frame length) and allows the viewer to instantly switch from a compression ratio lower than IVIA / t / ZUZZ / UII Doó generates a longer video summary that provides clearer views of individual activity components at a much higher compression ratio, resulting in a more concise video summary that displays denser activity patterns. As the time compression control bar 304 is adjusted, the duration of the video summary, indicated by the video summary duration indicator 308, also changes. GUI 300 may also include other video control functions that allow the user to, for example, zoom in, zoom out, play, pause, rewind, and / or fast-forward a video summary. In example modes, a Motion Sensitivity Control Bar 306 allows a user to achieve a desired balance between activity detection and interference suppression by dynamically adjusting the Motion Sensitivity threshold value (the ω parameter). For example, a lower Motion Sensitivity threshold value results in greater motion or activity detection, but may also result in false activity detection caused by shadows or other small changes in pixel values ​​that do not represent actual activity in the video frame. Conversely, a higher Motion Sensitivity threshold value eliminates interference and many false activity detections, but may miss portions of actual activity within a frame. By using the Motion Sensitivity Control Bar 306, the user can adjust the sensitivity between sensitive and insensitive to achieve the desired balance. Figures 4A to 4C are screenshots of compressed summary videos that can be viewed using GUI 400 according to the example modes described herein. In these example modes, once the pixel activity levels of the source video have been pre-computed, a summary video can be generated at a desired compression ratio as needed for viewing. Figure 4A is a screenshot of GUI 400 showing a 402 window displaying a summary video that compressed the source video to one-eighth of its original length. In other words, an eight-to-one compression ratio was applied to the source video to produce the 402 summary box. Similarly, Figure 4B is a screenshot of GUI 400 showing a 404 window displaying a summary video with a compression ratio of sixteen to one. Figure 4C is a screenshot of GUI 400 showing a 406 window displaying a summary video with a compression ratio of 32:1. As can be seen, there is greater overlap of activity in summary videos with a higher compression ratio. In example modes, the time compression ratio can be adjusted in real time using a GUI control function, such as the time compression control bar 304 shown in Figure 3. IVIA / t / ZUZZ / UII Doo κ c κ N In summary videos with a higher compression ratio, such as the one shown in Figure 4C, pixel overlap can occur, resulting in a lack of clarity in the summary video. To avoid visual confusion where pixel overlap is present, a pixel with higher contrast than the background image can be given more weight in the summary video. Pixel overlap can also be mitigated by the fact that the operator has dynamic (i.e., real-time) control over the time compression ratio and, therefore, the desired activity density. In some modes, the active pixels of the source video (which has a number of frames N₀) can be remapped with respect to the compressed summary video (which has a number of frames N±) in blocks of consecutive frames. In an alternative mode, to allow for minimizing activity overlap, slightly smaller blocks with a length of N₀ < N₀ consecutive frames can be transferred, leaving room for translation in the mapping. Each block of frames can then start at any frame from 0 to -N₂ - 1 on the summary timeline. The starting frame of the A₀ block can be indicated by the delay variable Lk, which represents a degree of freedom in the optimization process. This is equivalent to the method described in the previous section, where A₀ equals A₁ and all Lk values ​​are restricted to zero.To describe this modified mapping approach, an indicator function equal to 1 is introduced if block k contributes any foreground pixel to the summary box t, according to the set of delay variables:. (5) 5(k, t) = \(Lk< t <N2+ Lkde otro modo Therefore, the counting function of equation (2) can be rewritten according to equation (6) below: (6) Similarly, the mapped images calculated in the preceding equation (4) can be rewritten according to the equation (7) below: «o-tx=Σ^1SCk,t)-BMIx,kN2-Lk +t) Where the Imss image is a shorthand notation for the product of the image sequence and its activity mask, calculated according to equation (8) below: (8) — I(x, t) · M(Xt) The relative temporal changes of the mapped activity blocks provide an effective mechanism for reducing overlaps. The values ​​of Lo...LK can be optimized to minimize the sum of all overlapping foreground pixels in the video summary, using an iterative scaling optimization procedure (for some standard variant, such as simulated tempering, which is less likely to converge to a locally deficient minimum of the cost function). As a result, this alternative approach allows for a reduction in activity overlap in the video summary at the cost of the additional computation required to run the optimization procedure. Figures 5A to 5B illustrate screenshots of a sample GUI that can be generated according to the example methods in this description to access a portion of a source video from a summary video. As described earlier, a summary video can be used as a visual index to the original source video, with each active pixel linked to its corresponding frame in the source video. For example, a user can navigate between a summary video and a source video by selecting an object of interest using, for example, a touchscreen, mouse, or other pointing device, and access the relevant portion of the source video that displays the selected object. Thus, a summary video can function as a navigation tool to more easily find and access activities within a source video. Figure 5A shows a window 502 displaying a video summary, along with various GUI controls, as described earlier with reference to Figure 3. Specifically, a video summary duration indicator 510 displays the length of the video summary shown in window 502, the time compression control bar 504 displays the time compression ratio used for this specific video summary, the motion sensitivity control bar 506 displays the motion sensitivity threshold used for this specific video summary, and cursor 508 displays the GUI cursor. The video summary displayed in window 502 shows a number of activities at a street intersection, and cursor 508 is positioned over a large semi-trailer truck. Selecting the truck allows the user to access the portion of the source video where the truck is located.In other words, by selecting the truck in the summary video, the user can access the part of the source video where the truck's active pixels were detected. Figure 5B shows the relevant source video in window 512 after the user selected the truck using cursor 508. As can be seen, the activity shown in source video window 512 is less dense than that shown in window 502, and the timestamp indicator of source video 514 shows the timestamp of the source frame displayed in window 512. Therefore, by using the summary video as an index to the source video, a user can select an activity and jump to the corresponding moment in the source video to examine the activity in more detail. Figure 6 is a diagram of an example of a rewind process provided by a forensic video rewind tool, according to the modality examples in this description. The example rewind process can be implemented using the 948 Rewind Module in Figure 9. The Rewind Module addresses the need to quickly assess, for example, in real time, the context surrounding an abandoned or otherwise suspicious item, a common task for video operators. In some modalities, given a stationary object in place, the 948 Rewind Module automatically jumps to the point in the video when the object first appeared. In other modalities, given a stationary object in place, the 948 Rewind Module can alert a human operator, who can then initiate a rewind task (process) to jump to the point in the video when the object first appeared.At this point, the operator can assess the context of the situation and respond accordingly. The backtracking task can mimic how a human operator would approach this task. Instead of simply rewinding the video, which can take a considerable amount of time, a user can directly go back in time using a more accurate estimate of when the object was likely placed. Since the object may have been present for only a few minutes or potentially for days, the algorithm(s) described herein adapt by searching nonlinearly. For example, in the first backtracking phase, the algorithm(s) go back in time with exponentially increasing time deltas until the object is deemed not to be present. In the second backtracking phase, the algorithm(s) employ a divide-and-conquer approach to refine the estimate of the time period around the object's first appearance. In some forms, the backtracking process 600 may include a first phase (phase 1 601) and a second phase (phase 2 603). In phase 1 601, the backtracking process 600 may receive user contributions corresponding to a user sketching a bounding box 607 around or partially around an image corresponding to object 605 in image segment 609. In this case, object 605 may be a suitcase. Image segment 617 may function as a "pointer" or reference frame from which edge, color, and intensity information is extracted. A segment may be a cropped subset of an original image. The segment may be the content of the original image within the bounding box 607 that corresponds to image segment 617. As detailed in Algorithm 1, the algorithm then backtracks from image segment 609 for an initial duration Jfy and performs the following operations. Algorithm 1 Backtrack until object is not present tmcxv indicator character box time tivsí frame time until evaluation while object is present if there is no movement in region of interest then score- ::-:5-5-. -.--25-5Q5-.H '-·5-ι·25>: ?:<· ce if score > threshold then set object is not present back Tantes· ' end if end if end while The first operation is motion estimation. After rewinding time, the motion within the bounding box is evaluated over a short period. In some modalities, this period may be less than 10 seconds. For example, motion within the bounding box can be determined based on the pixel activity levels determined in step 106 of Figure 1. If there is no motion within the bounding box during this period, a comparison is made between the object at or during the short time period and the object at or during a time period when the bounding box was first placed around the object. The video frame corresponding to the time period when the bounding box was first placed around the object can be called the indicator reference frame. The video frame corresponding to the short time period can be called the test frame.The time period in which each boundary marker was first placed around the object is after the short time period. The comparison, described below, can be performed to see if the object is still present. If the amount of movement exceeds a certain threshold, the object is considered to be temporarily blocked by the movement of people walking in front of it. The backtracking algorithm(s) (illustrated as Algorithm 1 above) do not perform a comparison and continue working backward in time. The second operation can be a comparison of image features. To determine if the object is present, image features are extracted, including border content, color, and pixel intensity. The evaluated frame is compared to a reference frame. The algorithm looks for a significant change in shape and color rather than the explicit presence of an object against a known background, which would be a more computationally expensive approach. The known background can correspond to the background of the summary video. Edge information can be extracted using a Canny edge detection algorithm that produces a binary edge mask. The percentage of edge overlap is calculated by summing the pixel overlap between the edge masks of the indicator box and the evaluated box and dividing by the number of edge pixels in the indicator box. This provides a metric for determining how closely the contents of the test box match the shape of the object in the indicator box. Similarly, a pixel-to-pixel difference is calculated by subtracting the indicator frame from the test frame for each of the three color channels. Initially, the images are normalized to remove any uniform changes in lighting. The absolute element-to-element difference is then calculated, summed, and normalized. This metric indicates a significant change in color or intensity and produces a high value when the content of a test frame does not match the indicator frame. Algorithm 1 calculates a weighted difference, or score, between factors indicating no change and factors indicating significant change. Significant change can be a pixel-to-pixel difference between a scoreboard and the test box that exceeds a user-defined threshold, thus indicating that an object is no longer in the position it was in the scoreboard. One factor that can indicate no change is the percentage of pixel overlap between the border masks of the scoreboard and the test box. A factor that can indicate significant change is a change in color intensity determined by the absolute element-to-element difference described above. If this score exceeds a particular threshold, the object is considered not to be present. Otherwise, the object is still present, and the algorithm continues backtracking with an exponential scaling factor, N.The scaling factor (increasingly larger movements) provides a balance between calculation time and robustness during both short and long periods of downtime. The indicator image is updated periodically to minimize discrepancies caused by slowly changing lighting over extended periods. In phase 2, the time period can be adjusted. In the second processing phase (see Algorithm 2), Algorithm 2 moves forward in time, from "not present" to the first moment it "is still present." This divide-and-conquer approach is repeated until the estimated time for the object's appearance is reduced to a reasonably short interval (e.g., 10 seconds). In some modalities, phase 2 may not occur. For example, if Tafter-Tbefore is not greater than the permissible time period after the backtracking algorithm moves backward in time, then phase 2 would never take place. Algorithm 2 Refmar the tender period of appearance of the object time before the appearance of the object Tiespues re^pc after the appearance of ctepc. nca zade a while í Tana. Tiespim] > permissible time period to evaluate in the middle of the time period íewi · I Tañes - Tie$pues\:2 if the object is not present then Tañes · te va / de lo contrario el objeto sigue presente T / espues · teva! finish if finish while The user interface for viewing the results of the backtracking algorithms is shown in Figure 7. The algorithms result in several representative samples over the time period of the object's appearance. The backtracking algorithm (Algorithm 1) often needs only a few seconds to find the first occurrence of an object, but it can take tens of seconds if the object was present for many hours or days. The rate at which frames can be extracted from the video management system (VMS) significantly affects the total processing time. The parameters for Jf and the exponential scaling factor / V can be adjusted depending on the VMS characteristics and the operator's needs. The 948 reversal module reduces the workload of human operators who receive numerous reports of abandoned items, whether from observant passengers or other video monitoring systems. In addition to stationary objects, the reversal tool can also be used as a general change detector for everyday investigative tasks; for example, determining when an object disappeared (e.g., a stolen bicycle) or when an object's appearance changed (e.g., graffiti on a wall). Figure 7 shows a screenshot of a 700 graphical user interface for displaying and reviewing a back-up function, as per the modality examples in this description. The 700 graphical user interface can begin playing a video a few seconds before the object comes to a stop. The graphical user interface 700 may include one or more menus (path reconstruction menu 701, video summary menu 703, and rewind menu 705). The path reconstruction menu 701 may display icons such as anchor points or markers that, when activated, allow the user to view a path traveled by an object in and through the IVIA / I 1300 multiple fields of view corresponding to different cameras. For example, the path reconstruction menu 701 may display screens similar to those illustrated by the screenshots in Figures 8, 13, 14, and 16, and may be implemented as a result of executing one or more instructions in the path reconstruction module 950. The video summary menu 703 can be a menu that displays a sample GUI 300 to show a video summary, as illustrated in Figure 3. In some modes, the video summary menu 703 can display screens similar to those illustrated in Figures 4A to 4C, which display compressed versions of a video. The video summary menu 703 can be implemented as a result of executing one or more instructions in the video summary creation module 934. The rewind menu 705 may be a menu that displays one or more controls (video speed decrease icon 713, rewind icon 715, play icon 717, fast forward icon 723, video speed increase icon 721, playback speed 719, and scroll bar 725). The rewind menu 705 may also include an icon for executing the rewind operation, which a user can activate by clicking the icon. This may cause one or more processors to execute instructions in a rewind module 948, thereby causing the rewind menu 705 to generate multiple frames 729, at least one of which includes an object of interest. For example, each of the multiple frames 729 may include object 709, which are extracted over a period of time. The rightmost box includes a clear image of object 709 similar to the one shown in indicator box 711, without the delimiter box 707.Object 709 in the leftmost frame of the multiple 729 frames is obscured by a person and therefore may correspond to a frame where the backspace tool determines that object 709 is no longer present due to a pixel-by-pixel comparison as described above. The frames from rightmost to leftmost are frames selected by the backspace tool that correspond to test frames given before indicator frame 711 and can be compared to indicator frame 711 as the backspace tool selects the multiple 729 frames. The lower portion of the interface displays several image clips (729) from the time period surrounding the object's appearance. In Figure 7, the time period is 11:56:04–12:56:04, which is one hour. Although the time period illustrated in Figure 7 is one hour, it can be shorter or longer than one hour. In some modes, this time period can be configured by the user. In other modes, one or more of the modules in Figure 9 can determine a suitable time period based on the user's observation patterns. Figure 8 illustrates example transition zones according to the example modalities described herein. Large-scale video surveillance systems provide methods for viewing multiple camera feeds simultaneously. However, tools are lacking that allow an operator to track activity from one camera view to another, especially when the cameras' fields of view do not overlap. A path reconstruction module 950, as described herein, addresses this need by allowing human operators to add comments about the activity in the video. As a result, an operator or the path reconstruction module 950 can seamlessly traverse the camera views (801, 803, 805, 807) using transition zones (809, 813, 815, 821, and 811) and reconstruct video evidence by automatically combining fragments, derived from anchor points, from multiple cameras.The anchor points are described with respect to Figure 14. The 950 Path Reconstruction module incorporates several user interface capabilities that assist human operators during multi-camera video investigations. Transition zones 809, 813, 815, 821, and 811 are selectable regions overlaid on the video that direct a user to views of nearby cameras, eliminating the need to pause the video to search for a specific camera from a list or menu. While the 950 Path Reconstruction module is agnostic to the object type (person, vehicle, bag), for ease of explanation, the following description focuses on the task of tracking a person of interest across multiple cameras. Transition zones define connections between two camera views. In some configurations, each zone is defined by its shape (polygon coordinates), the camera's unique identification number, and the zone (in another camera) to which it links. In other configurations, each zone is defined by a subset of either its shape (polygon coordinates), the camera's unique identification number, and the zone (in another camera) to which it links. Transition zones are placed at each main entrance or exit area, or both, as depicted in Figure 8. For indoor installations, the shapes and locations often correspond to main pedestrian traffic routes.For example, transition zone 811 can be placed near an exit where pedestrians leave a first area, thus leaving the field of view of a first camera corresponding to camera view 807, and enter a second area, appearing in the field of view of a second camera corresponding to camera view 805 and transition zone 821. These transition zones can be linked. Similarly, transition zones 819 and 813 can also be linked. In some modes, the transition zones can have different colors. For example, transition zone 819 can be a different color than transition zone 813. The color of transition zone 819 can indicate to an operator that clicking on it allows the operator to view the camera view corresponding to the area from which an object originates.For example, an operator might be looking at camera view 805 and want to determine how an object of interest entered camera view 805. The operator can click on transition zone 819, which closes the feed corresponding to camera view 805 and opens a feed corresponding to camera view 801, which might be the object's previous camera view. The operator can then determine how the object appeared in camera view 805 after leaving camera view 801. In some modes, a green transition zone icon might indicate the next camera in whose field of view the object will be when it leaves the field of view of a previous camera and enters the field of view of the next camera. A yellow transition zone icon might indicate the opposite.This means that a yellow transition zone icon can indicate the previous camera in whose field of view the object was located before entering the field of view of the camera the user is currently viewing. Because the transition zone icons are color-coordinated, this makes it easy for the user to quickly switch between cameras in whose field of view the object may have entered as the object moves through a specific area within the cameras' field of view. When a transition zone is selected or clicked, the path reconstruction tool closes the current camera feed on the screen, opens the linked camera feed, and begins displaying the new camera view in the video player. The transition zone corresponding to the previous camera is displayed in a different color to give the user contextual information and a way to go back, if needed. The nearby camera can be previewed by hovering over the transition zone; a static thumbnail of the camera view is displayed in response. The thumbnail may be a pop-up message that appears briefly while the mouse hovers over the green arrow. It then disappears when the user moves the mouse away from the green arrow. In some modalities, defining transition zones is a one-time offline configuration step, and manually creating each transition zone can be time-consuming. To reduce this burden, in some modalities, transition zones can be estimated algorithmically, using a pedestrian detection algorithm, and then confirmed or edited by the end user. In some modalities, an accumulated channel feature (ACF) algorithm can be used for multi-resolution object detection. The algorithm's input can be an image frame, and the output can be a bounding box around the object and a confidence score for each detected object. The person's appearance and location information can be monitored over time. In practice, a color association algorithm is used to eliminate detections of other people who may be present.Once high-confidence detections of a single person are compiled, the entry and exit zones for each camera view are estimated based on the person's time and location. In some modalities, a high-confidence detection can be defined based on a statistical parameter such as a confidence interval. In some modalities, the person leaves one camera and reappears in the next camera after a short time; the transition zones linking the two cameras are placed at the location of the last / first sighting, respectively. Consequently, the exit zone might be the last sighting of the person within the field of view of the first camera, and the entry zone might be the first sighting of the person within the field of view of the second camera.Transition zones can be placed where the person disappears from the first camera's field of view and reappears in the second camera's field of view. Furthermore, transition zones can be linked together. In some modalities, additional logic is used in cases where camera density increases (causing the person to appear in multiple cameras simultaneously) or where there are significant gaps (the person is not in view for an extended period). In some modes, anchor points, or markers, are set by the operator as a person enters and exits the camera's field of view. An anchor point consists of a bounding box (x, y coordinates in the upper left corner, along with the width and height), the time in milliseconds since the date (e.g., January 1970 UTC), and the camera's unique identification number. In modes where a human operator marks anchor points as the person of interest enters and exits each camera view, the algorithm switches videos at a point midway between the last anchor point in the current camera view and the first anchor point in the next camera view. In some modes, where closely spaced camera anchor points may overlap over time, the Path Reconstruction module 950 alternates between camera views to display all sightings marked by the operator. In some modes, additional logic is executed for cases where two observations are separated by a large time gap or where the resolution is not uniform across camera views. In some modes, the final reconstructed video can be exported to an MPEG-4 video file or another standard compression format. In some modes, the name, date, and time of the source camera can be overlaid below the video. Calculating the video streams and time periods to be included in the composite video has negligible computation time. Any latency is often due to communication costs with the video management system (VMS) and acquiring new camera streams. ML / E / ZuZz / uΊ 1300 Exporting the video to a file can take a few minutes or more depending on the length of the video, the number of camera views, the cost of VMS communication, and the resolution of each camera. The 950 Path Reconstruction module is useful for tracking a person of interest's activity across multiple camera views and producing a composite video that concisely illustrates the activity. This composite video can then be used to collaborate with other investigators. Additionally, annotated metadata (camera ID numbers, timestamps) can be stored for later reference. In some applications, the 950 Path Reconstruction module is used for other tasks, such as vehicle tracking. The example tracing function described above can be implemented using the path reconstruction module 950 in Figure 9. Example computing devices Figure 9 is a block diagram of a sample 900 computer device that can be used to perform any of the methods provided in the sample modes. The 900 computer device includes one or more non-transient, computer-readable media for storing one or more computer-executable instructions or software for implementing sample modes. The non-transient, computer-readable media may include, but are not limited to, one or more types of hardware memory, non-transient tangible media (for example, one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), and the like. The 906 memory may include computer system memory or random-access memory, such as DRAM, SRAM, EDO RAM, and the like. The 906 memory may also include other types of memory or combinations thereof.For example, the memory 906 included in the computing device 900 can store computer-readable and computer-executable instructions or software to implement example modes detailed herein. The computing device 900 also includes a processor 902 and associated core 904, and may include one or more configurable processors 902' and one or more associated cores 904' (for example, in the case of computing systems having multiple processors / cores), to execute computer-readable and computer-executable instructions or software stored in memory 906 and other programs to control the system hardware. The processor 902 and one or more processors 902' may each be a single-core processor or a multi-core processor (904 and 904'). Virtualization can be used on the 900 computing device so that the infrastructure and resources on the device can be dynamically shared. A 914 virtual machine can be provided to handle a process running on multiple processors so that the process appears to be using only one computing resource instead of multiple computing resources. Multiple virtual machines with one processor can also be used. A user can interact with the computing device 900 through a visual display device 918, such as a touchscreen or computer monitor, which can display one or more user interfaces 920 that can be provided according to example modalities, for example, the example interfaces illustrated in Figures 3, 4A to 4C, 5A to 5B, 7, 8, 13, and 16. The visual display device 918 can also display other aspects, elements, and / or information or data associated with example modalities, for example, database views, maps, tables, graphs, diagrams, and the like. The computing device 900 can include other I / O devices for receiving input from a user, for example, a keyboard or any suitable multi-point touch interface 908 and / or a pointing device 910 (for example, a pen, stylus, mouse, or touchpad).The keyboard and / or pointing device 910 can be electrically coupled to the visual display device 918. The computing device 900 can include other suitable peripheral I / O devices. The 900 computing device may include a 912 network interface configured to interact through one or more 922 network devices with one or more networks, e.g., local area network (LAN), wide area network (WAN), or the Internet through a variety of connections including, but not limited to, standard telephone lines, LAN or WAN links (e.g., 802.11, TI, T3, 56kb, X.25), broadband connections (e.g., ISDN, frame relay, ATM), wireless connections, controller area network (CAN), or some combination of any or all of the above.The 912 network interface may include an integrated network adapter, a network interface card, a PCMCIA network card, a card bus network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for interconnecting the 900 computing device with any type of network capable of communicating and performing the operations described herein. Furthermore, the 900 computing device may be any computing system, such as a workstation, desktop computer, server, laptop computer, handheld computer, tablet (e.g., the iPad® tablet), mobile computing or communication device (e.g., the iPhone® communication device), or any other form of computing or telecommunications device capable of communication and having sufficient processing power and memory capacity to perform the operations described herein. The 900 computing device can run any 916 operating system, such as any version of the Microsoft® Windows® operating system, different versions of Unix and Linux operating systems, any version of MacOS® for Macintosh computers, any embedded operating system, any real-time operating system, any open-source operating system, any proprietary operating system, any operating system for ML / E / ZuZZ / u 1300 mobile computing devices or any other operating system capable of running on the computing device and performing the operations described in this document. In example modes, the operating system 916 can run in native mode or emulated mode. In one example mode, the operating system 916 can run on one or more cloud machine instances. The computing device 900 can include one or more video input devices 924, such as one or more video cameras that a user can use to provide one or more video input streams. The computing device 900 may also include one or more storage devices 926, such as a hard disk, CD-ROM, or other computer-readable media, for storing computer-readable data and instructions and / or software that implements example modes as detailed herein or parts thereof. The storage 926 includes a video editing system 925. The video editing system 925 includes an active pixel detection module 928, a background pixel detection module 930, user interfaces 920, a video summary creation module 934, and / or the summary box creation module 946, in example modes. User interfaces 920 may include a GUI 932 that can be rendered by the visual display device 918. The GUI 932 may be a GUI that corresponds to one or more of the example interfaces illustrated in Figures 3, 4A to 4C, 5A to 5B, 7, 8, 13 and 16.In other words, GUI 932 can be a GUI that triples and / or displays GUI 300, 400, 500, 700, 800, 1300 and 1600. In example modes, the active pixel detection module 928 can detect active pixels within each source frame by comparing each pixel value with the corresponding pixel in the background image. As described earlier, a motion sensitivity threshold value can be used to determine which pixels are active versus which are simply part of the background image. The background image pixels, as well as the adaptive background model that characterizes the background over a local time lag of the source video, can be generated using the background pixel detection module 930. In example modes, the summary frame creation module 946 creates the frames of a summary video by overlaying the active pixels onto the background image; and the summary video creation module 934 creates the summary video by adding the summary frames in the appropriate order. As described earlier, the number of frames included in a summary video can be determined based on a compression ratio that can be dynamically adjusted by a user via GUI 932 in some modes. These modules can be logically or physically separated or combined into one or more modules. IVIA / t / ZUZZ / UII Doo A sample database 945 can store one or more additional databases, such as the detection storage database 936 or the archived video database 944, to store any relevant information needed to implement sample modes. The archived video database 944 can store, for example, the original video source and / or video data related to previously created video summaries. The video management system can store video frames in the archived video database 944 for extraction by the summary video creation module 934, the summary frame creation module 946, or the path reconstruction module 950.The Path Reconstruction Module 950 can execute computer-executable instructions that cause the module to extract frames from the archived video database 944 and combine, or concatenate, the extracted frames to create a composite video. In example modes, the detection storage database 936 may include the active pixel storage 938 to store information about the active pixels within the source video, the background pixel storage 940 to store information about the pixels that make up the background image, and / or a summary video storage 942 to store the summary video once it is created. The rewind module 948 may use data stored in the active pixel storage 938 and the background pixel storage 940 to determine whether to continue rewinding in time as described above with reference to Algorithm 1 and Algorithm 2.The back-up module 948 may include computer-executable instructions that cause the module to perform operations on algorithm 1 and algorithm 2 using data stored in active pixel storage 938 and background pixel storage 940. The database 945 may be provided on the computing device 900 or provided separately or remotely from the computing device 900. Example network environments Figure 10 is a diagram of an example network environment suitable for a distributed implementation of example modes. The 1000 network environment may include one or more 1002 and 1004 servers, which may include the active pixel detection module 1028, the background pixel detection module 1030, the summary box creation module 1046, the video summary creation module 1034, the backtracking module 1048, the path reconstruction module 1050, the detection storage 1036, or other items described with reference to Figure 10. In example modes, the 1004 server may include the active pixel detection module 1028, the background pixel detection module 1030, the summary box creation module 1046, the video summary creation module 1034, the backtracking module 1048, and the path reconstruction module 1050; while the 1002 server ML / E / ZuZZ / uΊ 1300 includes detection storage 1036. As will be seen, various distributed or centralized configurations can be implemented, and in some modes, a single server can be used. The network environment may also include a computing device 1000 and one or more video input devices 1024 and / or other elements described with reference to Figure 10. In example modes, servers 1002 and 1004, computing device 1000, and video input device(s) 1024 can communicate with each other via a communication network 1012. This communication network 1012 can include, but is not limited to, the internet, an intranet, a LAN (local area network), a WAN (wide area network), a MAN (metropolitan area network), a wireless network, an optical network, and similar networks. In example modes, in response to user input commands on computing device 1000, a user can dynamically configure video summary parameters, such as the motion sensitivity threshold and / or the video summary compression ratio. Once a video summary is created in the video summary creation module 1034, the video summary can be transmitted to computing device 1000 and displayed to a user via a GUI.The backtracking module 1048 includes executable code and other code for implementing Algorithm 1 and Algorithm 2 described above. The path reconstruction module 1050 includes executable code and other code for implementing one or more of the tracing functions as described with reference to Figure 8 above. Figure 11 is a flowchart illustrating an example method for reviewing video recordings when an object was placed within a camera's field of view, according to the modalities described herein. In block 1102, one or more instructions in backspace module 948 may be executed by a processor to detect a bounding box around part or all of the object in a first frame at a first moment in the video. For example, the processor may detect a bounding box, similar to bounding box 607 in Figure 6, around object 605 in segment 609, in response to executing the instructions in backspace module 948. As noted earlier, segment 609 may be referred to as the indicator reference frame. In block 1104, the processor can identify a second frame in the video, corresponding to a second moment. With respect to Figure 6, the processor can search through one or more segments that occurred in the past (also called rewinding) and identify a frame, or segment 615, at a time before segment 609. Segment 615 can also be referred to as the test frame. Segment 615 can be called the test frame because it is the frame against which rewind module 948 compares, or tests, the indicator reference frame. In block 1106, the processor can determine whether there is movement within the bounding box of the second frame. For example, the processor can execute one or more instructions in rewind module 948 that cause the processor to perform an operation to evaluate movement within a bounding box 615. In block 1108, the processor can compare at least one first edge piece, one first color piece, and one first intensity piece in the first frame to a second edge piece, a second color piece, and a second intensity piece in the second frame. For example, the processor can also execute instructions in backspace module 948, causing the processor to compare intensity and color piece 617 corresponding to segment 609 to intensity and color piece (not shown) corresponding to segment 615. The processor can also perform the operation of comparing edge piece 619 corresponding to segment 609 to edge piece 615. In block 1110, the processor can generate a score based at least in part on comparing at least the first edge information with the second edge information, at least the first color information with the second color information, or at least the first intensity information with the second intensity information. In block 1112, the processor can determine that the score does not exceed a first threshold. The threshold can be user-defined in some modes. In other modes, the threshold can be determined based at least in part on backspace module 948. In block 1114, the processor can determine an estimated time frame for the object's first appearance. For example, the processor can execute one or more instructions associated with backspace module 948, which causes the processor to estimate a time frame during which the object first appeared. More specifically, the processor can perform operations associated with Algorithm 2, described earlier, which refines, or limits, the time frame around which the object first appeared. The processor can perform these operations until the time difference between when the object is first seen and when it is first determined that the object is not present is less than the allowable time frame. The allowable time frame can be a user-defined amount of time (e.g., 10 seconds).In other modes, the permissible time period may be less than 10 seconds and may ultimately be determined based on the configuration in which the 948 recoil module is used and / or the movement activity in a location. In block 1116, the processor can cause multiple frames to be displayed within the time period on a screen connected to at least one processor. The at least one processor can execute instructions that cause the visual display device 918 to display multiple frames that fall within the estimated time period. ML / I 1303 Figure 12 is a flowchart illustrating an example method for tracking an object across multiple cameras, according to the modalities described herein. In block 1202, a processor can detect the location and appearance of an object in the field of view of a first camera. The processor can execute one or more instructions associated with path reconstruction module 950 to detect the object's location and appearance. For example, the processor can detect object 709 in Figure 7 by executing one or more instructions in path reconstruction module 950. In block 1204, the processor can generate one or more anchor point annotations around the object in the first camera's field of view over a period of time. The processor can execute one or more instructions associated with path reconstruction module 950 that cause the processor to generate the one or more anchor points. For example, one or more instructions can cause the processor to activate GUI 932, thereby displaying one or more anchor points (e.g., anchor points 729). In block 1206, the processor can use at least one transition zone to cause the display of a second camera's field of view. The processor can execute one or more instructions associated with path reconstruction module 950, thereby causing the processor to send a signal to the second camera, which can be one of input devices 924, to display the second camera's field of view. In block 1208, the processor can detect the object's location and appearance within the second camera's field of view. The processor detects the object's location and appearance based, at least in part, on user input to generate a bounding box around the object. When a bounding box is created around the object, a timestamp and a camera ID can be stored as an anchor point. The processor can then execute one or more instructions associated with the path reconstruction module 950, which causes the processor to determine the object's location relative to a spatial reference frame and its appearance. In block 1210, the processor can generate one or more anchor point annotations around the object in the second camera's field of view over a period of time. The processor can execute one or more instructions associated with path reconstruction module 950 that cause the processor to generate the one or more anchor points. For example, one or more instructions can cause the processor to activate GUI 932, thereby displaying one or more additional anchor points that are different from the anchor points in 729. In block 1212, the processor may cause multiple frames to be displayed in which the object is present in the field of view of the first camera, followed by multiple frames in which the object is present in the field of view of the second camera. IVIA / t / ZUZZ / UII Doo chronologically. The processor can execute one or more instructions associated with the path reconstruction module 950, which causes the processor to generate multiple frames. For example, with reference to FIG. 14, the leftmost anchor point of the 1408 anchor points corresponds to a first frame, in the chronological order of the multiple frames. The first frame can be a frame generated by the first camera and can be a frame associated with the field of view of the first camera. The rightmost anchor point of the 1408 anchor points corresponds to a last frame, in the chronological order of the multiple frames. The second frame can be a frame generated by the second camera and can be a frame associated with the field of view of the second camera. Figure 13 illustrates a screenshot of an overlay graphical user interface 1300 for navigating a camera network, according to embodiments of the present invention. Figure 13 depicts a field of view from a first camera with an overlay graphical user interface that includes transition zone icons 1304, 1306, and 1308. Transition zone icon 1308 (double arrow) can switch to a view that is identical, but a mirror image, of what the user is seeing. In Figure 13, this would correspond to the user viewing the same area from another camera located on the other side of the railway tracks. The overlay graphical user interface can be used to navigate from the field of view of the first camera to the field of view of a second camera (not shown). The overlay graphical user interface may include one or more transition zone icons 1304 and 1306.In some modes, transition zone icons 1304 and 1306 may be arrow-shaped. In other modes, transition zone icons 1304 and 1306 may be squares or rectangles. A user can interact with the transition zone icons using a cursor 1302 to select one. When a user selects a transition zone icon, the overlay graphical user interface displays the field of view of the other camera associated with that transition zone icon. Figure 14 represents a screenshot of a graphical user interface 1400 for selecting an object of interest when it is in the field of view of multiple cameras and generating anchor points, according to embodiments of the present invention. Figure 14 illustrates an individual of interest 1406 in the field of view of one camera. The graphical user interface can generate a bounding box 1404 around the individual of interest 1406 in response to user input via a cursor 1402. The graphical user interface can also display multiple anchor points 1408. The anchor points 1408 can be specified by the user to identify the individual of interest or object in the field of view of other video cameras. In some embodiments, this can allow the user to quickly review, or play back, the video recording forwards and backwards in time if they lose sight of the individual of interest. ML / I 1303 Composite video, rewind composite video recording, fast forward composite video recording, increase or decrease the playback speed of the composite video recording using a plus and minus sign respectively. The playback speed can also be changed using a slider. Video segments 1610, 1612, 1614, and 1616 correspond to portions of the composite video recording in which an object of interest is within the field of view of one or more cameras. Images 1602, 1604, 1606, and 1608 are composite video frames from the recording in which the object of interest is within the field of view of a camera capturing video during segments 1610, 1612, 1614, and 1616 respectively. Figure 17 represents a timeline 1700 corresponding to segments of a composite video, where each segment is a portion of a video stream from a different camera, according to the modalities described herein. The path reconstruction module 950 may include instructions that, when executed by a processor, can cause the processor to combine one or more video stream segments, each produced by a corresponding camera. Part 1702 may be a segment generated by a first camera (Cam1), part 1704 may be a segment produced by a second camera (Cam2), and part 1706 may be a segment produced by a third camera (Cam3). The times t1 and t2 correspond to the times of the earliest and latest anchor points generated by Cam1. The times t3 and t4 correspond to the times of the earliest and latest anchor points generated by Cam2.The ts and te times correspond to the times of the earliest and latest anchor points generated by Cam3. The times corresponding to limits A and B are the start and end of the segment for Cam1. The Path Reconstruction Module 950 may include instructions that cause the processor to determine which portion of the Cam1 video should be used in the composite video. The times corresponding to limits B and C are the start and end of the segment for Cam2. The Path Reconstruction Module 950 may include instructions that cause the processor to determine which portion of the Cam2 video should be used in the composite video. The times corresponding to limits C and D are the start and end of the segment for Cam3. The Path Reconstruction Module 950 may include instructions that cause the processor to determine which portion of the Cam3 video should be used in the composite video. In some modes, limit B is calculated by finding the midpoint between t2 and te.Likewise, the midpoint between t4 and te is equal to the limit time C. Limits A and D are generally equal to ti and te, respectively. The limits are sometimes extended or restricted so that they do not last more than 30 seconds before / after the anchor point, if there is a large gap between the times. Figure 18 represents a screenshot of a reconstruction tool 1300 path detection tool that detects an object moving in the field of view of two cameras and a logic diagram indicating the location of the transition zone icons that join the two cameras, according to the modalities of this description. The example path reconstruction tool 1800 may include the field of view 1802 of a first camera and the field of view 1804 of a second camera. The path reconstruction module 950 may include instructions that cause a processor to detect a person 1826 at some point in time when the object 1826 enters the field of view 1802. In some modalities, the person 1826 may be an object. The processor may detect the person 1826 and include a bounding box 1806 around the person 1826.The processor can maintain the same bounding box (represented by 1808, 1810, and 1812) around person 1826 as they move through the area associated with field of view 1802. 1808, 1810, and 1812 can represent the same bounding box that moves along with person 1826 as person 1826 traverses the area associated with field of view 1802. When person 1826 leaves field of view 1802 and enters field of view 1804, the processor can detect when person 1826 enters field of view 1804 and can add a bounding box 1814 that corresponds to when the processor first detects person 1826 in field of view 1804. The same bounding box (represented by 1816 and 1818) can be found around person 1826 as they move through the area associated with field of view 1804. A user instruction, or instructions in the 950 path reconstruction module, can cause the processor to place transition zone icons 1824 and 1822 at the rightmost edge of field of view 1802 and the leftmost edge of field of view 1804. Transition zone icons 1824 and 1822 may be arrow-shaped and, in some modes, may take other forms. Clicking transition zone icon 1824 will terminate a video stream associated with camera and field of view 1802 and initiate a video stream associated with camera and field of view 1804. In some modes, transition zone icon 1824 may be green. The user can also click on the transition zone icon 1822 and a video stream associated with camera and field of view 1804 will end and a video stream associated with camera and field of view 1802 will begin.In some modes, the transition zone icon 1824 may be yellow. A link 1820 can be established between a camera corresponding to transition zone icon 1824 and a camera corresponding to transition zone icon 1822. The link 1820 can be determined between the two cameras based on the presence of a person 1826 or an object entering or leaving the field of view 1802 or 1804. Figure 19 illustrates a graphical representation of experimental results of the ML / I 1303 The amount of video recording reviewed over a period of time for a first participant. Figure 19 is a graph comparing the amount of video recording (amount of video reviewed (min) 1906) that a user reviews in a given period of time (time spent reviewing video (min) 1908) when tracking an object. Curve 1904 is a line of best fit of the amount of video recording reviewed compared to the amount of time it takes a user to review the amount of video recording when the user uses a video player to locate an object of interest. The amount of video recording reviewed compared to the amount of time it takes a user to review the amount of video recording is expressed as a ratio of the former to the latter.The speed (ratio of the amount of video recording reviewed to the amount of time it takes a user to review that amount of video recording) for curve 1904 is 0.25. This means that for every minute of video recording available for viewing, it took the user approximately four minutes to review the video recording. For example, it takes the user twenty minutes to review four minutes of video recording. When the same user utilizes the transition zone and anchor point functions provided by the 950 path reconstruction module, the amount of video footage reviewed compared to the time it takes a user to review the video footage for curve 1906 is 0.47. This means that for every minute of video footage available for viewing, it took the user slightly more than two minutes to review the same video footage. For example, it takes the user 20 minutes to review nine minutes of the same video footage they viewed without the anchor points and transition zones. Figure 20 illustrates a graphical representation of experimental results for the amount of video recording reviewed over a period of time for another participant. Figure 20 is a graph comparing the amount of video recording (amount of video reviewed (min) 2006) that another user reviews in a given period of time (time spent reviewing the video (min) 2008) to track an object. The curve 2004 is a line of best fit of the amount of video recording reviewed compared to the amount of time it takes a user to review the amount of video recording, when the user uses only a video player to track the detection of an object of interest. The amount of video recording reviewed compared to the amount of time it takes a user to review the amount of video recording is expressed as a ratio of the former to the latter.The speed (ratio of the amount of video recording reviewed to the amount of time it takes a user to review that amount of video recording) for the 2004 curve is 0.18. This means that for every minute of video recording available for viewing, it took the user approximately five minutes to review the video recording. For example, it takes the user twenty minutes to review four minutes of video recording. When the same user utilizes transition zone and anchor point functions provided by the 950 path reconstruction module, however, the user speed increases to 0.35. The speed (ratio of the amount of video recording reviewed to the amount of time it takes a user to review that amount of video recording) for curve 2002 is 0.35. This means that for every minute of video recording available for viewing, it took the user approximately two minutes to review the same video recording. For example, it takes the user ten minutes to review five minutes of the same video recording they viewed without the anchor points and transition zones. Figure 21 is a graphical representation of the amount of time an object of interest is within the field of view of multiple cameras. This represents the time a first user spent reviewing the video recording of the object of interest across multiple cameras and performing different parts of the video recording on the multiple cameras using a video player without the aid of transition zones and anchor points provided by the 950 Path Reconstruction module. Figure 21 illustrates a user tracking an object of interest as the object moves through the field of view of multiple cameras. Block 2102 represents when the object of interest is within the field of view of a camera. Figure 21 shows the data for a total of twenty cameras and when the object of interest is visible to each of the multiple cameras.For example, between times t1 and t2, the object of interest is visible to cameras nine and ten. At time t1, the user marked a portion of the video recording from camera 10 in which the user noticed the object of interest using marker 2106. Also starting at time t1, the user began viewing the camera 10 recording of the object of interest. Block 2140 represents a portion of the camera recording reviewed by a user. The length of block 2140 indicates the amount of time a user spends reviewing the recording from a given camera. For example, the amount of time the user spends reviewing the camera 10 recording, between t1 and t2, corresponds to block 2110. The length of block 2110 is equal to t10, which in this example could be approximately one minute of video recording from camera 10. Figure 22 illustrates a screenshot of the amount of time an object of interest is within the field of view of multiple cameras, the time one user spent reviewing the video recording of the object of interest across multiple cameras, and the time another user spent reviewing different parts of the video recording across multiple cameras using a video player with the help of transition zones and anchor points provided by the Path Reconstruction Module 950. Compared to Figure 21, Figure 22 includes more video recording markers across multiple cameras when an object of interest is within their field of view. Blocks 2202, 2204, and 2206 have the same meaning as blocks 2102, 2104, and 2106, respectively.It should be noted that markers 2106 and 2206 correspond to anchor points that the user added to the video recording generated by a given camera. There are more markers in Figure 22 than in Figure 21 because the user, whose performance is captured in Figure 22, was using transition zones and anchor points to track the object across the fields of view of different cameras. Transition zones make it easier to locate and track an object of interest across multiple video streams, allowing the user to follow an object since they can click on the transition zone to begin viewing the recording from a camera linked to the stream of the camera they are currently viewing. Without transition zones, a user would have to locate the appropriate camera from among potentially several dozen or more than one hundred cameras, causing the user to lose track of the object as it moves from one camera's field of view to another's.This not only allows the user to track the object more effectively, but also increases the amount of video footage the user can review. Anchor points also help users quickly determine which cameras may have captured an object in the past. For example, 1408 anchor points can provide a user with the ability to quickly review footage from any number of cameras to determine a path the object has taken. Because anchor points can be shared between different users, if one user has viewed and created an anchor point for the same object a second user is tracking, the second user can access the anchor points created by the first user to determine a path the object took before the second user lost track of it. As a result, a user can review more video footage and more precisely mark when an object was seen.When a user uses transition zones and anchor points, they can more accurately mark (create anchor points) when an object of interest enters the camera's field of view and leaves the camera's field of view. In describing exemplary embodiments, specific terminology is used for the sake of clarity. For descriptive purposes, each specific term is intended to encompass at least all technical and functional equivalents that operate similarly to achieve a similar purpose. Furthermore, in some cases where a particular exemplary embodiment includes multiple system elements, device components, or method steps, those elements, components, or steps may be replaced by a single element, component, or step. Likewise, a single element, component, or step may be replaced by multiple elements, components, or steps that serve the same purpose. Moreover, although exemplary embodiments have been shown and described with reference to particular embodiments thereof, those skilled in the art will understand that various substitutions and alterations in form and details may be made without departing from the scope of the invention.Furthermore, other aspects, functions, and advantages are also within the scope of the invention. Sample flowcharts are provided herein for illustrative purposes and are not exhaustive examples of the methods. A person skilled in the art will recognize that the sample methods may include more or fewer steps than those illustrated in the sample flowcharts and that the steps in the sample flowcharts may be performed in a different order than shown in the illustrative flowcharts.

Claims

1. A system for locating a detected object in a video, the system comprising: a screen, a memory that stores executable instructions, and at least one processor programmed to execute the instructions contained in the memory to: detect a bounding box around an object in a first frame at a first time in the video; identify a second frame in the video that corresponds to a second time; determine that there is no movement within the bounding box of the second frame; compare at least one of a first edge piece of information, or first color piece of information, or first intensity piece of information associated with one or more pixels within the bounding box in the first frame, to a corresponding second edge piece of information, or corresponding second color piece of information, or corresponding second intensity piece of information associated with one or more pixels within the bounding box in the second frame;generate a score based, at least in part, on the comparison; determine based on the score that the object is not present in the second frame; and determine an estimated time period when the object first appeared in a video stream.

2. The system of claim 1, wherein the comparison is based at least in part on a pixel overlap between an edge mask of the at least one first edge information and an edge mask of the at least one second edge information.

3. The system of claim 1, wherein the at least one processor is further programmed to execute the instructions contained in this memory to: cause multiple frames to be displayed on the screen within the time period.

4. The system of claim 1, wherein the at least one processor is further programmed to execute instructions corresponding to a backspace module, thereby causing the at least one processor to: determine whether the object is present in one or more frames prior to the first frame in the video transmission.

5. A system for tracking an object in multiple cameras, the system comprising: a display; a memory that stores computer-executable instructions; and at least one processor programmed to: cause a transition zone overlay to be displayed on an image on a display, the transition zone overlay being selectable by a user to navigate between a first field of view of a first camera and a second field of view of a second camera; and in response to the user's selection of the transition zone overlay, cause the second field of view of the second camera to be displayed as the object moves between the first field of view and the second field of view.

6. The system of claim 5, wherein the at least one processor is further programmed to: generate one or more first anchor points in the first field of view, the one or more first anchor points representing the identification of an object of interest in the first field of view.

7. The system of claim 6, wherein the at least one processor is further programmed to: generate one or more second anchor points in the second field of view, the one or more second anchor points representing the identification of the object of interest in the second field of view.

8. The system of claim 7, wherein the at least one processor is further programmed to: generate a composite video based at least in part on one or more first anchor points and one or more second anchor points over time.

9. The system of claim 7, wherein the at least one processor is further programmed to execute instructions corresponding to a trajectory reconstruction module, thereby causing the at least one processor to: reconstruct a trajectory that the object travels through the field of view of multiple cameras.

10. A non-transient, computer-readable medium that stores computer-executable instructions stored therein, which, when executed by at least one processor, cause at least one processor to perform the operations of: detecting a bounding box around an object in a first frame at a first time in the video; identifying a second frame in the video that corresponds to a second time; determining that there is no movement within the bounding box of the second frame; comparing at least one of a first edge piece of information, or first color piece of information, or first intensity piece of information associated with one or more pixels within the bounding box in the first frame, to a corresponding second edge piece of information, or corresponding second color piece of information, or corresponding second intensity piece of information associated with one or more pixels within the bounding box of the second frame;generate a score based, at least in part, on the comparison; determine based on the score that the object is not present in the second frame; and determine an estimated time period when the object first appeared in a video stream.

11. The non-transient, computer-readable medium of claim 10, wherein the comparison is based at least in part on a pixel overlay between an edge mask of the at least first edge information and an edge mask of the at least second edge information. ML / I 1303 12. The non-transient computer-readable medium of claim 10, wherein the at least one processor is further programmed to perform the operations of: causing multiple frames to be displayed within a period of time on a screen.

13. A non-transient, computer-readable medium that stores computer-executable instructions stored therein, which, when executed by at least one processor, cause at least one processor to perform the operations of: causing a transition zone overlay to be displayed over an image on a screen, the transition zone overlay being selectable by a user to navigate between a first field of view of a first camera and a second field of view of a second camera; and in response to the user's selection of the transition zone overlay, causing the second field of view of the second camera to be displayed as the object moves between the first field of view and the second field of view.

14. The non-transient computer-readable medium of claim 13, wherein the at least one processor is further programmed to perform the operations of: generating one or more first anchor points in the first field of view, the one or more first anchor points representing the identification of an object of interest in the first field of view.

15. The non-transient computer-readable medium of claim 13, wherein the at least one processor is further programmed to perform the operations of: generating one or more second anchor points in the second field of view, the one or more second anchor points representing the identification of the object of interest in the second field of view.

16. The non-transient computer-readable medium of claim 13, wherein the at least one processor is further programmed to perform the operations of: generating a composite video based at least in part on one or more first anchor points and one or more second anchor points over time.

17. A method for locating a detected object in a video, the method comprising: detecting a bounding box around an object in a first frame at a first time in the video; identifying a second frame in the video corresponding to a second time; determining that there is no movement within the bounding box of the second frame; comparing at least one piece of first edge information, or first color information, or first intensity information associated with one or more pixels within the bounding box in the first frame, to a corresponding second piece of second edge information, or corresponding second color information, or corresponding second intensity information associated with one or more pixels within the bounding box of the second frame; generating a score based, at least in part, on the comparison; and determining based on the score that the object is not present in the second frame.and determine an estimated time period when the object ML / E / ZuZZ / u 1300 first appeared in a video transmission.; 18. The method of claim 17, wherein the comparison is based at least in part on a pixel overlap between an edge mask of the at least one first edge information and an edge mask of the at least one second edge information.

19. The method of claim 17, wherein the method further comprises: causing multiple frames to be displayed within the time period on a screen.

20. A method for tracking an object across multiple cameras, the method comprising: causing a transition zone overlay to be displayed on an image on a screen, the transition zone overlay being selectable by a user to navigate between a first field of view of a first camera and a second field of view of a second camera; and in response to the user's selection of the transition zone overlay, causing the second field of view of the second camera to be displayed as the object moves between the first field of view and the second field of view.

21. The method of claim 20, the method further comprises: generating one or more first anchor points in the first field of vision, the one or more first anchor points representing the identification of an object of interest in the first field of vision.

22. The method of claim 21, the method further comprises: generating one or more second anchor points in the second field of vision, the one or more second anchor points representing the identification of the object of interest in the second field of vision.

23. The method of claim 22, the method further comprises: generating a composite video based at least in part on one or more first anchor points and one or more second anchor points over time.