Methods and systems for video-based object tracking using isochrones
Patent Information
- Authority / Receiving Office
- AU · AU
- Patent Type
- Applications
- Current Assignee / Owner
- FUSUS LLC
- Filing Date
- 2024-12-12
- Publication Date
- 2026-07-30
AI Technical Summary
Existing systems lack the capability to efficiently track objects across a geographic area monitored by multiple video sources in real-time or near real-time, particularly in emergency response and crime center dispatch scenarios.
The system employs isochrones to determine the geographic area reachable by a tracked object based on its speed and direction, allowing for the identification of relevant video sources and continuous tracking of the object's path forward or backward in time.
This approach enables real-time or near real-time tracking of objects across multiple video sources, improving response times and decision-making in emergency and crime center dispatch operations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
METHODS AND SYSTEMS FOR VIDEO-BASED OBJECT TRACKINGUSING ISOCHRONESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 609,106 entitled “METHODS, SYSTEMS, AND COMPUTER PROGRAM PRODUCTS FOR ARTIFICIAL INTELLIGENCE (AI)-ASSISTED VIDEO BASED MULTI-OBJECT TRACKING” filed on December 12, 2023; U.S. Provisional Patent Application No. 63 / 609,051 entitled “METHODS, SYSTEMS, AND COMPUTER PROGRAM PRODUCTS FOR VIDEO CAMERA SELECTION ISOCHRONES” filed on December 12, 2023; U.S. Provisional Patent Application No. 63 / 609,025 entitled “METHODS, SYSTEMS, AND COMPUTER PROGRAM PRODUCTS FOR ISOCHRONE-BASED HISTORICAL PATH MAPPING” filed on December 12, 2023; and U.S. Provisional Patent Application No. 63 / 609,038 entitled “METHODS, SYSTEMS, AND COMPUTER PROGRAM PRODUCTS FOR LOSS PREVENTION USING ARTIFICIAL INTELLIGENCE VIDEO TRACKING” filed on December 12, 2023, the entire contents of which are incorporated by reference herein.BACKGROUND Field of the Description
[0002] The present description relates to methods and systems for video-based object tracking using isochrones.Description of Related Art
[0003] A need exists for improved methods and systems, particularly in emergency response and crime center dispatch, for automatically locating, identifying, and tracking an object across a geographic area monitored by multiple video sources in real-time or near real-time.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] FIG. 1 is a diagram overview of the system according to an embodiment of the subject matter described herein.
[0005] FIG. 2 is a diagram illustrating functional components of a processing hub according to an embodiment of the subject matter described herein.
[0006] FIG. 3 is a flow chart of a method performed by the system according to an embodiment of the subject matter described herein.
[0007] FIG. 4 is a map showing 15- and 30-minute isochrones from an initial location according to an embodiment of the subject matter described herein.
[0008] FIG. 5 is a map showing isochrones adjusted based on the speed and direction of travel of the tracked object according to an embodiment of the subject matter described herein.
[0009] FIG. 6 is a map showing the path and corresponding isochrones of a tracked object observed by video sources according to an embodiment of the subject matter described herein.
[0010] FIG. 7 is a map showing isochrones for multiple tracked objects having different speeds and directions according to an embodiment of the subject matter described herein.
[0011] FIG. 8 is a bar chart showing a cumulative value over time for an object as compared to a threshold according to an embodiment of the subject matter described herein.
[0012] FIG. 9 is a composite image of a view captured by a video source before and after an incident that may be provided in a user interface according to an embodiment of the subject matter described herein.
[0013] FIG. 10 is an illustration of an object tracking user interface that may be generated and presented to an emergency or crime center dispatcher according to an embodiment of the subject matter described herein.DETAILED DESCRIPTION
[0014] The subject matter described herein includes methods and systems for video-based object tracking using isochrones in real-time or near real-time. One computer-implemented method includes detecting a trigger event and, in response, determining an initial time and a location from the trigger event. Video for the time and location is then obtained from video sources and an object to be tracked is identified based on analysis of the video. A speed and direction of the object is then determined, and additional video sources are identified using isochrones that account for the speed and direction of the tracked object. Video is then obtained from additional video sources along the predicted path of the tracked object to locate the object’ s earlier or later location in order to continue to track the object’s path forward or backward in time.
[0015] The subject matter described herein alternate or additionally includes systems and methods to support users, such as emergency dispatchers, when responding to an incident by tracking objects in real-time or near real-time. The system may provide a central hub or repository for gathering live video and data streams and processing them together to provide a user interface that can be displayed on a client device.
[0016] The system described herein addresses problems with prior systems by combining multiple streams of video and data, and thereby providing for automatically tracking an object across a geographic area covered by one or more video sources in real-time or near real-time. Artificial intelligence (Al) algorithms, including multimodal and hybrid Al models, can be implemented to enable video and data to be analyzed faster and to create automation that was not previously available.
[0017] Particularly, the system processes information gathered from a network of cameras to determine a location and orientation of each camera. The system then uses the location of the object at a given time to determine which cameras will have a view of the object at a later or earlier time.
[0018] The system may produce an isochrone-map, as part of a user interface, centered on an initial time and location of the object to identify potentially useful camera feeds. According to one aspect, the system al s o generates an interface that can be presented to a user. The interface may include icons that, when selected, display a camera feed associated with the icon. The interface may be updated in real-time or near-real-time to display camera icons near the object at a given time using the time, location, and direction of travel of the object.
[0019] The determination of potentially useful cameras using an isochrone map may be performed using Al-assisted methods for calculating the geographic area reachable by a tracked object andautomatically updating the isochrone map based on the updated location of the tracked object. Because the system receives and knows the location of all cameras and their orientations, the system can be configured to generate a set of cameras that are in the vicinity of the current or active camera's location. This enables the system to follow a person, vehicle, or other tracked object moving across cameras.
[0020] If the object is identified in a video feed, the system updates the location of the object, updates the isochrone map based on the location, and updates the user interface to indicate the projected path of the object. In this manner, the user can capture images of the object at various times and locations.
[0021] The system may include one or more software modules for determining a speed and a direction of travel of the tracked object. The speed and direction of travel of an object may be determined by comparing changes in the object’s location over time. For example, the system may receive emergency call for service information reporting an initial location of the object. The emergency call for service may comprise an emergency call for service and the emergency call for service information may comprise, for example, 9-1-1 call information. In another example, the system may analyze video feeds from multiple cameras having known locations to automatically update the object’s location. Regardless of how the object’s location is updated, the system may update a user interface (e.g., an isochrone map view) to identify the geographic area reachable by the object within a configured time frame based on assumptions or calculations about the object’s speed or method of travel.
[0022] The following is a discussion of a system for tracking objects in real-time or near real-time using the networked data gathered from the video sources.
[0023] FIG. 1 is an overview diagram of the system according to an embodiment of the subject matter described herein. The system 100 may include one or more video sources 102, one or more data sources 104, one or more client devices 106, and a processing hub 108 connected via communications network 110. The system 100 may be implemented as a client / server type architecture but may also be implemented using other architectures, such as cloud computing, software as a service model (SaaS), a mainframe / terminal model, a stand-alone computer model, and a plurality of lines of code on a non-transitory computer readable medium.
[0024] Video information may be generated by the video source(s) 102. Each video source 102 may include a camera 112 for capturing a video stream or video feed 120. The video sources 102 may take a variety of forms such as drones, traffic cameras, private cell phone video, building securitycameras, user-utilized robots with cameras, and so on.
[0025] The camera 112 may be associated with a known location 114 and a known orientation 116. The location 114 and the orientation 116 of the cameras 112 may be used by the system to determine whether the camera 112 is predicted to be near the projected or historical path of a tracked object.
[0026] The camera 112 may be associated with a camera appliance 118. The camera appliance 118 may mediate communication between the camera 112 and the network 110. In one embodiment, the camera 112 may provide a video stream but not provide any image processing, data compression, or other processing. The camera appliance 118 may perform these additional functions to augment the capabilities of the camera 112. Thus, the combination of the camera 112 and its associated camera appliance 118 may be referred to as a video source 102.
[0027] The camera appliance 118 may be equipped with Al at the edge-type code / software to turn the camera 1 12 into a smart, cloud-connected device capable of analyzing data as close as possible to the source. In some embodiments of system 100, video data may be processed by the camera appliance 118 to generate, alter, or compress video information which may be transmitted as video 120 to the processing hub 108.
[0028] The camera appliance 118 may or may not be physically located in proximity to the camera 112. For example, the functions performed by multiple camera appliances 118 may be executed in the cloud (e.g., as part of processing hub 108). In another embodiment, the camera appliance 118 may be integrated with the camera 112. The camera appliance 118 may pre-process video from the camera 112 to compress a video stream 120 provided to other devices. The camera appliance 118 may also multiplex or demultiplex video streams, process or add metadata to the video stream 120, and / or perform Al-based image processing (e.g., object detection) on the video stream 120.
[0029] The data sources 104 may include a source of information other than the video sources 102. Notable data sources 104 include sources of map data and sources of trigger events. Each of these types of data sources are described in greater detail below.
[0030] The map data source 104 may be a cloud-based system that stores and provides the map data 122 to the processing hub 108. The processing hub 108 may request updated map data 122 periodically, subscribe to updated map data 122, or receive push updates including map data 122 from the map data source 104. The map data 122 may include map tiles, geospatial coordinates (e.g., latitude and longitude, GPS, street intersections) and topographic details about terrain, elevation, and landforms. The map data 122 may also include directions, traffic flow, and connectivity of streets androads. The map data 122 may also include administrative boundaries outline the borders of cities, states, countries, and other political or administrative regions.
[0031] If the data source 104 is a source of trigger event data, it may be referred to as a trigger event data source 104 and may provide data describing the occurrence of a trigger event. A trigger event is a condition that, when satisfied, may cause the processing hub 108 to begin object tracking.
[0032] In one embodiment, trigger event data may be a signal produced by an environmental sensor, such as an audio sensor. The audio sensor may be configured to detect gunshots and determine a time and an approximate location of the gun shot. When a gunshot is detected, the trigger data source 104 may generate a signal including the trigger data 122 and transmit the signal including the trigger data 122 to the processing hub 108. It is appreciated, however, that various other types of environmental sensors may be used as trigger data sources 104 and are not limited to audio gunshot sensors.
[0033] In another embodiment, the trigger data source 104 may include an emergency dispatch system and the trigger event data 122 may include data gathered by an emergency dispatch system related to a emergency call for service. The emergency dispatch system may comprise an emergency dispatch system and the emergency call for service may comprise a 9-1-1 call. When a 9-1-1 call is made in the United States, the Enhanced 9-1-1 (E911) system automatically collects specific information to assist emergency services. The data gathered may include the Automatic Number Identification (ANI) (i.e., the caller's phone number) and the Automatic Location Identification (ALI) (i.e., the caller's location). For landline phones, this location is the registered address associated with the phone number. For mobile phones, location information is provided in two phases: Phase I transmits the caller's phone number and the location of the cell tower handling the call, while Phase II offers more precise location data using GPS or network-based triangulation to determine the caller's geographic coordinates. Additionally, a timestamp records the date and time when the call was initiated. The E911 system also uses the Emergency Service Routing Key (ESRK) or Emergency Service Routing Digit (ESRD), codes employed to route the call to the appropriate Public Safety Answering Point (PSAP).
[0034] This information may be transmitted to the PSAP, where it is displayed on the call taker's system. The data is integrated into the PSAP's telephony and Computer-Aided Dispatch (CAD) systems for use in coordinating emergency response. Call details may be stored in secure PSAP databases and logging systems, to maintain records for operational purposes, legal compliance, andauditing. The call details stored by PSAP may then be provided to the processing hub 108 as data 122.
[0035] The client device 106 may take the form of a computing device that may communicate directly or indirectly with the processing hub 108 over the network 110 and may take the form of a desktop computer, mobile phone, or dedicated terminal. The client device 106 may include a display 124 (e.g., a touchscreen or monitor screen) operable to display or present a user interface 126 to a user.
[0036] The user interface 126 may include one or more data and / or video layers generated and transmitted by the processing hub 108 during operations of the system 100. Display 124 may be operable to display the user interface 126 with one or more layers of video, data, or combinations thereof generated and transmitted by the processing hub 108.
[0037] The processing hub 108 may pull in real-time data, such as video 120 and data 122 from a variety of sources and display it in a user interface 126 (e.g., map-based format). Primary users of the user interface 126 may be real-time crime center personnel and emergency call for service operators using the client devices 106 and officers, emergency users, SWAT leaders, event and incidence response coordinators, and the like using the user interface 126 to direct unfolding situations.
[0038] FIG. 2 is a diagram illustrating functional components of a processing hub according to an embodiment of the subject matter described herein. The processing hub 108 may take the form of one-to-many computing and data storage devices that are cloud-based or accessible via the Internet or other communications network 110. For case of explanation, though, the processing hub 108 is shown to include a processor 200 that manages input / output (VO) devices 202 for communications over the network 110. The processor 200 also manages storage and retrieval of information from data storage / memory 204, which may be a data storage device such as a server accessible directly or over the network 110 by the processor 200.
[0039] The camera orientation data 116 associated with the camera 112 may be provided to the processing hub 108 as part of the data stream 122. The processing hub 108 may store the received camera orientation data 208 in data storage / memory 204. In some embodiments, the received camera orientation data 208 may be determined when the camera 112 was installed. For example, if the camera 112 included a fixed, pole-mounted surveillance camera, the camera’s orientation may be recorded when it was installed. In other embodiments, the received camera orientation data 208 maybe inferred from other information. For example, knowing the street layout and locations of landmarks from the received map data, the processing hub 108 may analyze the video feed 120 received from the camera 112 to identify the camera’s location and orientation.
[0040] The received map data 210 may include map data received from the data source 104. Map data refers to a collection of geographic information used to create and display maps (e.g., in digital applications like Google Maps). The map data 210 encompasses more than just visual representations and may also include metadata that enhances the functionality and utility of the maps.
[0041] The map data 210 may include geospatial coordinates (e.g., latitude and longitude, GPS, street intersections) and topographic details about terrain, elevation, and landforms. The map data 210 may also include directions, traffic flow, and connectivity of streets and roads. Traffic and transit information may include real-time or historical data on traffic conditions, public transportation routes, and schedules. The map data 210 may also include administrative boundaries outline the borders of cities, states, countries, and other political or administrative regions.
[0042] Cloud-based systems may store and provide the map data 210 to the processing hub 108 through distributed databases and scalable architectures. Application Programming Interfaces (APIs) can be used by the processing hub 108 to request specific services such as geocoding, routing information, or map tiles, and receive real-time responses from the data source 104.
[0043] The received trigger event data 212 may include information received from the data sources 104 indicating that an event has occurred that triggers the processing hub 108 to begin object tracking. As mentioned above, the received trigger data 212 may include emergency call for service call data provided by an emergency dispatch system. The data gathered may include an ANI and an ALI which provides the caller's phone number and location.
[0044] Once the trigger event data 212 is received, the trigger module 214 may interpret, analyze, convert, or reformat the trigger event data 212 into another format for further processing. If the trigger event data 212 includes emergency call for service information indicating a time of the call, but not a location, the trigger module 214 may automatically attempt to determine the location. For example, if the trigger event data 212 includes a live audio stream or audio recording of a 911 call, the trigger module 214 may automatically analyze the transcript for indications of a location. For example, if the caller described their location with reference to street names or nearby landmarks, the trigger module 214 may automatically identify these indications of the caller’s location and combine it with the time information that was automatically captured by the E911 system.
[0045] The initial camera set determination module 216 may be configured to determine an initial set of cameras. The initial camera set determination module 216 may compare the time, location, or other information generated by the trigger module 214 with the camera location data 206 and the camera orientation data 208 to determine the cameras having a view of the incident location.
[0046] A default time window and geographic range may be used when determining which cameras to include in the initial camera set 218. For example, a default time window of + / - 30 seconds and a range of 100 meters may be used. These values may be configurable by the user or may be automatically updated by the system.
[0047] The initial set of cameras determined by the initial camera set determination module 216 may be stored as initial camera set 218 from among the set of cameras 112.
[0048] As a first selection criterion, the camera prediction module 232 may use the camera location 206 to identify or select a predicted set of cameras 240. The camera prediction module 232 may determine which of the cameras 112 are in a vicinity surrounding the geographical space captured in the video stream of the active camera. The size of this space may be determined relative to various criteria. For example, the vicinity may comprise a 360-degree ring with an inner diameter set by the edge of geographic space captured by the active camera and an outer diameter set to provide a ring thickness or depth of 50 to 1000 feet or more. Alternately, this space may be determined to cover 1 to 12 city blocks or some other predefined vicinity or one that may be varied based on a speed at which the tracked object 222 is traveling with larger sized spaces being useful for faster moving objects and smaller sized spaces useful for slower objects such as person on foot or on a bike or scooter.
[0049] As a second selection criterion, the module may also verify that the camera 112 has the proper orientation 116 to have a field of vision to capture images of the tracked object 222. For example, a camera in a useful location may not be selected for the set of cameras if it is pointed away from an area of interest. Specifically, if the tracked object 222 is a vehicle, the cameras with orientations 116 useful for capturing streets or spaces where a vehicle may travel would be selected for the set whereas a different set of cameras may be selected when the object is a person on foot (e.g., store video may be useful in such situations and cameras oriented to capture inside spaces may be included in the set of cameras as the person may travel into such spaces).
[0050] The video stream 120 from the video source 102 may be stored in memory 204 as received video 120. The data stream 122 from the data source 104 may be stored in memory 204 asreceived data 122. The received data 122 may include location data 114 and orientation 116 for the video camera 112. The data stream 122 may be provided as part of, or separately from, the video stream 120.
[0051] Next, the received video 120 may be analyzed by the object identification module 220. Object detection may include identifying and localizing objects within the frames of a video sequence. A goal of object detection is to detect the presence of objects and determine their positions within the frame. Al-based image processing leverages machine learning algorithms to analyze pixel data and extract features relevant to object detection. Convolutional neural networks (CNNs) may be particularly effective in this domain due to their ability to automatically learn hierarchical feature representations from raw images. CNNs consist of multiple layers that apply convolutional filters to the input data, capturing spatial hierarchies and patterns that indicate the presence of objects. Object detection may also include pattern matching techniques that compare segments of the image to predefined templates or feature descriptors to identify regions that match known object patterns.
[0052] Once objects are detected, object identification, or classification, assigns specific labels to the detected objects based on their features. This process involves analyzing characteristics such as size, shape, texture, and movement to determine the object's category . For example, if one or more frames from a video show a person, object detection algorithms first identify the regions containing the object. Object identification algorithms then process these regions to recognize that the object is a person by evaluating features like the human silhouette, facial structures, or typical movement patterns associated with walking or running. Techniques such as deep learning classifiers, including CNNs trained on labeled datasets, arc commonly used for this purpose. The combination of detection and identification enables systems to locate objects within video frames and to understand what those objects are.
[0053] Meta-attributes of an object may be extracted from the video 120 by applying specialized algorithms to images of the tracked object. The process may begin with region of interest extraction, where the system isolates the area containing the individual to focus the analysis. Visual features pertinent to the meta-attributes may then be extracted from this region using the same CNNs to learn hierarchical feature representations from image data.
[0054] Depending on the meta-attribute, different models may be employed for meta- attribute classification. Gender and race recognition, for example, may use pre-trained CNN classifiers to analyze facial features and morphological characteristics using large, labeled datasets to capturesubtle differences in facial structures. Height estimation may be achieved by analyzing the object’s scale relative to known references in the scene or by utilizing camera metadata when available. Techniques such as pose estimation can also be used to model skeletal structure for calculating an object’s height.
[0055] Determining the color of an object’s clothing may involve segmenting a region of interest to isolate items like shirts and pants. The system may then perform color space analysis, often in the HSV (Hue, Saturation, Value) color space, to identify the dominant colors of these clothing items while accounting for valuations in lighting conditions. To detect whether the person is wearing a hat, object detection algorithms based on CNNs, such as Faster R-CNN or YOLO (You Only Look Once), may be used to scan the head region for features indicative of headwear.
[0056] Temporal analysis of multiple frames from the video may be used to aggregate information over time, improving the robustness of attribute detection. This approach may mitigate issues like occlusion, motion blur, or momentary changes in appearance. The outputs from multiple classifiers within the object identification module 220 may be integrated to form a comprehensive profile of the detected person, with probabilistic models or ensemble methods combining predictions and handling uncertainties.
[0057] The system may also apply confidence thresholds to ensure reliability, accepting attribute predictions if the confidence score exceeds a predefined limit. Any anomalies or conflicting information may trigger re-analysis or flag the data for human review. By utilizing advanced machine learning techniques and decomposing the task into specialized components targeting specific attributes, the Al image processing system can accurately determine meta-attributes like gender, height, race, clothing colors, and the presence of accessories such as hats. This modular approach allows for scalability and adaptability, enabling the system to incorporate additional attributes or enhance existing ones as more advanced models and data become available.
[0058] As used herein, an isochrone is a line or boundary on a map or graph that connects points of equal time, typically representing areas or locations that can be reached within the same time interval from a specific origin or event. Isochrones can be calculated based on various modes of travel, such as walking, driving, or public transportation, or other temporal phenomena. Depending on the mode of travel, isochrones may consider variables like terrain, road networks, traffic patterns, or other accessibility constraints. Isochrones may look forward in time, identifying areas or points that can be reached within a specified time window from an origin, or backward in time, determining the locationsor origins that could have contributed to a specific point within a given time frame.
[0059] As used herein, an isochrone map is a visual representation that uses isochrone lines or areas to depict zones of equal travel time or temporal reach from a given starting point. Isochrone maps show areas reachable from a specific location within defined time intervals. Isochrone maps can also illustrate the possible origin of phenomena based on historical time data.
[0060] Storing an isochrone includes capturing spatial and temporal information in a way that preserves geometric and attribute data. Isochrones may be represented as lines (boundaries) or polygons (enclosed areas) and stored with metadata such as the origin point, time intervals, and mode of travel. Several formats can be used.
[0061] GeoJSON is one format for storing isochrones that can store the geometric shape of the isochrone as a polygon or multipolygon, along with additional properties such as the origin and time interval. For example, a GeoJSON file might define a walking isochrone as a polygon with coordinates outlining the reachable area within a 15-minute interval.
[0062] Shapefiles are another example of a format for storing isochrones that may be used. Shapefiles include multiple files (.shp, .shx, .dbf, etc.) and support complex geometries and attributes. In contrast to some other isochrone formats, shapefiles are not human-readable and require specialized software like QGIS or ArcGIS for interpretation.
[0063] For storing the isochrones, databases such as PostGIS specialize in managing isochrone data. For example, PostGIS stores isochrone polygons directly as spatial data in a PostgreSQL database. Queries can then be performed on the stored data for analysis or visualization. Alternatively, NoSQL databases like MongoDB can store isochrone data in JSON format for easy integration with web applications.
[0064] Custom JSON or CSV formats can also be used. These formats can include time-based attributes and coordinates but may lack advanced spatial processing capabilities. For raster representations of isochrones, a format like GeoTIFF may be used because it encodes travel time as pixel values, which is useful for high-resolution, continuous travel-time surfaces.
[0065] It may be appreciated that isochrones 234 may include future or predicted path isochrones 236 and past or historical path isochrones 238. For example, starting at the time determined by the trigger module 214, a 30-second isochrone can describe the area in which the tracked object 222 may be found within the next 30 seconds or, alternatively, the 30-second isochrone can describe the area where the tracked object 222 might have been in the previous 30 seconds. It may be important todistinguish between future isochrones 236 and past isochrones 238 because the path of the tracked object is likely different before and after the incident.
[0066] The predicted camera set 240 includes any number of cameras 112 that are within a given isochrone 234. A predicted camera set 240 may be generated when the location of the tracked object 222 is updated and a new isochrone 234 may be determined. Once the boundary of an isochrone 234, any cameras 112 located within that boundary are identified as a predicted cameras set 240 and are associated with the isochrone.
[0067] Because there may be many isochrones 234 for the tracked object 222, there may also be many predicated camera sets 240. As the tracked object 222 moves, the prediction of which cameras the tracked object 222 is likely to be near changes. In this way, every predicted camera set 240 is associated with a particular time period. By scrubbing through successive time periods, whether in the future or the past, the user can see the changes in the predicated camera set 240. This process is illustrated and described later in FIGS. 5 and 6.
[0068] In some embodiments, the object tracking module 230 may automate the process of tracking cumulative metrics for the same tracked object 220 over time. The object tracking module 230 may evaluate these metrics against set thresholds and initiate appropriate actions in response. By doing so, the object tracking module 230 enables timely responses to significant events — such as substantial losses due to shoplifting — and supports informed decision-making within operational contexts.
[0069] The object tracking module 230 monitors and tracks cumulative values associated with specific entities over time, determining when a predefined threshold is met for the tracked object. Its primary purpose is to aggregate relevant data points for an object — such as a shoplifter — and compare the accumulated value against a set threshold. When the threshold is reached or exceeded, the module triggers predefined actions, like generating notifications or compiling detailed reports.
[0070] The object tracking module 230 receives input data related to individual events or transactions associated with the tracked entities. This data includes unique identifiers for the entities and quantifiable values pertinent to the cumulative total. The object tracking module 230 also maintains an ongoing cumulative total for the tracked object by aggregating these incoming values over time, updating the stored total whenever new data is ingested.
[0071] The object tracking module 230 continuously evaluates the cumulative total for each entity against the predetermined threshold. Upon detecting that the total meets or exceeds the threshold, it initiates predefined actions. These actions might involve generating alerts to notify relevantpersonnel, compiling comprehensive summaries of the entity's activities, or interfacing with other systems to escalate the response.
[0072] In addition to monitoring and evaluation, the object tracking module 230 handles data management tasks and securely stores historical data for each entity, including timestamps, event details, and cumulative totals, facilitating auditing, reporting, and further analysis. The object tracking module 230 also manages entities by allowing the addition of new ones and updating existing records, ensuring that tracking remains accurate and up-to-date.
[0073] Administrators can configure the object tracking module 230 by defining thresholds, specifying the actions to be taken when thresholds are breached, and adjusting other operational parameters to suit the organization's needs. The object tracking module 230 may include mechanisms for error handling and data validation to ensure the integrity of the tracking process. The object tracking module 230 may also incorporate security measures to protect sensitive information and control access according to established policies.
[0074] The object tracking module 230 is thus configured for tracking a cumulative value over time and comparing it against a predetermined threshold to trigger a specific action when that threshold is reached or exceeded. This may be used to monitor ongoing activities, aggregate relevant data points, and initiate responses based on the accumulated value. The cumulative value can represent any measurable quantity or metrics. The threshold serves as a control limit that, once met, prompts the system to execute predefined actions like generating alerts.
[0075] Finally, it may be appreciated that in some embodiments, the information received and processed as described above may be displayed in a user interface, such as the user interface 126 displayed on the client device 106. The user interface module 242 may include one or more subroutines or callable applications to create a common operating picture for users as the user interface 126. For example, these subroutine / applications may operate to provide additional data views to video 120 and data 122 and to provide controls that can be stored within a tab in the user interface 126. The users may also be allowed to configure (and pre-configure via profiles) specific user views provided by the user interface module 242.
[0076] The user interface 126 may provide an isochrone-based camera selection interface (see FIG. 5 for example) that is generated by the user interface module 242 via calls to the camera set selection module. During operations of the system 100, an object may be identified in received video 120 that is to be tracked, and data file may be stored for this object in memory 204 as shown at. The directionof travel may be determined for the tracked object 222 by the module or another module in hub 108 along with its current location and a set of Al-drivcn meta- attributes.
[0077] The user interface 126 generated by the user interface module 242 may include a main viewing or active camera viewing window that displays the received video stream 120 from the active camera. The user interface 126 may further include icons corresponding to the location of cameras on the map. Views of the user interface 126 may be configurable by the user interface module 242 based on default or user- modified interface profiles. Profiles can be used by users to cause the user interface module 242 to bring in various video elements 120 and data elements 122 as needed and may be provided in user- selectable or default data / video layers.
[0078] The user interface 126 may further be configured to include a tab or button that the user can select to initiate Al-based tracking of the tracked object 222 by the module. In response, the module may utilize the Al-driven meta-attributes as well as other data (such as the last known location and / or direction of travel) to process the received video 120 of all or a subset of the cameras 112 to detect the tracked object 222.
[0079] The predicted camera set 240 can be provided to the user interface module 242 for use in a section or portion of the user interface 126. For example, a map-based portion of the interface 126 may provide a map with a centrally-located icon representing the current location of the object. The direction of travel of the object may be displayed, such as with an arrow leading away from the active camera icon, along with icons corresponding to the cameras in the predicted camera set 240 of cameras along the direction of travel.
[0080] FIG. 3 is a flow chart of the method according to an embodiment of the subject matter described herein. The method 300 may begin with implementing the system 100 by providing the processing hub 108 on the network 110 and communicatively linking the client device 106 to the processing hub 108 and achieving a data link with data sources 104 and video sources 102 so as to make them cloud-based (e.g., by providing the appliance on cameras 112 and / or networks of cameras 112). Thus, the method 300 may start by configuring a mobile phone, laptop computer, dispatch center system, crime center system, or the like with software as discussed above to process feeds from a network of video cameras to track a moving object across a geographic space.
[0081] Step 302 includes receiving a trigger event and determining a time and a location of an incident indicated by the trigger event. For example, a trigger event may include a call for service received by an emergency or crime center dispatch system.
[0082] In another example, the trigger event may include receiving a signal generated by an environmental sensor. For example, an audio sensor may detect gun shots and provide a time, and an approximate location of the gun shot.
[0083] Step 304 includes determining the time and the location of the incident based on the trigger event.
[0084] In one embodiment, the first time and the first location of the incident may be determined based on location information received in the signal from the environmental sensor.
[0085] In another embodiment, the first time and the first location of the incident may be based on processing the emergency call for service using Natural Language Processing (NLP) by an LLM- based Al model. This processing may include automatically generating a transcript of the emergency call for service and analyzing the transcript for indications of the first time and the first location of the incident. An NLP-based analyzer may perform NLP (and / or other analysis) of the dispatcher call and / or the dispatcher incident narrative of the call, which is stored in the memory 204. The analysis of the call and / or the dispatcher incident narrative may be used to determine a location (e.g., latitude and longitude) of the incident to which the call is related.
[0086] Step 306 includes obtaining video from an initial set of cameras associated with the time and location of the incident. For example, a module may be conf igured for generating an incident- specific camera set from the available cameras 112 (or video sources 102). The cameras 112 determined to be in the initial camera set 218 may be public and / or private cameras. The initial camera set 218 may also be associated with a geographical region associated with the location of the incident. In some cases, the region or area used to locate useful cameras 112 for the set is defined by the visible range circumference about the incident location. The initial camera set 218 in some cases includes cameras with an orientation 116 that would allow that camera 112 to capture some or all of the scene associated with the incident. In some cases, object detection in the video streams 120 is used for determining the initial camera set 218, e.g., is a detected object, such as a particular individual or vehicle, associated with the incident found in the video stream 120, and, if so, include that camera 112 in the initial camera set 218.
[0087] The automated use of camera orientation 116 allows the processing hub 108 to detect the cameras 112 within the visible range circumference of, and oriented on, the location of the incident. To capture all possibly useful cameras "oriented on" may be defined as some angular offset from the line of focus of the camera 112 (e.g., the line of focus may be plus or minus 30 degrees (or less) frombeing orthogonal to the location (latitude and longitude, for example) of the incident). The locations of the cameras 112 in the set arc indicated via camera icons on the user interface 126.
[0088] Step 308 includes identifying an object to be tracked from a video. Object identification may include using one or more Al-based image processing techniques to determine one or more metaattributes of the object. These meta-attributes may include, but are not limited to, the color, size, shape, gender, race, or other characteristics of the object. For example, a man wearing red pants and blue shirt may be associated with meta-attributes: male, red pants, blue shirt. The meta- attribute indicating that it is pants that are red instead of something else may be inferred based on the location of the red color in the lower half of the bounding box of the man.
[0089] Various Al-based image processing techniques, algorithms, or models may be used to identify one or more meta-attributes of an object from a video. The Al-based image processing techniques used to identify meta-attributes of objects in a video may include several components and steps. These may include computer vision and deep learning methodologies that detect and classify objects based on their attributes across different video sequences.
[0090] For example, meta- attribute extraction may begin with a video pre-processing stage that includes frame extraction, where individual frames are extracted from the video at a suitable frame rate for analysis. Image pre-processing techniques such as normalization, resizing, and enhancement may also be applied to prepare frames for consistent analysis. Data augmentation enhances the robustness of models by simulating variations like rotations, scaling, and lighting changes during training.
[0091] In the object detection phase, algorithms like YOLOv5, Faster R-CNN, or SSD may be used and a selected Al model may then be trained or fine-tuned. For training, a dataset may be prepared by collecting and labeling images containing the objects of interest, such as persons, cars, and guns. Transfer learning fine-tunes pre-trained models on the custom dataset to improve detection accuracy. During inference, the detection model runs on each frame to locate objects, yielding bounding boxes and class probabilities. Post-processing steps like Non-Maximum Suppression (NMS) eliminate duplicate detections, and detections are filtered based on confidence thresholds.
[0092] The meta- attribute extraction phase identifies several types of attributes. For color detection, regions of interest (RO I) may be extracted by taking the pixels within the bounding boxes for detected objects. Images are converted from RGB to HSV color space for better color segmentation. Color classification is performed using histogram analysis or by training a color classification model usinga convolutional neural network to determine the dominant colors. For example, cars can be classified into predefined color categories like red, blue, or black.
[0093] For clothing attribute recognition, deep learning models which are trained to recognize clothing attributes may be used. Features related to clothing, such as style, type, and color may be extracted, and attributes such as shirt color, pant color, and clothing type (e.g., T-shirt, jacket) are predicted.
[0094] In demographic attribute classification, face detection and alignment are performed using models like MTCNN to detect and align faces within person bounding boxes. Attribute prediction is carried out by implementing models like VGGFace or ResNet, trained on demographic datasets to predict gender and race. It is important to be mindful of biases in datasets and models and ensure compliance with legal and ethical standards.
[0095] Feature representation and embedding involve deep feature extraction using CNNs to extract high-dimensional feature vectors representing the object's appearance. Dimensionality reduction techniques like Principal Component Analysis (PCA) or t-SNE can be applied for visualization and efficiency. Feature vectors are normalized to have unit length for consistent similarity measurements.
[0096] For object tracking within the video, multi-object tracking (MOT) algorithms like DeepSORT or ByteTrack are implemented to assign IDs to objects across frames. Data association uses motion (e.g., Kalman filters) and appearance features to link detections over time. Occlusion handling incorporates re-identification strategies to handle temporary object occlusions.
[0097] Cross-video object re-identification (Re-ID) may include global feature matching, where extracted features arc compared using distance metrics like Euclidean or Cosine Similarity. Attributebased filtering uses meta-attributes like color, clothing, and demographics to narrow down potential matches. Re-ID models are trained specifically for Re-ID tasks, such as Triplet Loss Networks or Siamese Networks, to improve matching accuracy. Features and attributes are stored in an indexed database, such as using FAISS for efficient similarity search.
[0098] Evaluation and metrics involve assessing detection performance using metrics like Precision, Recall, and mAP (mean Average Precision). Attribute classification accuracy may be evaluated to assess detection performance.
[0099] Continuous learning and maintenance include continuously retraining models with new data to improve accuracy and adapt to new conditions. A feedback loop may incorporate user feedback or manual corrections to refine the models.
[0100] It may be appreciated that changes in lighting, weather, and camera angles may affect detection and classification. Additionally, crowded scenes and other occlusions where objects may be partially visible or overlap with others may also make detection and classification more difficult. These challenges may be addressed by training models on diverse datasets and applying data augmentation techniques.. Robust tracking algorithms and Re-ID models are used to maintain object identities.
[0101] An example workflow might proceed as follows. A frame is processed, and a person and a car are detected using YOLOv5. For the person, clothing colors are identified, such as a red shirt and blue jeans, and gender is classified as male. For the car, the color is identified as black. Feature vectors for both the person and the car are extracted using a CNN. The person and car are tracked across subsequent frames, maintaining their IDs. In another video, a person with similar attributes and feature vector is detected. The system computes the similarity, and if above a threshold, flags it as the same individual. The findings are stored in a database, and alerts or reports are generated as needed.
[0102] Step 310 includes determining a speed and direction of the tracked object. In one embodiment, the object’s speed and orientation may be determined using video information obtained from a single camera as the object moves within the view of the camera. When initially determining the speed and direction of the object, the system may compare multiple frames of the initial video information to determine the object’s movement.
[0103] In other embodiments, the system may determine the object’s speed and direction over a longer timespan as the object traverses the views of multiple cameras. For example, the speed and direction of the object may be determined by comparing the object’s location relative to the known locations of fixed landmarks. These landmarks may include, but are not limited to, the locations of different video cameras and the distance between blocks. So, if the object was near a first video camera with a known location and ten seconds later was near a second video camera also having a known location, the system can determine how far the object traveled in ten seconds and in which direction.
[0104] Step 312 includes predicting a next set of cameras based on the speed and direction of travel of the object using one or more isochrones.
[0105] While a simple isochrone may represent a geographic area in which the tracked object may be located within a given time period since the location of the object was last determined, an isochroneaccording to the subject matter described herein may represent a more targeted geographic area in which the tracked object is more likely to appear based on the speed and direction of the object. For example, a basic isochrone may encompass 360 degrees around the tracked object. The camera set resulting from this basic isochrone would include all cameras located near the object.
[0106] In contrast to a simple isochrone, an intelligent isochrone that may be generated by the system described herein may take into account the speed and direction of the object. For example, for a vehicle as the tracked object, the system may elongate the isochrone in the direction of travel of the vehicle. Moreover, the isochrone may be reduced in the direction opposite to the direction of travel, as well as reduced in directions or areas which do not include roads and / or which include obstacles to the vehicle’s ability to move at a given speed. The result is a smaller isochrone that includes fewer predicted cameras in which the tracked object is likely to appeal-.
[0107] In some embodiments, the method 300 may also include generating an isochrone map. An isochrone map may indicate the locations of cameras, the path of object(s) after leaving an initial time and location or the isochrone map may indicate the historical path of the object before arriving at the initial time and location. These past and future paths may include actual determined locations, predicted locations, or a combination to provide the user with a real-time or near real-time view of the tracked object’s movements before, during, and after a particular time and location (incident). Exemplary isochrones and isochrone maps that may be generated by the user interface module 242 are described in greater detail below with respect to FIGS. 4-7.
[0108] Step 314 includes determining whether a condition to stop tracking the object has been satisfied. For example, video or still images of the tracked object may be presented to a user on a client device, such as a mobile phone, to a first responder, or may be displayed on a crime-center dispatch system to a dispatcher. The user may then confirm by visual inspection or otherwise, whether the system should continue tracking the identified object. This can save valuable computing resources by avoiding unnecessary object tracking.
[0109] FIG. 4 is an isochrone map showing geographic areas reachable by walking within 15 minutes and 30 minutes of a location according to an embodiment of the subject matter described herein. Referring to FIG. 4, in response to a trigger event identifying a time and location (of an alleged incident), the system may identify camera 402 as the nearest camera. Video may then be obtained from the camera 402 and analyzed to identify any objects within view of the camera 402 during the target timeframe.
[0110] Isochrone 406 illustrates a boundary encompassing a geographic area. The size of the isochrone 406 represents the area in which the object 404 may travel at an assumed maximum speed within a predetermined time. Isochrone 406 is shaped the way that it is because the system does not know the speed or direction of travel for the object 404.
[0111] To generate the isochrone 406, the system makes an assumption about the speed of the object 404. Here, the system does not identify the object or attempt to determine its actual speed. Instead, the system may use a predetermined value for the speed of the object 404. It will be appreciated, however, that the object 404 may be capable of moving faster than the assumed speed, or may be incapable of moving as fast as the assumed speed. This may be the case if the object 404 was instead a vehicle or a bicycle.
[0112] When generating the isochrone 406, the system does not make any assumptions about the future direction or path of the object 404. In other words, the object 404 may travel in any direction from its initial location. As a result, the isochrone 406 is largely symmetrical. However, rather than being a circle, some portions of the isochrone 406 reach farther from the initial location than do other portions of the isochrone 406. This may be the result of, for example, differences in terrain, which alter the speed of the object 404. In one embodiment, the object 404 may travel faster on paved roads than on other surfaces. Moreover, buildings may be located within the perimeter of the isochrone 406 which require the object 404 to circumnavigate.
[0113] The isochrone 406 may include a number of video sources. Here, the isochrone 406 includes the initial camera 402, as well as camera 408 and camera 410.
[0114] Location 412 and location 414 represent two possible future locations of the object 404 within the time period associated with the isochrone 406. Black object icons represent the object at a determined or known location and time. For example, the object 404 is represented by a block icon because it is confirmed to be at that location and time by the camera 402.
[0115] Black camera icons in image 400 may represent cameras from which video information was actually obtained and used to determine the location of a tracked object at a particular time. For example, video information provide by the camera 403 may confirm that the object 404 was located within the viewable area of the camera 402 at the time of the incident.
[0116] White icons represent possible or predicted future locations of the object 404. Here, the object 404 may be at location 412, location 414, or any other location within the isochrone 406. It may be appreciated that the distance and direction from the initial location may vary and is notknown to the system until the object 404 is observed by one of the cameras 402, 408, 410 at a time subsequent to the time of the incident.
[0117] Gray icons with no outline represent possible or predicted past locations of the object 404. Here, the object 404 may have been located at position 415 at some point within the predetermined time associated with the isochrone 406.
[0118] Isochrone 416 is like isochrone 406, but represents a longer period of time corresponding to a larger geographical area. Like when generating isochrone 406, the system may assume a speed and time period when generating isochrone 416. In the example illustrated in FIG. 4, the isochrone 406 may represent a time period of + / - 15 minutes from the incident and the isochrone 416 may represent a time period of + / - 30 minutes from the incident. The isochrone 416 may also be associated with a set of one or more additional cameras 418.
[0119] The isochrones 406 and 416, however, have a drawback. Because the isochrones 406 and 416 are not based on the observed speed or the direction of the object 404, the isochrones 406 and 416 may encompass a large geographic area for a given time period. As a result, the number of cameras 408, 410, 418 from which video information must be obtained and analyzed may prevent the system from tracking one or more objects in real-time or near real-time. In order to address this limitation, the system intelligently reshapes and resizes isochrones based on the observed speed and direction of the object 404. The process of determining the speed and direction of an object is discussed below.
[0120] FIG. 5 is a map showing isochrones adjusted based on the speed and direction of travel of the tracked object according to an embodiment of the subject matter described herein. In this example, the system determines that the object 504 (e.g., a person) is traveling North at a rate of 1 meter per second (m / s). Using this information, the system can significantly shrink the isochrones and reduce the number of video sources from which video is obtained and analyzed.
[0121] Here, for example, an isochrone 506 represents the predicted area in which the object 504 is likely to be located within 15 minutes after an incident occurring at a determined time and location (i.e., 504). This isochrone 506 predicts an area North of the incident location. The isochrone 506 encompasses just one camera, camera 508, within the predicted set of cameras for the isochrone 506.
[0122] Continuing after the incident, an isochrone 516 represents the predicted area in which the object 504 is likely to be located within 30 minutes after the incident and is also located North ofthe previous location. This indicates that the object 504 is expected to continue traveling North. The isochrone 516 includes just an additional camera, camera 526, within the predicted set of cameras for the isochrone 516. Thus, it may be appreciated that as the time frame increases, the length of the isochrone lengthens, but does not spread in all directions equally. This is because the isochrones shown in FIG. 5 assume that the object 504 will continue traveling at the same speed and direction as initially determined from the video of the incident.
[0123] Conversely, an isochrone 529 represents a geographic area which the object 504 is predicted to be located 15 minutes before the incident. In this case, the predicted past location 530 is South of the incident location. The geographic area encompassed by the isochrone 529, like the isochrone 529, includes a single camera 510.
[0124] Continuing into the past leading up to the incident, an isochrone 532 represents a geographic area which the object 504 is predicted to have been located 30 minutes before the incident. In this case, the geographic area encompassed by the isochrone 532 includes an additional camera 534.
[0125] It may be appreciated that videos are obtained from fewer cameras in FIG. 5 than in FIG. 4 due to the more targeted isochrones taking into account the observed speed and direction of the object. This allows the system to perform object tracking, including multi-object tracking, for one or more objects in real-time or near real-time.
[0126] FIG. 6 is a map showing the path and corresponding isochrones of a tracked object observed by video sources according to an embodiment of the subject matter described herein. In FIG. 6, the initial video obtained from the camera 602 from before the incident indicates that the object 604 is traveling North at walking speed (c.g., 1 m / s). Therefore, a relatively small isochrone 606 may be generated North of the incident location that predicts, with a statistically significant likelihood, that the object 604 will subsequently be located within the area encompassed by the isochrone 606. Indeed, in this example, the camera 608 captures video of the object 612.
[0127] Notably, neither the speed nor the direction is constant over time. Instead, based on the video from the camera 608, it may be determined that the person 612 has changed speed and direction. Here, the object 612 increases their speed (e.g., from walking to running) and turns East. As a result, the shape and size of the next isochrone - isochrone 616 - is longer than the isochrone 606 and encompasses a geographic area East of the last known position of the object 612. In this example, the isochrone 616 encompasses two more cameras than the isochrone 606, where the camera 638 captures video of the object 628 within its view during the target time frame.
[0128] Gray icons with a black outline may represent locations of the object 604 observed before the incident at its initial time and location 604. Here, for example, the isochrone 629 is similar to the isochrone 606, indicating that the object was traveling in a similar direction and speed before and after the incident (e.g., walking). Also like the predicted future path of the object, the object 630 in the isochrone 629 is observed transitioning from running to walking speed and changing direction from West to North (indicating a left turn). Thus, like the isochrone 616, the isochrone 632 may be elongated to account for the increased speed of the object and the isochrone 632 may be directed to the West to account for the change in direction of the object. As the system searches video information retrieved from cameras further forward or backward in time from appropriate locations, the path of the object both leading up and after the initial time and location determined from the trigger event may be determined.
[0129] F FIG. 7 is a map showing isochrones for multiple tracked objects having different speeds and directions according to an embodiment of the subject matter described herein. In this example, three objects are identified in the initial video obtained at the location of an incident by camera 702. The object 704 is a person running East from the incident. Based on the direction and speed of the object 704, an isochrone 742 may be generated that covers an area in which the person is likely to appear, as indicated by the predicted location(s) 746. Notably, the area within isochrone 742 does not include any additional video sources. As such, the time window of the isochrone 742 may be extended to include additional cameras, such as the camera 744 that is in the object’s predicted path.
[0130] Next, a bicycle object 748 may be identified based on the initial video information. In response to the real-time or near real-time analysis of the video information (e.g., frame comparison), it may be determined that the speed and direction of the object 748 is different than that of the object 704. This is indicated by predicted future positions 754 and 756. Based on the predicted future path of the object 748, an isochrone 750 may be determined that includes video source 752. Video information may be obtained from video source 752 and analyzed to determine whether the object 748 appears within view during a target time period. In this case, it may be appreciated that the isochrone 750 is larger than the isochrone 742 resulting from the object 748 traveling faster than the object 704.
[0131] Finally, a car object 758 may be identified based on the initial video information. In response to the real-time or near real-time analysis of the video information, it may be determined that the speed and direction of the object 758 is faster than the speed of the object 704 and the object 748.This is indicated by predicted future position 770. Based on the predicted future path of the object 758, an isochrone 760 may be determined that includes video sources 762 and 768. Video information may be obtained from video sources 762 and 768 and analyzed to determine whether the object 758 appears within view during a target time period. In this case, it may be appreciated that the isochrone 760 is larger than either the isochrone 742 or the isochrone 750 due to the object 758 traveling faster than either the object 704 or the object 748.
[0132] Thus, FIG. 7 illustrates an exemplary multi-object tracking scenario where three different types of objects are identified for tracking and where each object has a different speed and direction. In some embodiments, each object may be tracked (either in the future to the present or in the past to a given time) until an indication to stop tracking is received. This indication to stop tracking an object may include a signal produced by a user in response to viewing each tracked object and assessing its importance.
[0133] FIG. 8 is a bar chart showing a cumulative value over time for an object as compared to a threshold according to an embodiment of the subject matter described herein. Applying this to a shoplifting loss prevention scenario, a retail chain may wish to track the total value of merchandise stolen by an individual over time across multiple store locations. The retailer may wish to record each theft incident with details including the time, date, location, and the monetary value of the items stolen. The system may then aggregate these values to maintain a running total of the losses attributed to that individual.
[0134] When the cumulative stolen amount reaches a predetermined threshold, such as $5,000, the system may trigger a notification or other action. The notification may include relevant details for all of the incidents associated with a particular individual, as well as any relevant surveillance footage of the thefts. This information can enable the retailer to recognize patterns, assess the severity of the losses, and take appropriate actions such as enhancing security measures or involving law enforcement to address the issue.
[0135] In many jurisdictions, $5,000 is the threshold for felony theft. Rather than charging a shoplifter each incident separately, none of which individually are for more than $5,000, the retailer may wish to wait until it is a felony, which carries greater penalties. By automating the process of tracking cumulative metrics for specific entities, evaluating these metrics against set thresholds, and initiating appropriate actions when necessary, the system enables timely responses to significant events — such as substantial losses due to shoplifting.
[0136] FIG. 9 is a composite image of a view captured by a video source before and after an incident that may be provided in a user interface according to an embodiment of the subject matter described herein. Referring to FIG. 9, the composite image of the view 900 may be captured by a camera 602 located near the scene of a shooting incident. In response to receiving an emergency call for service indicating the time (fo) and location of the incident, the camera 602 was identified and a video clip retrieved for the 10 seconds before to and the 10 seconds after to. The system may then automatically analyze the video clip for any objects. In this example, the system identifies the object 604 to be tracked. In other examples, multiple objects (if any) may be automatically tracked by default.
[0137] Once the tracked object’ s initial location at to is identified, the system may attempt to identify the tracked object at other locations by comparing each frame of the video clip. In this example, the object 604 may be seen running North in the ten seconds before and after the incident. This is represented by multiple images of the object being displayed at different locations indicated by different times. In the composite image of the view 900, the object 630 appears larger in the foreground at to -10 seconds and the object 612 appears smaller in the background at to +10 seconds. It may, therefore, be appreciated that objects 63, 604, and 612 represent the same object at different times and locations within view of the camera during the target time frame.
[0138] From this video clip, the system may determine both the existence of the object, as well as its speed and direction of travel. In other embodiments, as mentioned previously, additional information, such as meta-attributes of the object (e.g., red pants, white shirt) and the mode of travel of the object (e.g., walking or miming on foot) may also be determined from the video clip shown in the composite image of the view 900.
[0139] It also may be appreciated that the composite image of the view 900 includes a four-way intersection between an East-West road and a North-South road and, from the middle of the intersection, one road leads in each one of these four cardinal directions. The composite image of the view 900 also shows a perspective showing true North following the road but displayed as being curved toward the upper right comer of the composite image of the view 900. This may lead the viewer to incorrectly assume that, as object 612 moves along the road, the object 612 is traveling North-East rather than North. However, the system may combine its knowledge of the street’s layout and orientation with its knowledge of the camera’s location and orientation to determine the direction of the road shown in the composite image of the view 900 (here shown with arrowsindicating to North, South, East, West). Moreover, the system may be configured in some embodiments to supplement the received camera location, camera orientation, and map data to identify any landmarks having known locations by analyzing the composite image of the view 900.
[0140] FIG. 10 is an illustration of an object tracking user interface that may be generated and presented to an emergency or crime center dispatcher according to an embodiment of the subject matter described herein.
[0141] In FIG. 10, a user interface 1000 displays different camera views obtained from different cameras when tracking the object 604. For example, in view 1002, the object 604 can be seen shooting the victim at the scene of the incident (from FIG. 9). In views 1004 and 1006, the object is identified as object 630 because these views correspond to the location of the object before the incident. In view 1004, the object 630 is shown walking toward the camera, while in view 1006, the object is shown walking from left to right across the image. However, in both the view 1004 and the view 1006, the object 630 is walking in the same direction (i.e., North). This can be determined by the system as described above to account for differences in the location or orientation of the different cameras. Finally, in view 1008, the object is identified as object 612 because this view corresponds to the location of the object after the incident. In view 1008, the object 612 is shown running (no longer walking) North away from the incident. The views 1002, 1004, 1006, and 1008 may be automatically generated by the system in response to receiving a trigger event that initiates object tracking.
[0142] An emergency or crime center dispatcher may evaluate the views 1002, 1004, 1006, and 1008 to confirm that the tracked objects 604, 612, and 630 arc the same object and that object tracking should continue to be performed. Additionally, the dispatcher may select one of the views 1002, 1004, 1006, or 1008 for enlargement as the selected view. Here, for example, the view 1008 is the selected view.
[0143] The user interface 126 may also include a map view 1010 showing the isochrones and the predicted and historical paths of the tracked object. The map view 1010 synthesizes the various location, object identification, predicted object path, and map data to provide the dispatcher with a visual representation of the object location during a time window before and after the incident. Different isochrone colors or shading may be used in the map view 1010 to indicate, for example, different time periods or modes of travel.
[0144] The user interface 126 may also include an information summary view 1012 showing various(non-video, non-map) details related to the incident. These details may include time and location information gleaned from the trigger event information provided by an E911 system or ShotSpottcr environmental sensor, as well as meta-attributes determined by the system based on an analysis of the video information retrieved from the cameras. Here, for example, the information summary view 1012 indicates that the incident occurred at 1 pm at 100 Main Street and that currently (5 minutes after the incident), the object identified by the system is traveling North on foot and wearing a red shirt. This information may be communicated to users of other client devices, such as mobile phones or laptops of police officers, to aid in apprehending the tracked object.
[0145] Selected camera view 1008 may include a live, current, real time, or near real time video feed. It is appreciated that selected camera view 1008 is not limited to a live view. It is also appreciated that selected camera view 1008 may be updated to include a new or different view of video information based on the camera selected by the user.
[0146] The user interface 1000 includes a plurality of camera icons representing different camera sets associated with tracking one or more objects. The system can obtain both historical video recordings and live video streams being captured by a camera at or near the location of the icon on map.
[0147] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the embodiments described. The terminology used herein was chosen to explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
[0148] Aspects of the subject matter described herein may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit,” “module” or “system.”
[0149] Aspects of the subject matter described herein may take the form of a computer program product embodied in a non-transitory computer readable medium having computer executable program code embodied thereon. A computer readable storage medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system,apparatus, or device. A computer readable medium may include, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD- ROM), an optical storage device, a magnetic storage device, or a suitable combination thereof.
[0150] The subject matter described herein may be implemented as software-as-a-service (SaaS) that is centrally hosted in the cloud and licensed on a subscription basis. One or more users may communicate with the system over a communications network, such as the Internet. Cloud storage may store or manage information using a public or private cloud where the physical storage spans multiple servers (sometimes in multiple locations). System services may be accessed through a colocated cloud computing service, a web service API, or by applications that utilize the API.
[0151] The subject matter described herein may be implemented machine learning (ML) or other artificial intelligence models. ML uses computer algorithms that can improve automatically through experience and using data. Machine learning algorithms build a model based on sample data (training data) to make predictions or decisions without being explicitly programmed to do so.
[0152] In certain embodiments, the system may perform some or all the functions using machine learning or artificial intelligence. The system may be operable for identifying one or more data sources and extracting data from the identified data sources. Machine learning-based software may load the data in an unstructured format and automatically determine relationships between the data. Machine learning-based software may identify relationships between data in an unstructured format, assemble the data into a structured format, evaluate the correctness of the identified relationships and assembled data, and / or provide machine learning functions to a user based on the extracted and loaded data, and / or evaluate the predictive performance of the machine learning functions (e.g., “learn” from the data). Machine learning-based software may assemble data into an organized format using one or more supervised or unsupervised learning techniques to identify relationships between data elements in an unstructured format.
[0153] Machine learning, as used herein, may be performed using one or more modules, computer executable program code, logic hardware, and / or other entities configured to learn from or train on input data, and to apply the learning or training to provide results or analysis for subsequent data. Machine learning-based software may include a model generator, a training data module, a modelprocessor, a model memory, and a communication device. In some embodiments, machine learningbased software may generate decision trees. Machine learning-based software may also calculate coefficients and hyper parameters of a decision tree based on the training data set. In yet other embodiments, machine learning-based software may use association rule mining, artificial neural networks, and / or deep learning algorithms to develop models. In some embodiments, machine learning-based software may utilize hardware optimized for machine learning functions, such as an FPGA or GPU.
[0154] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0155] The flowchart and block diagrams in the figures may illustrate and / or represent the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0156] Although the invention has been described and illustrated with a certain degree of particularity, it is understood that the present disclosure has been made only by way of example, and that numerous changes in the combination and arrangement of parts can be resorted to by those skilled in the art without departing from the spirit and scope of the invention, as hereinafter claimed.
Claims
WE CLAIM:
1. A computer- implemented method for video-based object tracking using isochrones comprising instructions stored on a non-transitory computer-readable storage medium and executed on a computing device provided with a hardware processor and a memory, the method comprising: detecting a trigger event; determining a first time frame and a first location based on the trigger event; obtaining first video information from a first video source, wherein the first video source has a view of the first location, and wherein the first video information is associated with the first time frame; identifying an object to be tracked in the first video information; determining a first speed and a first direction of the object; predicting a second set of video sources using an isochrone based on the first speed and the first direction of the object; and obtaining second video information from the second set of video sources, wherein the second video information is associated with a second time frame; and identifying the object in the second video information.
2. The method of claim 1, wherein the detecting the trigger event includes one of: receiving information from an environmental sensor; and receiving emergency call for service information.
3. The method of claim 2, wherein the environmental sensor is an audio sensor having a fixed location that is configured to detect gunshot sounds and provide an approximate time and location of a detected gunshot.
4. The method of claim 2, further comprising performing natural language processing (NLP) on the emergency call for service information using a large language model (LLM) configured to automatically generate a transcript of the emergency call for service and analyze the transcript for indications of the first time and the first location.
5. The method of claim 1, wherein identifying the object in the first video information includesdetermining a meta- attribute of the object using one or more artificial intelligence (Al) -based image processing techniques, wherein the meta-attribute is at least one of a size, a shape, and a color of the object.
6. The method of claim 1 , wherein determining the first speed and the first direction of the object includes comparing two or more frames in the first video information.
7. The method of claim 1 , wherein determining the first speed and the first direction of the object includes comparing two or more locations of the object relative to a known location of the first video source.
8. The method of claim 1, wherein the isochrone indicates a geographic area based on: the first location of the object; the first speed of the object; the first direction of the object; and a prediction time window.
9. The method of claim 8, wherein the prediction time window is associated with at least one of a period of time before the first time frame and a period of time after the first time frame.
10. The method of claim 1, wherein multiple objects are identified to be tracked in the first video information and wherein video sources are individually predicted for each of the multiple objects based on the speed and first direction individually determined for each of the multiple objects.
11. A system for video-based object tracking using isochrones, the system comprising: a processing hub operable to: detect a trigger event; determine a first time frame and a first location based on the trigger event; obtain first video information from a first video source, wherein the first video source has a view of the first location, and wherein the first video information is associated with the first time frame; identify an object to be tracked in the first video information; determining a first speed and a first direction of the object; predict a second set of video sources using an isochrone based on the first speed and the first direction of the object; andobtain second video information from the second set of video sources, wherein the second video information is associated with a second time frame; and identify the object in the second video information.
12. The system of claim 11, wherein the trigger event includes one of: information from an environmental sensor; and emergency call for service information.
13. The system of claim 12, wherein the information from the environmental sensor provides an approximate time and location of a detected gunshot.
14. The system of claim 12, wherein the processing hub is further operable to perform natural language processing (NLP) on the emergency call for service information using a large language model (LLM) configured to automatically generate a transcript of the emergency call for service and analyze the transcript for indications of the first time and the first location.
15. The system of claim 11, wherein the processing hub is further operable to determine a metaattribute of the object using one or more Al-based image processing techniques, wherein the meta-attribute is at least one of a size, a shape, and a color of the object.
16. The system of claim 11, wherein the processing hub is further operable to determine the first speed and the first direction of the object by comparing two or more frames in the first video information.
17. The system of claim 11, wherein the processing hub is further operable to determine the first speed and the first direction of the object by comparing two or more locations of the object relative to a known location of the first video source.
18. The system of claim 11, wherein the isochrone indicates a geographic area based on: the first location of the object; the first speed of the object; the first direction of the object; and a prediction time window.
19. The system of claim 18, wherein the prediction time window is associated with at least oneof a period of time before the first time frame and a period of time after the first time frame.
20. The system of claim 11, wherein the processing hub is further operable to identify multiple objects to be tracked in the first video information.