Multi-sensor body tracking of monitored environments for real-time and near-real-time systems and applications
Through the sub-second processing and hierarchical cluster matching algorithm of the Real-Time Location System (RTLS), the problems of insufficient real-time and accuracy in multi-camera tracking are solved, and efficient multi-subject tracking is achieved, which is suitable for environments such as warehouses, factories, and retail places.
Patent Information
- Application Number
- CN202510292102.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-14
- Filing Date
- 2025-03-12
- Publication Date
- 2025-09-16
AI Technical Summary
Existing multi-camera tracking technologies have difficulty achieving accurate multi-agent tracking in real-time and high-density environments, especially when processing a large number of behavior detections and trajectory links that exceed the capabilities of computing resources, resulting in insufficient real-time performance and accuracy.
A real-time positioning system (RTLS) is used to define initialization anchor points, utilize sub-second processing capabilities, perform behavior detection and trajectory continuity analysis based on machine learning models, and combine hierarchical clustering and matching algorithms to achieve real-time data processing and subject tracking of multiple sensors.
It achieves real-time tracking of multiple subjects in high-density environments with sub-second response time, provides smoother ID assignment and more continuous trajectory tracking, improves spatiotemporal scene understanding, and supports efficient multi-camera multi-subject tracking.
Smart Images

Figure CN120655685A_ABST
Abstract
Description
Background Art
[0001] Multi-subject multi-camera (MTMC) tracking is a computer vision-based technique that simultaneously monitors and tracks the motion of numerous objects (subjects) across multiple camera views—taking input from video feeds captured by multiple (potentially non-overlapping) cameras and applying algorithms and / or machine learning techniques to analyze the video streams to track and identify subjects of interest. MTMC tracking can be used in applications such as security and surveillance, vehicle traffic monitoring, monitoring transportation, activities in factories and warehouses, for retail analytics to monitor customer behavior in retail stores, and / or for crowd management / public safety at events, gatherings, or public venues. Summary of the Invention
[0002] Embodiments of the present disclosure relate to multi-sensor subject tracking of monitored environments for real-time and near-real-time systems and applications. The systems and methods disclosed herein can be used to track multiple subjects within a monitored area across the field of view of multiple image sensors to generate tracking solutions for the multiple subjects within a second of capturing image data.
[0003] Compared to existing MTMC tracking technology, the systems and methods described herein provide a real-time location system (RTLS) that can perform multi-subject tracking using real-time (e.g., live) or near-live streaming data from multiple sensors viewing a monitored area. The monitored area can include any area in which a subject of interest may be traveling, such as, but not limited to, warehouses, factories, retail locations, office buildings, security facilities, arenas, public transportation stations, parks, public places, etc.
[0004] RTLS subject tracking can be based on defining a set of initialization anchor points, where a separate anchor point is initialized for each individual subject identified from the live streaming data. The live streaming data can be processed in micro-batches (e.g., a synchronized set of streaming data from multiple sensors across sub-second durations) so that the rendered subject tracking data is up to date (e.g., current within one second) relative to the actual position of the tracked subject. RTLS can receive a set of synchronized optical image stream data that includes, for example, video image feeds from multiple optical image sensors. The video image feeds can be synchronized such that the streaming data includes individual image feeds from different optical image sensors that capture image data simultaneously. The video image feeds can include timestamps that can be used to align the simultaneously captured image data from multiple feeds.
[0005] In some embodiments, anchor point and associated behavior state can be initialized to represent the subject that can be detected from stream data.Each initialized anchor point can be maintained in RTLS state.The anchor point of tracked subject can be used to track or determine the behavior data of tracked subject in the field of view of multiple optical image sensors.In stable state operation, anchor point can be tracked based on trajectory continuity analysis, and this trajectory continuity analysis associates the previous behavior state of the live stream data from previous micro batch with the representation (such as, behavior embedding) derived from the live stream data of current micro batch, to identify the active anchor point that can propagate (such as, continue) forward from the live stream data of previous micro batch.The remaining representation that cannot be tracked based on trajectory continuity analysis may represent the subject that is not represented in the live stream data of the most recent previous micro batch.In some cases, one or more of these representations can be associated with a dormant anchor point, such as, the anchor point of the subject previously initialized, which has not yet appeared in the live stream data of the most recent micro batch. To identify when the remaining representations are associated with a dormant anchor, the RTLS may perform a hierarchical clustering process to group the remaining behaviors into a set of one or more clusters based on similarity of the representations.
[0006] Once clustering is performed, the RTLS may perform a matching algorithm to determine whether one or more remaining representations associated with the cluster match (e.g., are similar) to a representation associated with a dormant anchor. If the cluster is determined to match the dormant anchor, the RTLS may perform behavioral state management, which reclassifies the dormant anchor as an active anchor, and the anchor's behavioral state may be updated based on the current representations that formed the cluster. If the cluster of remaining representations is determined not to match the dormant anchor, the cluster may represent behavior associated with a subject not previously observed by the RTLS, or may have been previously associated with a dormant anchor that has expired (e.g., deleted based on age and / or other data staleness criteria). In this case, the RTLS may initialize a new active anchor associated with the cluster, initialize a new global ID assigned to the new anchor, and initialize the behavioral state of the new anchor based on the current representations that formed the cluster.
[0007] The subject tracking data derived from the behavioral states of the active anchors can be used as input, for example, by one or more subject assessment systems. For example, the subject tracking data can be used by a rendering system to display a comprehensive view of the monitored area and visually render the position and trajectory tracking data of each subject in the area and / or other applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The present system and method for multi-sensor subject tracking of monitored environments for real-time and near real-time systems and applications is described in detail below with reference to the accompanying drawings, wherein:
[0009] Figure 1 is a data flow diagram illustrating a real-time location system according to some embodiments of the present disclosure;
[0010] Figure 2 is a data flow diagram illustrating a behavior extraction process according to some embodiments of the present disclosure;
[0011] Figure 3 is a data flow diagram illustrating a real-time location system state management process according to some embodiments of the present disclosure;
[0012] Figure 4A and Figure 4B is a diagram illustrating an example user interface display generated by a real-time locating system according to some embodiments of the present disclosure;
[0013] Figure 5 is a flow chart illustrating a method for a real-time location system according to some embodiments of the present disclosure;
[0014] Figure 6 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0015] Figure 7 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] The disclosed systems and methods relate to multi-sensor subject tracking of monitored environments for real-time and near real-time systems and applications.
[0017] Historically, camera-based tracking techniques have involved single object tracking (SOT) or multiple object tracking (MOT) based on image data from a single camera. More recently, multi-camera-based tracking has leveraged the development of multi-camera networks to apply techniques for providing comprehensive monitoring of an area from multiple viewpoints. However, multi-camera tracking introduces complexities in tracking subjects across multiple camera views—including performing accurate synchronization and calibration between cameras, addressing changes in the subject's appearance that may vary with viewpoint, and fusing data from different video streams to unambiguously define the subject and associated track.
[0018] While object re-identification (Re-ID) techniques and other advanced matching algorithms have been developed that can take the outputs of object detection algorithms and match them to compute individual trajectories, these techniques face difficulties in scalability and supporting real-time multi-subject tracking (MTMC) for live video streams. For example, if a monitored area has five cameras tracking ten subjects moving around the area, then up to 50 different behavior detections may need to be processed and individually linked to the subjects appearing in the video stream, and their respective trajectories may need to be tracked over time. If a monitored area has five cameras tracking 100 subjects moving around the area, then up to 500 different behavior detections may need to be processed and individually linked to the subjects appearing in the video stream, and their respective trajectories may need to be tracked over time. For monitored areas with densely populated objects (e.g., a grocery store during peak shopping hours), the number of different behavior detections that need to be processed simultaneously for different subjects (including their respective trajectory data) can quickly exceed the capabilities of the available computing resources used by various algorithms to support real-time multi-subject tracking.
[0019] Compared to existing MTMC tracking technology, the systems and methods described herein provide a real-time location system (RTLS) that can perform multi-subject tracking using real-time (e.g., live) or near-live streaming data from multiple sensors viewing a monitored area. The monitored area can include any area in which a subject of interest may be traveling, such as, but not limited to, warehouses, factories, retail locations, office buildings, security facilities, arenas, public transportation stations, parks, public places, etc.
[0020] As described herein, RTLS subject tracking can be based on defining a set of initialization anchor points, wherein for each individual subject identified from live streaming data, a respective anchor point is initialized. Live streaming data can be processed in micro-batches (e.g., a set of synchronized streaming data from multiple sensors spanning sub-second durations) so that the rendered subject tracking data is up-to-date (e.g., up-to-date within one second) with respect to the actual position of the tracked subject. In some embodiments, unique advantages of the RTLS embodiments described herein include, for example, but not limited to, real-time or near real-time processing capabilities with sub-second response times, smoother and more continuous ID assignment to tracked subjects compared to existing MTMC tracking techniques, and / or providing enhanced spatiotemporal scene understanding.
[0021] As described herein, an RTLS can receive a set of synchronized optical image stream data comprising, for example, video image feeds from multiple optical image sensors. The video image feeds can be synchronized such that the stream data comprises individual image feeds from different optical image sensors that simultaneously capture image data. In some embodiments, the video image feeds can include a timestamp that can be used to align the simultaneously captured image data from the multiple feeds. The RTLS can process the synchronized optical image stream data into continuous micro-batches (e.g., wherein each micro-batch defines a different sub-second time frame of the synchronized optical image stream data). In some embodiments, the synchronized optical image stream data may include live streaming feeds from multiple optical image sensors. In some embodiments, the synchronized optical image stream data may include previously recorded batches of live streaming feeds from the multiple optical image sensors.
[0022] In order to initialize RTLS, the live streaming data of the initial micro batch can be processed to generate a set of initialized anchor points. Each individual anchor point initialized by RTLS represents a subject that can be detected from the live streaming data. In some embodiments, each initialized anchor point remains in the RTLS state. As described below, the initialization process may include processing the stream data to calculate the representation of the behavior based on the features of the detected subject. In one or more embodiments, the representation of the behavior is implemented as behavior embedding, and the initialization process includes clustering of the behavior embedding to associate the grouping (e.g., clustering) of the behavior embedding with different subjects, and assigning an anchor point to each cluster. When the anchor point is initialized, the behavior state of the anchor point is also defined and initialized based on the behavior data (e.g., appearance data and spatiotemporal data) of the corresponding different subjects represented by the behavior embedding contained in the cluster assigned by it. The anchor point of the tracked subject can be used to track or determine the behavior data of the tracked subject within the field of view of multiple sensors that generate live streaming data. The term "active anchor point" as used herein refers to a previously initialized anchor point that represents a subject that can be observed from the live streaming data of the current micro-batch, which has been propagated from the live streaming data of the most recent previous micro-batch. In contrast, a "dormant anchor point" can refer to a previously initialized anchor point associated with a subject that can previously be detected from the live streaming data, but the subject does not seem to appear in the live streaming data of the last previous micro-batch. In some embodiments, the RTLS state can associate the anchor point with its corresponding behavior state and maintain historical behavior data associated with each initialized anchor point. When a new anchor point is initialized, and / or when a new behavior embedding successfully matches the previous behavior state of the previously initialized anchor point, the RTLS state can be updated. In some embodiments, the historical data from the RTLS state can be used to increase the reliability of behavior matching for time spans exceeding the time frame of the current micro-batch.
[0023] To generate a set of initialized anchor points for initializing the RTLS, computer-based perception can be used to process feeds from various cameras to perform behavior detection, perform single-camera tracking (e.g., subject behavior localization), and generate a behavior embedding (e.g., computing a vector representing the subject) for each detected behavior. The behavior embedding can include encoded behavior data that includes appearance data and / or spatiotemporal data that captures features of the tracked subject. The behavior embedding can be calculated based on the detected behavior, for example, using a machine learning model trained to recognize features of the expected subject (e.g., a model trained to recognize and extract features associated with features of people, vehicles, machines, animals, and / or other subjects of interest). In some embodiments, the RTLS can implement multiple parallel live stream data processing paths to process feeds from multiple cameras to generate behavior embeddings. For example, in some embodiments, the RTLS can implement a set of parallel processing paths (e.g., using multiple processing threads and / or multiple core processors) to calculate behavior embeddings, where each processing path generates a behavior embedding based on the live stream feed of the corresponding camera. The behavior embedding associated with an active anchor point can be used to define and / or update the behavior state of the active anchor point, as described below.
[0024] Based on a set of behavior embeddings, RTLS can map the spatiotemporal data from each behavior embedding to a global image coordinate system (e.g., based on external camera calibration parameters associated with each individual camera) and perform initial clustering based on the similarity of the behavior embeddings to group the behavior embeddings into clusters. For example, a subject (e.g., a person) in a monitored area can be represented as a detected subject in a micro-batch of live streaming data and as a first behavior embedding derived from the feed of a first camera. Similarly, a second behavior embedding can be derived from the feed of a second camera for the detected subject. For each camera, a set of external calibration parameters (e.g., rotation and translation transformations) can be used to map the position information of the behavior extracted from the two-dimensional (2D) image data captured by the camera to a global coordinate system associated with the monitored area. Since in this example the first behavior embedding and the second behavior embedding are both representations of the same subject (person) within the same time frame, the first behavior embedding and the second behavior embedding should share very similar behavior data (e.g., appearance data and spatiotemporal data). Therefore, these behavior embeddings will form clusters that can be uniquely associated with different subjects (e.g., detected people). In addition, for each additional camera that observes the subject within the time frame, the behavioral embeddings derived from these camera feeds should also be clustered with the behavioral embeddings derived from the first camera and the second camera. Other different clusters can be formed in the same manner based on the behavioral embeddings associated with other subjects represented in the micro-batch of live streaming data. In some embodiments, when the RTLS performs clustering, hierarchical clustering can be performed to group the behavioral embeddings into clusters based on the similarity of the behavioral embeddings. In some embodiments, the RTLS can apply one or more hierarchical clustering algorithms, such as, but not limited to, the Balanced Iterative Reduction and Clustering using Hierarchy (BIRCH) algorithm, the Ordering Points to Identify Cluster Structures (OPTICS) algorithm, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN), the Hierarchical DBSCAN (HDBSCAN*), and / or other clustering algorithms.
[0025] Each cluster that explicitly represents a subject can be used to initialize an active anchor point, and a unique global identifier (global ID) can be assigned to the anchor point. In addition, for each anchor point, the RTLS can initialize and maintain a behavioral state that can be updated using subsequent micro-batch live streaming data (e.g., as long as the anchor point continues to be an active anchor point). In some embodiments, the global ID can be an anonymous identifier and / or can be associated with a more personal identifier, such as the subject's name, employee number, customer number, account number, student number, and / or other identifier associated with the tracked subject.
[0026] After initializing the RTLS and the first set of anchor points, for subsequent micro-batches of live streaming data (e.g., subsequent time frames), active anchor points can be identified based on the streaming data and their respective corresponding behavior states can be updated. For example, computer-based perception can be used to process the feeds from each camera to perform behavior detection, single-camera tracking, and create behavior embeddings for each detected behavior. The behavior embeddings can be calculated by a machine learning model based on the behaviors detected from the live streaming data and the spatiotemporal data from each behavior embedding mapped to the global image coordinate system. Given the short sub-second time intervals that occur between each consecutive micro-batch of live streaming data, the RTLS can use trajectory continuity analysis to associate the behavior data from the previous behavior state with the behavior embedding derived from the live streaming data of the current micro-batch to identify active anchor points that can be propagated forward (e.g., continued) from the previous micro-batch of live streaming data. For example, the previous position and trajectory data from the previous behavior state of the anchor point can be used to accurately predict the current position and trajectory expected in the time elapsed between iterations. Those behavior embeddings derived from the current micro-batch with spatiotemporal data corresponding to the predicted current position and trajectory (e.g., within a threshold distance) can be used to propagate the anchor and update the behavior state data for the activity anchor. In some embodiments, trajectory continuity analysis can determine the distance (e.g., Euclidean distance) between the predicted current position and trajectory and the position and trajectory represented by one or more behavior embeddings to associate the behavior embedding with the activity anchor. The similarity of the appearance data from the behavior embedding can be used to further verify the association between the behavior embedding and the activity anchor.
[0027] Advantageously, an active anchor can be propagated forward from a previous micro-batch without performing clustering of the behavior embedding to compute an updated behavior state for the anchor. Instead, behavior state management can be performed using current behavior data (e.g., appearance and / or spatiotemporal data) extracted from the current behavior embedding and the previous behavior state of the anchor to update the previous behavior state based on the current behavior data. During propagation, incremental changes in the appearance, position, and / or trajectory of the tracked subject can be incorporated into the behavior state data maintained by the RTLS for the anchor. In some embodiments, the updated behavior state data can include a statistically weighted fusion of previous and current appearance and spatiotemporal data. In some embodiments, the behavior state of the active anchor can represent the subject for a moving time window. For example, the behavior state of the active anchor can represent appearance and spatiotemporal data based on the behavior embedding obtained within the past "t" seconds (e.g., within the past 5 or 10 seconds) so that the trajectory and rendered tracking data can be calculated based on the subject's most recent motion period.
[0028] In some embodiments, a set of current behavior embeddings from current micro-batch may include one or more behavior embeddings that do not conform to the trajectory prediction continuity from previous micro-batch. These remaining behavior embeddings can represent a subject that does not exist (e.g., is not represented) in the live streaming data of the most recent previous micro-batch. In some cases, one or more of these behavior embeddings can be associated with a dormant anchor point, that is, a previously initialized anchor point that does not appear in the live streaming data of the most recent micro-batch. For example, a subject with an initialized anchor point can move to a position that is temporarily unable to be observed by a set of multi-camera views available through RTLS (e.g., due to leaving the monitored area and / or due to occlusion), and after a period of time, the subject may return to a position where they can again be observed from a set of multi-camera views.
[0029] To determine whether the remaining behavior embeddings are associated with a dormant anchor, the RTLS may perform an assignment process that includes a hierarchical clustering process followed by a matching algorithm. The hierarchical clustering process may evaluate the remaining behaviors (e.g., behavior embeddings that could not be propagated as active anchors based on trajectory continuity analysis) to group the remaining behaviors into a set of one or more clusters based on the similarity between their behavior embeddings. Different clusters may be formed from the remaining behavior embeddings that are not associated with the current active anchor. Once clustering is performed, the RTLS may perform a matching algorithm to determine whether a cluster in the remaining behavior embeddings matches (e.g., is similar to) a behavior embedding associated with a dormant anchor. If a cluster is determined to match a dormant anchor, the RTLS may perform behavior state management to reclassify the dormant anchor as an active anchor and update the anchor's behavior state based on the current behavior embedding that formed the cluster. If a cluster of the remaining behavior embeddings cannot be matched with a dormant anchor, the cluster may represent behaviors associated with a subject not previously observed by the RTLS, or behaviors that may have been previously associated with a dormant anchor that may have expired (e.g., behaviors that were deleted based on age and / or other data staleness criteria). In this case, the RTLS can initialize a new active anchor associated with the cluster, initialize a new global ID assigned to the new anchor, and initialize the behavior state of the new anchor based on the current behavior embeddings that form the cluster. Through the matching algorithm, each cluster from the remaining behaviors not associated with the dormant anchor can similarly represent a new subject and be used to initialize the new anchor. In some embodiments, when the RTLS determines that the number of remaining unmatched behaviors after executing the matching algorithm exceeds a threshold (e.g., the ratio of unmatched behavior embeddings to matched behavior embeddings), the RTLS can reinitialize (e.g., perform system initialization as described above) to generate a new set of initialized anchors.
[0030] In some embodiments, the RTLS may include a machine learning model trained to perform behavior detection and tracking and derive a behavior embedding. The machine learning model may include, for example, a re-identification (Re-ID) embedding model that encodes the appearance of each tracked subject. Optical image data comprising an image of the subject may be used to generate an enclosing shape (e.g., a box) for the subject, and the optical image data may be cropped to the enclosing shape. The Re-ID embedding model may output a behavior embedding for the cropped image, the behavior embedding comprising an embedding vector. The behavior embedding may encode one or both of the appearance data and spatiotemporal data (e.g., position and / or trajectory) that characterizes the subject. In some embodiments, the machine learning model may include a ResNet50 backbone architecture or other deep neural network architecture.
[0031] In some embodiments, when the RTLS performs clustering of behavioral embeddings, the RTLS performs a two-step hierarchical clustering process. As described above, clustering can be performed on behavioral embeddings that have not been found to be associated with activity anchors based on the continuity of the trajectory data. In the first step of the clustering process, the set of behavioral embeddings is clustered based on the similarity of the behavioral embeddings. This first cluster is used to predict a number (e.g., a predicted number of subjects) representing how many clusters are associated with the actual subject represented by the behavioral embedding. The first step of the clustering process can be based on applying a fine-tuned clustering threshold parameter to the behavioral embeddings. That is, the clustering threshold parameter can define the degree to which a set of behavioral embeddings need to be tightly clustered (e.g., how close the distance is) in order for the set of behavioral embeddings to be considered a cluster. In some embodiments, the clustering threshold parameter may include a cluster quality threshold (QT) that specifies a threshold distance between cluster members and / or a minimum number of behavioral embeddings that are within a threshold distance of each other for a set to be considered a cluster. In the second step of the clustering process, the set of behavioral embeddings is again used and clustering is performed based on the similarity of their behavioral embeddings, and the clustering is further constrained based on the number of agents predicted from the first step of the clustering process. That is, the second step of the clustering process is constrained to cluster the behavioral embeddings into a number of clusters corresponding to the number of agents predicted to be present in the mini-batch of live streaming data. Because each cluster resulting from the second step of the clustering process is expected to contain behavioral embeddings representing the same distinct agent, the appearance data and / or spatiotemporal data encoded in the behavioral embeddings of each cluster should be highly similar—and easily distinguishable from the appearance data and spatiotemporal data provided by the behavioral embeddings of other clusters representing other distinct agents in the monitored area. Advantageously, using two-step clustering, the process of initializing new anchors from the resulting clusters and / or reclassifying dormant anchors as active anchors can benefit from the high confidence that the behavioral embeddings of each cluster correspond to a distinct agent that is distinguishable from other agents.
[0032] In some embodiments, when RTLS performs a two-step hierarchical clustering process, the process can use ReID appearance data from behavioral embeddings without integrating spatiotemporal data. Spatiotemporal information can then be incorporated into the reassignment phase through a matching process, where both spatiotemporal and ReID distances can be normalized and weighted for decision making. In each iteration of the matching process, the clustered appearance data and spatiotemporal data (e.g., location) can be updated and optimized.
[0033] As described above, in some embodiments, the clustering process may be followed by a matching process to determine whether the behavior embedding associated with the cluster matches (e.g., is similar to) the behavior embedding associated with a previously initialized active anchor that is subsequently classified as a dormant anchor. In some embodiments, the matching may be performed using an iterative matching combinatorial optimization algorithm (e.g., a matching algorithm that solves the assignment problem by matching agents to tasks). In some embodiments, the matching process may apply a Hungarian matching algorithm (which may be referred to as a Kuhn-Munkres algorithm or a Munkres assignment algorithm) to assign the clusters generated by the clustering process to the current dormant anchor by matching the cluster's behavior embedding with the previously initialized behavior state associated with the dormant anchor (e.g., based on similarity). In some embodiments, the iterative refinement may include sufficient iterations to achieve optimal matching accuracy saturation (e.g., approximately 10 iterations in the case of the Hungarian matching algorithm). If a cluster is assigned (matched) to the previous behavior state of the dormant anchor, the RTLS may reclassify the dormant anchor as an active anchor and apply behavior state management to update the behavior state of the now active anchor based on the appearance data and spatiotemporal data provided by the current behavior embedding from the cluster. If the matching algorithm determines that a cluster cannot be assigned (matched) to a dormant anchor, a new active anchor can be initialized for that cluster, whose initial behavior state is defined based on the appearance data and spatiotemporal data provided by the current behavior embedding from the cluster.
[0034] As described above, RTLS can apply state management to update the behavior state of the active anchor point based on the appearance data and / or spatiotemporal data provided by the current behavior embedding associated with the active anchor point. For example, in some embodiments, the RTLS can determine the current behavior embedding associated with the active anchor point and the previous behavior embedding represented by the previous behavior state of the active anchor point from the live streaming data of the current micro-batch, and update the behavior state of the active anchor point based on the current behavior embedding. In this way, the behavior state of the active anchor point can be updated based on the current appearance data of the subject and / or based on the current position and / or trajectory data of the subject. In some embodiments, the update can include fusing or splicing the behavior represented by the current and previous / historical behavior embeddings of the anchor point. Therefore, the behavior state of the anchor point can include both current and historical behavior data. In some embodiments, the appearance data and / or spatiotemporal data provided by the current behavior embedding can have different weights than the appearance data and / or spatiotemporal data from the behavior state representing the previous behavior embedding. For example, a set of historical appearance, position and / or trajectory data of the active anchor point can be considered more reliable than the most recent appearance, position and / or trajectory data. In some embodiments, when the behavioral state is updated, historical appearance, location, and / or trajectory data may be given a higher weight (e.g., 0.9) while current appearance, location, and / or trajectory data may be given a relatively lower weight (e.g., 0.1) because the data is combined to update the behavioral state of the active anchor. Conversely, when the RTLS is updating the behavioral state of a dormant anchor that has been reclassified as an active anchor, the current appearance, location, and / or trajectory data may be more accurate than the historical appearance, location, and / or trajectory data, depending on how long the anchor has been a dormant anchor. In this case, the historical appearance, location, and / or trajectory data may be given a lower weight (e.g., 0.1) while the historical appearance, location, and / or trajectory data may be given a relatively higher weight (e.g., 0.9) because the data is combined to update the behavioral state of the anchor. In some embodiments, the behavioral state may retain historical appearance data and / or spatiotemporal data for a predetermined duration and purge historical data from the behavioral state based on age and / or other data staleness criteria. For example, in some embodiments, the behavior state may maintain the previous 5 or 10 seconds of historical appearance, position, and / or trajectory data for the active anchor.
[0035] In some embodiments, the RTLS may use a configurable state retention time to determine the duration that a dormant anchor may remain eligible for reclassification back to an active anchor, rather than the RTLS initializing a new anchor. If a dormant anchor remains dormant for longer than the state retention time, it may expire. For example, given a configurable state retention time of 10 minutes, the RTLS will trigger the initialization of a new anchor for a tracked subject that reappears after being dormant for 15 minutes. After an anchor has been dormant for longer than the configurable state retention time, the RTLS may delete the expired dormant anchor (e.g., making it impossible to reactivate as an active anchor). In some embodiments, historical behavior data and / or behavior state associated with the deleted anchor may be saved in a database for later analysis (e.g., using the QBE functionality described below).
[0036] For example, one or more subject assessment systems can use subject tracking data derived from the behavioral states of active anchors as input. A rendering system can use the subject tracking data to display a comprehensive view of the monitored area and visually render the position and trajectory tracking data of each subject in the area, wherein the rendering system can benefit from real-time spatiotemporal knowledge of the position and motion of the subjects of interest within the monitored area. This can include switching between real and / or virtual camera views when tracking subjects passing through the area to locate the current position of a particular subject within the area (e.g., based on a global ID or other query).
[0037] In some embodiments, the RTLS can use the behavioral states of active anchors to render a user interface (UI) on a human-machine interface display. The UI can present a comprehensive computer vision-based view of the monitored area and / or a view of a selected portion of the monitored area. For example, the RTLS can use the position and / or trajectory data of each active anchor to display the predicted position and / or tracking information associated with each subject. In some embodiments, the RTLS can generate a display of a computer vision-based top-down (bird's-eye) view of the monitored area—displaying the position and / or tracking of each tracked subject. The RTLS can generate a display of a computer vision-based environment corresponding to the viewpoints of one or more real image sensors providing real-time live streaming data, and / or a display of the viewpoints of one or more virtual camera views instantiated in the computer vision-based environment. In some embodiments, a subject can be selected through the user interface and its movement tracked through the computer vision-based environment based on the behavioral states of the subject's associated anchors. A global ID (or other corresponding identifier) can be displayed for one or more subjects tracked by the RTLS in the monitored area.
[0038] In some embodiments, the RTLS may generate an output representing data representing the behavioral state of one or more subjects being tracked in the monitored area. The output may be used as input to other systems. For example, in some embodiments, a navigation control system of a mobile machine (e.g., an autonomous mobile robot (AMR) and / or a self-driving machine) may dynamically guide the mobile machine away from a path that may be congested due to the presence of tracked subjects (e.g., people, obstacles) on the path. In some embodiments, the RTLS implements an RTLS application programming interface (API) to facilitate AMR integration for communicating tracked subject motion to the AMR control system, and / or displaying tracked subjects associated with the AMR on a user interface (UI), as described below. In some embodiments, the security system may use the output from the RTLS to track the person of interest in real time as the person of interest passes through the monitored area.
[0039] In some embodiments, RTLS can integrate a query by example (QBE) function that supports long-term QBE queries spanning a selected time period, leveraging global IDs and behavior embeddings generated by the RTLS and stored in a database (e.g., a Milvus database, a vector database, and / or other vector database management systems (VDBMS)). The QBE function can operate based on Representational State Transfer (REST) application programming interface (API) inputs, which may include (but are not limited to) object IDs, sensor IDs, timestamps, and optional parameters (e.g., time range, match score threshold, and / or top K matches) to retrieve similar behaviors. The QBE function can normalize the behavior embeddings before searching the database for similar behaviors, and matching behaviors can have IDs used by the REST API to retrieve behavior metadata using a search engine (e.g., Elasticsearch). This search capability of the QBE function can provide accurate and efficient retrieval of relevant tracked subject behavior patterns over a longer period of time.
[0040] Because RTLS initializes an active anchor for each subject of interest and updates the behavioral state of the active anchor on a micro-batch-by-micro-batch basis based on tracking trajectory continuity, the execution of complex clustering and / or matching algorithms can be limited to cases where the live streaming data of a micro-batch contains behavioral embeddings that have not yet been associated with an active anchor, and / or the computation is limited to applying the clustering and / or matching algorithms only to those embeddings that have not yet been associated with an active anchor. Therefore, RTLS can process each micro-batch of live streaming data quickly enough to support real-time multi-camera multi-subject tracking applications in high-density environments, for example, where the position and trajectory tracking data of multiple subjects can be simultaneously computed and / or displayed within a short period of time (e.g., within one second or less) of the captured micro-batch of live streaming data.
[0041] Figure 1 is an example data flow diagram for a real-time locating system (RTLS) 100 according to some embodiments of the present disclosure. It will be understood that this arrangement and other arrangements described herein are presented as examples only. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by one or more processors (e.g., processing units, processing circuits) executing instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may be implemented using Figure 6 The example computing device 600 and / or Figure 7 The example data center 700 may be implemented with similar components, features, and / or functions.
[0042] like Figure 1 As shown, the RTLS 100 can receive image data 104 from a plurality of optical image sensors 102 (e.g., cameras) to track the positions and trajectories of a plurality of subjects within a monitored area based on initializing anchor points assigned to respective tracked subjects and updating the behavioral states of these anchor points based on behavioral data extracted from the image data 104. In some embodiments, the image data 104 includes synchronized optical image stream data based on real-time streaming feeds from the plurality of optical image sensors 102. In some embodiments, the image data 104 includes synchronized optical image stream data based on batches of previously recorded live streaming feeds from the plurality of optical image sensors 102.
[0043] The image data 104 may include individual video image feeds that are synchronized such that the streaming data includes individual image feeds from different optical image sensors 102 that captured the image data simultaneously. The video image feeds may include timestamps that may be used to align the simultaneously captured image data from the multiple feeds. The RTLS 100 may process the image data 104 into successive micro-batches, e.g., where each micro-batch defines a different sub-second time frame of the image data 104. In some embodiments, the optical image sensor may include, for example, one or more RGB cameras, one or more IR cameras, one or more RGB-IR cameras, and / or one or more cameras, as discussed below with respect to the computing device 600.
[0044] The image data 104 can be processed using behavior extraction 110 to extract behavior embeddings 120 that represent individual behaviors of tracked subjects represented in the image data 104. That is, each behavior embedding 120 can include a representation of a tracked subject that appears in a single feed of image data 104 associated with a different one of the optical image sensors 102. The representation provided by the behavior embedding 120 can include behavior data that includes appearance data (e.g., characterizing the appearance of the tracked subject) and / or spatiotemporal data (e.g., characterizing the position and / or trajectory of the tracked subject relative to the monitored area). As described herein, each optical image sensor 102 can capture image data 104 based on its particular view of the monitored area, which depends at least in part on extrinsic camera calibration parameters 112 (e.g., one or more rotational-translational transformations) associated with each individual optical image sensor 102. Thus, the camera calibration parameters 112 can be used to map the spatiotemporal data represented by the behavior embedding 120 to a global coordinate system associated with the monitored area, thereby facilitating associating the spatiotemporal data from the behavior embedding 120 for clustering, matching, trajectory tracking, and / or other purposes.
[0045] Figure 2 is a data flow diagram illustrating the operation of behavior extraction 110 for the process of extracting behavior embedding 120 from image data 104. Figure 2 As shown, behavior extraction 110 can implement multiple parallel data processing paths 228, each of which can independently process a feed of image data 104 associated with one of the optical image sensors 102. Within each path 228, behavior extraction 110 can perform behavior detection 222, behavior tracking 224, and / or behavior encoding 226. Behavior detection 222 can be used to identify and extract features corresponding to the behavior of tracked subjects (e.g., people, vehicles, machines, animals, and / or other subjects of interest) observed in each processing path 228. In some embodiments, behavior detection 222 can use the optical image data 104 containing subject images to generate a bounding shape (e.g., a box) for one or more subjects and generate subject-specific optical image data cropped to the bounding shape for each subject. Behavior tracking 224 can be used to calculate spatiotemporal features associated with each detected behavior associated with a tracked subject, such as position and trajectory (speed and direction of movement) relative to the local image frame of the sensor capturing the image data. The behavior encoding 226 can generate an encoding of the appearance and spatiotemporal data associated with each behavior detected from the image data, which is used to generate the behavior embedding 120. Figure 2As shown, since each processing path 228 is associated with a different optical image sensor 102 that can capture image data 104 based on its specific view of the monitored area, in some embodiments, the behavior extraction 110 can also include a behavior mapping 230 that applies the camera calibration parameters 112 to map the spatiotemporal data associated with each behavior to a global coordinate system associated with the monitored area. The resulting output from the behavior extraction 110 can include a plurality of behavior embeddings 120, wherein each individual behavior embedding 240 can include behavior data (e.g., appearance data 242 and / or spatiotemporal data 244) representing the tracked subject as the tracked subject appears in a separate feed of image data 104 associated with a different one of the optical image sensors 102.
[0046] In some embodiments, behavior extraction 110 may use a behavior encoding model 220 (e.g., a machine learning model) to perform one or more of behavior detection 222, behavior tracking 224, and / or behavior encoding 226. The behavior encoding model 220 may include, for example, a re-identification (Re-ID) embedding model that encodes the appearance of each tracked subject appearing in the image data 104 carried by the processing path 228. In some embodiments, based on the optical image data 104 including an image of the tracked subject, the behavior encoding model 220 may generate a bounding shape (e.g., a box) around the detected behavior and crop the optical image data to the bounding shape. For example, the image of the subject's detected behavior may be cropped and the cropped image may be resized, for example, to a 256x128 pixel image. The behavior encoding model 220 may output a behavior embedding 240 corresponding to the cropped image. In some embodiments, the behavior embedding 240 may be constructed in the form of an embedding vector. The embedding vector may include, for example, a vector with a configurable size, ranging from 1 to 2048 elements. In some embodiments, the embedding vector may have a default size of 256. The behavior embedding 240 calculated by the behavior encoding model 220 can encode one or both of the appearance data 242 and the spatiotemporal data 244 that characterize the tracked subject. In some embodiments, the behavior encoding model 220 may include a ResNet50 backbone architecture or other deep neural network (DNN) architecture. The behavior encoding model 220 can be trained with a combination of triplet loss, center loss, and ID loss (e.g., cross entropy loss). Re-ID features can be used to calculate the triplet loss, thereby minimizing the embedding distance of behaviors associated with the same subject (positive samples) while maximizing the distance of other behaviors associated with different subjects (negative samples).
[0047] Back to Figure 1 ,like Figure 1As shown, the RTLS 100 may include an RTLS state management 130 that generates and maintains anchor points (and their corresponding behavioral states) for tracking subjects observable within a monitored area based on behavioral embeddings 120 generated from a sequence of image data 104 over a series of time frames. More specifically, the RTLS state management 130 assigns clusters of behavioral embeddings 120 to anchor points, where each anchor point represents a different tracked subject within the monitored area. Each anchor point has a corresponding behavioral state, where behavioral data for different tracked subjects can be maintained, associated with a global ID, and tracked over time. In some embodiments, the RTLS state management 130 may include an RTLS state 132, which includes one or more anchor points 136 and corresponding previous behavioral states 134 associated with the tracked subject. As described above, the anchor points 136 maintained by the RTLS state 132 may include active anchor points 137 and dormant anchor points 138. In each iteration of the operation of RTLS state management 130, RTLS state 132 may be updated to reflect the newly initialized anchor point and / or current activity data derived from the current time frame in which the activity is embedded.
[0048] In some embodiments, when the RTLS state management 130 receives a set of behavior embeddings 120, the RTLS state management 130 performs trajectory tracking 140. In some embodiments, the trajectory tracking 140 can use trajectory continuity analysis to correlate previous behavior states 134 and activity anchors 137 from previous time frames with the behavior embeddings 120 derived from the live streaming data of the current micro-batch to identify tracked activity anchors 144 that can be propagated forward and updated as activity anchors 137.
[0049] For example, further reference Figure 1 , Figure 3is a data flow diagram illustrating the operation of the RTLS state management 130 according to some embodiments. The trajectory tracking 140 can perform trajectory analysis, such as subject tracking continuity 300, based on the behavioral data contained in the behavior embedding 120. For each previous behavioral state 134 associated with the active anchor point 137, the trajectory tracking 140 can use the previous position and trajectory data (e.g., direction vectors) to calculate a predicted (e.g., expected) current position and trajectory (direction of motion) for a given elapsed time between iterations - and evaluate tracking continuity based on the predicted current position and trajectory. If the magnitude of the direction vectors indicates motion of the tracked subject, the trajectory tracking 140 can normalize the direction vectors and calculate the difference between their angles. This normalized angular difference is then used as a factor to adjust the final spatiotemporal distance calculation. The behavior embedding 120, which includes position and trajectory data within a threshold distance (e.g., a Euclidean distance) of the expected current position and trajectory of the previous behavior state 134 of the activity anchor 137 at the previous time frame (e.g., the immediately previous micro-batch of image data 104), can be assigned to the anchor 136 of the previous behavior state 134 and used to define a set of tracked activity anchors 144. The tracked activity anchors 144 can then be processed using the behavior state propagation 156 to update the previous behavior state 134. That is, the behavior state propagation 156 can generate an updated behavior state 158 associated with the tracked activity anchor 144, which includes the corresponding current behavior data from the behavior embedding 120. In some embodiments, the similarity of appearance data from the behavior embedding 120 can be used to further verify the association between the behavior embedding 120 and the tracked activity anchor 144. Because the RTLS 100 updates at least a portion of the previous behavior state 134 of the activity anchor point 137 based on the tracking trajectory continuity, executing complex clustering and / or matching algorithms may be avoided at least for the tracked activity anchor point 144 .
[0050] In some embodiments, trajectory tracking 140 may identify one or more of the behavior embeddings 120 that do not have tracking continuity with the previous behavior state 134 associated with the active anchor point 137. These remaining behavior embeddings may be classified by trajectory tracking 140 as non-associated embeddings 146, which represent subjects detected by behavior extraction 110 that are not associated with the active anchor point 137. For example, non-associated embeddings 146 may represent behaviors captured in the current time frame of the image data 104 that were not present in the previous micro-batch of image data 104. In some cases, one or more of these behavior embeddings may be associated with an anchor point 136 that has been reclassified as a dormant anchor point 138. A dormant anchor point 138 (and its corresponding previous behavior state 134) may represent a subject that was previously observable by the optical image sensor 102, but was not observable in the most recent previous time frame of the image data 104 (e.g., due to leaving the monitored area and / or due to occlusion). To determine whether one or more non-associated embeddings 146 can be associated with a dormant anchor 138 (e.g., in the event that the subject has returned to the monitored area or is otherwise again observable by the optical image sensor 102 ), the RTLS state management 130 may selectively apply an assignment process that includes behavioral embedding clustering 148 followed by one or more matching algorithms 150 .
[0051] The behavioral embedding clustering 148 can cluster the non-associative embeddings 146 based on similarities in behavioral data (e.g., appearance data and / or spatiotemporal data). Thus, behavioral embeddings containing representations of the same person within the same time frame should share similar behavioral data so that these behavioral embeddings will form clusters that can be uniquely associated with different subjects. Other different clusters can be formed in the same manner based on behavioral embeddings associated with other subjects represented in the non-associative embeddings 146. In some embodiments, the behavioral embedding clustering 148 can apply one or more hierarchical clustering algorithms, such as, but not limited to, the Balanced Iterative Reduction and Clustering using Hierarchy (BIRCH) algorithm, the Ordering Points to Identify Cluster Structures (OPTICS) algorithm, the Density-Based Application Space Clustering with Noise (DBSCAN), the Hierarchical DBSCAN (HDBSCAN*), the agglomerative clustering algorithm, and / or other clustering algorithms.
[0052] In some embodiments, the behavior embedding clustering 148 performs a two-step hierarchical clustering process. As described above, the clustering process can be performed on non-associated embeddings 146 that are not found to be associated with activity anchors 137 based on trajectory tracking 140. In the first step of the clustering process, the behavior embedding clustering 148 can cluster the non-associated embeddings 146 based on, for example, the similarity of the behavior data. This first step of the clustering process can generate a number of clusters for determining the number of subject predictions. The number of subject predictions indicates how many clusters generated from the non-associated embeddings 146 are related to the actual subject. The clustering of the first step can be based on applying a fine-tuned clustering threshold parameter to the non-associated embeddings 146 - to define how closely a set of behavior embeddings need to be clustered (e.g., how close they are) to be considered a cluster. In some embodiments, the clustering threshold parameter may include a cluster quality threshold (QT) that specifies a threshold distance between cluster members and / or a minimum number of behavior embeddings that are within a threshold distance of each other for a set to be considered a cluster.
[0053] In some embodiments, in a second step of the clustering process, behavioral embedding clustering 148 performs a second clustering of the non-associative embeddings 146, which includes constrained clustering of the non-associative embeddings 146 based on the predicted number of subjects derived from the first step of the clustering process. That is, behavioral embedding clustering 148 constrains the clustering process to cluster the behavioral embeddings into a number of clusters corresponding to the number of subjects predicted to be present in the non-associative embeddings 146. Therefore, since each cluster generated by behavioral embedding clustering 148 from the second step of the clustering process is expected to contain behavioral embeddings representing the same different subjects, the appearance data and / or spatiotemporal data encoded in the behavioral embeddings within each cluster should be highly similar - and easily distinguishable from the appearance data and spatiotemporal data provided by behavioral embeddings of other clusters representing other different subjects in the monitored area. In some embodiments, when behavioral embedding clustering 148 performs behavioral embedding clustering 148, the process may use the ReID appearance data from the non-associative embeddings 146 without integrating spatiotemporal data. The spatiotemporal information can then be incorporated by the matching algorithm 150 in the reassignment phase, where both the spatiotemporal and ReID distances can be normalized and weighted for decision making. In each iteration of the matching algorithm 150, the clustered appearance data and / or spatiotemporal data (e.g., location) can be updated and optimized.
[0054] The clusters generated by the behavioral embedding clustering 148 can be used as input to the matching algorithm 150 to determine whether one or more clusters match a previous behavioral state 134 associated with a dormant anchor 138 (e.g., a previously initialized active anchor that has since been reclassified as a dormant anchor). The active anchor 137 may have been reclassified as a dormant anchor 138, for example, because the subject of the anchor ceased to appear in one or more time frames of the image data 104 for at least a threshold duration. In some embodiments, the matching algorithm 150 may associate clusters with the dormant anchor 138 using an iterative matching combinatorial optimization algorithm (e.g., a matching algorithm that solves an assignment problem by matching agents to tasks). As an example, the matching algorithm 150 may apply a Hungarian matching algorithm, which may be referred to as a Kuhn-Munkres algorithm or a Munkres assignment algorithm, to the clusters. The matching algorithm 150 can assign the cluster generated by the behavior embedding clustering 148 to the current dormant anchor 138 by matching the clustered behavior embeddings with the behavior data represented by the previous behavior state 134 associated with the dormant anchor 138 (e.g., based on similarity). The output from the matching algorithm 150 can include a set of matched dormant anchors 154 (e.g., a set of one or more dormant anchors 138 that have been successfully matched to the clusters generated by the behavior embedding clustering 148). The matched dormant anchors 154 can then be processed by the behavior state propagation 156 to update the previous behavior state 134. For example, the behavior state propagation 156 can reclassify the dormant anchors 138 included in the matched dormant anchors 154 as active anchors 137. In addition, the behavior state propagation 156 can generate one or more updated behavior states 158 associated with the matched dormant anchors 154. Based on the updated behavior state 158, the previous behavior state 134 associated with the matched dormant anchor 154 can be updated to include the corresponding current behavior data based on the behavior embedding 120 represented by the assigned cluster. In some embodiments, the iterative refinement performed by the matching algorithm 150 can include enough iterations to obtain optimal matching accuracy saturation (e.g., approximately 10 iterations in the case of the Hungarian matching algorithm).
[0055] When the matching algorithm 150 determines that one or more clusters generated by the behavioral embedding clusters 148 cannot be assigned (matched) to the dormant anchor point 138, the RTLS state management 130 can update the RTLS state 132 to include one or more newly initialized anchor points 152, each corresponding to an unmatched cluster. In addition, for each newly initialized anchor point 152, the RTLS state management 130 can update the RTLS state 132 to include an associated new behavioral state that will define one of the previous behavioral states 134 to process the next time frame of the image data 104. In some embodiments, each new behavioral state can include a representation of behavioral data associated with a different tracked subject based on the behavioral embedding from the cluster.
[0056] In some embodiments, the RTLS state management 130 can use a configurable state retention time to determine the duration that a dormant anchor 138 can remain eligible for reclassification back to an active anchor 137. For example, given a configurable state retention time of 10 minutes, the RTLS state management 130 will trigger the initialization of a new initialized anchor 152 for a tracked subject that reappears 15 minutes after being dormant. After an anchor point has been dormant for longer than the configurable state retention time, the RTLS state management 130 can delete the dormant anchor from the list of dormant anchors 138 so that it cannot be reactivated as an active anchor 137. In some embodiments, the RTLS state management 130 can maintain historical behavior data and / or behavior state associated with the deleted anchor 136 in the RTLS state 132 in a database for later analysis (e.g., using the QBE function).
[0057] Back to Figure 1 , the RTLS state 132 may be updated based on the newly initialized anchor point 152 to include the new anchor point 136, and / or update the previous behavioral state 134 based on the updated behavioral state.
[0058] For example, one or more subject evaluation functions 162 (e.g., analysis, query, and / or rendering systems) can use subject tracking data 160 derived from the RTLS state 132 and / or behavioral data represented by the previous behavioral state 134 as input. For example, the subject evaluation function 162 can include a query-by-example (QBE) function that utilizes the global ID and behavioral embeddings generated by the RTLS 100. The behavioral embeddings and / or other behavioral data associated with one or more tracked subjects can be stored in a database 164 (e.g., a Milvus database, a vector database, and / or other vector database management system (VDBMS)) to support long-term QBE queries spanning a selected time period. In some embodiments, the QBE function can operate based on Representational State Transfer (REST) application programming interface (API) inputs, which can include (but are not limited to) one or more of an object ID, a sensor ID, a timestamp, and optional parameters (e.g., a time range, a match score threshold, and / or a top-K match) to obtain similar behaviors. The QBE function can normalize the behavior embeddings before searching for similar behavior embeddings in the database 164, and the matching behaviors can have IDs used by a REST API to obtain behavior metadata using a search engine (e.g., Elasticsearch). This search capability of the QBE function can provide accurate and efficient retrieval of relevant tracked subject behavior patterns over longer periods of time. In some embodiments, subject tracking data 160 can be generated based on queries performed using the QBE function. In some embodiments, one or more subject evaluation functions 162 can use the subject tracking data 160 (and / or other data representing behavior data from the RTLS 100) to control the operation of one or more machines and / or systems. For example, the subject evaluation function 162 can control one or more operations of the AMR and / or self-machine based on the location of one or more tracked subjects represented by the subject tracking data 160.
[0059] In some embodiments, one or more subject assessment functions 162 may include a rendering system to display a comprehensive view of the monitored area via a UI and visually render the position and trajectory tracking data for each subject within the area as indicated by the subject tracking data 160. The rendering system may allow a user to switch between real and / or virtual camera views when tracking subjects moving through the monitored area, to locate the current position of a particular subject within the area (e.g., based on its global ID), and / or other applications that may benefit from real-time spatiotemporal knowledge of the position and motion of subjects of interest within the monitored area.
[0060] For example, Figure 4A and Figure 4Bis a diagram illustrating an example UI that may be displayed using a human-machine interface coupled to RTLS 100 (e.g., one or more presentation components 618 of example computing device 600). Figure 4A and Figure 4B In the example of , RTLS 100 may include a set of optical image sensors 102 including eight cameras arranged to view different portions of a monitored area 405 (e.g., an environment). Here, the monitored area 405 is shown in the form of a grocery store, but it should be understood that in various embodiments, the monitored area 405 may include any area where subject tracking is desired, such as, but not limited to, warehouses, factories, retail locations, hospitals, office buildings, security facilities, arenas, public transportation stations, parks, public places, etc. Figure 4A , UI display 410 presents multiple views 415 of monitored area 405, referred to as image sensors 1-8. In some embodiments, each view 415 may correspond to an image feed from a different sensor in optical image sensor 102. In some embodiments, RTLS 100 may generate one or more views 415 as a display of a computer vision-based environment, the views 415 corresponding to viewpoints of one or more optical image sensors 102 and / or viewpoints of one or more virtual camera views instantiated within the computer vision-based environment.
[0061] like Figure 4A As shown, each view 415 can present multiple behaviors 422, each corresponding to a tracked subject within the view of a corresponding optical image sensor 102. Furthermore, for each behavior 422, RTLS 100 can extract behavioral data and generate a corresponding one of the behavioral embeddings 120. Given that 20 tracked subjects are present and observable within monitored area 405 across eight image sensors, a maximum of 160 behavioral embeddings 120 can be expected to be generated by behavior extraction 110. However, in practice, the number of behavioral embeddings 120 can be expected to be fewer, as it is unlikely that each of the 20 tracked subjects will be observed by all eight image sensors simultaneously. In this example, 6 behaviors are observable from image sensor 1, 6 behaviors are observable from image sensor 2, 6 behaviors are observable from image sensor 3, 6 behaviors are observable from image sensor 4, 9 behaviors are observable from image sensor 5, 6 behaviors are observable from image sensor 6, 5 behaviors are observable from image sensor 7, and 8 behaviors are observable from image sensor 8. Thus, in total, the behavior extraction 110 processing these image feeds can detect 52 different behaviors 422 and generate 52 behavior embeddings 120 accordingly. In other embodiments, the RTLS 100 can obtain image data 104 from a greater or lesser number of optical image sensors 102 and display the image data from the image feeds. Figure 4AA view of one or more views of the image data is shown.
[0062] As described herein, similarities between behaviors 422 can be used to associate behaviors with respective tracked entities. For example, image sensor 3 captures behavior 424 at a cargo display area 425 at the end of an aisle that has a similar appearance to corresponding behavior 426 captured by image sensor 4 at the cargo display area 425 at the end of an aisle. Therefore, the behavior embeddings for behavior 424 and behavior 426 are similar in appearance data and / or spatiotemporal data and will be clustered together based on similarity. Therefore, the behavior embeddings for behavior 424 and behavior 426 can both be assigned to the same anchor point, behavior state, and / or global ID. In some embodiments, the RTLS 100 can display the global ID (or other separate identifier) assigned to the anchor point of the tracked entity next to the behavior of the tracked entity displayed in one or more views 415. In some embodiments, the UI display 410 may include an information field 420 that displays information about the observed tracked subjects, such as the number of tracked subjects presented in the plurality of views 415 and / or the number of optical image sensors 102 contributing optical image data 104 to the RTLS 100 to produce the views 415 .
[0063] Figure 4B is an example UI display 440 presenting an overhead view of monitored area 405. In some embodiments, UI display 440 presents virtual subjects 460 at the locations of various tracked subjects corresponding to anchor points generated by RTLS 100 and based on behavioral data represented by behavioral states corresponding to these anchor points. Figure 4A Each of the 52 behavioral embeddings 120 generated from multiple image sensor data feeds shown in is clustered by the RTLS state management 130 and associated with an activity anchor 137 and is Figure 4B 460. The spatiotemporal data for each behavior 424 from view 415 has been mapped to the global coordinate system of monitored area 405 so that the position of the tracked subject can be presented relative to a common reference system within UI display 440 using virtual subject 460.
[0064] As the RTLS state 132 is updated, the RTLS 100 may update the location and / or appearance of each virtual subject 460 representation of the tracked subject within the UI display 440. For each virtual subject 460, the UI display 440 may include a global ID 464 assigned to the anchor point of the corresponding tracked subject. The global ID 464 may be an anonymous identifier and / or may be associated with a more personal identifier, such as the subject's name, employee number, customer number, account number, student number, and / or other identifier associated with the tracked subject. In some embodiments, for one or more virtual subjects 460, the RTLS 100 may update a path indicator 462 that shows the past location and / or movement of the tracked subject over a selectable duration.
[0065] As described herein, subject assessment functionality 162 can control one or more operations of an AMR and / or a self-machine based on the position of one or more tracked subjects represented by subject tracking data 160. In some embodiments, UI display 440 can display the position and / or movement of a virtual subject 460 representing tracked subjects in monitored area 405 relative to the position and / or movement of one or more AMRs 470. In some embodiments, UI display 440 can display alternative routes 471 and 472 for navigating AMR 470 through the monitored area that avoid congestion and / or avoid interaction with tracked subjects, and / or display route change events due to congestion.
[0066] Figure 5 is a flow chart illustrating a method 500 for multi-sensor real-time subject tracking in a monitored environment according to some embodiments of the present disclosure. Figure 5 The features and elements described in the method 500 may be used in conjunction with, combined with, or substituted for elements of any other embodiment discussed herein, and vice versa. Furthermore, it should be understood that Figure 5 The functionality, structure, and other descriptions of the elements of the embodiments described in the specification may apply to the same or similarly named or described elements in any of the figures and / or embodiments described herein, and vice versa. Each block of the method 500 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, the various functions may be performed by one or more processors that include processing circuitry to execute instructions stored in memory. The method may also be embodied as computer-usable instructions stored on a computer storage medium. The method may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Additionally, the method 500 may be provided with respect to any of the figures and / or embodiments described herein. Figure 1The RTLS 100 is described by way of example. However, the method may additionally or alternatively be performed by any system or any combination of systems, including but not limited to the systems described herein.
[0067] As discussed in greater detail herein, in some embodiments, the method can include generating a behavior state representing behavior data of one or more tracked subjects within the monitored area based on a plurality of behavior embeddings computed from synchronized optical image stream data representing the monitored area. The behavior state is updated based on using the plurality of behavior embeddings based on at least one or both of the following: associating a first set of individual behavior embeddings in the plurality of behavior embeddings with one or more previous behavior states in a first plurality of previous behavior states computed from the synchronized optical image stream data based on at least the trajectory tracking data; and selectively applying an assignment process to assign a second plurality of previous behavior states to a second set of individual behavior embeddings in the plurality of behavior embeddings based on at least one of the plurality of behavior embeddings not having an associated previous behavior state from the first plurality of previous behavior states.
[0068] At block B502, method 500 includes calculating multiple representations corresponding to the behavior of one or more subjects within an environment based on a first time frame of streaming sensor data, wherein the streaming sensor data includes behavioral data of the one or more subjects within the environment captured from multiple optical sensors. The streaming sensor data may include behavioral data captured from multiple optical sensors and may represent one or more tracked subjects within a monitored area. Multiple representations (e.g., behavioral embeddings) may individually represent behavioral data of a single tracked subject from the one or more tracked subjects. As described above, RTLS 100 may include behavior extraction 110 that operates to generate behavior embeddings 120 based on image data 104, wherein the image data 104 represents the monitored area captured by multiple optical image sensors 102. The image data may include synchronized optical image stream data based on live streaming feeds from the multiple optical image sensors. The streaming sensor data may include synchronized optical image stream data that includes individual image feeds from the multiple optical sensors, wherein the multiple optical sensors are synchronized to capture the individual image feeds simultaneously. In some embodiments, the image data includes synchronized optical image stream data based on a previously recorded batch of live streaming feeds from the multiple optical image sensors. RTLS can process image data into successive micro-batches, e.g., where each micro-batch defines a different sub-second time frame of image data. The image data can be processed using behavior extraction to extract behavior embeddings representing individual behaviors of the tracked subject represented in the image data.
[0069] Each behavioral embedding may include a representation of a tracked subject as it appears in a separate feed of image data associated with different ones of the optical image sensors. The representation provided by the behavioral embedding may include behavioral data including appearance data (e.g., characterizing the appearance of the tracked subject) and / or spatiotemporal data (e.g., characterizing the position and / or trajectory of the tracked subject relative to the monitored area). The multiple behavioral embeddings may be mapped to a global image coordinate system based on camera calibration parameters associated with the multiple optical sensors. In some embodiments, the method may include generating the multiple behavioral embeddings using a machine learning model that is trained to detect features representing one or more tracked subjects based on the streaming sensor data. For example, the behavioral embeddings may be computed using a behavioral encoding model (e.g., a machine learning model). The behavioral encoding model may include, for example, a re-identification (Re-ID) embedding model that encodes the appearance of each tracked subject that appears in the image data.
[0070] In some embodiments, the method may define one or more anchor points to associate individual previous behavior states from at least one of the first plurality of previous behavior states and the second plurality of previous behavior states with a global identifier (ID). For example, the RTLS may include an RTLS state management function that generates and maintains anchor points (and their corresponding behavior states) for tracking subjects observable within the monitored area based on behavior embeddings generated from a sequence of image data over a series of time frames. The RTLS state management function may assign clusters of behavior embeddings to anchor points, where each anchor point represents a different tracked subject within the monitored area. Each anchor point has a corresponding behavior state, where behavior data for different tracked subjects can be maintained, associated with a global ID, and tracked over time.
[0071] Method 500 includes, at block B504, associating one or more of a first plurality of prior behavior states with a plurality of representations based at least on trajectory tracking data of one or more subjects. The first plurality of prior behavior states may be derived using a second time frame of streaming sensor data that occurs prior to the first time frame. For example, as described herein with respect to Figure 1 and Figure 3 As discussed, the behavior embeddings may be received by RTLS state management including trajectory tracking functionality. In some embodiments, trajectory tracking may use trajectory continuity analysis to correlate previous behavior states and activity anchors from previous time frames with behavior embeddings derived from the current micro-batch live stream data to identify tracked activity anchors that may be propagated forward and updated as activity anchors.
[0072] Trajectory tracking can include subject tracking continuity analysis based on behavioral data contained in the behavior embedding. For each previous behavior state associated with an active anchor, trajectory tracking can use the previous position and trajectory data to calculate the expected current position and trajectory for the elapsed time between given iterations. A behavior embedding comprising position and trajectory data within a threshold distance of the expected current position and trajectory calculated based on the previous behavior state of the active anchor for the previous time frame (e.g., the immediately previous micro-batch of image data) can be assigned to the anchor of the previous behavior state and used to define a set of tracked active anchors. The tracked active anchors can then be processed using behavior state propagation to update the previous behavior state.
[0073] Method 500 includes, at block B506, assigning a second plurality of previous behavior states to the plurality of representations based on at least one of the plurality of representations not having an associated previous behavior state from the first plurality of previous behavior states. The second plurality of previous behavior states may be derived using one or more time frames occurring prior to the first time frame of the streaming sensor data. For example, to determine when the remaining behavior embeddings are associated with the dormant anchor point, the RTLS may perform an assignment process that includes a hierarchical clustering process followed by a matching algorithm. As described above, the method may selectively apply the assignment process to determine when one or more non-associated embeddings 146 may be associated with the dormant anchor point 138. The RTLS state management 130 may apply an assignment process that includes behavior embedding clustering 148 followed by one or more matching algorithms 150. For those instances where there are still behavior embeddings that cannot be associated with an active anchor point, the RTLS state management may selectively perform a complex clustering and / or matching algorithm of the assignment process and apply the assignment process to those remaining behavior embeddings. Therefore, by first matching the current behavior embedding to the existing activity anchors using less computationally intensive trajectory tracking, and then selectively applying the assignment process to those behavior embeddings that could not be matched using trajectory tracking, the processing resources of the RTLS system can be more efficiently utilized. As a result, the RTLS can process each micro-batch of live streaming data quickly enough to support real-time multi-camera multi-subject tracking applications in high-density environments, for example, where the position and trajectory tracking data of multiple subjects can be simultaneously calculated and / or displayed within the short period (e.g., one second or less) of live streaming data being captured for the micro-batch.
[0074] Thus, in some embodiments, the assignment process may cluster at least one behavior embedding from the plurality of behavior embeddings that does not have an associated previous behavior state from the first plurality of previous behavior states to generate one or more clusters based at least on similarity; and selectively apply a matching algorithm to assign the second plurality of previous behavior states to the one or more clusters. Based on at least one cluster from the one or more clusters not being assigned at least one of the second plurality of previous behavior states by the matching algorithm, a new anchor point and corresponding behavior state may be initialized for the at least one cluster.
[0075] In some embodiments, clustering of behavior embeddings without associated previous behavior states may include a hierarchical clustering process that performs the following operations: determining a subject prediction number based on a first clustering process that clusters at least one of a plurality of behavior embeddings that does not have an associated previous behavior state based at least on similarity of the behavior embeddings; and applying a second clustering process to at least one of the plurality of behavior embeddings that does not have an associated previous behavior state, the second clustering process being constrained to cluster at least one of the plurality of behavior embeddings that does not have an associated previous behavior state based at least on the subject prediction number. In addition, the matching algorithm may include at least one of the following: an iterative matching combinatorial optimization algorithm; an algorithm for solving an assignment problem by matching agents to tasks; a Hungarian matching algorithm; a Kuhn-Munkres algorithm; or a Munkres assignment algorithm.
[0076] At block B508, method 500 includes updating at least one of the first plurality of previous behavior states or the second plurality of previous behavior states based on the plurality of representations to generate an updated behavior state. For example, behavior state propagation may generate an updated behavior state associated with the tracked active anchor point, which includes corresponding current behavior data from the behavior embedding. Behavior state propagation 156 may also or alternatively generate one or more updated behavior states associated with the matched dormant anchor point. Based on the updated behavior state, the RTLS state may be updated to include corresponding current behavior data based on the current behavior embedding.
[0077] For example, one or more subject assessment functions (e.g., analysis, query, and / or rendering systems) can use subject tracking data derived from the RTLS state and / or the behavioral state of the active anchor as input. As described herein, the subject assessment function can include a query by example (QBE) function that utilizes a global ID and behavioral embedding generated by the RTLS. The behavioral embedding and / or other behavioral data associated with one or more tracked subjects can be stored in a database (e.g., a Milvus database, a vector database, and / or other vector database management system (VDBMS)) to support long-term QBE queries across a selected time period. Subject tracking data can be generated based on queries executed using the QBE function. In some embodiments, one or more subject assessment functions can use the subject tracking data (and / or other data representing behavioral data from the RTLS) to control the operation of one or more machines and / or systems. For example, the subject assessment function can control one or more operations of an AMR and / or a self-machine based on the position of one or more tracked subjects represented by the subject tracking data. In some embodiments, the one or more subject assessment functions 162 may include a rendering system to display a comprehensive view of the monitored area through the UI and visually render the position and trajectory tracking data of each subject in the area indicated by the subject tracking data, such as the position and trajectory tracking data of each subject in the area. Figure 4A and Figure 4B As discussed, the UI may include a computer vision-based view of one or more tracked subjects of at least a portion of the monitored area based at least on the updated behavioral state.
[0078] The systems and methods described herein can be used for various purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, generative AI, and / or any other suitable application.
[0079] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems implementing one or more language models - such as one or more large language models (LLMs), systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems.
[0080] Example computing device
[0081] Figure 6 FIG6 is a block diagram of an example computing device 600 suitable for implementing some embodiments of the present disclosure. Computing device 600 may include an interconnect system 602 that directly or indirectly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, I / O components 614, a power supply 616, one or more presentation components 618 (e.g., a display), and one or more logic units 620. In at least one embodiment, computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more GPUs 608 may include one or more vGPUs, one or more CPUs 606 may include one or more vCPUs, and / or one or more logic units 620 may include one or more virtual logic units. Thus, computing device 600 may include discrete components (e.g., a complete GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof. In some embodiments, one or more functions of RTLS 100 disclosed herein may be implemented as code executed by computing device 600, for example, using CPU 606 and / or GPU 608.
[0082] although Figure 6The various blocks of are shown as being connected via an interconnect system 602 having wires, but this is not intended to be limiting and is provided for clarity only. For example, in some embodiments, a presentation component 618 such as a display device may be considered an I / O component 614 (e.g., if the display is a touch screen). As another example, the CPU 606 and / or the GPU 608 may include memory (e.g., the memory 604 may represent a storage device in addition to the memory of the GPU 608, the CPU 606, and / or the other components). Thus, Figure 6 The term computing device is illustrative only. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all are considered within the Figure 6 within the range of computing devices.
[0083] Interconnect system 602 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. Interconnect system 602 can include one or more links or bus types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standard association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 606 can be directly connected to memory 604. In addition, CPU 606 can be directly connected to GPU 608. In the case where there is a direct or point-to-point connection between components, interconnect system 602 can include a PCIe link to perform the connection. In these examples, it is not necessary to include a PCI bus in computing device 600.
[0084] Memory 604 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 600. Computer-readable media can include volatile and non-volatile media and removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0085] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by computing device 600. When used herein, computer storage media does not include the signals themselves. In some embodiments, database 164 may be implemented at least in part using memory 604.
[0086] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information into the signal. By way of example and not limitation, computer storage media may include wired media such as a wired network or a direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0087] The CPU 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. The CPU 606 can include any type of processor and can include different types of processors, depending on the type of computing device 600 implemented (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers). For example, depending on the type of computing device 600, the processor can be an Advanced RISC (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 600 can also include one or more CPUs 606 in addition to one or more microprocessors or supplementary coprocessors such as math coprocessors.
[0088] In addition to or in place of the CPU 606, the GPU 608 may also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. One or more GPUs 608 may be integrated GPUs (e.g., with one or more CPUs 606) and / or one or more GPUs 608 may be discrete GPUs. In embodiments, one or more GPUs 608 may be coprocessors for one or more CPUs 606. The computing device 600 may use the GPU 608 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 608 may be used for general-purpose computing on a GPU (GPGPU). The GPU 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 may generate pixel data for outputting an image in response to a rendering command (e.g., a rendering command received from the CPU 606 via a host interface). The GPU 608 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory can be included as part of memory 604. GPU 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or through a switch (e.g., using NVSwitch). When combined, each GPU 608 can generate pixel data or GPGPU data for different portions or different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.
[0089] In addition to or in lieu of the CPU 606 and / or GPU 608, the logic unit 620 may be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more methods and / or processes described herein. In embodiments, the CPU 606, GPU 608, and / or logic unit 620 may perform any combination of methods, processes, and / or portions thereof, either separately or in combination. The one or more logic units 620 may be part of and / or integrated within the one or more CPUs 606 and / or GPUs 608 and / or the one or more logic units 620 may be discrete components of or otherwise external to the CPU 606 and / or GPU 608. In embodiments, the one or more logic units 620 may be processors of the one or more CPUs 606 and / or GPUs 608. In some embodiments, the behavioral coding model 220 may be implemented, at least in part, using the one or more GPUs 608.
[0090] Examples of logic unit 620 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0091] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that allow the computing device 600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communications. The communication interface 610 may include components and functionality that allow communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or InfiniBand communications), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 620 and / or the communication interface 610 may include one or more data processing units (DPUs) to transmit data received over the network and / or through the interconnect system 602 directly to one or more GPUs 608 (e.g., their memories).
[0092] The I / O ports 612 can allow the computing device 600 to be logically coupled to other devices including I / O components 614, presentation components 618, and / or other components, some of which can be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include a microphone, a mouse, a keyboard, a joystick, a game pad, a game controller, a satellite dish, a browser, a printer, a wireless device, and the like. The I / O components 614 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological input. In some instances, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 600 (as described in more detail below). The computing device 600 can include a depth camera such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof for gesture detection and recognition. Additionally, computing device 600 may include an accelerometer or gyroscope to enable motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 600 to render immersive augmented or virtual reality.
[0093] The power supply 616 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 616 may provide power to the computing device 600 to enable the components of the computing device 600 to operate.
[0094] The presentation component 618 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 may receive data from other components (e.g., the GPU 608, the CPU 606, the DPU, etc.) and output the data (e.g., as images, video, sound, etc.). In some embodiments, one or more presentation components 618 may be used to display one or more UIs used in conjunction with the RTLS 100.
[0095] Sample Data Center
[0096] Figure 7 An example data center 700 is shown, which may be used in at least one embodiment of the present disclosure. The data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0097] like Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources ("node CRs") 716(1)-716(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 716(1)-716(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules and cooling modules, etc. In some embodiments, one or more of the node CRs 716(1)-716(N) may correspond to a server having one or more of the above-mentioned computing resources. In addition, in some embodiments, the node CRs
[0098] 716(1)-716(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of the node CRs 716(1)-716(N) may correspond to virtual machines (VMs). In some embodiments, one or more functions of the RTLS 100 disclosed herein may be implemented as code executed by the computing device 600, for example, using one or more of the node CRs 716(1)-716(N).
[0099] In at least one embodiment, the computing resources 714 of the grouping can include a separate grouping (not shown) of the node CR716 housed in one or more racks, or many racks (also not shown) housed in the data center of each geographical location. The separate grouping of the node CR716 in the computing resources 714 of the grouping can include computing, network, memory or storage resources that can be configured or assigned to support the grouping of one or more workloads. In at least one embodiment, several node CR716 comprising CPU, GPU, DPU and / or other processors can be grouped in one or more racks to provide computing resources to support one or more workloads. One or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.
[0100] Resource coordinator 712 may configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may comprise a software design infrastructure ("SDI") management entity for data center 700. Resource coordinator 712 may comprise hardware, software, or some combination thereof.
[0101] In at least one embodiment, Figure 7 As shown, the framework layer 720 may include a job scheduler 728, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 may include a framework that supports the software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application 742 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be, but is not limited to, a free and open source software network application framework, such as Apache Spark, which can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 728 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and a distributed file system 738 for supporting large-scale data processing. The resource manager 736 can manage clustered or grouped computing resources mapped to or allocated to support the distributed file system 738 and the job scheduler 728. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 714 at the data center infrastructure layer 710. The resource manager 736 can coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources. In some embodiments, one or more applications 742 and / or software 732 can be used to implement one or more functions of the RTLS 100 disclosed herein.
[0102] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0103] In at least one embodiment, the one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0104] In at least one embodiment, any of the configuration manager 734, resource manager 736, and resource coordinator 712 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. The self-modification actions can relieve the data center operator of the data center 700 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0105] The data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 700. In at least one embodiment, by using weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 700 may be used to infer or predict information using a trained machine learning model corresponding to one or more neural networks, such as but not limited to those described herein. For example, in some embodiments, the data center 700 may be used, at least in part, for training and / or implementation.
[0106] In at least one embodiment, the data center 700 may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or reasoning using the aforementioned resources. In addition, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.
[0107] Sample network environment
[0108] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to: Figure 6 The backend device 700 may be implemented on one or more instances of the computing device 600 of the embodiment of the present invention—for example, each device may include similar components, features and / or functions of the computing device 600. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of the data center 700, examples of which are described herein with respect to FIG. Figure 7 Describe in more detail.
[0109] The components of the network environment can communicate with each other through the network, which can be wired, wireless, or both. The network can include multiple networks, or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.
[0110] Compatible network environments may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment), and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to the server may be implemented on any number of client devices.
[0111] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or application may include network-based service software or application programs, respectively. In an embodiment, one or more client devices may use network-based service software or application programs (e.g., by accessing the service software and / or application programs via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software network application framework that may, for example, use a distributed file system for large-scale data processing (e.g., "big data").
[0112] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Any of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across a state, region, country, global, etc.). If the connection to the user (e.g., client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0113] Client devices may include Figure 6 The client device 600 may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality head-mounted display, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a watercraft, an aircraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, an in-vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these described devices, or any other suitable device.
[0114] The present disclosure can be described in the general context of machine-usable instructions or computer code executed by a computer or other machine such as a personal digital assistant or other handheld device, including computer-executable instructions such as program modules. Generally, program modules including routines, programs, objects, components, data structures, etc. refer to code that performs a specific task or implements a specific abstract data type. The present disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0115] As used herein, the statement "and / or" with respect to two or more elements should be interpreted as referring to only one element or combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B and C. In addition, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0116] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways to include steps that are different from the steps described herein in conjunction with other current or future technologies, or combinations of similar steps. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be interpreted as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
Claims
1. One or more processors, including processing circuitry configured to: Computing a plurality of representations corresponding to behavior of one or more agents within the environment based on a first time frame of the streaming sensor data, wherein The streaming sensor data includes behavioral data of one or more subjects within an environment captured from a plurality of optical sensors; associating one or more previous behavioral states of a first plurality of previous behavioral states with the plurality of representations based at least on trajectory tracking data of the one or more subjects; assigning a second plurality of previous behavior states to the plurality of representations based on at least one representation of the plurality of representations not having an associated previous behavior state from the first plurality of previous behavior states; as well as At least one of the first plurality of previous behavior states or the second plurality of previous behavior states is updated based on the plurality of representations to generate an updated behavior state.
2. The one or more processors of claim 1 , wherein the processing circuit is further configured to: One or more anchor points are defined to associate an individual previous behavior state from at least one of the first plurality of previous behavior states and the second plurality of previous behavior states with a global identifier ID.
3. The one or more processors of claim 1 , wherein the processing circuitry is further configured to perform a dispatch process to: clustering at least one representation of the plurality of representations that does not have the associated previous behavior state from the first plurality of previous behavior states to generate one or more clusters based at least on similarity; and The second plurality of previous behavior states are assigned to the one or more clusters using a matching algorithm.
4. The one or more processors of claim 3, wherein the processing circuit is further configured to: Based on at least one cluster of the one or more clusters not being assigned at least one of the second plurality of previous behavior states by the matching algorithm, a new anchor point and corresponding behavior state are initialized for the at least one cluster.
5. The one or more processors of claim 3 , wherein the processing circuitry is further configured to cluster at least one of the plurality of representations that does not have the associated previous behavior state based on a hierarchical clustering process that: determining a subject prediction number based on at least similarities of representations in the plurality of representations based on a first clustering process that clusters at least one representation in the plurality of representations that does not have the associated prior behavioral state; and Based at least on the predicted number, a second clustering process is applied to the at least one representation of the plurality of representations that does not have the associated previous behavioral state to cluster the at least one representation of the plurality of representations that does not have the associated previous behavioral state.
6. The one or more processors of claim 3, wherein the matching algorithm comprises at least one of: Iterative matching combination optimization algorithm; Algorithms that solve the assignment problem by matching agents with tasks; Hungarian matching algorithm; Kuhn-Munkres algorithm; or Munkres assignment algorithm.
7. The one or more processors of claim 1, wherein the behavioral data comprises at least one of appearance data or spatiotemporal data represented by the plurality of behavioral embeddings.
8. The one or more processors of claim 1 , wherein the processing circuitry is further configured to: The plurality of representations is generated using a machine learning model trained to detect features representative of the one or more subjects based at least on the streaming sensor data.
9. The one or more processors of claim 1 , wherein the streaming sensor data comprises synchronized optical image stream data comprising individual image feeds from the plurality of optical sensors, wherein: The plurality of optical sensors are synchronized to capture the individual image feeds simultaneously.
10. The one or more processors of claim 1 , wherein the processing circuitry is further configured to: The plurality of representations are mapped to a global image coordinate system based at least on camera calibration parameters associated with the plurality of optical sensors.
11. The one or more processors of claim 1 , wherein the processing circuitry is further configured to cause display of a computer vision-based view of the one or more agents of at least a portion of the environment based at least on the updated behavioral state.
12. The one or more processors of claim 1 , wherein the one or more processors are included in at least one of: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; a system for performing simulation operations; Systems for performing digital twin operations; a system for performing light transport simulations; A system for performing collaborative content creation of three-dimensional assets; Systems for performing deep learning operations; systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system that implements one or more large language models (LLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
13. A system comprising one or more processors, wherein the one or more processors are configured to: generating a plurality of representations computed based on the optical image stream data, the plurality of representations representing one or more tracked subjects within the monitored area; associating, based at least on the trajectory tracking data, a first set of individual representations of the plurality of representations with one or more previous behavior states of a first plurality of previous behavior states computed from the optical image stream data; assigning, using an assignment process, a second plurality of previous behavior states to a second set of individual representations in the plurality of representations based on at least one representation in the plurality of representations not having an associated previous behavior state from the first plurality of previous behavior states; as well as At least one of the first plurality of previous behavior states and the second plurality of previous behavior states is updated based on the plurality of representations to generate an updated behavior state.
14. The system of claim 13, wherein the one or more processors are further configured to execute the assignment process to: clustering at least one representation of the plurality of representations that does not have an associated previous behavioral state from the first plurality of previous behavioral states to generate one or more clusters based at least on similarity; and A matching algorithm is selectively applied to assign the second plurality of previous behavior states to the one or more clusters.
15. The system of claim 14, wherein the one or more processors are further configured to: Based on at least one cluster of the one or more clusters not being assigned at least one of the second plurality of previous behavior states by the matching algorithm, a new anchor point and corresponding behavior state are initialized for the at least one cluster.
16. The system of claim 14, wherein the one or more processors are further configured to cluster the at least one of the plurality of representations that does not have the associated previous behavior state based on a hierarchical clustering process that: determining a subject prediction number based on at least similarity of the representations based on a first clustering process that clusters at least one of the plurality of representations that does not have the associated prior behavior state; and Based at least on the number of predictions, a second clustering process is applied to the at least one representation in the plurality of representations that does not have the associated previous behavioral state, the second clustering process being constrained to cluster the at least one representation in the plurality of representations that does not have the associated previous behavioral state.
17. The system of claim 14, wherein the matching algorithm comprises at least one of the following: Iterative matching combination optimization algorithm; Algorithms that solve the assignment problem by matching agents with tasks; Hungarian matching algorithm; Kuhn-Munkres algorithm; or Munkres assignment algorithm.
18. The system of claim 13, wherein the one or more processors are further configured to: The plurality of representations is generated using a machine learning model trained to detect features representative of the one or more tracked subjects based at least on the optical image stream data.
19. The system of claim 13, wherein the system is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; a system for performing simulation operations; Systems for performing digital twin operations; a system for performing light transport simulations; A system for performing collaborative content creation of three-dimensional assets; Systems for performing deep learning operations; systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system that implements one or more large language models (LLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
20. A method comprising: generating a behavioral state representing behavioral data of one or more subjects within the environment based on a plurality of representations computed from optical image stream data representing the environment, wherein the behavioral state is updated using the plurality of representations based on at least one of: associating, based at least on the trajectory tracking data, a first set of individual representations of the plurality of representations with one or more previous behavior states of a first plurality of previous behavior states computed from the optical image stream data; and An assignment process is applied to assign a second plurality of previous behavior states to a second set of individual representations in the plurality of representations based on at least one representation in the plurality of representations not having an associated previous behavior state from the first plurality of previous behavior states.