Multi-Object Tracking Using Correlation Filters in Video Analysis Applications

By processing the scaled search area and image area in parallel in GPU memory, and learning related filters using focus windows and occlusion maps, the inefficiency problem of multi-object tracking method in crowded environments and occlusion situations is solved, and efficient multi-object tracking is achieved.

CN113950702BActive Publication Date: 2025-07-18NVIDIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080040489.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-03
Filing Date
2020-05-29
Publication Date
2025-07-18
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

The existing multi-object tracking method is inefficient in handling crowded environments and occlusions, making it difficult to effectively track multiple objects, and has high computing and storage requirements, which limits the number of objects that can be tracked simultaneously.

Method used

Using a multi-object tracking method of related filters, the scalable search area and image area are processed in parallel in the graphics processing unit (GPU) memory, and the focus window and occlusion diagram are used to learn related filters, reducing background learning and improving processing efficiency.

Benefits of technology

It improves the computing efficiency and storage efficiency of multi-object tracking, and can effectively track multiple objects in crowded environments and occlusions, reducing the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113950702B_ABST
    Figure CN113950702B_ABST
Patent Text Reader

Abstract

In various examples, image regions can be extracted from a batch of one or more images, and the image regions can be scaled to one or more template sizes in batches. In cases where the image region includes a search region for object localization, the scaled search region can be loaded into the Graphics Processing Unit (GPU) memory and processed in parallel for localization. Similarly, in cases where the image region is used for filter updates, the scaled image region can be loaded into the GPU memory and processed in parallel for filter updates. The image regions can be processed in batches from any number of images and / or from any number of single-object and / or multi-object trackers. Other aspects of the present disclosure provide methods for associating locations using correlation response values, methods for enabling a correlation filter to learn in object tracking based at least on focus windowing, and methods for enabling a correlation filter to learn in object tracking based at least on an occlusion map.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Efficient and effective object tracking is a key task in the visual perception pipeline as it bridges inference results across video frames, enabling temporal analysis of objects of interest. Tracking multiple objects is a key issue for many applications such as surveillance, animation, activity recognition, or vehicle navigation. Conventional multi-object trackers can be implemented using independent single-object trackers that operate on the full frame of a video and track objects by associating bounding boxes between frames. Tracking is typically performed on a single video stream and is divided into localization and data association. For localization, each single-object tracker can independently estimate the position of the detected objects in the frame - and for data association - the estimated object positions from the trackers can be linked across frames to form complete trajectories. Discriminative correlation filters (DCFs) have recently been used for localization in object tracking. A DCF-based tracker can define a search region around the object of interest, where the best correlation filter is learned so that the object can be localized in the next frame as the peak position of the correlation response within the search region.

[0002] Single-object trackers can each analyze and generate non-homogeneous data across the trackers, such as image regions of the video and correlation filters (in the case of DCF-based trackers) of different sizes and shapes. The non-homogeneous data of the trackers is processed serially by the trackers and then combined in data association. In an example of implementing conventional methods in a system where many single-object trackers are running and possibly many multi-object trackers (one per video stream), the processing and data storage requirements may limit the number of objects that can be tracked simultaneously. Conventional methods also have difficulty tracking objects in crowded environments and / or environments where occlusions are common, making tracking a challenging task. Summary of the Invention

[0003] Embodiments of the present disclosure relate to multi-object tracking using correlation filters. Systems and methods are disclosed that can improve the computational and storage efficiency of multi-object trackers (such as those implemented using correlation filters). Additional aspects of the present disclosure relate to various improvements to embodiments of correlation filters for object tracking in, for example, multi-object trackers.

[0004] Compared with conventional systems (such as those described above), image regions can be extracted from a batch of one or more images, and the image regions can be scaled to one or more template sizes in batches. In cases where the image region includes a search region for object localization, the scaled search region can be loaded into the graphics processing unit (GPU) memory and processed in parallel for localization. Similarly, in cases where the image region is used for filter updates, the scaled image region can be loaded into the GPU memory and processed in parallel for filter updates. Image regions can be processed in batches from any number of images and / or from any number of single-object and / or multi-object trackers. Further aspects of the present disclosure provide methods for associating positions using correlation response values, methods for making correlation filters learn in object tracking based at least on focus windowing, and methods for making correlation filters learn in object tracking based at least on occlusion maps. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The present system and method for a multi-object tracker using a correlation filter are described in detail below with reference to the accompanying drawings, in which:

[0006] Figure 1 is a diagram showing an example of an object tracking system according to some embodiments of the present disclosure;

[0007] Figure 2 is a diagram showing an example of how an object tracking system according to some embodiments of the present disclosure can implement multi-object tracking over multiple frames; Figure 1

[0008] Figure 3 is a diagram showing an example of batch processing that can be used to implement multi-object tracking according to some embodiments of the present disclosure;

[0009] Figure 4 is a diagram showing an example of processing a batch of one or more images according to some embodiments of the present disclosure;

[0010] Figure 5 is a flowchart showing a method for batch processing search regions of an object tracker to determine an estimated object position according to some embodiments of the present disclosure;

[0011] Figure 6 is a flowchart showing a method for batch processing image regions of an object tracker to initialize or update a correlation filter according to some embodiments of the present disclosure;

[0012] Figure 7 is a flowchart showing a method for batch cropping and scaling search regions of an object tracker to determine an estimated object position according to some embodiments of the present disclosure;

[0013] ​Figure 8 is a diagram showing an example of associating positions based on relevant response values according to some embodiments of the present disclosure;

[0014] Figure 9 is a flowchart showing a method for associating positions based on relevant response values at least according to some embodiments of the present disclosure;

[0015] Figure 10 is a diagram showing an example of applying a focus window to an image to enable a correlation filter to learn using the focus window according to some embodiments of the present disclosure;

[0016] Figure 11 is a flowchart showing a method for applying a focus window to an image to enable a correlation filter to learn using the focus window according to some embodiments of the present disclosure;

[0017] Figure 12A is a diagram showing an example of an image, an occlusion map of the image, and a target model of a correlation filter learned using the image and the occlusion map according to some embodiments of the present disclosure;

[0018] Figure 12B is a diagram showing an example according to some embodiments of the present disclosure of Figure 12A the correlation response of a correlation filter and the estimated object position determined using the correlation filter;

[0019] Figure 13 is a flowchart showing a method for applying an occlusion map to an image to enable a correlation filter to learn according to some embodiments of the present disclosure;

[0020] Figure 14 is a diagram showing an example of a correlation response having multiple modes according to some embodiments of the present disclosure; and

[0021] Figure 15A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0022] Figure 15B is for according to some embodiments of the present disclosure for Figure 15A an example of the camera positions and fields of view of an example autonomous vehicle;

[0023] Figure 15C is for according to some embodiments of the present disclosure for Figure 15A a block diagram of an example system architecture of an example autonomous vehicle;

[0024] Figure 15D is for according to some embodiments of the present disclosure of Figure 15A a system diagram of the communication between a cloud-based server and an example autonomous vehicle; and

[0025] Figure 16 A block diagram of an example computing device suitable for implementing some embodiments of the present disclosure. Detailed implementation

[0026] Systems and methods related to multi-object trackers using correlation filters are disclosed. Systems and methods are disclosed that can improve the computational and storage efficiency of multi-object trackers (such as those implemented using correlation filters). Additional aspects of the present invention relate to various improvements in the implementation of correlation filters for object tracking, for example, in multi-object trackers.

[0027] The disclosed embodiments can be implemented in a variety of different perception-based object tracking and / or recognition systems, such as in automotive systems, robots, aerial systems, boating systems, intelligent area monitoring, simulation, and / or other technical fields. The disclosed methods can be used for any perception-based control, analysis, monitoring, tracking, and / or behavior modification of machines and / or systems.

[0028] For intelligent area monitoring, various disclosed embodiments can be incorporated into the systems and / or methods described in U.S. Non-Provisional Application No. 16 / 365,581, entitled "Smart Area Monitoring with Artificial Intelligence," filed on March 26, 2019, the entire content of which is hereby incorporated by reference.

[0029] For simulation, various disclosed embodiments can be incorporated into the systems and / or methods described in U.S. Non-Provisional Application No. 16 / 366,875, entitled "Training, Testing, and Verifying Autonomous Machines Using Simulated Environments," filed on March 27, 2019, the entire content of which is hereby incorporated by reference.

[0030] For locomotive systems, although the present disclosure may be described with respect to the exemplary autonomous vehicle 1500 (or, referred to herein as "vehicle 1500" or "this vehicle 1500," examples of which are described herein), this is not intended to be limiting. For example, the systems and methods described herein can be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), robots, warehouse vehicles, off-road vehicles, airships, boats, and / or other vehicle types. Figures 15A to 15D For example, the systems and methods described herein can be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), robots, warehouse vehicles, off-road vehicles, airships, boats, and / or other vehicle types.

[0031] Compared with conventional systems (such as those described above), image regions can be extracted from a batch of one or more images, and the image regions can be scaled to one or more template sizes in batches. In doing so, the size and shape of the scaled image regions, as well as the correlation filters and correlation responses in the correlation filter-based method, can be made more uniform. This can reduce the storage size and improve the processing efficiency, while allowing the image regions to be analyzed and processed efficiently and effectively in parallel, for example, using one or more GPU threads. For example, in the case where the image region contains a search region for object localization, the scaled search region can be loaded into the GPU memory and processed in parallel for localization. Similarly, in the case where the image region is used for filter update, the scaled image region can be loaded into the GPU memory and processed in parallel for filter update. Using the disclosed method, image regions can be batched from any number of images and / or from any number of single-object and / or multi-object trackers, allowing parallelization across single-object trackers and / or video stream processing.

[0032] A further aspect of the present disclosure provides a method for associating positions using correlation response values. A correlation filter having a correlation response at an estimated position can be used to determine the estimated position of an object. When determining whether to associate a position with the estimated position, the value of the correlation response can be determined for the position and used as the visual feature for the determination. Thus, it is not necessary to separately extract visual features from the image at that position.

[0033] An additional aspect of the present invention provides for learning a correlation filter in object tracking based at least on focus windowing. When the correlation filter is learning from an image region, a focused window can be applied to the image region (e.g., one or more of its channels), which blurs the background of the target object, where the blurring increases based on the distance from the target. The focused window can be applied to one or more color and / or feature channels of the image using a Gaussian filter. The correlation filter can learn from the blurred image, reducing learning from the background while still allowing the background to provide learning of the context around the target object. Additionally, when the image region is a search region for localizing the target object, a larger search region can be used without the risk of overlearning the background.

[0034] The present disclosure also provides for learning a correlation filter in object tracking based at least on an occlusion map. When learning the correlation filter from an image region, the occlusion map can be applied to mask, exclude, and / or blur the occluded image region of a target object. The correlation filter can learn from the modified image, thereby reducing or eliminating learning from occlusions while still allowing learning of the target object from the exposed portions. The occlusion map can be generated using a machine learning model, such as a Gaussian mixture model (GMM) trained using the target object as the background (e.g., using the image region for learning the correlation filter), such that the occlusion is detected as foreground.

[0035] Figure 1 FIG. is an illustration showing an example of an object tracking system 100 according to some embodiments of the present disclosure. It should be understood that such and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) can be used in addition to or in place of the shown arrangements and elements, and some elements can be omitted altogether. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. The different functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, the different functions can be implemented by a processor executing instructions stored in a memory. By way of example, the object tracking system 100 can be implemented on one or more instances of a Figure 16 computing device 1600.

[0036] The object tracking system 100 can include, among other things, one or more sensors 102, one or more object detectors 104, one or more object trackers 106, one or more data associators 108, and a batch manager 110.

[0037] One or more sensors 102 may be configured to generate sensor data, such as image data representing one or more images (e.g., image 102A, 102B, or 102C), which may be frames of one or more video streams. One or more object detectors 104 may be configured to detect objects in the sensor data, which may include detected object locations, such as one or more points of a bounding box or other shape within an image. One or more object trackers 106 may be configured to analyze the sensor data and, in some examples, the detected object locations to locate the detected objects in a frame based at least on determining one or more estimated object locations. One or more data associators 108 may be configured to manage object tracking across frames and / or video streams based at least on linking the estimated object locations across frames (e.g., to generate a track 132). This may include, for example, associating and / or assigning the estimated object locations to one or more objects and / or detected objects based at least on assigning an object identifier (ID) to the estimated object locations. A batch manager 110 may be configured to manage batching in the object tracking system 100 for implementations that employ batching.

[0038] The object tracker 106 may include an object locator 112. In implementations that employ a correlation filter, the object tracker 106 may also include a filter initializer 114 and a filter updater 116. The object locator 112 may be configured to use one or more machine learning models to locate the detected objects. For example, the machine learning model may be implemented using a correlation filter learned using the filter initializer 114 and the filter updater 116. The filter initializer 114 may be configured to, for example, initialize a correlation filter for a newly tracked and / or detected object (e.g., by the object detector 104). The filter updater 116 may be configured to, for example, update a correlation filter for a previously tracked object (e.g., previously located by the object tracker 106). In some examples, each object tracker 106 may be responsible for tracking and maintaining the state of a single corresponding object.

[0039] The data associator 108 may include an object matcher 118, a tracker state manager 120, a location aggregator 122, a tracker instantiator 124, and a tracker terminator 126. The object matcher 118 may be configured to match the estimated object locations from the object tracker 106 with one or more detected object locations (e.g., from the object detector 104) and / or object IDs. The tracker state manager 120 may be configured to manage the state of an object tracker, such as the object tracker 106. For example, the tracker state manager 120 may manage the state of the object tracker based at least on the matching results of the object matcher 118.

[0040] The tracker state manager 120 may use a location aggregator 122 that is configured to aggregate, combine, and / or merge object locations. These locations may include detected object locations from the object detector 104 and estimated object locations from the object tracker 106, where the estimated object locations of the object tracker 106 are matched to the detected object locations by the object matcher 118. The tracker state manager 120 may assign the aggregated locations to the state of the object tracker 106 for an image and / or frame. The tracker state manager 120 may also use a tracker instantiator 124 that is configured to instantiate a new object tracker 106. The tracker state manager 120 may instantiate an object tracker 106 for detected object locations from the object detector 104 for which the object matcher 118 cannot match an estimated object location and / or a previously tracked object. The tracker state manager 120 may further use a tracker terminator 126 that is configured to terminate an existing object tracker 106. The tracker state manager 120 may terminate the object tracker 106 for an estimated object location from such an object tracker 106 where the object matcher 118 cannot match it to a detected object location from the object detector 104, an estimated object location from a previous frame, and / or a previously tracked object (e.g., where localization fails and / or is below a threshold level of confidence). In some examples, the tracker state manager 120 may further be used to re-identify a tracked object across frames and / or a video stream (e.g., merge detections of the same object across a video stream and / or activate a tracker for a re-emerging object). In some examples, each data associator 108 may be responsible for multi-object tracking within a respective video stream and / or feed.

[0041] As described herein, one or more sensors 102 may be configured to produce sensor data, such as image data representing one or more images (e.g., image 102A, 102B, or 102C), which may be frames of one or more video streams. In some examples, the sensor data may be produced by any number of sensors, such as one or more image sensors of one or more cameras. Other examples of sensors that may be employed include LIDAR sensors, RADAR sensors, ultrasonic sensors, microphones, and / or other sensor types. The sensor data may represent one or more fields of view and / or sensing fields of the sensors 102, and / or may represent the perception of the environment by the one or more sensors 102.

[0042] Sensors, such as image sensors (e.g., of a camera), LIDAR sensors, RADAR sensors, SONAR sensors, ultrasonic sensors, etc., may be referred to herein as perception sensors or perception sensor devices, and sensor data generated by the perception sensors may be referred to herein as perception sensor data. In some examples, instances of sensor data may represent an image captured by an image sensor, a depth map generated by a LIDAR sensor, etc. LIDAR data, SONAR data, RADAR data, and / or other sensor data types may be related or associated with image data generated by one or more image sensors. For example, image data representing one or more images may be updated to include data related to LIDAR sensors, SONAR sensors, RADAR sensors, etc., such that the sensor data used by object detector 104 and / or object tracker 106 may be more informative or detailed than the image data alone. Thus, this additional information from any number of perception sensors may be used to perform object tracking.

[0043] In various embodiments, filter initializer 114 and / or filter updater 116 may cause a correlation filter (e.g., DCF) to learn, which generates and / or is used to identify a peak correlation response for a target from an image region. For example, the correlation filter may learn to generate a peak correlation response at the center of a target in an image region. Object locator 112 may use the correlation filter to locate a target in a search region based at least on determining the peak location of the correlation response, which may correspond to an estimated object location. In each frame, an optimal filter may be created that generates a peak correlation response at the target location on a per-frame basis. Filter updater 116 may update the correlation filter using positive and / or negative samples, which may be based on the object location determined using object locator 112. Filter updater 116 may use an image region (e.g., a search region) corresponding to the object location and find a filter that maximizes the correlation response for positive samples and minimizes the correlation response for negative samples (e.g., for DCF). Filter updater 116 may update the target model of the correlation filter at least in part based on an exponential moving average (EMA) of the optimal filter generated at each frame. This method may be used for temporal consistency across frames. For example, the correlation filter F may be calculated at frame N using equation (1):

[0044] F(N) = α * observation+(1 - α) * F(N - 1) (1)

[0045] where α may represent a learning rate and observation may represent the correlation filter created for frame N.

[0046] In various embodiments, the learning rate α can be based at least on the relevant response signal-to-noise ratio (SNR). In embodiments that employ confidence scores, the ratio between the confidence scores and / or relevant response values described herein can be used to determine the learning rate α (e.g., the learning rate can be a function of the confidence scores and / or ratios or be proportional to the confidence scores and / or ratios). For example, in cases where the ratio and / or confidence score is low, a lower learning rate can be used when updating the target appearance model of the correlation filter, while a higher learning rate can be used when the ratio and / or confidence score is high.

[0047] Different types of correlation filters are contemplated within the scope of the present disclosure. A correlation filter can refer to a class of classifiers that are configured to produce a peak in a correlation output or response, e.g., to achieve accurate localization of a target in a scene. Examples of suitable correlation filters include Kernelized Correlation Filters (KCF), Discriminative Correlation Filters (DCF), Correlation Filter Neural Networks (CFNN), Multi-Channel Correlation Filters (MCCF), kernel correlation filters, adaptive correlation filters, and / or other filter types. KCF is a variant of DCF that uses the so-called "kernel trick" when solving the internal optimization to find the global minimum during the filter update phase. All other workflows can be the same as in a typical DCF.

[0048] One or more machine learning models (MLMs) can be used to implement the correlation filter. The MLMs described herein can take various forms, e.g., but not limited to, the MLMs can include any type of machine learning model, such as machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid machines, etc.) and / or other types of machine learning models.

[0049] In at least one embodiment, the position of the search region in the frame can be based at least in part on a previously determined position associated with the object. For example, the search region can be offset from the previously determined position and can include the previously determined position. Examples of previously determined positions include the detection position of the object determined by the object detector 104 (e.g., for a previous frame), the estimated position of the object determined by the object tracker 106 for a previous frame, and / or combinations thereof. For example, in the case where the search region is based on a combination of the estimated position and the detection position, the position aggregator 122 can aggregate, combine, fuse, and / or merge the estimated position with the detection position to produce a combined position. The aggregated position can be the estimated position, the detection position, or a different position based on those positions (e.g., a statistical combination).

[0050] Now refer to Figure 2 , Figure 2 which is a diagram showing an example of how an object tracking system 100 according to some embodiments of the present disclosure Figure 1 can implement multi-object tracking over multiple frames. Figure 2 By way of example, frames 0, 1, and 2 are shown, which can belong to the same video stream. The video stream can be analyzed using the Figure 1 object tracking system 100 to generate trajectory data of the trajectories of one or more objects detected in the video stream and / or to track the trajectories over multiple frames. For example, trajectory data 130 of a trajectory 132 of an object 160 (e.g., a vehicle) can be generated and / or the trajectory 132 can be tracked over multiple frames. The numbers used to label the frames are intended to indicate the temporal relationship between the frames and do not necessarily mean that the frames are consecutive (although they may be). For example, frame 0 can be before frame 1 in the video, and frame 2 can be after frame 1, but there can be intermediate frames.

[0051] In frame 0, the object detector 104 can analyze one or more portions of the image data and / or sensor data (e.g., of frame 0 and / or temporally related to frame 0) to detect the position of one or more objects (if any) in the field of view of the sensor 102. As shown, the object detector 104 can determine the position 202 of object 160, the position 206 of object 208, the position 210 of object 212, and the position 214 of object 216. The data associator 108 can instantiate an object tracker 106 for each detected object position in frame 0 and assign a new object ID to each object tracker 106.

[0052] The instantiated object tracker 106 can initialize the object locator 112 with its corresponding location from the object detector 104. This can include learning the visual features of the object associated with the location. For example, in the case where a correlation filter is used to learn the visual features of the object 160, the filter initializer 114 can initialize the correlation filter using the image data and / or sensor data corresponding to the location 202 in the image. In an example, the filter initializer can extract an image region from frame 0, which can at least include the location 202 (e.g., the region within the bounding box), and in different examples, include a larger region that can correspond to the size of the search region used by the object locator 112 to estimate the object location. The detected object location 202 can be used as the object location 138 of the track 132 in the state data of the object tracker 106 for frame 0. Other object trackers 106 can similarly use the locations 206, 210, and 214.

[0053] In frame 1, each object tracker 106 can analyze one or more portions of the image data and / or sensor data (e.g., of frame 1 and / or temporally related to frame 1) to estimate the location of an object (if any) in the field of view of the sensor 102. As shown, the object tracker 106 can estimate the location 220 of the object 160, the location 222 of the object 208, the location 224 of the object 212, and the location 226 of the object 216. As an example, to track the object 160, the object tracker 106 can define a search region at least partially based on the previous location of the object 160, and determine the estimated object location 220 by searching for the object 160 within the search region at least based on the visual features learned during initialization.

[0054] For example, the location of the search region can be based on the location 202 of the object 160 detected in frame 0. In the case where a correlation filter is used, the correlation filter can be applied to the search region to compare the locations within the search region with the learned visual features. The estimated object location 220 can be determined at least based on the correlation response of the correlation filter applied to the search region, and can be centered at the location of the peak of the correlation response or based on the location of the peak of the correlation response, or otherwise determined from it. The locations 222, 224, and 226 can be estimated similarly.

[0055] In the example of Frame 1, the detected object location may not be used to determine the object location of Frame 1. For example, the object detector 104 may not analyze Frame 1 to calculate the detected object location for association with the estimated object location. In this case, one or more of the estimated object locations 220, 222, 224, and 226 may be used as the object location (e.g., the bounding object) in the state data of the corresponding object tracker 106 for Frame 1, and / or the tracker terminator 126 may terminate or deactivate one or more of the object trackers 106. For example, the detected object location 220 may be used as the object location 140 of the track 132 in the state data of the object tracker 106 for Frame 0 and Object 160. The object location may also be used by the filter updater 116 of the object tracker 106 to update the relevant filter (e.g., using the image region based at least on the object location).

[0056] Frame 2 is an example where both the estimated object location and the detected object location are used to determine the object location of the frame. In different embodiments, the detected object location may or may not be used to determine the object location of any given frame. For example, the detected object location may be used per frame, may be used periodically (e.g., every N frames, where N is an integer), or may be used based on evaluating different criteria. When the detected object location is not used for a frame, the object detector 104 does not need to be in operation, thus saving computational resources. Similarly, the object tracker 106 may run every Z frames, where Z is an integer. For example, the object tracking system 100 may operate in the case of N = 2 and Z = 1 or N = 2 and Z = 2. The filter initializer 114 and / or the filter updater 116 may operate in a similar spanning manner to save computational resources.

[0057] For Frame 2, each object tracker 106 may determine the estimated locations 230, 232, 234, and 236 using the corresponding object locator 112 - similar to that already described for Frame 1. The object detector 104 may also determine the detected locations 240, 242, and 244, similar to that already described for Frame 1. The object matcher 118 may attempt to match the estimated locations 230, 232, 234, and 236 with the detected object locations and / or previously tracked objects among the locations 230, 242, and 244.

[0058] For example, the estimated position 230 may match the detected position 240, and the estimated position 232 may match the detected position 242. Thus, the position aggregator 122 may aggregate, combine, and / or merge the detected position 240 with the estimated position 232 to determine the position of the corresponding object in frame 2, thereby determining the position of the corresponding object in frame 2 for the corresponding object tracker 106. For example, the aggregated position may be used as the object position 142 of the track 132 in the state data of the object tracker 106 for frame 2 and object 160. The object position may also be used by the filter updater 116 of the object tracker 106 to update the relevant filter (e.g., using an image region based at least on the object position). Similarly, the position aggregator 122 may aggregate, combine, and / or merge the detected position 242 with the estimated position 232 to determine the position of the corresponding object in frame 2 for the corresponding object tracker 106. The object position may also be used by the filter updater 116 of the object tracker 106 to update the relevant filter (e.g., using an image region based at least on the object position).

[0059] The estimated positions 234 and 236 may not match any detected positions and / or tracked objects. As a result, the tracker terminator 126 may terminate and / or deactivate the tracking of the corresponding object. The detected position 244 may also not match any estimated positions and / or tracked objects. Thus, the tracker instantiator 124 may instantiate an object tracker 106 for a new object, as described herein. For example, the detected position 244 may be used by the filter initializer 114 to initialize the relevant filter of the object tracker 106.

[0060] As with respect to Figure 2 As can be seen, in the object tracking system 100, particularly in cases where many object trackers 106 and / or video streams are employed, a large number of image regions may be extracted, processed, and analyzed for object tracking. Similarly, a large number of relevant filters may need to be initialized, updated, and / or applied for object tracking. In different embodiments, the batch manager 110 may manage the batching of any combination of this various data in the object tracking system 100 to allow for effective and efficient parallelization of the processing performed by the object trackers 106. For example, batching may be used to parallelize one or more of the object locator 112, filter initializer 114, or filter updater 116 of the object trackers 106 for different objects.

[0061] Examples of batching

[0062] In various examples, the batch manager 110 may be used to extract image regions from a batch of one or more images. Now referring to Figure 3 , Figure 3FIG. is an illustration showing an example of a batch process that can be used to implement multi-object tracking according to some embodiments of the present disclosure. In some examples, the batch manager 110 can use the Figure 3 method for the object locator 112 of the object tracker 106, the filter initializer 114 of the object tracker 106, or the filter updater 116 of the object tracker 106, respectively, to parallelize the processing of those components across objects and / or video streams.

[0063] For example, the image regions of the filter initializer 114 of the object tracker 106 can be batched and processed, followed by the batching and processing of the image regions of the object locator 112 of the object tracker 106, followed by the batching and processing of the image regions of the filter updater 116 of the object tracker 106. For subsequent batch processing, this order can be repeated. In other examples, batches can be formed from image regions processed by any combination of these components, and the processing does not need to be sequential across each object tracker 106 (e.g., when processing a batch, the filter can be updated for one object tracker 106 while the object is being located for another object tracker 106).

[0064] Figure 3 FIGS. show batches 0, 1, 2, and 3, which can be generated by the batch manager 110, and each of which can contain one or more images and / or frames. In the example shown, the batch manager 110 forms each batch from multiple sources (e.g., Src 0, Src 1, Src 2). Each source can correspond to a video stream, sensor 102, camera, video feed, multi-object tracker, and / or tracked object, and can provide multiple images (e.g., frames) to the batch manager 110. In the example shown, each source corresponds to a respective video stream from a respective camera and provides a sequence of frames (e.g., as they are generated and / or available). For example, Src 0 can correspond to Figure 2 the video stream, and Src 1 and Src 2 can correspond to other video streams.

[0065] The batch manager 110 can generate each batch, for example, at least based on the time when image data is received from the corresponding source. For example, the batch manager 110 can batch the frames received within a time window (or, in some examples, before the time window expires). Thus, frame 0 from Src 0, frame 0 from Src 1, and frame 0 from Src 2 can each be received within the time window of batch 0. The time window for each batch can be the same or different. In some examples, the time window is configured such that the object tracking system 100 has completed the processing of the previous batch. For example, in the case where the time window is dynamic, the endpoints of the time window can be based at least on the completion of the processing of the previous batch and / or the completion of one or more intermediate processing steps of the processing pipeline of the previous batch.

[0066] The batch manager 110 may operate using a maximum batch size that may be based on, for example, the number of frames to be extracted from and / or processed in a batch (e.g., one per source), the number of sources, and / or the number of image regions. For example, the image regions to be extracted for a video stream and / or frame extraction may correspond to the number of locations to be processed using the object locator 112, filter initializer 114, and / or filter updater 116 of the object tracker 106. A frame may be excluded from a batch, for example, if the image region containing the frame would cause the number of image regions processed for the batch to exceed a threshold (e.g., which may be limited based on the available memory size).

[0067] For different batches, the maximum batch size may be the same or different. As shown, batch 0 may be a full batch because each source has provided frames within the time window. However, batch 1 is a partial batch because Src 1 has not provided a frame before the time window for the batch has expired. Batch 2 is also a partial batch because Src2 has not provided a frame before the time window for the batch has expired. Batch 3 is likewise a full batch because each source has frames available for processing within the time window. The batch manager 110 may process each frame it receives from a source or may discard one or more frames, for example, to maintain real-time object tracking.

[0068] In addition to forming batches of one or more images, the batch manager 110 may also manage the processing of each batch. Now referring to Figure 4 , Figure 4 is a diagram illustrating an example of processing a batch of one or more images according to some embodiments of the present disclosure. At 402, the processing of a batch may include extracting a batch of image regions from a batch of one or more images and scaling the extracted image regions to one or more template sizes. Extracting an image region from an image may include cropping the image region from the image, and the cropped image data may be scaled to the template size. The batch manager 110 may receive a list of bounding boxes and determine the corresponding image regions to be extracted and scaled for each bounding box. The cropped and scaled image regions may be arranged (e.g., consecutively) in a texture memory (e.g., texture cache) and may belong to a texture object or reference. The arrangement of the scaled image regions may be based at least on mapping the bounding box indices from the list of bounding boxes to the coordinates of one or more grids in a memory grid. Cropping and scaling may use normalized coordinates, bilinear interpolation, and image boundary handling that may be provided by the texture memory hardware. The image regions and / or the scaled image regions may have any suitable color format, such as NV12. Although bounding boxes are described herein, the bounding boxes may more generally be referred to as bounding shapes.

[0069] Figure 4An example showing the scaled image regions 402A and 402B is presented. The batch manager 110 can cause the scaled image regions 402A and 402B to be stored into a texture that may belong to a texture object. In the example shown, a single template size with width W and height H is used. For P image regions (or objects), at least WxHxP pixels of texture memory (e.g., Compute Unified Device Architecture texture memory) may be required for storage. However, any number of template sizes can be used for one or more image regions. Although the image regions can have various shapes and sizes in the image from which the image regions are extracted, the shown scaled image regions 402A and 402B are each of the template size (in some examples, the template size can be selected and / or configured to be less than or equal to the size of each of the image regions associated with the template size). In some embodiments, using texture memory to process batches allows for free scaling using hardware interpolation and free image boundary handling (e.g., image regions partially falling outside the image boundary can be filled).

[0070] In embodiments in which the object locator 112, filter initializer 114, and / or filter updater 116 of the object tracker 106 analyze image features, at 404, one or more feature channels of the image region can be generated from one or more of the scaled image regions. For example, the batch manager 110 can cause each scaled image region to be analyzed in parallel through the processing of the texture object to generate one or more corresponding feature channels. Figure 4 Three feature channels for the image region are shown by way of example. Thus, three feature regions 404A can be generated from the scaled image region 402A. The batch manager 110 can cause the texture containing the scaled image regions 402A and 402B to be processed in parallel to produce an image (e.g., in a new texture or an existing texture) that may contain at least the feature regions of each feature channel and the scaled image regions. In the example, each feature region of the scaled image region can have the same size as the template size of the scaled image region or can be a different template size. For M feature channels, at least MxWxHxP pixels of texture memory may be required for storage. In some embodiments, stacked composite images in the texture memory can be used to perform batch feature extraction. Feature images can be generated in the memory for each feature channel, where the feature regions are arranged based at least on mapping the indices of the scaled image regions from the stacked composite image to the coordinates of one or more grids in the memory grid of the memory.

[0071] At 406, the batch manager 110 may load textures from the texture memory into the device memory. For example, the texture memory may be off-chip and the device memory may be on-chip. The device memory may have one or more parallel processing units (PPUs), such as one or more GPUs. In different embodiments, the PPU may correspond to, for example, one or more of the GPUs 1508, GPU 1520, GPU 1584, GPU 1608, and / or the logic unit 1620 described herein. In some embodiments, the loading of the textures may be performed by one or more CPUs (such as CPU 1506, CPU 1518, CPU 1580, and / or CPU 1606). Loading the textures into the device memory may further include rearranging one or more of the scaled image regions and / or feature regions. For example, the feature region may be arranged by the scaled image region (e.g., an object) when copying data into the device memory. In some examples, the device memory may be pre-allocated at least based on the maximum number of objects and / or image regions per video stream and / or batch. The PPU cores may run only on the memory blocks in use. In some examples, memory blocks may be reserved for each tracked object and / or object tracker 106 and used batch by batch. When the tracker terminator 126 terminates the tracking of an object, the reserved memory may be released and made available for the objects and / or object trackers 106 instantiated by the tracker instantiator 124. Thus, the pre-allocated memory may be reused across batches, allowing for efficient and low-overhead memory management. The memory may be further allocated continuously for efficient batching. The feature region corresponds to various types of visual image features. As an example, the feature region may correspond to the grayscale representation of the image region, the histogram of oriented gradients (HOG), ColorName, etc.

[0072] At 408, the scaled image region and the extracted feature region may be processed by the PPU. For example, worker threads may operate on one or more scaled image regions and / or associated feature regions in parallel to perform any combination of the functions of the object locator 112, filter initializer 114, and / or filter updater 116 of the object tracker 106. This may result in an input for generating and / or processing the image regions of subsequent batches. Due to batching and scaling the image regions, the size and shape of the scaled image regions—and the correlation filters and correlation responses in the related filter-based methods—may be made more uniform. This may reduce the storage size and improve the processing efficiency while allowing for efficient and effective parallel analysis and processing of the image regions, such as using the threads of one or more graphics processing units (GPUs).

[0073] Now refer to Figure 5, each block of method 500 and other methods described herein include computational processes that may be executed using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory. The method may also be embodied as computer-usable instructions stored on a computer storage medium. By way of example only, the method may be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with other hosted services), or a plug-in to another product. Additionally, by way of example, a system for Figure 1 is described for the method. However, the method may alternatively or additionally be executed by any one system or any combination of systems, including but not limited to those described herein.

[0074] Figure 5 is a flow chart showing method 500 for batch processing search regions of an object tracker to determine an estimated object location according to some embodiments of the present disclosure. Method 500 includes, at block B502, extracting a batch of search regions from a batch of one or more images. For example, batch manager 110 may cause image data representing a batch of search regions corresponding to detection locations of objects (e.g., Figure 3 detection locations 202, 206, 210, and / or 214 of one or more of objects 160, 208, 212, or 216) in one or more images of one or more videos (e.g., Figure 2 batch 0, 1, or 2) to be extracted from image data representing a batch of one or more images of one or more videos.

[0075] Method 500 includes, at block B504, generating scaled search regions based at least on scaling the batch of search regions to one or more template sizes. For example, batch manager 110 may cause image data representing scaled search regions having a template size (e.g., Figure 4 scaled image regions 402A and 402B) to be generated from image data representing a batch of image regions. The generation may be based at least on scaling the batch of search regions to one or more template sizes. Batch manager 110 may also cause image data representing one or more features of the scaled search regions (e.g., Figure 4 feature region 404A of scaled image region 402A in

[0076] Method 500 includes, at block B506, determining an estimated object location within the scaled search regions. For example, batch manager 110 may cause the scaled search regions (and in some embodiments the feature regions) to be loaded into the PPU. Figure 1The object locator 112 of the object tracker 106 of one or more multi-object trackers can determine data representing an estimated object position within the scaled search region (and in some embodiments, from image data representing a feature region) from the image data representing the scaled search region (e.g., Figure 2 the estimated positions 220, 222, 224, and / or 226).

[0077] Method 500 includes, at block B508, generating an association between one or more estimated object positions and one or more objects. For example, the data associator 108 can associate one or more estimated object positions with one or more previously and / or newly tracked objects and / or trajectories (e.g., using object IDs). In some embodiments, the estimated object positions can be aggregated and / or fused with one or more other positions by the position aggregator 122 (e.g., as described with respect to Figure 2 frame 2) before being associated with an object.

[0078] Figure 6 is a flow chart of a method 600 for batching image regions of an object tracker to initialize or update a correlation filter according to some embodiments of the present disclosure. Method 600 includes, at block B602, extracting a batch of image regions from a batch of one or more images. For example, the batch manager 110 can cause the extraction of image data representing a batch of image regions corresponding to the detection positions (e.g., Figure 3 of the batch 0, 1, or 2) of one or more images representing one or more videos, of one or more objects (e.g., Figure 2 the detection positions 202, 206, 210, and / or 214 of one or more of the objects 160, 208, 212, or 216).

[0079] Method 600 includes, at block B604, generating scaled image regions based at least on scaling the batch of image regions to one or more template sizes. For example, the batch manager 110 can cause the generation of image data representing scaled image regions having a template size (e.g., Figure 4 the scaled image regions 402A and 402B) from the image data representing the batch of image regions. The generation can be based at least on scaling the batch of image regions to one or more template sizes. The batch manager 110 can also cause the generation of image data representing one or more features of the scaled image regions (e.g., Figure 4 the feature region 404A of the scaled image region 402A in ) from the image data representing the scaled image regions. The features of the scaled image regions can be represented as feature regions having the template size of the scaled image regions.

[0080] Method 600 includes, at block B606, determining a correlation filter from a scaled image region. For example, batch manager 110 may cause the scaled image region (and in some embodiments, the feature region) to be loaded into the PPU. Figure 1 The filter initializer 114 of object tracker 106 of one or more multi-object trackers of Figure 1 may determine data representing a correlation filter (which, in some embodiments, may include one or more feature channels) from image data representing the scaled image region (and in some embodiments, from image data representing the feature region). Additionally or alternatively, for one or more image regions, Figure 1 The filter updater 116 of object tracker 106 of one or more multi-object trackers of Figure 1 may determine data representing an updated correlation filter (which, in some embodiments, may include one or more feature channels) from image data representing the scaled image region (and in some embodiments, from image data representing the feature region).

[0081] Method 600 includes, at block B608, using the correlation filter to determine one or more estimated object positions. For example, Figure 1 The object locator 112 of object tracker 106 of one or more multi-object trackers of Figure 1 may determine data representing one or more estimated object positions (e.g., Figure 2 estimated positions 220, 222, 224, and / or 216 of Figure 2 ) from data representing the correlation filter.

[0082] Figure 7 is a flow chart showing a method 700 for batch cropping and scaling a search region of an object tracker to determine an estimated object position according to some embodiments of the present disclosure. Method 700 includes, at block B702, cropping and scaling a batch of one or more images to generate a batch of scaled search regions. For example, batch manager 110 may cause image data representing a batch of one or more images (e.g., Figure 3 batches 0, 1, or 2 of Figure 3 ) of one or more videos to extract image data representing cropped and scaled search regions having one or more template sizes (e.g., Figure 4 scaled image regions 402A and 402B of Figure 4 ). Batch manager 110 may also cause image data representing one or more features of the scaled search region (e.g., Figure 4 feature region 404A of scaled image region 402A in Figure 4 ) to be generated from the image data representing the scaled search region. The features of the scaled search region may be represented as a feature region having the template size of the scaled search region.

[0083] Method 700 includes, at block B704, determining an estimated object position within a scaled search region. For example, the batch manager 110 may cause the scaled search region (and in some embodiments, the feature region) to be loaded into the PPU. Figure 1 The object locator 112 of the object tracker 106 of one or more multi-object trackers may determine data representing an estimated object position within the scaled search region (and in some embodiments, from image data representing the feature region) from the image data representing the scaled search region (e.g., Figure 2 the estimated positions 220, 222, 224, and / or 226 of

[0084] Method 700 includes, at block B706, generating an assignment between one or more estimated object positions and one or more object IDs. For example, the data associator 108 may assign one or more estimated object positions to the object IDs of existing and / or newly tracked objects and / or trajectories (e.g., using object IDs). In some embodiments, the estimated object positions may be aggregated and / or fused with one or more other positions by the position aggregator 122 (e.g., as described with respect to Figure 2 frame 2 of

[0085] Example of data association using correlation response values

[0086] Aspects of the present disclosure provide data association in object tracking at least in part based on correlation response values. These methods may be implemented by Figure 1 the object tracking system 100 or a different object tracking system, which may employ object tracking techniques different from those of the object tracking system 100. The disclosed methods may enable data association to be performed between an estimated object position and one or more other positions based on visual similarity without generating additional correlation responses and / or features of the estimated object position and / or other positions.

[0087] Data association may be used to link estimated object positions from an object tracker to positions within a frame and / or across frames (e.g., detected object positions and / or estimated object positions). For example, as described herein, Figure 1 the object matcher 118 of Figure 2The estimated position 230). For example, the object locator 112 may apply a correlation filter to the search region to determine the estimated object position. Thus, a correlation response for the estimated object position (which may include one or more channels, one or more of which may include feature channels) may have been generated. Embodiments of the present disclosure may enable this correlation response to be reused in data association. For example, the correlation response may be located in memory when used to determine the estimated object position, and the correlation response may be looked up in memory (e.g., at the same location used for localization) for data association. Thus, the correlation response that has already been computed for localization can be used to associate the estimated object position with one or more other positions, and there is no need to generate additional correlation responses and / or visual features for performing data association (although they may be generated in some embodiments).

[0088] According to the disclosed method, the value of the correlation response of the estimated object position corresponding to another position (e.g., of a detected bounding box) can be used to compare that another position with the estimated object position. Based at least on the value of the correlation response (e.g., the sum of regions and / or a single value or aggregation of values of the correlation response channels), that another position may or may not be associated with the estimated object position. In some embodiments, the comparison may also be based at least on the value of the correlation response corresponding to the estimated object position (e.g., the sum of regions and / or a single value or aggregation of values of the correlation response channels). For example, the comparison may include determining a ratio between the value associated with another position and the value associated with the estimated object position. Using this method for different correlation responses and positions can be used to normalize this factor when comparing different positions.

[0089] In various examples, the value of the correlation response for estimating the object position may be based at least on the peak correlation response value of the correlation response (and / or one value or multiple values used by the object locator 112 to select the estimated object position). In some examples, the peak correlation response value may be at the center of the bounding box corresponding to the estimated object position. The value of the correlation response for another position may be based at least on the correlation response value of the correlation response at that another position, such as at the center of the bounding box corresponding to that position or otherwise within the bounding box.

[0090] In various examples, comparisons between values can be used to compute a confidence value that quantifies a level of similarity between locations and / or a likelihood that the locations correspond to the same object. Confidence values (which may also be referred to as confidence scores) between different locations can be used to associate the locations with each other. For example, any suitable matching algorithm can be used to match locations (e.g., an estimated object location and a detected object location) at least based on the confidence value. Examples of suitable matching algorithms include global matching algorithms, greedy algorithms, or non-greedy algorithms such as those using the Hungarian method. For example, a bipartite graph can be formed that relates locations (e.g., a set of estimated object locations and a set of detected object locations) to weights corresponding to confidence values. The bipartite graph can be formed, for example, by minimizing the cost between location nodes. The association between locations can then correspond to the linked nodes. Additionally or alternatively, in some embodiments, locations can be associated with each other at least based on corresponding confidence values that exceed a threshold.

[0091] In addition to or instead of the correlation response value, the confidence score for associating locations can also be based on other factors. For example, the confidence value of a location can be at least based on the intersection over union (IoU) between locations (e.g., between bounding boxes). In some examples, the confidence score of a location can be at least based on spatio-temporal data. For example, an estimated object location can be computed from a previous location of the object (e.g., in a previous frame). The aggregation of the locations being compared (e.g., between bounding boxes) can be computed by the location aggregator 122 and compared with the previous location as a factor in computing the confidence score. In some embodiments, the confidence score can be at least based on the IoU between the aggregation of the locations and the previous location (e.g., between bounding boxes). A lower IoU can correspond to a higher confidence score. Another factor for the confidence score can be at least based on the speed between the aggregation of the locations and the previous location. The speed can be computed based on the distance between bounding boxes and / or based on measured or inferred speed information. A lower speed can correspond to a higher confidence score. Although the different factors are described as being used to compute the confidence score, additionally or alternatively, any of these factors can be used as a threshold to prevent a match when the corresponding value exceeds the threshold (e.g., when the speed is greater than the threshold).

[0092] In one or more embodiments, during a first stage, a correlation filter can be determined by the filter initializer 114 or the filter updater 116, and then the correlation filter can be applied by the object locator 112 in a next frame to obtain a correlation response. The peak correlation response can correspond to an estimated location of the object being tracked, which is located and applied to the next frame.

[0093] The correlation response generated by the correlation filter can cover the entire search area. If the detection location (e.g., bounding box) from the object detector 104 for data association is within the search area, instead of extracting new features from the detection location, the correlation response value associated with the detection location (e.g., the center of the bounding box) can be used. The correlation response value can be based on the correlation filter learned for the object being tracked and can indicate a confidence level from the perspective of the object tracker 106. If there are multiple object trackers 106 whose search areas include the same detection location, then there can be multiple correlation values corresponding to the same detection location (e.g., bounding box).

[0094] Now refer to Figure 8 , Figure 8 which is a diagram showing an example of associating positions based on correlation response values according to some embodiments of the present disclosure. By way of example, it is described for data association between an estimated object position and a detected object position Figure 8 . However, the positions being compared can generally be any positions associated with the object.

[0095] Figure 8 The example of Figure 2 shows the correlation response 802 of the estimated position 230 of Figure 8 and also shows Figure 2 the correlation response 806 of the estimated position 232 of Figure 2 and the image regions 804 and 808 that can be used to generate the correlation responses 802 and 806 (although it can be scaled, it is not shown as scaled). The object matcher 118 of the data associator 108 can use the correlation responses 802 and 806 and the correlation responses of the estimated positions 234 and 236 in

[0096] As described herein, these confidence scores can be based at least on the values of the correlation responses (e.g., used by object locator 112 to determine the estimated location). For example, a table 810 is shown, in which each cell represents a confidence score between a corresponding estimated location and a detected location. The confidence score 0.9 can be calculated based at least on the values 840 and 830 of the correlation response 802. The value 830 can be at the center of the bounding box of the estimated location 230 and / or can be the peak correlation response value. The value 840 can be at the center of the bounding box of the detected location 240 or can otherwise correspond to the detected location 240. These values can be a single value from one or more channels of the correlation response or a combination of values therefrom (e.g., statistical). In some examples, the values can be derived (e.g., statistically) from the values in a region of one or more channels (such as within the corresponding bounding box). Since the values 830 and 840 are similar, the ratio between the values can be high, resulting in a high confidence score.

[0097] Similarly, the confidence score 0.3 can be calculated based at least on the values 842 and 832 of the correlation response 802. The value 832 can be at the center of the bounding box of the estimated location 232 and / or can be the peak correlation response value. The value 842 can be at the center of the bounding box of the detected location 240 or can otherwise correspond to the detected location 240. Since the values 832 and 842 are not as similar as the values 830 and 840, the ratio between the values can be lower, resulting in a lower confidence score. Thus, the detected location 240 can be matched with the estimated location 230 rather than with the estimated location 0.3. Confidence scores can be calculated similarly between each combination of detected and estimated locations, and the object matcher 118 can use these confidence scores for object matching, as described herein (e.g., using the Hungarian matching or greedy matching algorithm with a bipartite graph). In some embodiments, the confidence scores can be calculated without extracting visual features from the detected locations 240, 242, and 244 of frame 2, thereby reducing the computation and storage required for data association.

[0098] Now refer to Figure 9 , Figure 9 is a flow chart showing a method 900 for associating locations based at least on correlation response values according to some embodiments of the present disclosure. Method 900 includes, at block B902, determining an estimated object location based at least on the correlation response. For example, the object locator 112 of the object tracker 106 can determine the estimated location 230 based at least on the correlation response 802 of the correlation filter generated using the filter initializer 114 and / or the filter updater 116.

[0099] Method 900 includes, at block B904, determining at least one value of a correlation response corresponding to a detected object location. For example, object matcher 118 may determine value 840 of correlation response 802 corresponding to detected location 240.

[0100] Method 900 includes, at block B906, associating the detected object location with an estimated object location based on at least one value of the correlation response. For example, object matcher 118 may use the at least one value to calculate a confidence score (e.g., Figure 8 0.9 in ), and associate detected location 240 with estimated location 230 based at least on the confidence score. Location aggregator may aggregate estimated location 230 and detected location 240 based on the association, and the aggregated location may be assigned to an object ID associated with object tracker 106 of frame 2.

[0101] Example of using focus-plus-windowing for correlation filter learning

[0102] Aspects of the present disclosure provide, in part, for correlation filter learning in object tracking based at least on focus-plus-windowing. These methods may be implemented by Figure 1 object tracking system 100 or a different object tracking system, which may employ object tracking techniques different from those of object tracking system 100. The disclosed approach may enable correlation filter learning while focusing on a target region without target segmentation and without excluding background from training.

[0103] As described herein, filter initializer 114 and / or filter updater 116 may cause a correlation filter that produces a correlation response based on a target in an image region to learn. Object locator 112 may use the correlation filter to locate the target in a search region based at least on the correlation response, which may correspond to an estimated object location. The size of the search region may present certain tradeoffs. A large search region may allow object tracker 106 to track an object target even with large displacements between frames. However, a correlation filter that learns using a large search region will include more background. This may cause the filter to inadvertently and undesirably learn to detect and / or track background rather than the target object. Conversely, when using a smaller search region, less background is learned, but at the cost of an increased frequency of tracking failure when the object being tracked undergoes large displacements across consecutive frames.

[0104] A conventional method segments a target object from an image and uses the segmented target to exclude the background when training a correlation filter. Since the background is excluded from learning, this method can work equally well for large or small search regions. However, this method is not without problems. For example, in order for segmentation to work well, complex segmentation algorithms are required. Object segmentation itself is typically a computationally intensive task, and thus the conventional method uses an efficient / simple segmentation algorithm to fit within the computational budget. Specifically, the conventional method uses a Markov Random Field (MRF) on color likelihoods for segmentation. This may compromise the quality of the segmentation, resulting in parts of the target object being excluded from learning or parts of the background being included in learning. Additionally, even with high-quality segmentation, learning only from the target object can result in a large number of false positives, especially when the background is cluttered or has similar color components.

[0105] Additional aspects of the present disclosure provide for learning a correlation filter in object tracking based at least on focus windowing. When learning a correlation filter from an image region, a focus window can be applied to the image region, the focus window blurring the background of the target object, where the blurring is enhanced based on the distance from the target. Applying the focus window to the image can refer to applying the focus window to one or more channels of the image, such as one or more color channels or feature channels extracted from the image (e.g., from one or more of the color channels). Embodiments of focus windowing can be considered a crude, approximate object segmentation.

[0106] Now refer to Figure 10 , Figure 10 which is a diagram showing an example of applying a focus window to an image to use the focus window for learning a correlation filter according to some embodiments of the present disclosure. Figure 10 A search region 1000 is shown that can be used for learning a correlation filter for a target object 1002. Figure 10 A feature region 1004 is also shown, which can be a feature channel of the search region 1000 extracted from the search region 1000. In this example, the focus windowing can be applied to at least the feature region 1004, resulting in a focused feature region 1006. The focus windowing can be similarly applied to one or more other channels of the search region 1000. As can be seen, the focused feature region 1006 reduces learning from the background while still allowing the background to provide learning of the context around the target object. Thus, the search region 1000 can be made larger without the risk of overlearning the background.

[0107] In some embodiments, a focus windowing may be applied to each image region using a blurring filter, such as a Gaussian filter. The blurring filter may be configured to increase blurring based on the pixel distance from a target object (e.g., estimated object location) and / or the center of the image region. In some embodiments, the same blurring filter with the same set of blurring parameters may be applied to each image region to enable the correlation filter to learn from that region. In other embodiments, the blurring parameters of the blurring filter may be dynamically adjusted, e.g., based on image region attributes. For example, the size of the blurring filter may be adjusted (e.g., to cover the entire region) at least based on the size of the image region. Additionally or alternatively, the width, height, slope speed, and / or other dimensions or characteristics of the impulse response of the blurring filter may be adjusted at least based on the size of the image region and / or the target object. Although segmentation of the target object need not be performed, in some embodiments, segmentation and / or bounding box detection of the target object may be derived from the image region and used to determine and / or adjust one or more parameters of the blurring filter (e.g., one or more dimensions of the impulse response).

[0108] Now referring to Figure 11 , Figure 11 FIG. is a flow chart of a method 1100 for applying focus windowing to an image to enable a correlation filter to learn using the focus windowing according to some embodiments of the present disclosure. The method 1100 includes, at block B1102, determining an image region based at least on a location associated with an object. For example, a filter initializer 114 or a filter updater 116 may determine Figure 10 a search region 1000 based at least on a detected object location from an object detector 104, an estimated object location from an object locator 112, and / or an aggregated location from a location aggregator 122 associated with a target object 1002.

[0109] At block B1104, the method 1100 includes generating a focused image region based at least on applying a focus window to the image region. For example, a filter initializer 114 or a filter updater 116 may apply a focus windowing to one or more channels of Figure 10 the search region 1000 to generate a focused image region that includes the focus windowing in one or more channels. A focused feature region 1006 is an example of a feature channel of the focused image region corresponding to the search region 1000.

[0110] At block B1106, method 1100 includes making a correlation filter learn based at least on a focused image region. For example, filter initializer 114 or filter updater 116 can make the correlation filter of target object 1002 learn from at least focused feature region 1006 and / or other channels of the focused image regions generated from search region 1000, which may or may not include focused windowing. In some examples, the correlation filter is a multi-channel correlation filter, and each channel can learn from one or more corresponding channels of the focused image region (e.g., the histogram of oriented gradients channel of the correlation filter can learn from the histogram of oriented gradients channel of the focused image region). In different embodiments, the channel weights of the correlation filter can be calculated using the per-channel contributions to the correlation response.

[0111] Example of making a correlation filter learn using an occlusion map

[0112] The present disclosure also provides making a correlation filter learn in object tracking based at least on an occlusion map. These methods can be implemented by Figure 1 object tracking system 100 or a different object tracking system, which may employ different object tracking techniques than object tracking system 100. The disclosed methods can enable making the correlation filter learn while reducing and / or eliminating learning from occlusions.

[0113] Conventional methods of making a correlation filter of a target object learn do not consider which pixels are part of the target object and which are part of an occlusion. When there is partial / complete occlusion of the target object, such methods increase the risk of including the background in the target modeling / learning, even when segmentation or focused windowing is applied.

[0114] According to aspects of the present disclosure, when making a correlation filter learn from an image region, an occlusion map can be applied to mask, exclude, and / or blur the occluded image regions of the target object. The correlation filter can learn from the modified image, thereby reducing or eliminating learning from occlusions while still allowing learning from the exposed portions of the target object. For example, the occlusion map can be used to detect partial occlusions and to exclude those pixels when using a moving average (such as using equation (1)) or other temporal learning algorithms to update the model. Pixels in the target object region and pixels in the background region of the target object have different characteristics, but the moving average scheme treats them the same. In the proposed method, the background learned by GMM background can be directly used as the learned target model, which means that considering the variance of pixel values, the target pixel values can be estimated from GMM.

[0115] Some embodiments may include using an adaptive GMM history to train a GMM based on the target object state. When the target object state is stable (e.g., non-occluded), the appearance of the target object may be relatively static, resulting in a high peak correlation response with little intensity variation. Thus, for learning purposes, the GMM history may be set high. When the target object state is partially occluded, this may result in a lower peak correlation response, possibly with higher sidelobes (over a short time period). Again, the GMM history may be set high, but may not be used to update the target appearance model. When the target state includes rapid changes, this may result in lower correlation peaks (over a longer time period). In these cases, the GMM history may be set low to quickly adopt the latest changes.

[0116] Now refer to Figure 12A and Figure 12B , Figure 12A FIG. is an illustration showing an example of an image region 1200, an occlusion map 1202 of the image region 1200, and a target model 1204 of a correlation filter learned using the image region 1200 and the occlusion map 1202, in accordance with some embodiments of the present disclosure. Figure 12B is an illustration showing an example of Figure 12A the correlation response 1206 of a correlation filter and an estimated object position 1208 determined using the correlation filter, in accordance with some embodiments of the present disclosure.

[0117] The occlusion map 1202 can be generated from the image region 1200, which can be, for example, the search region used by the filter initializer 114 or the filter updater 116 to enable the target model 1204 of the relevant filter to learn. The occlusion map 1202 is shown as having an identified occlusion region 1210. In different examples, the occlusion map 1202 can be applied to the image region 1200 (e.g., one or more channels of an image), such as by masking, blurring, or otherwise adjusting the image region 1200 and / or the relevant filter learning model to reduce or eliminate learning from one or more of the occlusion regions 1210. In various embodiments, the occlusion map 1202 can be used to create a modified version of the image region 1200, which the filter initializer 114 or the filter updater 116 can use to enable the target model 1204 of the relevant filter to learn. In embodiments that also employ focus windowing, the image region 1200 can be further modified using focus windowing (before or after adjustment based on the occlusion map 1202). In some examples, the occlusion map 1202 can be represented by the output data of a machine learning model trained to generate the occlusion map 1202. The occlusion map 1202 can optionally be further processed and applied to the image region 1200 as a mask. For example, when applying the occlusion map 1202 to the image region 1200, one or more of the occlusion regions 1210 can be removed, merged, combined, or otherwise adjusted (e.g., at least based on the distance from the center of the image region 1200 or other locations of the target object 1212).

[0118] The occlusion map 1202 can be generated using a machine learning model, such as a Gaussian mixture model (GMM) or other MLM, that is trained using the target object 1212 as the background over multiple frames (e.g., using the image region used for relevant filter learning) such that the occlusion is detected as the foreground. For example, the machine learning model can be trained in parallel with the relevant filter of the target object 1212 from the same source image (e.g., the image region 1200).

[0119] Generate an occlusion map for each frame and / or based on occlusion detection. The proposed method can be used to minimize the damage to the target model 1204 caused by occlusions in the image region 1200. In some embodiments, as long as the occlusion region 1210 of the target object 1212 is detected, the detection of occlusion regions in the non-target region may be irrelevant because the pixels in the non-target region change as the target object 1212 moves over multiple frames. Using the disclosed method, the target model 1204 can remain undamaged as the target object 121 moves past the occluder 1214 over multiple frames. Thus, the correlation response 1206 can reliably indicate the estimated object position 1208. For example, as shown, by considering the occluder 1214, the shape of the peak of the correlation response can become slightly elongated in the horizontal direction, forming an elliptical shape that follows the target object 1212. In contrast, without considering the occluder 1214, the peak of the correlation response may be located on the occluder 1214 at the center of the image, causing the object tracker to get stuck.

[0120] Now refer to Figure 13 , Figure 13 FIG. is a flowchart of a method 1300 for applying an occlusion map to an image for correlation filter learning according to some embodiments of the present disclosure. Method 1300 includes, at block B1302, determining an image region based on a position associated with an object. For example, the filter initializer 114 or the filter updater 116 can determine the image region 1200 of FIG. 12 based at least on the detected object position from the object detector 104, the estimated object position from the object locator 112, and / or the aggregated position from the position aggregator 122 associated with the target object 1002.

[0121] At block B1304, method 1300 includes generating an occlusion map associated with the image region. For example, the filter initializer 114 or the filter updater 116 can use a machine learning model (such as, a GMM) to generate the occlusion map 1202 of FIG. 12.

[0122] At block B1306, method 1300 includes causing a correlation filter to learn based at least on the occlusion map and the image region. For example, the filter initializer 114 or the filter updater 116 can use the occlusion map 1202 to cause the target model 1204 of FIG. 12 to learn from the image region 1200 to exclude, remove, and / or ignore learning from the pixels corresponding to one or more of the occlusion regions 1210.

[0123] Example considering multiple mode correlation response filters

[0124] For different reasons, multiple modes may appear in a correlation response filter, such as in the case where the correlation filter is applied to an area that occludes a target object, or in the case where there are objects with a similar appearance nearby. Now refer to Figure 14 , Figure 14 which is a diagram illustrating an example of a correlation response 1400 having multiple modes according to some embodiments of the present disclosure. The correlation response 1400 includes a mode 1402 and a mode 1404, which may be caused, for example, by an occluder 1214 to Figure 12B the target object 1212.

[0125] Aspects of the present invention provide methods that can be used to better estimate the position of an object when there are multiple modes in a correlation response. In different embodiments, a particle filter may be applied to the correlation response to determine and / or select a peak correlation response value of the correlation response. The particle filter may be based at least on an expected response function, where the expected response function has a single node. In some embodiments, the response function may correspond to the function of a blur filter applied during focus windowing. In any example, the particle filter may be based on a Gaussian response function having a single node. When there are multiple modes, the modes may be fitted to an expected distribution and the fitted distribution may define an estimated object position.

[0126] Example autonomous vehicle

[0127] Figure 15AIllustrated is an example autonomous vehicle 1500 in accordance with some embodiments of the present disclosure. The autonomous vehicle 1500 (alternatively, referred to herein as "vehicle 1500") can include, but is not limited to, passenger vehicles such as cars, trucks, buses, first responder vehicles, shuttle vehicles, electric or electric bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, underwater vehicles, drones, and / or another type of vehicle (e.g., a driverless and / or one that accommodates one or more passengers). Autonomous vehicles are generally described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016 - 201806, issued June 15, 2018; Standard No. J3016 - 201609, issued September 30, 2016; and prior and future versions of this standard). Vehicle 1500 may be capable of implementing functions that conform to one or more of Levels 3 - 5 of the autonomous driving levels. For example, depending on the embodiment, vehicle 1500 may be capable of implementing conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).

[0128] Vehicle 1500 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. Vehicle 1500 may include a propulsion system 1550, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. The propulsion system 1550 may be connected to the driveline of vehicle 1500, which may include a transmission, in order to effect the propulsion of vehicle 1500. The propulsion system 1550 may be controlled in response to receiving a signal from the throttle / accelerator 1552.

[0129] A steering system 1554, which may include a steering wheel, may be used to steer vehicle 1500 (e.g., along a desired path or route) while the propulsion system 1550 is operating (e.g., while the vehicle is in motion). The steering system 1554 may receive a signal from a steering actuator 1556. For fully autonomous (Level 5) functionality, the steering wheel may be optional.

[0130] A brake sensor system 1546 may be used to operate the vehicle brakes in response to receiving a signal from a brake actuator 1548 and / or a brake sensor.

[0131] may include one or more system - on - chips (SoCs) 1504 (Figure 15C ) and / or one or more controllers 1536 of one or more GPUs may provide signals (e.g., signals representing commands) to one or more components and / or systems of the vehicle 1500. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 1548, to operate the steering system 1554 via one or more steering actuators 1556, and to operate the propulsion system 1550 via one or more throttle / accelerators 1552. One or more controllers 1536 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 1500. One or more controllers 1536 may include a first controller 1536 for autonomous driving functions, a second controller 1536 for functional safety functions, a third controller 1536 for artificial intelligence functions (e.g., computer vision), a fourth controller 1536 for infotainment functions, a fifth controller 1536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1536 may handle two or more of the above functions, two or more controllers 1536 may handle a single function, and / or any combination thereof.

[0132] One or more controllers 1536 may provide signals for controlling one or more components and / or systems of the vehicle 1500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example and without limitation, global navigation satellite system sensors 1558 (e.g., global positioning system sensors), RADAR sensors 1560, ultrasonic sensors 1562, LIDAR sensors 1564, inertial measurement unit (IMU) sensors 1566 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1596, stereo cameras 1568, wide-angle cameras 1570 (e.g., fisheye cameras), infrared cameras 1572, surround cameras 1574 (e.g., 360-degree cameras), remote and / or mid-range cameras 1598, speed sensors 1544 (e.g., for measuring the rate of the vehicle 1500), vibration sensors 1542, steering sensors 1540, brake sensors (e.g., as part of a brake sensor system 1546), and / or other sensor types.

[0133] One or more of the controllers 1536 may receive inputs (e.g., represented by input data) from the instrument cluster 1532 of the vehicle 1500 and provide outputs (e.g., represented by output data, display data, etc.) via the human-machine interface (HMI) display 1534, an audible annunciator, a speaker, and / or via other components of the vehicle 1500. These outputs may include information such as vehicle speed, rate, time, map data (e.g., Figure 15C the HD map 1522), location data (e.g., the location of the vehicle 1500 on a map, for example), direction, the location of other vehicles (e.g., occupancy grid), and information about objects and object states as perceived by the controller 1536, etc. For example, the HMI display 1534 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).

[0134] The vehicle 1500 also includes a network interface 1524 that may communicate over one or more networks using one or more wireless antennas 1526 and / or a modem. For example, the network interface 1524 may be capable of communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. One or more wireless antennas 1526 may also enable communication between objects (e.g., vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc. and / or one or more low-power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.

[0135] Figure 15B For an example of the camera positions and fields of view of an example autonomous vehicle 1500 in accordance with some embodiments of the present disclosure. The cameras and their respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on the vehicle 1500. Figure 15A For an example of the camera positions and fields of view of an example autonomous vehicle 1500 in accordance with some embodiments of the present disclosure. The cameras and their respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on the vehicle 1500.

[0136] The camera types for a camera can include, but are not limited to, digital cameras that can be adapted to be used with components and / or systems of a vehicle 1500. The camera can operate under an Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type can have any image capture rate, such as 60 frames per second (fps), 1520 fps, 240 fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array can include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as those with RCCC, RCCB, and / or RBGC color filter arrays can be used in an effort to increase light sensitivity.

[0137] In some examples, one or more of the cameras can be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can record and provide image data (e.g., video) simultaneously.

[0138] One or more of the cameras can be mounted in a mounting assembly such as a custom-designed (3-D printed) component to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the camera. Regarding the wing mirror mounting assembly, the wing mirror assembly can be custom 3-D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0139] A camera (e.g., a front camera) having a field of view that includes an environmental portion in front of the vehicle 1500 can be used for surround view to help identify forward paths and obstacles and, with the help of one or more controllers 1536 and / or a control SoC, assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning ("LDW"), adaptive cruise control ("ACC"), and / or other functions such as traffic sign recognition.

[0140] A variety of cameras can be used in a front-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imager. Another example can be a wide-angle camera 1570, which can be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 15B only one wide-angle camera is illustrated in, any number of wide-angle cameras 1570 can be present on the vehicle 1500. Additionally, a long-range camera 1598 (e.g., a long-range stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range camera 1598 can also be used for object detection and classification and basic object tracking.

[0141] One or more stereo cameras 1568 can also be included in a front-facing configuration. The stereo camera 1568 can include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (FPGA) with an integrated CAN or Ethernet interface on a single chip. Such a unit can be used to generate a 3-D map of the vehicle environment, including distance estimates for all points in the image. An alternative stereo camera 1568 can include a compact stereo vision sensor that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1568 can be used in addition to or in place of those described herein.

[0142] A camera (e.g., a side-view camera) having a field of view that includes an environmental portion on the side of the vehicle 1500 can be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, a surround camera 1574 (e.g., as Figure 15BThe four surround cameras shown in can be placed on vehicle 1500. The surround cameras 1574 can include wide-angle cameras 1570, fisheye cameras, 360-degree cameras, and / or the like. By way of example, four fisheye cameras can be placed at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 1574 (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a forward camera) as the fourth surround camera.

[0143] A camera having a field of view that includes an environmental portion at the rear of vehicle 1500 (e.g., a rearview camera) can be used to assist with parking, surround viewing, rear collision warning, and creating and updating occupancy grids. A variety of cameras can be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 1598, stereo cameras 1568, infrared cameras 1572, etc.).

[0144] Figure 15C FIG. is a block diagram of an example system architecture for an example autonomous vehicle 1500 in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be omitted altogether. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory. Figure 15A FIG. shows each of the components, features, and systems of vehicle 1500 in FIG. as being connected via bus 1502. Bus 1502 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as the “CAN bus”). CAN can be a network within vehicle 1500 used to assist in controlling various features and functions of vehicle 1500, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.

[0145] Figure 15C FIG.

[0146] Although the bus 1502 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or alternatively to the CAN bus, FlexRay and / or Ethernet can be used. Further, although a single line is used to represent the bus 1502, this is not intended to be limiting. For example, any number of buses 1502 can be present, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1502 can be used to perform different functions and / or can be used for redundancy. For example, a first bus 1502 can be used for a collision avoidance function and a second bus 1502 can be used for drive control. In any example, each bus 1502 can communicate with any component of the vehicle 1500, and two or more buses 1502 can communicate with the same component. In some examples, each SoC 1504, each controller 1536, and / or each computer in the vehicle can have access to the same input data (e.g., input from sensors of the vehicle 1500) and can be connected to a common bus such as a CAN bus.

[0147] The vehicle 1500 can include one or more controllers 1536, such as those described herein with respect to Figure 15A The controllers 1536 can be used for a variety of functions. The controllers 1536 can be coupled to any other different components and systems of the vehicle 1500 and can be used for the control of the vehicle 1500, the artificial intelligence of the vehicle 1500, the infotainment for the vehicle 1500, and / or the like.

[0148] The vehicle 1500 can include one or more system-on-chips (SoC) 1504. The SoC 1504 can include a CPU 1506, a GPU 1508, a processor 1510, a cache 1512, an accelerator 1514, a data store 1516, and / or other components and features not shown. In a variety of platforms and systems, the SoC 1504 can be used to control the vehicle 1500. For example, one or more SoC 1504 can be combined with an HD map 1522 in a system (e.g., a system of the vehicle 1500), and the HD map can obtain map refreshes and / or updates via a network interface 1524 from one or more servers (e.g., Figure 15D one or more servers 1578) of.

[0149] The CPU 1506 may include a CPU cluster or a CPU complex (alternatively referred to herein as a "CCPLEX"). The CPU 1506 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 1506 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 1506 may include four dual-core clusters, each of which has a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 1506 (e.g., CCPLEX) may be configured to support simultaneous cluster operations such that any combination of the clusters of the CPU 1506 can be active at any given time.

[0150] The CPU 1506 may implement power management capabilities including one or more of the following: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 1506 may further implement an enhanced algorithm for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this work is offloaded to the microcode.

[0151] The GPU 1508 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 1508 may be programmable and efficient for parallel workloads. In some examples, the GPU 1508 may use an enhanced tensor instruction set. The GPU 1508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 1508 may include at least eight streaming microprocessors. The GPU 1508 may use a compute application programming interface (API). Additionally, the GPU 1508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0152] In automotive and embedded use cases, the GPU 1508 can be power optimized for best performance. For example, the GPU 1508 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be restrictive, and the GPU 1508 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads leveraging a mix of compute and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to allow for more fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0153] The GPU 1508 can include, in some examples, high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900GB / s and / or a 16GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.

[0154] The GPU 1508 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thus improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 1508 to directly access the CPU 1506 page tables. In such an example, when the GPU 1508 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 1506. In response, the CPU 1506 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 1508. In this way, the unified memory technology can allow a single unified virtual address space for the memory of both the CPU 1506 and the GPU 1508, thus simplifying GPU 1508 programming and porting applications to the GPU 1508.

[0155] In addition, the GPU 1508 may include an access counter that can track how frequently the GPU 1508 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.

[0156] The SoC 1504 may include any number of caches 1512, including those described herein. For example, the cache 1512 may include an L3 cache that is available to both the CPU 1506 and the GPU 1508 (e.g., that is connected to both the CPU 1506 and the GPU 1508). The cache 1512 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.

[0157] The SoC 1504 may include an arithmetic logic unit (ALU) that can be utilized when performing processing for any one of the various tasks or operations regarding the vehicle 1500 (such as processing a DNN). In addition, the SoC 1504 may include a floating-point unit (FPU) or other math co-processor or digital co-processor type for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs that are integrated as execution units within the CPU 1506 and / or the GPU 1508.

[0158] The SoC 1504 may include one or more accelerators 1514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 1504 may include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement the GPU 1508 and offload some of the tasks of the GPU 1508 (e.g., freeing up more cycles of the GPU 1508 for performing other tasks). As an example, the accelerator 1514 can be used for targeted workloads that are stable enough to be easily accelerated (such as perception, convolutional neural networks (CNNs), etc.). As used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0159] The accelerator 1514 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for executing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations, as well as inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processor functions.

[0160] The DLA may execute neural networks, especially CNNs, quickly and efficiently for any of a variety of functions on processed or unprocessed data, such as, for example, and without limitation: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner recognition using data from a camera sensor; and / or CNNs for security and / or safety-related events.

[0161] The DLA may perform any function of the GPU 1508, and by using inference accelerators, for example, the designer may configure the DLA or the GPU 1508 for any function. For example, the designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 1508 and / or other accelerators 1514.

[0162] The accelerator 1514 (e.g., a hardware acceleration cluster) may include a Programmable Vision Accelerator (PVA), which may alternatively be referred to herein as a Computer Vision Accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example, and without limitation, any number of Reduced Instruction Set Computing (RISC) cores, Direct Memory Access (DMA), and / or any number of vector processors.

[0163] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0164] DMA can enable components of the PVA to access system memory independently of the CPU 1506. DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0165] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0166] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on sequential images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. In addition, a PVA may include additional error correcting code (ECC) memory to enhance overall system security.

[0167] Accelerator 1514 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high bandwidth, low latency SRAM for the accelerator 1514. In some examples, the on-chip memory may include at least 4MB SRAM consisting of, for example and without limitation, eight field configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high speed memory access to the PVA and the DLA. The backbone may include (e.g., using APB) an on-chip computer vision network that interconnects the PVA and the DLA to the memory.

[0168] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transfer. This type of interface may comply with the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0169] In some examples, the SoC 1504 can include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate real-time visualization simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulations, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.

[0170] The accelerator 1514 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving applications. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0171] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, algorithms based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3 - 5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0172] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.

[0173] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), outputs of an inertial measurement unit (IMU) sensor 1566 related to the vehicle 1500 orientation and distance, a 3D position estimate of an object obtained from the neural network and / or other sensors (such as a LIDAR sensor 1564 or a RADAR sensor 1560), etc.

[0174] The SoC 1504 can include one or more data stores 1516 (e.g., memory). The data store 1516 can be on-chip memory of the SoC 1504, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 1516 can be large enough in capacity to store multiple instances of the neural network. The data store 1512 can include an L2 or L3 cache 1512. References to the data store 1516 can include references to memory associated with the PVA, DLA, and / or other accelerators 1514 as described herein.

[0175] The SoC 1504 may include one or more processors 1510 (e.g., embedded processors). The processor 1510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 1504 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, assist system low power state transitions, SoC 1504 heat and temperature sensor management, and / or SoC 1504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1504 may use the ring oscillator to detect the temperature of the CPU 1506, GPU 1508, and / or accelerator 1514. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 1504 in a lower power state and / or place the vehicle 1500 in a driver safe stop mode (e.g., safely stop the vehicle 1500).

[0176] The processor 1510 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0177] The processor 1510 may also include an always-on processor engine, which may provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0178] The processor 1510 may also include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.

[0179] The processor 1510 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0180] The processor 1510 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0181] The processor 1510 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 1570, the surround camera 1574, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.

[0182] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where an image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.

[0183] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 1508 does not need to continuously render new surfaces, the video image compositor may be further used for user interface composition. Even when the GPU 1508 is powered on and active for 3D rendering, the video image compositor may be used to relieve the burden on the GPU 1508 to improve performance and responsiveness.

[0184] The SoC 1504 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 1504 may also include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.

[0185] SoC 1504 may also include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 1504 can be used to process data from cameras (connected via gigabit multimedia serial link and Ethernet), sensors (such as LIDAR sensor 1564, RADAR sensor 1560, etc. that can be connected via Ethernet), data from bus 1502 (such as the speed of vehicle 1500, steering wheel position, etc.), and data from GNSS sensor 1558 (connected via Ethernet or CAN bus). SoC 1504 may also include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and can be used to free CPU 1506 from routine data management tasks.

[0186] SoC 1504 can be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thus providing an integrated functional safety architecture for a platform that utilizes and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. SoC 1504 can be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with CPU 1506, GPU 1508, and data storage 1516, accelerator 1514 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0187] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.

[0188] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (such as GPU 1520) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may also include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.

[0189] As another example, as required for level 3, 4, or 5 driving, multiple neural networks can run simultaneously. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third deployed neural network on multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 1508.

[0190] In some examples, the CNNs for face recognition and owner recognition can use data from the camera sensors to recognize the presence of an authorized driver and / or owner of the vehicle 1500. The processing engine always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in the security mode, to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 1504 provides security against theft and / or carjacking.

[0191] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 1596 to detect and recognize emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect the sirens and manually extract features, the SoC 1504 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle operates as recognized by the GNSS sensor 1558. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 1562, the control program can be used to execute emergency vehicle safety routines to slow down the vehicle, drive it to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0192] The vehicle may include a CPU 1518 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 1504 via a high-speed interconnect (e.g., PCIe). The CPU 1518 may include, for example, an X86 processor. The CPU 1518 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 1504, and / or monitoring the status and health of the controller 1536 and / or the infotainment SoC 1530.

[0193] The vehicle 1500 may include a GPU 1520 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 1504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1520 may provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors of the vehicle 1500.

[0194] The vehicle 1500 may further include a network interface 1524, which may include one or more wireless antennas 1526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1524 may be used to enable wireless connections to the cloud (e.g., to the server 1578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the passenger's client device) via the Internet. To communicate with other vehicles, a direct link may be established between the two vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 1500 with information about vehicles approaching the vehicle 1500 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1500). This function may be part of the cooperative adaptive cruise control function of the vehicle 1500.

[0195] The network interface 1524 may include an SoC that provides modulation and demodulation functions and enables the controller 1536 to communicate via a wireless network. The network interface 1524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0196] Vehicle 1500 may also include a data store 1528 that may include off-chip (e.g., outside of SoC 1504) storage devices. The data store 1528 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.

[0197] Vehicle 1500 may also include a GNSS sensor 1558. The GNSS sensor 1558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1558 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0198] Vehicle 1500 may also include a RADAR sensor 1560. The RADAR sensor 1560 may be used by the vehicle 1500 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1560 may use CAN and / or a bus 1502 (e.g., to transmit data generated by the RADAR sensor 1560) for control as well as access to object tracking data and, in some examples, access Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 1560 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0199] The RADAR sensor 1560 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 1560 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surrounding environment of the vehicle 1500 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 1500.

[0200] As an example, a mid-range RADAR system can include ranges up to 1560 m (front) or 80 m (rear) and fields of view up to 42 degrees (front) or 1550 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots beside the vehicle.

[0201] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0202] Vehicle 1500 can also include ultrasonic sensors 1562. Ultrasonic sensors 1562 that can be placed in the front, rear, and / or sides of vehicle 1500 can be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 1562 can be used, and different ultrasonic sensors 1562 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 1562 can operate at the functional safety level of ASIL B.

[0203] Vehicle 1500 can include a LIDAR sensor 1564. The LIDAR sensor 1564 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1564 can be at the functional safety level of ASIL B. In some examples, vehicle 1500 can include multiple LIDAR sensors 1564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).

[0204] In some examples, the LIDAR sensor 1564 may be capable of providing a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 1564 can have, for example, an advertised range of approximately 1500 m, an accuracy of 2 cm - 3 cm, and support for a 1500 Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 1564 can be used. In such examples, the LIDAR sensor 1564 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 1500. In such examples, the LIDAR sensor 1564 can provide a field of view of up to 1520 degrees horizontally and 35 degrees vertically even for low-reflectivity objects, with a range of 200 m. The front-mounted LIDAR sensor 1564 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0205] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses the flash of a laser as the emission source to illuminate the vehicle's surrounding environment up to about 200 m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 1500. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a fan. The flash LIDAR device can use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR and because flash LIDAR is a solid-state device without moving parts, the LIDAR sensor 1564 can be less susceptible to motion blur, vibration, and / or shock.

[0206] The vehicle can also include an IMU sensor 1566. In some examples, the IMU sensor 1566 can be located at the center of the rear axle of the vehicle 1500. The IMU sensor 1566 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1566 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1566 can include an accelerometer, a gyroscope, and a magnetometer.

[0207] In some embodiments, the IMU sensor 1566 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1566 can enable the vehicle 1500 to estimate the heading by directly observing the change in velocity from the GPS to the IMU sensor 1566 and correlating them without the need for input from a magnetic sensor. In some examples, the IMU sensor 1566 and the GNSS sensor 1558 can be integrated into a single unit.

[0208] The vehicle can include a microphone 1596 placed in and / or around the vehicle 1500. Among other things, the microphone 1596 can be used for emergency vehicle detection and identification.

[0209] The vehicle may also include any number of camera types, including a stereo camera 1568, a wide-angle camera 1570, an infrared camera 1572, a surround camera 1574, a long-range and / or mid-range camera 1598, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 1500. The camera types used depend on the embodiment and the requirements of the vehicle 1500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1500. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 15A and Figure 15B is described in more detail.

[0210] The vehicle 1500 may also include a vibration sensor 1542. The vibration sensor 1542 can measure the vibration of components of the vehicle such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 1542 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-rotating axle).

[0211] The vehicle 1500 may include an ADAS system 1538. In some examples, the ADAS system 1538 may include a SoC. The ADAS system 1538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0212] The ACC system can use RADAR sensors 1560, LIDAR sensors 1564, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 1500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and recommends that the vehicle 1500 change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0213] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 1524 and / or the wireless antenna 1526 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., the vehicle immediately in front of vehicle 1500 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given the information of the vehicle in front of vehicle 1500, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.

[0214] The FCW system is designed to alert the driver to a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of, for example, audible, visual warnings, vibrations, and / or rapid braking pulses.

[0215] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.

[0216] The LDW system provides visual, audible, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 1500 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0217] The LKA system is a variant of the LDW system. If the vehicle 1500 starts to leave the lane, then the LKA system provides steering input or braking to correct the vehicle 1500.

[0218] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0219] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 1500 is in reverse. Some RCTW systems include AEB to ensure application of the vehicle brakes to avoid a crash. The RCTW system can use one or more rear RADAR sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0220] Conventional ADAS systems can be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether a safe condition truly exists and act accordingly. However, in an autonomous vehicle 1500, in the case of conflicting results, the vehicle 1500 itself must decide whether to heed the results from the primary computer or an auxiliary computer (e.g., the first controller 1536 or the second controller 1536). For example, in some embodiments, the ADAS system 1538 can be a backup and / or auxiliary computer for providing perception information to a redundant computer sanity module. The redundant computer sanity monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1538 can be provided to the supervisory MCU. If the outputs from the primary computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0221] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the direction of the host computer regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine the appropriate result.

[0222] The supervisory MCU may be configured to run a neural network that is trained and configured to determine, based on outputs from the host computer and the secondary computer, the conditions under which the secondary computer provides a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is in fact the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include components of the SoC 1504 and / or be included as a component of the SoC 1504.

[0223] In other examples, the ADAS system 1538 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, the diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0224] In some examples, the output of the ADAS system 1538 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 1538 indicates a forward collision warning due to an object being immediately in front, then the perception block can use that information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0225] The vehicle 1500 can also include an infotainment SoC 1530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 1530 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 1500. For example, the infotainment SoC 1530 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 1534, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 1530 can further be used to provide information (e.g., visual and / or auditory) to the vehicle's user, such as information from the ADAS system 1538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0226] The infotainment SoC 1530 can include GPU functionality. The infotainment SoC 1530 can communicate with other devices, systems, and / or components of the vehicle 1500 via a bus 1502 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1530 can be coupled to a supervisory MCU such that in the event of a failure of the main controller 1536 (e.g., the main and / or standby computer of the vehicle 1500), the GPU of the infotainment system can perform some self-driving functions. In such examples, the infotainment SoC 1530 can place the vehicle 1500 in a driver safe parking mode as described herein.

[0227] Vehicle 1500 may also include an instrument cluster 1532 (such as a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1532 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 1532 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine fault light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, and so on. In some examples, information may be displayed and / or shared between the infotainment SoC 1530 and the instrument cluster 1532. In other words, the instrument cluster 1532 may be included as part of the infotainment SoC 1530, or vice versa.

[0228] Figure 15D A system schematic diagram for communication between a cloud-based server and Figure 15A an exemplary autonomous vehicle 1500 according to some embodiments of the present disclosure. The system 1576 may include a server 1578, a network 1590, and vehicles including the vehicle 1500. The server 1578 may include multiple GPUs 1584(A)-1584(H) (collectively referred to herein as GPUs 1584), PCIe switches 1582(A)-1582(H) (collectively referred to herein as PCIe switches 1582), and / or CPUs 1580(A)-1580(B) (collectively referred to herein as CPUs 1580). The GPUs 1584, CPUs 1580, and PCIe switches may be interconnected by high-speed interconnects such as, for example, and without limitation, the NVLink interface 1588 developed by NVIDIA and / or PCIe connections 1586. In some examples, the GPUs 1584 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 1584 and the PCIe switches 1582 are connected via a PCIe interconnect. Although eight GPUs 1584, two CPUs 1580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 1578 may include any number of GPUs 1584, CPUs 1580, and / or PCIe switches. For example, each of the servers 1578 may include eight, sixteen, thirty-two, and / or more GPUs 1584.

[0229] Server 1578 can receive image data over network 1590 and from the vehicle, the image data representing an image showing an unexpected or changed road condition such as a recently started road work. Server 1578 can transmit neural network 1592, updated neural network 1592, and / or map information 1594 over network 1590 and to the vehicle, including information about traffic and road conditions. Updates to the map information 1594 can include updates to the HD map 1522, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 1592, updated neural network 1592, and / or map information 1594 can be represented and / or generated based on data received from new training and / or from any number of vehicles in the environment and / or experience from training performed at a data center (e.g., using server 1578 and / or other servers).

[0230] Server 1578 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by the vehicle, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). The training can be performed according to any one or more categories of machine learning techniques, including but not limited to, categories such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle over network 1590), and / or the machine learning model can be used by server 1578 to remotely monitor the vehicle.

[0231] In some examples, server 1578 can receive data from the vehicle and apply the data to the latest real-time neural network for real-time intelligent inference. Server 1578 can include a deep learning supercomputer powered by GPU 1584 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1578 can include a deep learning infrastructure of a data center powered only by a CPU.

[0232] The deep learning infrastructure of server 1578 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 1500. For example, the deep learning infrastructure can receive periodic updates from vehicle 1500, such as an image sequence and / or objects located in the image sequence that vehicle 1500 has localized (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 1500. If the results do not match and the infrastructure concludes that the AI in vehicle 1500 has failed, then server 1578 can transmit a signal to vehicle 1500 instructing the fail-safe computer in vehicle 1500 to take control, notify the passengers, and complete a safe parking operation.

[0233] For inference, server 1578 can include GPU 1584 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, a CPU, FPGA, and other processor-powered servers can be used for inference.

[0234] Example computing device

[0235] Figure 16 FIG. is a block diagram of an example computing device 1600 suitable for implementing some embodiments of the present disclosure. Computing device 1600 can include an interconnect system 1602 that directly or indirectly couples the following devices: memory 1604, one or more central processing units (CPUs) 1606, one or more graphics processing units (GPUs) 1608, a communication interface 1610, input / output (I / O) ports 1612, input / output components 1614, a power supply 1616, one or more presentation components 1618 (e.g., a display), and one or more logic units 1620.

[0236] Although Figure 16 the various boxes are shown as being connected via an interconnect system 1602 with lines, this is not intended to be restrictive and is for clarity only. For example, in some embodiments, a presentation component 1618, such as a display device, can be considered an I / O component 1614 (e.g., if the display is a touchscreen). As another example, CPU 1606 and / or GPU 1608 can include memory (e.g., memory 1604 can represent a storage device in addition to the memory of GPU 1608, CPU 1606, and / or other components). In other words, Figure 16The computing devices described are merely illustrative. No distinction is made among categories such as "workstations", "servers", "laptop computers", "desktop computers", "tablet computers", "client devices", "mobile devices", "handheld devices", "gaming consoles", "electronic control units (ECUs)", "virtual reality systems", and / or other device or system types, as all of these are considered within the scope of Figure 16 the computing devices.

[0237] The interconnect system 1602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1602 may include one or more types of buses or links, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 1606 may be directly connected to the memory 1604. Additionally, the CPU 1606 may be directly connected to the GPU 1608. In cases where there are direct or point-to-point connections between components, the interconnect system 1602 may include a PCIe link to perform the connection. In these examples, a PCI bus need not be included in the computing device 1600.

[0238] The memory 1604 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 1600. Computer-readable media can include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0239] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 1604 may store computer-readable instructions (e.g., which represent programs and / or program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies, CD-ROM, digital versatile disks (DVDs), or other optical disk storage devices, magnetic tape cassettes, magnetic tapes, magnetic disk storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1600. As used herein, computer storage media does not include signals per se.

[0240] A computer storage medium can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0241] The CPU 1606 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 1606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. The CPU 1606 can include any type of processor, and can include different types of processors depending on the type of computing device 1600 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 1600, the processor can be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary co-processors such as a math co-processor, the computing device 1600 can also include one or more CPUs 1606.

[0242] In addition to or in place of the CPU 1606, the GPU 1608 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1600 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 1608 can be an integrated GPU (e.g., integrated with one or more of the CPUs 1606) and / or one or more of the GPUs 1608 can be a discrete GPU. In an embodiment, one or more of the GPUs 1608 can be a coprocessor for one or more of the CPUs 1606. The GPU 1608 can be used by the computing device 1600 for rendering graphics (e.g., 3D graphics) or for performing general-purpose computing. For example, the GPU 1608 can be used for general-purpose computing on the GPU (GPGPU). The GPU 1608 can include hundreds or thousands of cores capable of simultaneously processing hundreds or thousands of software threads. The GPU 1608 can generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 1606 via a host interface). The GPU 1608 can include graphics memory, such as display memory, for storing pixel data or any other suitable data (such as GPGPU data). The display memory can be included as part of the memory 1604. The GPU 1608 can include two or more GPUs operating in parallel (e.g., via a link). The link can directly connect the GPUs (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 1608 can generate pixel data or GPGPU data for a different part of the output or for a different output (e.g., the first GPU for a first image and the second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0243] In addition to or instead of the CPU 1606 and / or the GPU 1608, the logic unit 1620 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1600 to execute one or more of the methods and / or processes described herein. In an embodiment, the CPU 1606, the GPU 1608, and / or the logic unit 1620 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 1620 can be part of and / or integrated in one or more of the CPU 1606 and / or the GPU 1608, and / or one or more of the logic units 1620 can be discrete components or otherwise external to the CPU 1606 and / or the GPU 1608. In an embodiment, one or more of the logic units 1620 can be a coprocessor of one or more of the CPU 1606 and / or the GPU 1608.

[0244] Examples of the logic unit 1620 include one or more processing cores and / or their components, such as tensor cores (TCs), tensor processing units (TPUs), pixel vision cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, etc.

[0245] The communication interface 1610 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 1600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1610 can include components and functions that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or InfiniBand communication), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0246] The I / O port 1612 can enable the computing device 1600 to be logically coupled to other devices including I / O components 1614, presentation components 1618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 1600. Exemplary I / O components 1614 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 1614 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 1600 (described in more detail below). The computing device 1600 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touch screen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 1600 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 1600 to render immersive augmented reality or virtual reality.

[0247] The power supply 1616 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 1616 can power the computing device 1600 so that the components of the computing device 1600 can operate.

[0248] The presentation component 1618 can include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 1618 can receive data from other components (e.g., the GPU 1608, the CPU 1606, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0249] Examples of suitable network environments

[0250] A network environment suitable for implementing embodiments of the present disclosure can include one or more client devices and / or servers. The client devices and / or servers can be implemented on one or more instances of the Figure 16 computing device 1600 (e.g., each device).

[0251] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. As an example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide a wireless connection.

[0252] A compatible network environment can include one or more peer-to-peer network environments (in which case, no server may be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functions described herein with respect to a server can be implemented on any number of client devices.

[0253] In at least one embodiment, the network environment can include one or more cloud-based network environments. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports software layers and / or one or more applications of an application layer. The software or application can respectively include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. One or more of the client devices can use web-based service software or applications. The framework layer can be, but is not limited to, a free and open-source software web application framework that can perform large-scale data processing (e.g., "big data") using a distributed file system, such as Apache Spark TM 。

[0254] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more of their parts) described herein. Any of these different functions can be distributed across multiple locations from a central or core server. If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization) or can be public (e.g., available to many organizations).

[0255] Client devices can include those described herein with respect to Figure 16At least some of the components, features, and functions of the example computing device 1600 described. By way of example, and not limitation, a client device may be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, spacecraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device.

[0256] The present disclosure may be described in the general context of machine - usable instructions or computer code, including computer - executable instructions such as program modules, that are executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general - purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0257] As used herein, the recitation of "and / or" with respect to two or more elements should be interpreted to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0258] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from those described herein or combinations of steps similar to those described in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is expressly described.

Claims

1. A computer-implemented method, comprising: identifying, from one or more first images of one or more videos, one or more image regions corresponding to one or more detection locations of one or more objects in the one or more videos; applying a focus window to the one or more image regions based at least on the identification, the focus window comprising: blurring the one or more regions corresponding to at least one background of the one or more objects in the one or more image regions based at least on a distance of the one or more regions from a target object in the one or more image regions; using the one or more image regions and updating one or more correlation filters based at least on the blurring; and generating, based at least on applying the one or more correlation filters to one or more search regions of one or more second images of the one or more videos, an estimated object location corresponding to the one or more search regions from the one or more correlation filters and based at least on the update.

2. The method according to claim 1, wherein the one or more image regions are at least based on a first search region of the one or more videos, the first search region being identified by tracking an object using a version of the correlation filter, and wherein determining the representation of the update comprises: Updating one or more of the versions of the one or more correlation filters.

3. The method according to claim 1, wherein the updating comprises: Initializing one or more of the correlation filters based at least on determining one or more of a newly detected object or a newly tracked object.

4. The method according to claim 1, wherein the blurring is performed using a blurring filter, and one or more characteristics of an impulse response of the blurring filter are adjusted based at least on a size of the target object.

5. The method according to claim 1, comprising: Generating an occlusion map from the one or more image regions, wherein each of the correlation filters is generated from one of the image regions in the one or more image regions using one of the occlusion maps in the occlusion map.

6. The method according to claim 1, further comprising: Extracting one or more feature channels of the one or more image regions from the one or more image regions, wherein the update additionally uses the one or more feature channels.

7. The method according to claim 1, further comprising: determining a first estimated object location using a correlation response of a version of one of the one or more correlation filters; and determining a confidence score for a detection location based at least on a correlation response value of the correlation response corresponding to one of the one or more detection locations of the one or more objects, wherein determining the one or more correlation filters comprises: updating a version of the correlation filter using a learning rate based at least on the confidence score.

8. A system, comprising: one or more processing units configured to perform operations comprising: determining one or more image regions corresponding to one or more locations of one or more objects in one or more first images; Apply one or more focus windows to the one or more image regions, at least based on the determination, the application including: making one or more regions corresponding to one or more backgrounds of one or more objects in the one or more image regions blurred at least based on a distance from a target object in the one or more image regions to generate one or more focused image regions; Learn one or more values of one or more correlation filters using the one or more focused image regions; and Generate one or more object positions at least based on applying the one or more correlation filters to the one or more second images.

9. The system according to claim 8, wherein the blurring increases the blurring at least based on a distance from the target object in the one or more image regions.

10. The system according to claim 8, wherein the blurring includes one or more of the following: Blurring one or more color channels of the one or more image regions; or Blurring one or more feature channels of the one or more image regions.

11. The system according to claim 8, wherein the blurring uses a blurring filter having one or more characteristics at least based on a size of the target object.

12. The system according to claim 8, wherein the one or more image regions include one or more scaled image regions scaled to a template size.

13. The system according to claim 8, wherein the operation comprises: Perform multi-object tracking across frames of one or more videos using the one or more object positions.

14. The system according to claim 8, wherein the system is included in at least one of the following: A control system for an autonomous machine or a semi-autonomous machine; A perception system for an autonomous machine or a semi-autonomous machine; A system for performing simulation operations; A system for performing real-time streaming; A system for performing deep learning operations; A system implemented using an edge device; A system implemented using a robot; A system for presenting at least one of virtual reality content or augmented reality content; A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.

15. A processor, comprising: One or more circuits, the one or more circuits for: Determine one or more object positions at least based on applying one or more correlation filters to one or more first images, the one or more correlation filters being updated at least partially using one or more image regions and at least based on applying a focus windowing to the one or more image regions, the focus windowing including making one or more regions corresponding to one or more backgrounds of one or more objects in the one or more image regions blurred at least based on a distance from a target object in the one or more image regions, the one or more image regions corresponding to one or more positions of the one or more objects in the one or more second images.

16. The processor according to claim 15, wherein the blurring is performed using a Gaussian filter.

17. The processor according to claim 15, wherein the blurring comprises one or more of the following: blurring one or more color channels of the one or more image regions; or blurring one or more feature channels of the one or more image regions.

18. The processor according to claim 15, further comprising determining one or more characteristics of a blurring filter that is used to perform the blurring based at least on the size of the target object.

19. The processor according to claim 15, wherein the one or more image regions comprise one or more scaled image regions that are scaled to a template size.

20. The processor according to claim 15, wherein the processor is included in at least one of the following: a control system for an autonomous machine or a semi-autonomous machine; a perception system for an autonomous machine or a semi-autonomous machine; a system for performing simulation operations; a system for performing real-time streaming; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for presenting at least one of virtual reality content or augmented reality content; a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Training, testing, and verifying autonomous machines using simulated environments

    US11436484B2

  • Smart area monitoring with artificial intelligence

    US20190294889A1

  • Multi-target detection tracking method, electronic equipment and storage medium

    CN108121945A