Adaptive target tracking machine learning model engine
The adaptive target tracking model engine addresses inaccuracies in surround environments by collecting and optimizing data from surround scenes, enhancing gaze estimation accuracy and robustness for diverse vehicle types and lighting conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2022-03-15
- Publication Date
- 2026-06-01
Smart Images

Figure 0007867822000001 
Figure 0007867822000002 
Figure 0007867822000003
Abstract
Description
[Background technology]
[0001] Recent advancements in target tracking technology have been developed for desktops, laptops, and tablets (i.e., non-surround or 2D environments where the user is viewing a small plane), and as such, conventional target tracking solutions are limited when applied to surround deployment applications (e.g., driving monitoring systems). This is because the requirements for target tracking in vehicles (i.e., in surround or 3D environments) differ from those in non-surround environments. For example, the expected range of the driver's head posture and gaze angle is typically much wider in the natural path of driving compared to desktop users, and can involve far more surfaces at various depths from the driver. In addition, target tracking in vehicles must operate robustly in a variety of lighting conditions—good and bad lighting conditions. For example, a vehicle may be operated at different times of the day with different environments and ambient light, and as such, target tracking must operate effectively across the entire range between daylight and darkness.
[0002] Conventional visual target tracking technologies include machine learning techniques. Specifically, they are limited by the type of data collected to train a machine learning model and the rigidity of the application of the machine learning model generated from the data. At a high level, a visual target tracking machine learning model is usually a classification model that formulates fixation area detection as a classification problem. In a typical implementation, a rough head pose direction and a face part including both eyes are identified and used to train a support vector machine (SVM) fixation classifier. The SVM fixation classifier then outputs one of eight predetermined fixation areas. In another example, facial features are identified and classified in these spatial configurations within six regions. However, the classification of facial features is often performed without using accurate fixation estimation. This spatial configuration-based technique may cause fluctuations in classification accuracy between and within objects. Furthermore, the spatial configuration-based technique does not acquire explicit pupil features, and as such, the accuracy and robustness of understanding pupil movement relative to fixation (rather than head movement) may decrease.
[0003] In a third conventional approach, a convolutional neural network (CNN) is trained using a facial descriptor to classify fixation areas. However, the effectiveness and accuracy of the fixation classification models based on these approaches are often limited to specific vehicle types with a fixed 3D shape. As such, a more comprehensive visual target tracking system with an alternative basis for performing machine learning operations can improve the computational operations and interfaces for visual target tracking systems.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
[0005] Embodiments of this disclosure relate to an adaptive target tracking machine learning model engine ("Adaptive Model Engine") for a target tracking system. Systems and methods are disclosed for providing an Adaptive Model Engine, including a target tracking or gaze tracking development pipeline ("Adaptive Model Training Pipeline") - the Adaptive Model Training Pipeline may perform the following training operations: collecting data, training, optimizing, and deploying a customized Adaptive Target Tracking Model based on an identified set of features of the deployment environment. The Adaptive Model Engine may support training an Adaptive Target Tracking Model for gaze vector estimation in a surround environment, and generating an Adaptive Target Tracking Model based on multiple target tracking deformation models and multiple face landmark neural network criteria.
[0006] In contrast to conventional systems such as those described above, data may be collected from a surround scene or surround deployment environment to generate an adaptive target tracking model. For example, surround scene data collection may be based on multiple sensors, multiple viewpoints, and data synchronization techniques. Data collection for the adaptive model engine may include augmented data variability (e.g., a set of surround scene data types) for ground truth values, including gaze direction vectors and gaze direction vector ranges. The adaptive model engine supports adaptive model training of the adaptive target tracking model based on a set of surround scene data types. The set of surround scene data types may represent different inputs that can be used during adaptive model training. Adaptive model training may include several data preparation and processing stages to manage the diversity of the set of surround scene data types used to train the adaptive target tracking model.
[0007] In some embodiments, adaptive model training of an adaptive model engine may include techniques for optimizing the adaptive target tracking model. Iterative refinement can be specifically performed on the adaptive target tracking model to generate a retrained adaptive target tracking model. For example, a first adaptive target tracking model for a vehicle can be generated and then optimized as a second adaptive target tracking model based on the vehicle type, vehicle size and shape, changes in the size and basic settings of the deep neural network ("DNN") model, and the focus and range of the DNN functionality.
[0008] The system and method for providing an adaptive target tracking machine learning model engine are described in detail below with reference to the attached diagrams. [Brief explanation of the drawing]
[0009] [Figure 1A] This is a diagram of an example system for providing an adaptive model engine, according to some embodiments of the present disclosure. [Figure 1B] This is a data flow diagram illustrating an example of a process that provides an adaptive target tracking model using an adaptive model engine, according to some embodiments of the present disclosure. [Figure 2A] This figure includes a visual representation of an example data acquisition component associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 2B] This figure includes a visual representation of an example data acquisition component associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 2C] This figure includes a visual representation of an example data acquisition component associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 3] This is a data flow diagram of the in-vehicle data acquisition process for adaptive model engines according to some embodiments of the present disclosure. [Figure 4A]This is an example data flow diagram of the gaze detection processing associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 4B] This is an example data flow diagram of the gaze detection processing associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 4C] This is an example data flow diagram of the gaze detection processing associated with an adaptive model engine, according to some embodiments of the present disclosure. [Figure 5] This flowchart illustrates a method for interacting with an adaptive model engine to provide an adaptive target tracking model, according to some embodiments of the present disclosure. [Figure 6] This flowchart illustrates a method for an adaptive model engine to provide an adaptive target tracking model, according to some embodiments of the present disclosure. [Figure 7A] Illustrations of exemplary autonomous vehicles according to some embodiments of the present disclosure. [Figure 7B] Figure 7A shows examples of camera positions and fields of view of an exemplary autonomous vehicle according to some embodiments of the present disclosure. [Figure 7C] Figure 7A is a block diagram of an exemplary system architecture of an exemplary autonomous vehicle according to some embodiments of the present disclosure. [Figure 7D] This is a system diagram of communication between a cloud-based server and the exemplary autonomous vehicle shown in Figure 7A, according to some embodiments of the present disclosure. [Figure 8] This is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure. [Figure 9] This is an exemplary data center block diagram suitable for use in implementing some embodiments of the present disclosure. [Modes for carrying out the invention]
[0010] Systems and methods relating to an adaptive target tracking approach using a machine learning model engine are disclosed. The disclosure may be described in relation to an exemplary autonomous vehicle 700 (referred to herein as "vehicle 700" or "ego vehicle 700" as otherwise provided herein, examples of which are illustrated with reference to Figures 7A–7D), but this is not intended to be limiting. For example, the systems and methods described herein may be used, without limitation, by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), piloted or unpiloted robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles towed to one or more trailers, flying ships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, submarines, drones, and / or other vehicle types. In addition, while this disclosure may describe autonomous driving and target tracking technologies, it is not intended to be limiting, and the systems and methods described herein may be used in any other technological space in which autonomous driving may be used, including augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technological space in which autonomous driving may be used.
[0011] Embodiments of this disclosure aim to provide an approach for adaptive target tracking using a machine learning model engine for target tracking systems ("Adaptive Model Engine"). The Adaptive Model Engine may include a target tracking or gaze tracking development pipeline ("Adaptive Model Training Pipeline") that supports collecting, training, optimizing, and deploying a customized adaptive target tracking model based on an identified set of features of the deployment environment. The Adaptive Model Engine supports training an adaptive target tracking model for gaze vector estimation in a surround environment, and generating an adaptive target tracking model based on multiple target tracking deformation models and multiple face landmark neural network criteria.
[0012] According to one or more embodiments, the adaptive model engine generates an adaptive target tracking model using surround scene data acquisition in a surround scene or surround deployment environment. The surround scene can refer to any multidimensional environment (e.g., a 2D or 3D environment) from which ground truth data is collected to enable the capture of adaptive model engine data (e.g., a set of adaptive model engine data) of ground truth values. In contrast to conventional systems with 2D scene data acquisition, data acquisition for the adaptive model engine can include augmented data variation of ground truth values (e.g., a set of surround scene data types) including gaze direction vectors and gaze direction vector ranges.
[0013] Surround scene data collection can be based on multiple sensors, multiple viewpoints, and data synchronization techniques. By collecting data in the surround scene, several advantages over conventional visual target tracking data collection techniques can be achieved. These advantages can include, without limitation: flexibility in camera placement, rapid acquisition of diverse data including the body, head, and eye movements of one or more persons, and robust modeling of the surround scene and scene shape. For example, surround scene data collection can be specifically used to train an adaptive visual target tracking model such that the adaptive visual target tracking model can be customized for different types of vehicles having different vehicle shapes, in a manner that can be used to train the adaptive visual target tracking model. Surround scene data collection can be based on a bench setup where participants are presented with target ground truth (GT) fixation points on a screen through a data collection interface. Surround scene data collection can be based on an in-vehicle setup where the vehicle is divided into several regions and surround scene data is collected using a signaling mechanism (e.g., a light emitting diode (LED) panel) in the regions.
[0014] The adaptive model engine supports adaptive model training of an adaptive visual target tracking model based on a set of surround scene data types. The surround scene data types can refer to categories of data retrieved during surround scene data collection in the surround scene. The set of surround scene data types can represent different inputs used during adaptive model training. The adaptive model training pipeline according to various embodiments of the present disclosure can include processing a set of surround scene data types corresponding to data retrieved during surround scene data collection, based at least on multiple sensors and data synchronization operations.
[0015] Adaptive model training can include several data preparation and processing stages to manage the diversity of sets of surround scene data types used to train an adaptive visual target tracking model. The data preparation stage can include preprocessing, data filtering, ground truth validation, and task-specific subsampling. Using the set of surround scene data types, as well as the data preparation and processing stages, to train an adaptive visual target tracking model supports adapting adaptive visual target tracking to changes in the deployment environment. For example, an adaptive visual target tracking model can, first, support different lighting, camera, and placement data types for supporting different vehicle visual target tracking systems, second, support different occupant data types (e.g., facial characteristics, ethnicity, eyewear, headgear, clothing, facial hair, partial occlusion, and body, head, and eye movements), and third, be trained to handle ground truth generation and labeling errors.
[0016] Adaptive model training can be based on one or more adaptive model training pipelines and, in particular, on the ability of deep neural networks ("DNNs") to achieve high-quality results when there is sufficient data to be captured so that a particular type of problem can be solved using DNNs. In this way, the adaptive model engine can support the derivation of customized adaptive target tracking models for target tracking systems in different deployment environments. Customized adaptive target tracking models can be generated using adaptive model engine data—specifically, adaptable target tracking models. Adaptable target tracking models can be maximum variation models. Adaptable target tracking models can be initially trained on a large pool of data (i.e., samples selected to ensure both extremes or a wide variety of input data and features) that focuses on maximum variation. Adaptable target tracking models can be stripped down (e.g., by dropping feature maps with small valued norms) to reduce computational complexity and discard redundant parameters that would lead to overfitting. The removed, adaptable target tracking models can then be retrained or fine-tuned to match the data distribution that matches the set of features in the deployed environmental data (e.g., car type and camera location / viewpoint).
[0017] Customized adaptive target tracking models can be generated by using adaptive model engine data, including multiple target tracking deformation models (e.g., comprehensive head norm, single-eye model, restricted head posture) and multiple face landmark neural network criteria (e.g., face landmark confidence, head posture, left / right eye appearance quality), as input to the adaptive model selection ensemble manager. Customized adaptive target tracking models support gaze estimation, which predicts where a person is looking by considering gaze vector estimation data (e.g., face, left eye, right eye, 2D / 3D landmarks). Gaze estimation can be based on 3D gaze vector estimation to predict gaze vectors.
[0018] Furthermore, adaptive model training of an adaptive model engine may include techniques for optimizing the adaptive target tracking model. Iterative improvement can be specifically performed on the adaptive target tracking model to generate a retrained adaptive target tracking model. For example, a first adaptive target tracking model for a vehicle can be generated and then optimized as a second adaptive target tracking model based on the vehicle type, vehicle size and shape, user-specific parameters, changes in the size and basic settings of the DNN model, and the focus and range of the DNN function. More generally, the adaptive target tracking model can be adapted to different types of application scenarios beyond automotive applications.
[0019] Adaptive target tracking machine learning model engine Referring to Figure 1A, Figure 1A is an example system 100 for an adaptive target-tracking machine learning model engine ("Adaptive Model Engine") according to some embodiments of the present disclosure. For example, system 100 may provide an adaptive model engine including a target-tracking or gaze-tracking development pipeline ("Adaptive Model Training Pipeline"). System 100 further supports three different types of model improvement operations, including: adaptive model selection and ensemble operations, training-based improvement operations, and iterative feedback operations. It should be understood that this and other arrangements described herein are described merely as examples. Other configurations and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) may be used in addition to or instead of those illustrated, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any appropriate combination and location. The various functions described herein as performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions can be performed by a processor that executes instructions stored in memory.
[0020] System 100 provides components, instructions, and operations for providing an adaptive model engine, including an adaptive model training pipeline. As shown in Figure 1A and described below, the components, instructions, and operations include: data acquisition 110, including bench setup 112 and in-vehicle setup 114 for data acquisition operations; data extraction 122, data filtering 124, data labeling 126; data preparation 130 for synthetic data generation 128; data preparation 130 for data subsampling 132; data preprocessing 134; and data isolation 136; DNN model training 140; multi-stage activation 150; feedback 160; quality assurance ("QA") 170; key performance indicators ("KPI") 172; model selection and deployment 180; and deployment environment 190.
[0021] Referring to Figure 1B, Figure 1B shows an example system 100 for adaptive target tracking according to several embodiments of the present disclosure. Figure 1B includes vehicle 102 (i.e., a first deployment environment having first deployment environment data) and vehicle 104 (i.e., a second deployment environment having second deployment environment data). The first and second deployment environments can represent different types of vehicles (e.g., vehicle type, size, and shape). For example, the first vehicle could be a sports utility vehicle ("SUV") and the second vehicle could be a sedan—each vehicle having different characteristics supported by the target tracking model. Different types of vehicles (e.g., the first vehicle and another vehicle) as well as non-automotive deployment environments are assumed by embodiments of the present disclosure.
[0022] Vehicles 102 and 104 may access deployment environment data corresponding to sets of features for their respective environments and provide this deployment environment data to system 100 (i.e., adaptive model engine) in Figure 1A. The sets of features may be used to identify the corresponding models for deployment. For example, the sets of features may include spatial configuration features (e.g., vehicle type, size, and shape), DNN baseline configuration features (e.g., DNN model size and performance baseline, DNN function focus and range), or gaze type configuration features (e.g., region-based gaze for driver distraction, saccade-based cognitive load estimation, gaze-based human-machine interaction in both automotive and non-automotive domains).
[0023] The set of features of the deployment environment may further correspond to surround scene data types, which refer to categories of data retrieved during data acquisition in the surround scene. The set of surround scene data types may represent different inputs used during model training. In this regard, the set of features may correspond to the types of inputs used during model training. For example, an adaptive target tracking model can be trained to, firstly, support different lighting, camera, and placement data types to support different vehicle target tracking systems; secondly, support different occupant data types (e.g., facial features, ethnicity, eyewear, headgear, clothing, facial hair, partial occupancy, and body, head, and eye movements); and thirdly, handle errors during ground truth generation and labeling. Such as the corresponding deployment environment data
[0024] Deployment environment data may include data generated by, in, or within the deployment environment. For example, the driving systems (e.g., driving system 106 and driving system 108) and / or other components of the vehicle may communicate one or more portions of the deployment environment data to the adaptive model engine (e.g., adaptive model selection and ensemble manager). As a further example, one or more computing devices 800 may communicate one or more portions of the deployment environment data to the adaptive model engine. As an example, the computing devices 800 may be used in the deployment environment for data collection as described herein, but may not be part of the deployment environment and / or the vehicle being deployed.
[0025] Deployment environment data may be generated based on the gaze type of the deployment environment. The deployment environment data generated based on the gaze type may be captured by one or more features used by the adaptive model engine to generate one or more adaptive target tracking models for deployment in the deployment environment. The gaze type may include region-based gaze detection for driver distraction, saccade-based cognitive load estimation, and gaze-based human-machine interactions in both the automotive and non-automotive domains. For example, a first vehicle configuration (e.g., vehicle type, size, and shape) may be associated with region-based gaze for driver distraction, while a second vehicle configuration may be associated with gaze-based human-machine interactions. A single vehicle configuration may include a first part of the vehicle having a first gaze type, and a second part of the vehicle having a second gaze type. As such, the deployment environment data may include gaze type information that can be used to identify one or more gaze types of a vehicle configuration, and the corresponding parts associated with one or more gaze types. Other variations and combinations of gaze types are assumed by embodiments of this disclosure.
[0026] In one or more embodiments, the adaptive model engine can be used to derive an adaptive target tracking model customized to a corresponding deployment environment, based on at least a set of features of the deployment environment, as discussed herein. In one or more embodiments, the set of features can be instructed to the adaptive model engine using a database and / or from manually entered or selected settings or components. In further examples, deployment environment data can be provided to the adaptive model engine using one or more schemas or other metrics from the set of features captured by the deployment environment data. Also in one or more embodiments, one or more features may be latent within the deployment environment data. Operationally, vehicle 102 or vehicle 104 can provide deployment environment data to the adaptive model engine and receive an adaptive target tracking model customized to the corresponding vehicle's deployment environment from the adaptive model engine. In one example, the adaptive model engine can be used to derive adaptive target tracking models for a first vehicle deployment environment, a second vehicle deployment environment, and a third non-vehicle deployment environment, each having a different shape. According to the embodiment, the first vehicle deployment environment may be associated with a first vehicle type, and the second vehicle deployment environment may be associated with a second vehicle type that may differ from the first vehicle type.
[0027] In another embodiment, second deployment environment data corresponding to a set of features of a second deployment environment is accessed, and an adaptive model engine is provided to generate a second adaptive target tracking model. The second adaptive target tracking model is customized to the second deployment environment using adaptive model engine data identified at least based on the features of the second deployment environment. The second adaptive target tracking model is provided for deployment in the second deployment environment.
[0028] In at least one embodiment, vehicle 102 or vehicle 104 may provide deployment environment data by transmitting data to an adaptive model engine. However, in other embodiments, at least some of the deployment environment data may be provided by another device or user within the deployment environment.
[0029] Using data identified based on at least a set of features of the deployment environment, the adaptive model engine (e.g., using the adaptive model selection and ensemble manager 180) derives a first customized adaptive target tracking model for vehicle 102 and a second customized adaptive target tracking model for vehicle 104. The corresponding adaptive target tracking models can be trained on at least gaze surround scene data and gaze vector estimation data. The customized adaptive target tracking models support gaze estimation to predict where a person is looking, taking into account gaze vector estimation data (e.g., face, left eye, right eye, 2D / 3D landmarks). Gaze estimation can be based on 3D gaze vector estimation to predict gaze vectors.
[0030] Adaptive model engine data may include multiple target tracking deformation models (e.g., comprehensive head norm, single-eye model, restricted head pose) and multiple face landmark neural network criteria (e.g., face landmark confidence, head pose, left / right eye appearance quality). The multiple target tracking deformation models and multiple face landmark neural network criteria are selected as an ensemble (e.g., using an ensemble machine learning method for improved predictive performance) and packaged together or combined, based at least on deployment environment data, to generate an adaptive target tracking model. For example, an adaptive model selection and ensemble manager may receive a first target tracking deformation model and a first face landmark neural network model to generate an adaptive target tracking model. An ensemble method may refer to a meta-algorithm that combines several machine learning techniques into a single predictive model to reduce variance (e.g., bagging), bias (e.g., boosting), or improve prediction (e.g., stacking).
[0031] The ensemble method can be performed using a sequential ensemble method in which base learners are generated sequentially (e.g., AdaBoost), or a parallel ensemble method in which base learners are generated in parallel (e.g., Random Forest). In one or more embodiments, machine learning ensemble may be based on spatial configuration features, DNN baseline configuration features, or gaze type configuration features. For example, decisions can be made from specific DNN baseline configuration features, DNN model size and performance baselines, and focus and scope of DNN functionality, such that selection and ensemble are based on DNN baseline configuration features. In this regard, generating an adaptive target tracking model may include selecting multiple face landmark neural network criteria and multiple target tracking deformation models. Multiple target tracking deformation models and multiple face landmark neural network criteria are combined to produce an output (e.g., via an adaptive target tracking model) that includes: gaze vectors, gaze points in camera space, or both. The output is aligned by DNN model size, performance preferences, and DNN functionality, focus, and scope, as identified in the deployment environment data.
[0032] In some embodiments, the deployment environment data received from vehicle 102 or vehicle 104 includes a gaze angle range associated with the deployment environment. The gaze angle range can be expressed using a predetermined gaze angle range criterion (e.g., eye height (0°), 25° up or down, eye height (10°), 30°) - other variations and combinations are assumed herein. The algorithm analyzes the relationship between the head and both eyes to determine the gaze angle and the corresponding gaze angle range. Customizing an adaptive target tracking model as such is at least based on identifying adaptive model engine data to generate an adaptive target tracking model based at least on the gaze angle range. The adaptive target tracking model includes a subset of data from the adaptive model engine data.
[0033] Vehicle 102 or Vehicle 104 receives an adaptive target tracking model customized for the deployment environment from the adaptive model engine (for example, via the adaptive model selection and ensemble manager 180) using adaptive model engine data identified at least on the characteristics of the deployment environment (for example, via its driving system). Vehicle 102 or Vehicle 104 uses the adaptive target tracking model for deployment in the deployment environment. The adaptive target tracking model can be deployed to efficiently identify the type of gaze in the deployment environment (e.g., region-based gaze for driver distraction, saccade-based cognitive load estimation, gaze-based human-machine interaction in both automotive and non-automotive domains).
[0034] Moving on to Figures 2A to 2C, Figures 2A to 2C illustrate data acquisition components, commands, and operations that support data acquisition. Data acquisition components may be used to generate at least some of the deployment environment data described herein. Data acquisition may be managed by a data collector 210, which may be implemented in any combination of the deployment environment 190 (e.g., a vehicle) and / or computing devices 800. Data acquisition components may include a bench configuration 220, an in-vehicle configuration 230, a ground truth display 240, a user recorder 250, a driveworks recorder 260, an animator 272, a point distribution 274, an LED panel 276, an LED controller 278, and a recorder / camera configuration 280. When data is generated using the deployment environment 190 and / or a device attached to the deployment environment 190 (e.g., an LED panel 276), the output may be provided to computing devices 800, such as a server, a cloud computing environment, a GPU device, etc.
[0035] In one or more embodiments, the computing device 800 may include an adaptive model selection and ensemble manager 450C, and / or the computing device 800 may provide (for example, transmit) at least some of the output data, deployment environment data, and / or data used to generate the deployment environment data to another device which may include the adaptive model selection and ensemble manager 450C. In one or more embodiments, data may be provided to the computing device 800 implementing one or more of the DNN model training 140, feedback 160, multi-stage activation 150, and / or data preparation 130.
[0036] The data acquisition process facilitates gaze data acquisition to capture large data variations. Data variations can include ambient lighting (indoor, outdoor), head movement in the 3D head box, head posture angle, pupil movement, eyewear (glasses, sunglasses, contact lenses), headwear (hats, caps, etc.), occlusion (face masks, drinking from cups), facial features (beards, mustaches) and makeup, as well as target demographics (age, sex, ethnicity, eye type).
[0037] The data acquisition operation can be based on a bench setup (e.g., on-bench configuration 220) or an in-vehicle setup (e.g., in-vehicle configuration 230). Advantageously, high throughput and practical data acquisition are possible for both bench and in-vehicle setups. Figure 2A includes building blocks that can be used for the bench and in-vehicle setups, respectively. For both setups, synchronized multi-camera (e.g., eight cameras) recording (i.e., data synchronization operation) can be used to achieve robustness against variations in camera location and viewpoint. This can provide the flexibility of in-vehicle camera placement, a highly desirable feature by automotive manufacturers.
[0038] Figure 2B shows a high-level architecture (e.g., gaze collection software) for gaze collection instructions and data. It is assumed that one or more instruction or data processors may be common to both on-bench and in-vehicle collection, but one or more instruction or data processors may be designed specifically for in-vehicle data collection. For example, the Animator, GazeExperimentGUI, Generator, GazeExperimentModel, and Listen are common to both on-bench and in-vehicle collection, while the GazeExperimentLED, LightGrid, and LED controller (PiLib) are designed specifically for in-vehicle data collection. In this regard, processing surround scene data of sets of surround scene data types from bench setups and in-vehicle setups is based on at least two or more common data processors, while at least one different data processor for the bench setup and the in-vehicle setup processes the surround scene data and performs data synchronization operations exclusively for the bench setup or the in-vehicle setup.
[0039] For the bench setup, participants may be presented with Target Ground Truth ("GT") gaze points on the screen via a data collection app. GT points may be generated pseudo-randomly for uniform data distribution on the screen. To function as a visual stimulus, the size of the circular target points may vary continuously. In addition, to ensure participants are paying attention while fixating on the target points, an L or R letter may be displayed in the center of each circle. Participants are then asked to right or left-click the indicated letter as appropriate to move to the next target.
[0040] Moving on to Figure 2C, which shows a car or vehicle for an in-car setup or vehicle setup. Figure 2C can be used to illustrate gaze definitions. The car is divided into seven areas for in-car data collection, including: “Front Center,” “Front Right,” “Information Cluster,” “Entertainment Console,” “Left External,” “Right External,” and “Off-Road.” The “Off-Road” area is defined as the interior of the car excluding the six areas. For the in-car setup, LED panels are mounted in each area of the car (e.g., Area 1, Area 2, Area 3, Area 4, Area 5, and Area 6). The position of the LED boards can be configured to change repeatedly to cover as many locations as possible across the entire shape of the car. As shown in Figure 3, a configurable set of points are highlighted across the LED panels on multiple gaze areas.
[0041] Figure 3 shows the in-vehicle data acquisition components, commands, and operations that support in-vehicle data acquisition. The in-vehicle data acquisition components include an in-vehicle grid 310, LEDs 230, an LED controller 330, a camera 340, and a user input interface 350. The LED controller 330 can be a breakout board connecting the in-vehicle grid 310, LEDs 230, LED controller 330, camera 340, and user input interface 350. A selected grid can be triggered to act as a visual stimulus so that user gaze data is collected based on the selected grid. As shown in the figure, the configuration of the LED boards supports gaze data acquisition. Each board is placed within a specific area, and individual LEDs are illuminated through the vehicle interface. The user or subject triggers the start and end of data capture through steering wheel control. Steering wheel control is connected to a network socket and a GPIO breakout. This ensures that data is collected when the user is looking at the LEDs. The steering wheel button activates the Start, Enable, and Record functions, corresponding to the JSON Pack Blue LED, JSON Pack Red LED, and JSON Pack Green LED on the network socket, as well as the SL Serial IDX / RGB, SL Serial IDX / RGB, and SL Serial IDX / RGB on the GPIO breakout.
[0042] Returning to Figure 1A, Figure 1A shows the data preparation components, instructions, and operations that support data preparation. The data preparation operations include data extraction 122, data filtering 124, data labeling 126, synthetic data generation 128, data subsampling 132, data preprocessing 134, and data isolation 136. Data extraction 122 involves accessing data acquisition records of raw video from the camera, along with files that store timestamps and GT gaze values for each target gaze point (via mouse clicks). These files are parsed to extract frames from the raw video and assign corresponding GTs. Data filtering 124 identifies the most useful frames from the extracted frames in a manner that optimizes data distribution. When video is recorded at 30 or 60 Hz, most extracted frames are visually very similar and therefore redundant for training. Also, if used for training, these frames should incur additional computational costs for data labeling and model training. To address this, the data filtering 124 may filter frames based on visual similarities, such as variations in head pose angle.
[0043] Data labeling (and activation) 126 includes generating a final list of frames annotated by a human labeler so that various attributes such as face and eye boundary boxes, face landmarks, eye state (open, half-open, slightly open, closed), glare / glint intensity and localization, etc., are used in model training. In addition, to reduce or eliminate human labeling errors, the activation step employs a ground truth activation mechanism that checks and validates given labels against algorithm-generated expected labels. These samples that deviate significantly from expected labels are discarded or sent back for another iteration of manual labeling.
[0044] Data subsampling 132 may be employed before training to generate task-specific models. For example, if a model is intended to be deployed for a specific type of car (compact, sedan, SUV, truck, etc.) or to function for a specific head pose or gaze angle range, the data may be subsampled as appropriate to feed into DNN model training 140. Data isolation 136 supports the DNN model training requirement to split the data into training, test, and enable sets. Data isolation 136 employs user-level data isolation (a unique anonymous user ID as the key) rather than frame-level isolation to achieve a high degree of model generalization capability across the entire subject. This ensures that no single user is included in more than one of the splits.
[0045] Data preparation 130 may also include data preprocessing 134 that supports data preprocessing operations including each of the following: glare detection, gaze data normalization, and manual eye-gaze feature extraction. Data preprocessing 134 may include glare detection for glare removal and corresponding masking. Glare detection may be used by data preparation 130 before training to achieve robustness against glare on eyeglasses and other reflective objects. Model training can leverage this in various ways, such as for glare removal using image processing on the input image, and to provide a glare mask as additional input to the trained network (which may also be provided to the model once trained).
[0046] Data preprocessing 134 may further include gaze data normalization. Gaze data normalization may aim to reduce the degrees of freedom of gaze estimation analysis. Gaze data normalization may use a normalized camera space in which the camera axis and the face or individual eye axis are aligned (i.e., aligning the gaze data axis from the camera with the face or eye in the gaze data from the camera). In such a normalized camera space, the DNN may learn more accurately the appearance variation of the entire aligned image. Normalizing gaze data may include not distorting the original frame (e.g., undistorted gaze data normalization operation), localization of face landmarks and 3D head pose estimation via PnP, 3D position estimation of the center of the eye and face, and image warping.
[0047] Data preprocessing 134 may further include hand-designed or manual eye-feature extraction. For each frame, the data preprocessing operation may include extracting hand-designed features correlated with gaze (e.g., 28 features), including the contours of the eyes and pupils, and the relative position of the pupils to the eyeballs. Along with the input image, these features may be explicitly provided to aid in the training process. In this way, data preprocessing 134 includes extracting a set of manually extracted gaze features, which are explicitly tagged for training operations in the adaptive model engine.
[0048] As shown in Figure 1A, the adaptive model engine includes DNN model training 140. Figures 4A and 4B include examples of DNN model training components, instructions, and operations that support DNN model training 140. With respect to Figure 4A, the DNN model training components may include input frame 402A, model input generator 404A, data augmentation 406A, input selection 408A, network architecture 410A, output 412A, output selection 414A, and final gaze 416A. The DNN model training components define a gaze model training pipeline through which training can be realized. During operation, an input frame (e.g., input frame 402A) can be received and processed through a model input generator (e.g., model input generator 404A). Potential network inputs (e.g., sets of surround scene data types) may include, but are not limited to, faces (original or normalized), left eye (original or normalized), right eye (original or normalized), channels with each eye + glare mask, channels with each eye + pupil mask, binocular stripes (rectangular crops covering the face that cover both eyes), 2D landmarks, 3D landmarks, face mask / grid 3D head pose angles, and / or handcrafted eye-gaze features. The input frames are processed by the model input generator 404A to generate input selections (e.g., input selections 408A) containing any number of these features or combinations thereof for a network architecture (e.g., neural network architecture 410A).
[0049] In some embodiments, data augmentation (e.g., data augmentation 406A) is further provided along with input selection, as will be discussed in more detail herein. The network architecture processes the inputs to generate outputs such as gaze points and gaze lines (e.g., output 412A), which are processed as output selections (e.g., output selection 414A) to generate a final gaze (e.g., final gaze 416A). It is assumed that different forms and combinations of model outputs are possible. For example, model outputs may include, but are not limited to, gaze vectors (Θ,Φ), gaze points in camera space (x,y,z), or combinations thereof (x,y,z,Θ,Φ). The gaze model training pipeline may operate on an end-to-end in-vehicle gaze tracking system, such as that shown in Figure 4B.
[0050] Figure 4B shows an example of a DNN model training component, including camera 402B, face detection 404B, detected face 406B, 3D gaze vector estimation 408B, face landmark detection 410B, 3D gaze origin estimation 412B, 3D vehicle shape 414B, projection-based in-vehicle gaze region mapping 416, and output gaze region 418B. During operation, the camera (e.g., camera 402B) records content processed using a face detection operation (e.g., face detection 404B) to identify detected faces (e.g., detected face 406B). The detected faces are further processed to identify face landmarks (e.g., face landmark detection 410B) and to determine a 3D gaze origin estimation (e.g., 3D gaze origin estimation 412B). The detected faces can also be processed to determine a 3D gaze vector estimation (e.g., 3D gaze vector estimation 408B).
[0051] Operationally, 3D gaze vector estimation may be provided to estimate the user's gaze direction or gaze vector. In some examples, 3D gaze vector estimation 4108B and 3D gaze origin estimation 412B determine the output gaze region 418B. The 3D gaze vector estimation, 3D gaze vector estimation, and 3D vehicle shape (e.g., 3D vehicle shape 414B) may be processed in combination with projection-based in-vehicle gaze region mapping (e.g., projection-based in-vehicle gaze region mapping 416B) to generate the output gaze region (e.g., output gaze region 418B).
[0052] Additional gaze generation operations are envisioned to complement the operations described above. In some embodiments, data augmentation (e.g., data augmentation 406A) may be provided for additional robustness in the gaze decision operation. For example, data augmentation may include Gaussian blur, gamma adjustment; occluding the eyes using shape (for occlusion invariance); and randomly dropping the eyes with some probability. Additional gaze generation operations may also provide support for data filtering and the use of loss functions for specific tasks. For example, data may be filtered based on head pose angles to generate a head pose constraint model, and loss functions (e.g., L1, L2, logcosh, cosine). Other variations and combinations of supplemental gaze generation operations are envisioned to be compatible with the embodiments of this disclosure.
[0053] The adaptive model engine further includes support for different types of model improvement operations, this example including adaptive model selection and ensemble, training-based improvement, and iterative feedback. Referring to Figure 4C, Figure 4C includes a data flow diagram for adaptive model selection and ensemble, which may be used to provide an adaptive target tracking model according to an embodiment of the present disclosure. Figure 4C includes adaptive model selection and ensemble components, commands, and operations that support generating an overall gaze. Figure 4C includes a face landmark network 410C, models 420C, 430C, 440C, and an adaptive model selection and ensemble manager 450C, which selects from models to generate an adaptive target tracking model 460C.
[0054] During operation, the Adaptive Model Selection and Ensemble Manager 450C is configured to receive inputs (e.g., face landmarks and deformation models) to generate an Adaptive Target Tracking Model 460C that determines (predicts) the overall gaze. As discussed herein, the target tracking models can be trained with a variety of inputs. Model training may include learning complementary information that is utilized when the target tracking models are combined. For example, an ensemble of a traditional joint model and a face-normalized theta-phi model can be generated by the Adaptive Model Selection and Ensemble Manager 450C.
[0055] Figure 4C shows various examples of deformable models that can be ensembled, or more generally, combined (e.g., Model #1: Comprehensive Head Norm; Model #2: Single Eye Model; to Model #N: Restricted Head Pose). These deformable models can be adaptively ensembled or selected based on a set of features of the deployment environment using various possible approaches. As described herein, features of the Adaptive Model Selection and Ensemble Manager 450C can be used to customize the Adaptive Target Tracking Model 460C, which can be instructed to the Adaptive Model Engine using a database and / or from manually entered or selected settings or configurations. For example, the Adaptive Model Selection and Ensemble Manager may have access to configuration and / or setting data that corresponds to the deployment environment and instructs a set of features. Settings or configurations can be entered and provided to the Adaptive Model Selection and Ensemble Manager 450C through a graphical user interface (GUI) and / or configuration or setting files.
[0056] As an addition or alternative, deployment environment data may be provided to the adaptive model engine using schemas or other metrics of one or more features captured by the deployment environment data. For example, settings, configurations, schemas, and / or other metrics may define or specify one or more characteristics of the deployment environment that the adaptive model selection and ensemble manager 450C can use to select one or more models for training, e.g., (models 420C, 430C, or 440C), how the models can be trained, how the models can be ensembled, etc. For example, if a characteristic indicates or specifies that the deployment environment provides inputs for one or more potential network inputs (e.g., surround scene data type), as illustrated with reference to Figure 4B, the adaptive model selection and ensemble manager 450C may select and / or collect one or more models in the ensemble that have the inputs for training. Similarly, if the characteristics indicate or specify that the deployment environment includes cameras that provide large head postures, the Adaptive Model Selection and Ensemble Manager 450C may select and / or collect an ensemble of one or more models having a single eye model, when it is possible to use such an ensemble when only one eye is visible.
[0057] Furthermore, in one or more embodiments, one or more features may be potential within the deployment environment data. For example, if the Adaptive Model Selection and Ensemble Manager 450C determines during training that the Adaptive Target Tracking Model 460C and / or one or more parts thereof are performing below a threshold level, the Adaptive Model Selection and Ensemble Manager 450C may automatically reconfigure the Adaptive Target Tracking Model 460C based, for example, on at least one or more other available configuration options. Other available configuration options may include, for example, one or more other models having inputs corresponding to the deployment environment data and / or corresponding to one or more characteristics of the deployment environment.
[0058] As a further example, in one or more embodiments, potential features may be extracted from one or more portions of deployment environment data or data otherwise corresponding to the deployment environment, for example, using one or more machine learning models and / or algorithms. In various embodiments, recent features may be extracted and applied to one or more criteria, at least based on data analysis. In one or more embodiments, the adaptive model engine may support (apply) various criteria, such as (for example, but not limited to) face landmark confidence, head pose, data normalization quality, and / or criteria for relating left / right eye appearance quality. As indicated in Figure 4C, data for criteria may be provided by the face landmark network 410C, for example, but not limited to. For left / right eye appearance quality deformation models, a full deformation model may be used if all inputs as indicated by the data are available; however, it is assumed that a single-eye model may be used if, for example, a large head pose results in only one eye being visible. Functionally, the head pose can be determined such that the corresponding deformation model is trained with limited data on the corresponding head pose.
[0059] Training-based improvement operations can be implemented using an iterative feedback mechanism that relies on, for example, QA170, KPI172, multi-stage enablement150, and feedback benchmarking and error analysis160 in Figure 1A to continuously improve a DNN model (e.g., identify and quantify failure cases). QA170 may provide operations to prevent errors and defects in the model or model change form to provide confidence that quality requirements are satisfied. KPIs may include a set of quantifiable measures used to measure the long-term performance of the model. Specifically, KPIs can help determine the performance of one model compared to the same model or a different model change form. Based on QA and KPIs, training-based improvement can be carried out iteratively using a multi-step policy (e.g., training, deactivating, and retraining). The policy can be employed to obtain a low-computation network that is robust to the full gaze angle range, with emphasis on the deployed environment. For example, in at least one embodiment, the deformed model is initially trained on a large pool of data that focuses on maximum variation. The deformed model is then stripped—by aggressively dropping feature maps with small norm values—to reduce computational complexity and discard redundant parameters that could lead to overfitting. In at least one embodiment, stripping may include—for each new model version deployed—large-scale benchmarking (e.g., benchmarking version) and an enumeration of the differences to previous versions of the model with respect to evaluation criteria (e.g., accurate gaze vector error across various data buckets, number of false gaze area detections per N hours of driving). The stripped model is then retrained and fine-tuned for a data distribution that matches the deployed environment, such as the type of car and the location or viewpoint of the camera.In this regard, training-based improvement operations may include receiving feedback data for the adaptive target tracking model and retraining the adaptive target tracking model, at least based on the adaptive model engine's multi-stage retraining pipeline. Retraining the adaptive target tracking model produces a retrained adaptive target tracking model.
[0060] Referring here to Figure 5, each block of Method 500 described herein includes a computation process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor that executes instructions stored in memory. The Method may also be performed as computer-usable instructions stored on a computer storage medium. The Method may be provided, to name a few, as a standalone application, a service or a hosted service (standalone or in combination with another hosted service), or as a plug-in to another product. In addition, Method 500 is described as an example with respect to System 100 in Figure 1. However, these Methods may be performed by any one system or any combination of systems, including but not limited to those described herein, in addition or by alternative means.
[0061] Figure 5 is a flowchart illustrating a method 500 for interacting with an adaptive model engine to provide an adaptive target tracking model, according to some embodiments of the present disclosure. Method 500 includes, in block B502, accessing deployment environment data corresponding to a deployment environment, based at least on the deployment environment (e.g., vehicle 102 or 104). The deployment environment data may include data on the features of the deployment environment corresponding to a set of surround scene data types, where the surround scene data type is a category of data retrieved during surround scene data acquisition.
[0062] Method 500 includes providing deployment environment data to an adaptive model engine in block B504 that supports deriving an adaptive target tracking model customized for one or more deployment environments, based on at least the characteristics of one or more deployment environments. For example, the deployment environment data may be provided to the adaptive model selection and ensemble manager 450C of the adaptive model engine for model selection and deployment 180.
[0063] Method 500 includes, in block B506, receiving from the Adaptive Model Engine an Adaptive Target Tracking Model 460C customized for the deployment environment using Adaptive Model Engine data identified at least on the characteristics of the deployment environment. For example, the Adaptive Model Selection and Ensemble Manager 450C may select, ensemble, and train one or more of Models 420C, 430C, and / or 440C and provide the resulting Adaptive Target Tracking Model 460C to one or more devices (e.g., the devices that provided the deployment environment data and / or requested the generation of the Adaptive Target Tracking Model 460C).
[0064] Method 500, in block B508, includes providing an adaptive target tracking model for deployment in a deployment environment. For example, a device that receives the adaptive target tracking model 460C may provide (e.g., transmit) the adaptive target tracking model 460C to a device in the deployment environment (e.g., vehicle 102 or 104).
[0065] Figure 6 is a flowchart illustrating a method 600 for an adaptive model engine to provide an adaptive target tracking model, according to some embodiments of the present disclosure. Method 600 includes, in block B602, receiving deployment environment data corresponding to a deployment environment in the adaptive model engine, which then supports deriving an adaptive target tracking model customized for one or more deployment environments based on at least the characteristics of one or more deployment environments. For example, the deployment environment data may be received by the adaptive model selection and ensemble manager 450C of the adaptive model engine for model selection and deployment 180.
[0066] Method 600 includes, in block B604, generating an adaptive target tracking model customized to the deployment environment using adaptive model engine data identified on at least the characteristics of the deployment environment, wherein the adaptive target tracking model is trained on at least gaze surround scene data and gaze vector estimation data. For example, the adaptive model selection and ensemble manager 450C may generate an adaptive target tracking model 460C customized to the deployment environment 190.
[0067] Method 600 includes, in block B606, communicating the adaptive target tracking model from the adaptive model engine. For example, the adaptive model selection and ensemble manager 450C may provide the resulting adaptive target tracking model 460C to one or more devices (e.g., the devices that provided the deployment environment data and / or requested the generation of the adaptive target tracking model 460C).
[0068] The systems and methods disclosed to provide an adaptive target tracking machine learning model engine are intended to be run using a processor. The processor may comprise one or more circuits. The processor may be provided in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system incorporating one or more virtual machines (VMs); a system at least partially implemented in a data center; or a system at least partially implemented using cloud computing resources.
[0069] Exemplary autonomous vehicle Figure 7A shows an exemplary autonomous vehicle 700 according to some embodiments of the present disclosure. The autonomous vehicle 700 (as otherwise referred to herein as "Vehicle 700") may include, but is not limited to, passenger vehicles such as cars, trucks, buses, first responder vehicles, shuttles, electric or motorized bicycles, motorcycles, fire engines, police vehicles, ambulances, boats, construction vehicles, submarines, drones, vehicles attached to trailers, and / or other types of vehicles (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of automation levels as defined by the National Highway Traffic Safety Administration (NHTSA), departments within the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (standard number J3016-201806, published June 15, 2018; standard number J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 700 may have the capability to perform functions at one or more of the autonomous driving levels from Level 3 to Level 5. For example, depending on the embodiment, Vehicle 700 may have the capability of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).
[0070] Vehicle 700 may include components such as the vehicle's chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components. Vehicle 700 may include a propulsion system 750, such as an internal combustion engine, a hybrid power unit, a fully electric engine, and / or another propulsion system type. The propulsion system 750 may be connected to the vehicle 700's drivetrain, which may include a transmission, to enable the vehicle 700 to propel itself. The propulsion system 750 may be controlled in response to receiving signals from a throttle / accelerator 752.
[0071] A steering system 754, which may include a steering wheel, may be used to steer the vehicle 700 (for example, along a desired course or route) when the propulsion system 750 is operating (for example, when the vehicle is moving). The steering system 754 may receive signals from the steering actuator 756. The steering wheel may also be an option for fully automated (level 5) functionality.
[0072] The brake sensor system 746 may be used to operate the vehicle brakes in response to receiving signals from the brake actuator 748 and / or the brake sensor.
[0073] The controller 736, which may include one or more system-on-a-chip (SoC) 704 (Figure 7C) and / or GPUs, can provide signals (e.g., expressions of commands) to one or more components and / or systems of the vehicle 700. For example, the controller can send signals to actuate the vehicle brakes via one or more brake actuators 748, actuate the steering system 754 via one or more steering actuators 756, and actuate the propulsion system 750 via one or more throttle / accelerators 752. The controller 736 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist the driver in driving the vehicle 700. The controller 736 may include a first controller 736 for autonomous driving functions, a second controller 736 for functional safety functions, a third controller 736 for artificial intelligence functions (e.g., computer vision), a fourth controller 736 for infotainment functions, a fifth controller 736 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 736 may handle two or more of the aforementioned functions, and two or more controllers 736 may handle a single function, and / or any combination thereof.
[0074] The controller 736 can provide signals for controlling one or more components and / or systems of the vehicle 700 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example and without limitation, global navigation satellite system sensors 758 (e.g., global positioning system sensors), RADAR sensors 760, ultrasonic sensors 762, LIDAR sensors 764, inertial measurement unit (IMU) sensors 766 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 796, stereo cameras 768, wide-view cameras 770 (e.g., fisheye cameras), infrared cameras 772, surround cameras 774 (e.g., 360-degree cameras), long-range and / or medium-range cameras 798, speed sensors 744 (e.g., for measuring the speed of a vehicle 700), vibration sensors 742, steering sensors 740, brake sensors (e.g., as part of a brake sensor system 746), and / or other sensor types.
[0075] One or more of the controllers 736 may receive inputs (represented, for example, by input data) from the instrument cluster 732 of the vehicle 700 and provide outputs (represented, for example, by output data, display data, etc.) via a human-machine interface (HMI) display 734, an audible annunciator, a loudspeaker, and / or other components of the vehicle 700. The outputs may include information such as vehicle velocity, speed, time, map data (e.g., HD map 722 in Figure 7C), location data (e.g., the location of the vehicle 700, such as on a map), direction, the location of other vehicles (e.g., occupied grid), and information about objects and the status of objects as perceived by the controller 736. For example, the HMI display 734 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., changing lanes now, exiting exit 34B within 3.22 km (2 miles), etc.).
[0076] Vehicle 700 further includes a network interface 724 that can communicate over one or more networks using one or more wireless antennas 726 and / or a modem. For example, the network interface 724 may have the capability to communicate over LTE, WCDMA®, UMTS, GSM, CDMA2000, etc. The wireless antennas 726 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth®, Bluetooth® LE, Z-Wave, ZigBee, and / or low-power wide-area networks (LPWAN) such as LoRaWAN, SigFox.
[0077] Figure 7B shows examples of camera positions and fields of view of the exemplary autonomous vehicle 700 of Figure 7A, according to several embodiments of the present disclosure. The cameras and their respective fields of view are exemplary embodiments and are not intended to limit the scope. For example, additional and / or alternative cameras may be included, and / or cameras may be placed in different locations on the vehicle 700.
[0078] The camera type may include, but is not limited to, a digital camera that can be used with components and / or systems of the vehicle 700. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or other ASILs. Depending on the embodiment, the camera type may have the capability of any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The camera may have the capability to use a roll shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras having RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to increase light sensitivity.
[0079] In some applications, one or more cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function mono-camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0080] One or more of the cameras may be mounted in custom-designed (3D-printed) mounting parts to eliminate stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capability. Referring to side mirror mounting parts, the side mirror parts may be custom 3D-printed so that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera may be integrated within the side mirror. For side-view cameras, the camera may also be integrated within four struts located at each corner of the cabin.
[0081] A camera having a field of view that includes a portion of the environment in front of the vehicle 700 (e.g., a forward-facing camera) may be used for surround view to help identify the forward path and obstacles and, with the help of one or more controllers 736 and / or control SoCs, to help provide information essential for generating an occupied grid and / or determining a preferred vehicle path. The forward-facing camera may also be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera may also be used for ADAS functions and systems, including other functions such as Lane Departure Warning ("LDW"), Autonomous Cruise Control ("ACC"), and / or traffic sign recognition.
[0082] Various cameras may be used in forward-facing configurations, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 770, which can be used to capture objects entering the view from the periphery (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in Figure 7B, any number of wide-view cameras 770 may be present in the vehicle 700. In addition, long-range cameras 798 (e.g., a long-view stereo camera pair) may be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. Long-range cameras 798 may also be used for object detection and classification, as well as basic object tracking.
[0083] One or more stereo cameras 768 may also be included in a forward-facing configuration. The stereo camera 768 may include an integrated control unit with an expandable processing unit that may provide a programmable logic (FPGA) and a multi-core microprocessor with an integrated CAN or Ethernet® interface on a single chip. Such a unit may be used to generate a 3D map of the vehicle's environment, including distance estimates for all points in the image. An alternative stereo camera 768 may include a compact stereo vision sensor comprising two camera lenses (one on the left and one on the right) and an image processing chip capable of measuring the distance from the vehicle to an object and activating autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 768 may be used in addition to or instead of those described herein.
[0084] A camera having a field of view including a portion of the environment on the sides of the vehicle 700 (e.g., a side-view camera) may be used for surround view, providing information used to create and update the occupancy grid and generate side impact collision warnings. For example, surround cameras 774 (e.g., four surround cameras 774 as shown in Figure 7B) may be positioned on the vehicle 700. The surround cameras 774 may include wide-view cameras 770, fisheye cameras, 360-degree cameras, and / or similar. For example, four fisheye cameras may be positioned in front of, behind, and on the sides of the vehicle. In an alternative configuration, the vehicle may use three surround cameras 774 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround view camera.
[0085] A camera having a field of view that includes a portion of the environment behind the vehicle 700 (e.g., a rear-view camera) may be used for parking assistance, surround view, rear collision warning, and creation and updating of the occupancy grid. A wide variety of cameras may be used, including, but not limited to, cameras suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 798, stereo cameras 768), infrared cameras 772, etc., as described herein.
[0086] Figure 7C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 700 of Figure 7A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted together. Furthermore, many of the elements described herein are functional entities that can be implemented as individual or distributed components or in combination with other components, and in any appropriate combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor that executes instructions stored in memory.
[0087] Each component, feature, and system of the vehicle 700 in Figure 7C is illustrated as being connected via a bus 702. Bus 702 may include a Controller Area Network (CAN) data interface (or referred to as the "CAN bus"). CAN may also be a network within the vehicle 700 used to help control various features and functions of the vehicle 700, such as the operation of brakes, acceleration, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0088] Bus 702 is described herein as a CAN bus, but this is not intended to limit it. For example, FlexRay and / or Ethernet® may be used in addition to, or as an alternative to, a CAN bus. In addition, a single line is used to represent bus 702, but this is not intended to limit it. There may be any number of buses 702, which may include, for example, one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 702 may be used to perform different functions and / or for redundancy. For example, a first bus 702 may be used for collision avoidance functions, and a second bus 702 may be used for operation control. In any example, each bus 702 may communicate with any of the components of vehicle 700, and two or more buses 702 may communicate with the same component. In some examples, each SoC 704, each controller 736, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 700) and may be connected to a common bus such as a CAN bus.
[0089] The vehicle 700 may include one or more controllers 736, such as those described herein with respect to Figure 7A. The controllers 736 may be used for a variety of functions. The controllers 736 may be connected to any of the various other components and systems of the vehicle 700 and may be used for the control of the vehicle 700, the artificial intelligence of the vehicle 700, infotainment for the vehicle 700, and / or the like.
[0090] Vehicle 700 may include a system-on-a-chip (SoC) 704. The SoC 704 may include a CPU 706, a GPU 708, a processor 710, a cache 712, an accelerator 714, a data store 716, and / or other components and features not shown. The SoC 704 may be used to control vehicle 700 in various platforms and systems. For example, the SoC 704 may be coupled in a system (e.g., a system of vehicle 700) that has an HD map 722 that can obtain map refreshes and / or updates via a network interface 724 from one or more servers (e.g., server 778 in Figure 7D).
[0091] The CPU 706 may include a CPU cluster or CPU complex (also referred to as "CCPLEX"). The CPU 706 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 706 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 706 may include four dual-core clusters, each cluster having its own dedicated L2 cache (e.g., 2MBL2 cache). The CPU 706 (e.g., CCPLEX) may be configured to support concurrent cluster operation, allowing any combination of the CPU 706 clusters to be active at any given time.
[0092] The CPU706 can implement power management capabilities that include one or more of the following features: individual hardware blocks may be automatically clock-gated when idle to conserve dynamic power; each core clock may be gated when a core is not actively executing instructions by executing WFI / WFE instructions; each core may be independently power-gated; each core cluster may be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster may be independently power-gated when all cores are power-gated. The CPU706 can further implement enhanced algorithms for managing power states, where acceptable power states and expected wake-up times are specified, and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing core may support a simplified power state input sequence in software where the work is offloaded to the microcode.
[0093] The GPU708 may include an integrated GPU (or, as referred to herein, "iGPU"). The GPU708 may be programmable and efficient for parallel workloads. In some embodiments, the GPU708 may be able to use an enhanced tensor instruction set. The GPU708 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU708 may include at least eight streaming microprocessors. The GPU708 may be able to use a Computation Application Programming Interface (API). In addition, the GPU708 may be able to use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0094] The GPU708 can be power-optimized for optimal performance in automotive and embedded use cases. For example, the GPU708 can be manufactured on a FinFET (Fin field-effect transistor). However, this is not intended to be a limitation, and the GPU708 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores divided into multiple blocks. Not limited to, for example, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In such an example, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In addition, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mixture of computation and addressing operations. A streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and coordination between concurrent threads. A streaming microprocessor may also include a combined L1 data cache and shared memory unit to simplify programming while improving performance.
[0095] In some examples, the GPU708 may include high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of 900 GB / s. In some examples, in addition to or instead of HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.
[0096] The GPU708 can incorporate unified memory technology, including access counters, to enable more precise movement of memory pages to the processor that most frequently accesses them, thereby improving the efficiency of shared memory ranges across processors. In some examples, address translation service (ATS) support may be used to allow the GPU708 to directly access the CPU706 page table. In such examples, when the GPU708 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU706. In response, the CPU706 can examine its page table for virtual-to-real-address mapping and send the translation back to the GPU708. As such, unified memory technology can enable a single, unified virtual address space for both the CPU706 and GPU708 memory, thereby simplifying GPU708 programming and porting of applications to the GPU708.
[0097] In addition, the GPU708 may include an access counter that can record how often the GPU708 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses that page most frequently.
[0098] The SoC704 may include any number of caches 712, including those described herein. For example, cache 712 may include an L3 cache available to both the CPU706 and the GPU708 (e.g., connected to both the CPU706 and the GPU708). Cache 712 may include a write-back cache that can record line states, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include 4MB or more, depending on the embodiment, although a smaller cache size may be used.
[0099] The SoC704 may include an arithmetic logic unit (ALU) that can be used to perform processing for any of the various tasks or operations of the vehicle 700 (for example, a processing DNN). In addition, the SoC704 may include a floating-point unit (FPU) (or other mass coprocessor or numerical coprocessor type) for performing mathematical operations within the system. For example, the SoC104 may include one or more FPUs integrated as execution units within the CPU706 and / or GPU708.
[0100] The SoC704 may include one or more accelerators 714 (e.g., a hardware accelerator, a software accelerator, or a combination thereof). For example, the SoC704 may include a hardware acceleration cluster that may include an optimized hardware accelerator and / or a large on-chip memory. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to complement the GPU708 and to offload some of the GPU708's tasks (e.g., to free up more cycles of the GPU708 to perform other tasks). As an example, accelerator 714 may be used for target workloads that are sufficiently stable to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs), etc.). In this specification, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (for example, as used for object detection).
[0101] Accelerator 714 (e.g., hardware acceleration cluster) may include a deep learning accelerator (DLA). A DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may also be an accelerator configured and optimized to perform image processing functions (e.g., CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inference. A DLA design can provide more performance per millisecond than a general-purpose GPU and significantly exceed CPU performance. The TPU can perform several functions, including, for example, single-instance convolutional functions supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.
[0102] DLA can quickly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, identification, and detection using data from microphones; CNNs for facial recognition and vehicle owner identification using data from camera sensors; and / or CNNs for security and / or safety-related events.
[0103] DLA can perform any function of GPU708, and by using inference accelerators, for example, a designer can target either DLA or GPU708 for any function. For example, a designer can focus on CNN and floating-point arithmetic processing on DLA, and leave other functions to GPU708 and / or other accelerators 714.
[0104] The accelerator 714 (for example, a hardware accelerator cluster) may include a programmable vision accelerator (PVA), which may also be referred to herein as a computer vision accelerator. A PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. A PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0105] A RISC core can interact with an image sensor (for example, the image sensor of one of the cameras described herein), an image signal processor, and / or similar devices. Each RISC core may contain any amount of memory. Depending on the embodiment, a RISC core may use one of several protocols. In some examples, a RISC core can run a real-time operating system (RTOS). A RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0106] DMA can enable PVA components to access system memory independent of the CPU 706. DMA can support any number of features used to bring optimizations to the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0107] A vector processor may also be a programmable processor that can be designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can act as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single-instruction, multiple-data (SIMD), or very-long instruction word (VLIW) digital signal processor. A combination of SIMD and VLIW can increase throughput and speed.
[0108] Each vector processor may include an instruction cache and be linked to dedicated memory. As a result, in some examples, each vector processor may be configured to run independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA can run the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA can run different computer vision algorithms simultaneously on the same image, or even run different algorithms sequentially on the image or parts of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In addition, a PVA may include additional error correction code (ECC) memory to enhance overall system safety.
[0109] The accelerator 714 (for example, a hardware accelerator cluster) may include a computer vision network on-chip and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 714. In some examples, the on-chip memory may include at least 4 MB of SRAM consisting of eight field-configurable memory blocks, which may be accessible by both the PVA and DLA, for example, and not limited to. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA can access the memory via a backbone that provides the PVA and DLA with high-speed access to the memory. The backbone may include a computer vision network on-chip that interconnects the PVA and DLA to the memory (for example, using an APB).
[0110] A computer vision network on-chip may include an interface that determines whether both the PVA and DLA are activatable and enable signals before any control signals / addresses / data are transmitted. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface may conform to ISO 26262 or IEC 61508 standards, but other standards and protocols may be used.
[0111] In some embodiments, the SoC704 may include a real-time ray tracing hardware accelerator, such as the one described in Patent Document 1. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the location and size of objects (e.g., in a world model) to generate real-time visualization simulations for RADAR signal interpretation, acoustic propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison to LIDAR data for localization and / or other functions, and / or other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.
[0112] The accelerator 714 (e.g., a hardware accelerator cluster) has diverse applications for autonomous driving. The PVA may also be a programmable vision accelerator that can be used in critical processing stages in ADAS and autonomous vehicles. The PVA's capabilities are suitable for areas of algorithms requiring predictable processing at low power and low latency. In other words, the PVA performs well in semi-high density or high density typical computations, even on small data sets, where predictable execution time is required along with low latency and low power. Therefore, because the PVA is efficient in object detection and integer computation, in relation to a platform for autonomous vehicles, the PVA is designed to run classic computer vision algorithms.
[0113] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. While semi-global matching-based algorithms may be used in some examples, this is not intended to be a limitation. Numerous applications for Level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions with input from two monocular cameras.
[0114] In some applications, PVA can be used to perform high-density optical flow by processing raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR data. In other applications, PVA is used for flight depth processing, for example, by processing raw flight data to provide processed flight data.
[0115] DLA can be used to run any type of network to enhance control and driving safety, for example, a neural network that outputs a confidence value for each object detection. Such confidence values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and consider only detections that exceed the threshold as true positives. In an automatic emergency braking (AEB) system, a false positive detection would cause the vehicle to automatically apply the emergency brakes, which is obviously undesirable. Therefore, only the most confident detection should be considered as a trigger for the AEB. DLA can run a neural network that regresses on the confidence values. The neural network can accept at least a subset of parameters as its input, such as bounding box dimensions, ground plane estimation acquired (e.g., from another subsystem), vehicle orientation, distance, inertial measurement unit (IMU) sensor output correlated with 3D position estimation of an object acquired from the neural network and / or other sensors (e.g., LIDAR sensor 764 or RADAR sensor 760), and others.
[0116] The SoC704 may include a data store 716 (for example, memory). The data store 716 may also be the on-chip memory of the SoC704 and can store neural networks that will run on the GPU and / or DLA. In some examples, the data store 716 may have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 712 may comprise an L2 or L3 cache 712. References to the data store 716 may include references to memory associated with the PVA, DLA, and / or other accelerators 714, as described herein.
[0117] The SoC704 may include one or more processors 710 (e.g., integrated processors). The processors 710 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and associated security enforcement. The boot and power management processor may also be part of the SoC704 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of SoC704 thermal and temperature sensors, and / or management of SoC704 power states. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC704 may use the ring oscillators to detect the temperatures of the CPU 706, GPU 708, and / or accelerator 714. If the temperature is determined to have exceeded a threshold, the boot and power management processor may enter a temperature fault routine, placing the SoC704 into a lower power state and / or putting the vehicle 700 into chauffeur safe shutdown mode (for example, safely shutting down the vehicle 700).
[0118] The processor 710 may further include a set of integrated processors that can perform the functions of an audio processing engine. The audio processing engine may also be an audio subsystem that enables full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0119] The processor 710 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timer and interrupt controllers), various I / O controller peripherals, and routing logic.
[0120] The processor 710 may further include a safety cluster engine, which includes a dedicated processor subsystem for handling safety management in automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operations.
[0121] The processor 710 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0122] The processor 710 may further include a high dynamic range signal processor, which may include an image signal processor, a hardware engine that is part of the camera processing pipeline.
[0123] The processor 710 may include a video image synthesizer, which may also be a processing block (for example, implemented on a microprocessor) that implements post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction on the wide-view camera 770, the surround camera 774, and / or the in-cabin surveillance camera sensors. The in-cabin surveillance camera sensors are preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify and appropriately respond to in-cabin events. The in-cabin system can perform lip-reading to activate cellular services and make phone calls, transcribe emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.
[0124] A video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs in the video, noise reduction reduces the weight of information provided by adjacent frames and appropriately weights the spatial information. If the image or part of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer can use information from previous images to reduce noise in the current image.
[0125] The video image synthesizer can also be configured to perform stereo rectification on the input stereo lens frame. Furthermore, the video image synthesizer can be used for user interface compositing when the operating system desktop is in use, so that the GPU708 is not required to continuously render new surfaces. Even when the GPU708 is powered on and actively performing 3D rendering, the video image synthesizer can be used to offload the GPU708 to improve performance and responsiveness.
[0126] The SoC704 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and associated pixel input functions to receive video and input from a camera. The SoC704 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0127] The SoC704 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC704 may be used to process data from cameras (connected, for example, via Gigabit Multimedia Serial Link and Ethernet®), sensors (e.g., Lidar sensor 764, Radar sensor 760, etc., which may be connected via Ethernet®), data from bus 702 (e.g., vehicle speed, steering wheel position, etc.), and data from GNSS sensor 758 (connected, for example, via Ethernet® or CAN bus). The SoC704 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to free up CPU 706 from routine data management tasks.
[0128] The SoC704 may also be an inter-terminal platform with a flexible architecture that spans automation levels 3-5, thereby providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS techniques for diversity and redundancy, and, together with deep learning tools, provides a platform for a flexible, reliable driving software stack. The SoC704 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when the accelerator 714 is coupled with the CPU 706, the GPU 708, and the data store 716 can provide a fast and efficient platform for autonomous vehicles at levels 3-5.
[0129] Therefore, this technology brings capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on a CPU, which can be configured using high-level programming languages such as the C programming language to execute a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time complex object detection algorithms, which are required for in-vehicle ADAS applications and actual Level 3-5 autonomous vehicles.
[0130] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technologies described herein enable multiple neural networks to run simultaneously and / or sequentially, and the results to be combined to enable Level 3–5 autonomous driving capabilities. For example, a DLA or a CNN running on a dGPU (e.g., GPU720) may include text and word recognition, enabling a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of identifying, interpreting, and providing a semantic understanding of signs, and passing that semantic understanding to a route planning module running on the CPU complex.
[0131] As another example, multiple neural networks may run simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of a flashing light and the text "Caution: Flashing light indicates frozen conditions" may be interpreted independently or collectively by several neural networks. The sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing light indicates frozen conditions" may be interpreted by a second deployed neural network that informs the vehicle's route planning software (preferably running on a CPU complex) that frozen conditions are present when flashing light is detected. The flashing light may be identified by informing the vehicle's route planning software of the presence (or absence) of the flashing light, and by operating a third deployed neural network through multiple frames. All three neural networks can run simultaneously within the DLA and / or on the GPU708, for example.
[0132] In some applications, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of the legitimate driver and / or owner of vehicle 700. An always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's side door, and in security mode, to stop the vehicle when the owner leaves the vehicle. In this way, SoC704 provides security against theft and / or vehicle hijacking.
[0133] In another example, a CNN for emergency vehicle detection and identification can detect and identify emergency vehicle sirens using data from microphone 796. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, SoC 704 uses a CNN for classifying environmental and urban sounds, as well as for classifying visual data. In a preferred embodiment, a CNN running on DLA is trained to identify the relative terminal velocity of emergency vehicles (for example, by using the Doppler effect). The CNN may also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by GNSS sensor 758. Thus, for example, when operating in Europe, the CNN would attempt to detect European sirens, and when in the United States, the CNN would attempt to identify only North American sirens. After an emergency vehicle is detected, a control program may be used, with the assistance of ultrasonic sensor 762, to perform emergency vehicle safety routines such as slowing down the vehicle, stopping it at the side of the road, parking the vehicle, and / or idling the vehicle until the emergency vehicle has passed.
[0134] The vehicle may include a CPU 718 (e.g., a separate CPU, or dCPU) which can be connected to the SoC 704 via a high-speed interconnect (e.g., PCIe). The CPU 718 may include, for example, an x86 processor. The CPU 718 may be used to perform any of a variety of functions, including, for example, mediating the consequences of a potential mismatch between ADAS sensors and the SoC 704, and / or monitoring the status and condition of the controller 736 and / or the infotainment SoC 730.
[0135] Vehicle 700 may include a GPU 720 (e.g., a separate GPU, or dGPU) which can be connected to SoC 704 via a high-speed interconnect (e.g., NVIDIA NVLINK). The GPU 720 can provide additional artificial intelligence capabilities, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based on input from sensors in Vehicle 700 (e.g., sensor data).
[0136] Vehicle 700 may further include a network interface 724 which may include one or more wireless antennas 726 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas and Bluetooth® antennas). The network interface 724 may be used to enable wireless connectivity to a cloud over the Internet (e.g., with a server 778 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct link may be established between two vehicles, and / or an indirect link may be established (e.g., over a network and over the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 700 information about vehicles in close proximity to vehicle 700 (e.g., vehicles in front of, beside, and / or behind vehicle 700). This functionality may also be part of the vehicle 700's joint adaptive cruise control functionality.
[0137] The network interface 724 may include an SoC that provides modulation and demodulation functions and enables the controller 736 to communicate over a wireless network. The network interface 724 may include a radio frequency front end for baseband-to-radio frequency upconversion and radio frequency-to-baseband downconversion. Frequency conversion can be performed through well-known processes and / or using superheterodyne processes. In some examples, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communication over LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth®, Bluetooth® LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0138] The vehicle 700 may further include a data store 728 which may include storage outside the chip (for example, outside the SoC 704). The data store 728 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0139] The vehicle 700 may further include GNSS sensors 758. The GNSS sensors 758 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS) sensors, etc.) assist in mapping, perception, occupy grid generation, and / or route planning functions. Any number of GNSS sensors 758 may be used, including, but not limited to, GPS using a USB connector with Ethernet® to a serial (RS-232) bridge.
[0140] Vehicle 700 may further include a RADAR sensor 760. The RADAR sensor 760 may be used by vehicle 700 for long-range vehicle detection, even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 760 may use CAN and / or bus 702 for control and to access object tracking data (for example, to transmit data generated by the RADAR sensor 760) using Ethernet® access for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 760 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor may be used.
[0141] The RADAR sensor 760 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, and short-range side coverage. In some examples, the long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view achieved by two or more independent scans, such as within a range of 250m. The RADAR sensor 760 can help distinguish between static and moving objects and may be used by ADAS systems for emergency brake assist and forward collision warning. The long-range RADAR sensor may include monostatic multimodal RADARs with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the four central antennas may create a focused beam pattern designed to record the area around the vehicle 700 at high speed with minimal interference from traffic in adjacent lanes. The other two antennas can widen the field of view, enabling rapid detection of vehicles entering or leaving the lane of the vehicle 700.
[0142] As an example, a medium-range RADAR system may include a range of up to 760m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 750 degrees (rear). A short-range RADAR system may include, but is not limited to, RADAR sensors designed to be mounted on both ends of the rear bumper. When mounted on both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and beside the vehicle.
[0143] Short-range radar systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0144] The vehicle 700 may further include ultrasonic sensors 762. Positioned on the front, rear, and / or sides of the vehicle 700, the ultrasonic sensors 762 may be used for parking assistance and / or for creating and updating the occupancy grid. A wide variety of ultrasonic sensors 762 may be used, and different ultrasonic sensors 762 may be used for detection at different ranges (e.g., 2.5m, 4m). The ultrasonic sensors 762 may operate at a functional safety level of ASIL B.
[0145] The vehicle 700 may include a LiDAR sensor 764. The LiDAR sensor 764 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 764 may also have a functional safety level of ASIL B. In some examples, the vehicle 700 may include multiple LiDAR sensors 764 (e.g., two, four, six, etc.) that can use Ethernet® (for example, to provide data to a Gigabit Ethernet® switch).
[0146] In some examples, the LIDAR sensor 764 may have the ability to provide a list of objects and their distances within a 360-degree field of view. A commercially available LIDAR sensor 764 may have an advertised range of approximately 700m, for example, with an accuracy of 2cm to 3cm and support for 700Mbps Ethernet® connectivity. In some examples, one or more non-protruding LIDAR sensors 764 may be used. In such examples, the LIDAR sensor 764 may be implemented as a small device that can be incorporated into the front, rear, side, and / or corners of a vehicle 700. In such examples, the LIDAR sensor 764 may have a range of 200m even for low-reflection objects and can provide a field of view up to 120 degrees horizontal and 35 degrees vertical. A front-mounted LIDAR sensor 764 may be configured for a horizontal field of view between 45 and 135 degrees.
[0147] In some applications, LiDAR technologies such as 3D flash LiDAR may also be used. 3D flash LiDAR uses a laser flash as a source to illuminate the area around the vehicle up to approximately 200m. The flash LiDAR unit includes a receptor that records the laser pulse travel time and reflected light on each pixel, sequentially corresponding to the range from the vehicle to the object. Flash LiDAR can enable the generation of high-precision and distortion-free images of the surroundings with every laser flash. In some applications, four flash LiDAR sensors may be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D steering array LiDAR cameras (e.g., non-scanning LiDAR devices) that have no moving parts other than a blower. The flash LiDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-documented intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 764 may be less susceptible to motion blur, vibration, and / or shock.
[0148] The vehicle may further include an IMU sensor 766. In some examples, the IMU sensor 766 may be positioned in the center of the rear axle of the vehicle 700. The IMU sensor 766 may include, but is not limited to, an accelerometer, magnetometer, gyroscope, magnetic compass, and / or other sensor types. In some examples, such as in a 6-axis application, the IMU sensor 766 may include an accelerometer and a gyroscope, while in a 9-axis application, the IMU sensor 766 may include an accelerometer, a gyroscope, and a magnetometer.
[0149] In some embodiments, the IMU sensor 766 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 766 may enable the vehicle 700 to estimate its direction of travel without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 766. In some embodiments, the IMU sensor 766 and the GNSS sensor 758 may be combined in a single integrated unit.
[0150] The vehicle may include a microphone 796 placed inside and / or around the vehicle 700. The microphone 796 may, among other things, be used for emergency vehicle detection and identification.
[0151] The vehicle may further include any number of camera types, including a stereo camera 768, a wide-view camera 770, an infrared camera 772, a surround camera 774, a long-range and / or medium-range camera 798, and / or other camera types. The cameras may be used to capture image data around the entire exterior of the vehicle 700. The type of camera used will depend on the embodiment and requirements of the vehicle 700, and any combination of camera types may be used to achieve the required coverage around the vehicle 700. In addition, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may, as an example, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet®. Each camera is described in more detail herein in relation to Figures 7A and 7B.
[0152] The vehicle 700 may further include a vibration sensor 742. The vibration sensor 742 can measure vibrations of vehicle components, such as axles. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 742 are used, the difference in vibration may be used to determine friction or slippage of the road surface (for example, when the difference in vibration is between a power-driven axle and a free-rotating axle).
[0153] Vehicle 700 may include ADAS system 738. In some examples, ADAS system 738 may include SoC. ADAS system 738 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0154] The ACC system may use a radar sensor 760, a lithium-ion sensor 764, and / or a camera. The ACC system may include longitudinal ACC and / or transverse ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of vehicle 700 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Transverse ACC performs distance maintenance and advises vehicle 700 to change lanes when necessary. Transverse ACC is related to other ADAS applications such as LCA and CWS.
[0155] CACC uses information from other vehicles that can be received from other vehicles via a wireless link through a network interface 724 and / or a wireless antenna 726, or indirectly via a network connection (e.g., via the Internet). Direct links may be provided by vehicle-to-vehicle (V2V) communication links, while indirect links may be infrastructure-to-vehicle (I2V) communication links. Generally, the V2V communication concept provides information about the vehicle immediately ahead (e.g., a vehicle in the same lane as vehicle 700, immediately in front of vehicle 700), while the I2V communication concept provides information about traffic further ahead. A CACC system may include either or both I2V and V2V information sources. Given information about vehicles ahead of vehicle 700, CACC can be more reliable, and CACC has the potential to make traffic flow smoother and reduce road congestion.
[0156] The FCW system is designed to warn the driver of hazards so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or radar sensor 760, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback such as a display, speaker, and / or vibration components. The FCW system can provide warnings in the form of audible, visual, vibration, and / or quick brake pulses.
[0157] An AEB system can detect an imminent forward collision with another vehicle or object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system may use a forward-facing camera and / or radar sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system may automatically apply the brakes as part of an effort to prevent, or at least mitigate, the impact of the anticipated collision. The AEB system may include techniques such as dynamic brake support and / or impending collision braking.
[0158] The LDW system warns the driver when the vehicle 700 crosses a lane marking by providing visual, audible, and / or tactile warnings, such as vibration of the steering wheel or seat. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn signal. The LDW system may use a forward-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration components.
[0159] The LKA system is a modified version of the LDW system. The LKA system provides steering input or braking to correct the vehicle 700 if it begins to drift out of its lane.
[0160] The BSW system detects and warns the driver of a vehicle in its blind spots. The BSW system can provide visual, audible, and / or tactile warnings to indicate that merging or changing lanes is unsafe. The system may provide additional warnings when the driver uses the turn signals. The BSW system may use a rear-facing camera and / or radar sensor 760, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0161] The RCTW system can provide visual, audible, and / or haptic notifications when an object is detected outside the range of the rear camera while the vehicle 700 is reversing. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a collision. The RCTW system may use one or more rear-facing RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0162] Conventional ADAS systems warn the driver and allow the driver to determine whether a safe condition truly exists and act accordingly. However, conventional ADAS systems have sometimes tended to produce misjudgments that, while not usually catastrophic, can be troubling and distracting to the driver. In the autonomous vehicle 700, however, if the results are contradictory, the vehicle 700 itself must decide whether to heed the results from the primary computer or the secondary computer (e.g., the first controller 736 or the second controller 736). For example, in some embodiments, the ADAS system 738 may also be a backup and / or secondary computer to provide perceptual information to a backup computer rationality module. The backup computer rationality monitor can run a variety of redundant software on hardware components to detect failures in perceptual and dynamic driving tasks. The output from the ADAS system 738 may be provided to the supervisory MCU. If the outputs from the primary and secondary computers are contradictory, the supervisory MCU must decide how to reconcile the contradiction to ensure safe operation.
[0163] In some implementations, a primary computer may be configured to provide a supervising MCU with a reliability score indicating the reliability of the primary computer in a selected outcome. If the reliability score exceeds a threshold, the supervising MCU may follow the primary computer's instructions, regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold, and the primary and secondary computers produce different results (e.g., conflicting results), the supervising MCU may mediate between the computers to determine an appropriate outcome.
[0164] The supervisory MCU may be configured to run a neural network trained and configured to determine, based on the outputs from the primary and secondary computers, when a secondary computer is providing a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer is reliable and when it is not. For example, when the secondary computer is a radar-based forward crossing (FCW) system, the neural network in the supervisory MCU can learn when the FCW is identifying metal objects that are not actually dangerous, such as sewer grates or manhole covers that trigger an alarm. Similarly, when the secondary computer is a camera-based lane departure warning (LDW) system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest operation. In embodiments involving a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running a neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise and / or be included as a component of the SoC704.
[0165] In other examples, ADAS system 738 may include a secondary computer that performs ADAS functions using conventional rules of computer vision. As such, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network within the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identities make the entire system more fault-tolerant, particularly to failures caused by software (or software-hardware interface) functions. For instance, if a software bug or error exists in the software running on the primary computer, and non-identical software code running on the secondary computer produces the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer did not cause a critical error.
[0166] In some examples, the output of the ADAS system 738 may be supplied to the perception block and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 738 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network, which is trained as described herein and therefore reduces the risk of misjudgment.
[0167] Vehicle 700 may further include an infotainment SoC 730 (for example, an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system does not have to be an SoC and may include two or more separate components. The infotainment SoC 730 may include a combination of hardware and software that can be used to provide vehicle 700 with audio (e.g., music, personal digital assistant, navigation commands, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assist, radio data system, fuel level, total mileage, brake fuel level, oil level, door open / close, air filter information, and other vehicle-related information). For example, the infotainment SoC 730 may also include wireless, disc player, navigation system, video player, USB and Bluetooth® connectivity, car computer, in-car entertainment, Wi-Fi, steering wheel audio control unit, hands-free voice control, heads-up display (HUD), HMI display 734, telematics device, control panel (for example, for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 730 may be further used to provide information (for example, visual and / or audible) to the vehicle user, such as information from the ADAS system 738, autonomous driving information such as planned vehicle operation, trajectory, surrounding environment information (for example, intersection information, vehicle information, road information, etc.), and / or other information.
[0168] The infotainment SoC 730 may include GPU functionality. The infotainment SoC 730 can communicate with other devices, systems, and / or components of the vehicle 700 via bus 702 (e.g., CAN bus, Ethernet®, etc.). In some examples, the infotainment SoC 730 may be coupled to a supervisory MCU so that the infotainment system's GPU can perform certain self-drive functions in the event of a primary controller 736 (e.g., the vehicle 700's primary and / or backup computer) failure. In such examples, the infotainment SoC 730 can put the vehicle 700 into a chauffeur-safe stop mode as described herein.
[0169] Vehicle 700 may further include an instrument cluster 732 (e.g., a digital dash, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 732 may include a controller and / or a supercomputer (e.g., a separate controller or supercomputer). The instrument cluster 732 may include a set of instruments such as a speedometer, fuel level indicator, oil pressure indicator, tachometer, odometer, turn signals, gear shift position indicator, seat belt warning light, parking brake warning light, engine fault light, airbag (SRS) system information, lighting control device, safety system control device, and navigation information. In some examples, information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. In other words, the instrument cluster 732 may be included as part of the infotainment SoC 730, and vice versa.
[0170] Figure 7D is a system diagram of communication between the cloud-based server of Figure 7A and an exemplary autonomous vehicle 700, according to some embodiments of the present disclosure. System 776 may include a server 778, a network 790, and a vehicle including the vehicle 700. Server 778 may include a plurality of GPUs 784(A) to 784(H) (collectively referred to herein as GPU 784), PCIe switches 782(A) to 782(H) (collectively referred to herein as PCIe switch 782), and / or CPUs 780(A) to 780(B) (collectively referred to herein as CPU 780). The GPUs 784, CPUs 780, and PCIe switches may be interconnected by high-speed interconnects, such as, for example, NVLink interfaces 788 and / or PCIe connections 786 developed by NVIDIA. In some examples, the GPU784 is connected via NVLink and / or NVSwitch SoCs, and the GPU784 and PCIe switch 782 are connected via PCIe interconnects. Eight GPU784s, two CPU780s, and two PCIe switches are illustrated, but this is not intended to be an limitation. Depending on the embodiment, each server 778 may contain any number of GPU784s, CPU780s, and / or PCIe switches. For example, server 778 may contain eight, sixteen, thirty-two, and / or more GPU784s, respectively.
[0171] Server 778 can receive image data from vehicles via network 790, representing images showing unexpected or altered road conditions, such as recently started road construction. Server 778 can transmit map information 794, including information about traffic and road conditions, to vehicles via network 790, including information about neural networks 792, updated neural networks 792, and / or map information 794. Updates to map information 794 may include updates to HD maps 722, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In some instances, neural networks 792, updated neural networks 792, and / or map information 794 may have arisen from new training and / or experience represented in data received from any number of vehicles in the environment, and / or based on training performed in a data center (for example, using server 778 and / or other servers).
[0172] Server 778 may be used to train a machine learning model (e.g., a neural network) based on training data. The training data may be generated by a vehicle and / or in a simulation (e.g., using a game engine). In some instances, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or otherwise pre-processed, while in other instances, the training data is not tagged and / or pre-processed (e.g., if the neural network does not require supervised learning). Training may be performed according to any one or more classes of machine learning techniques, including but not limited to the following: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including key component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine learning model has been traced, it may be used by the vehicle (for example, transmitted to the vehicle via network 790), and / or the machine learning model may be used by server 778 to remotely monitor the vehicle.
[0173] In some examples, Server 778 can receive data from vehicles and apply it to state-of-the-art real-time neural networks for real-time intelligent inference. Server 778 may include deep learning supercomputers and / or dedicated AI computers powered by GPU 784, such as the DGX and DGX Station Machines developed by NVIDIA. However, in some examples, Server 778 may include deep learning infrastructure that uses only CPU-powered data centers.
[0174] The deep learning infrastructure of server 778 can have the capability for high-speed real-time inference, which can be used to evaluate and verify the condition of the processor, software, and / or associated hardware within vehicle 700. For example, the deep learning infrastructure can receive periodic updates from vehicle 700, such as images of a sequence and / or objects located within images of that sequence (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them with objects identified by vehicle 700. If the results do not match and the infrastructure concludes that the AI within vehicle 700 is not functioning correctly, server 778 can send a signal to vehicle 700 instructing the vehicle's fail-safe computer to infer control, notify passengers, and complete a safe parking operation.
[0175] For inference, server 778 may include GPU 784 and one or more programmable inference accelerators (e.g., NVIDIA TensorRT). The combination of a GPU-powered server and inference accelerator can enable real-time responsiveness. In other examples, such as when performance is not a major requirement, a server powered by a CPU, FPGA, and other processors may be used for inference.
[0176] Exemplary computing devices Figure 8 is a block diagram of an example of a computing device 800 suitable for use in implementing some embodiments of the present disclosure. The computing device 800 may include an interconnection system 802 that indirectly or directly connects the following devices: memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply device 816, one or more presentation components 818 (e.g., a display), and one or more logical units 820. In at least one embodiment, the computing device 800 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). As an unrestricted example, one or more of the GPUs 808 may include one or more vGPUs, one or more of the CPUs 806 may include one or more vCPUs, and / or one or more of the logical units 820 may include one or more virtual logical units. As such, the computing device 800 may include individual components (e.g., an entire GPU dedicated to the computing device 800), virtual components (e.g., a portion of a GPU dedicated to the computing device 800), or a combination thereof.
[0177] The various blocks in Figure 8 are shown connected by lines via the interconnection system 802, but this is not intended to be restrictive and is simply for clarity. For example, in some embodiments, a presentation component 818, such as a display device, could be considered an I / O component 814 (for example, if the display is a touchscreen). As another example, the CPU 806 and / or GPU 808 may include memory (for example, memory 804 may represent a storage device in addition to the memory of the GPU 808, CPU 806, and / or other components). In other words, the computing devices in Figure 8 are merely illustrative. Categories such as “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “handheld device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types are all intended to fall within the scope of the computing devices in Figure 8 and are therefore not distinguished.
[0178] The interconnection system 802 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 802 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a VESA (video electronics standards association) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or other types of buses or links. In some embodiments, direct connections exist between components. For example, the CPU 806 may be directly connected to the memory 804. Furthermore, the CPU 806 may be directly connected to the GPU 808. Where direct or point-to-point connections exist between components, the interconnection system 802 may include PCIe links to implement the connections. In these examples, the PCI bus does not need to be included in the computing device 800.
[0179] Memory 804 may include any of various computer-readable media. The computer-readable media may be any available media accessible by the computing device 800. The computer-readable media may include both volatile and non-volatile media, and removable and non-removable media. For example, but not limited to, the computer-readable media may include computer storage media and communication media.
[0180] Computer storage media may include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 may store computer-readable instructions (e.g., representing programs and / or program elements), such as an operating system. Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store desired information and can be accessed by computing device 800. In this specification, computer storage media does not include signals themselves.
[0181] Computer storage media include any information distribution medium that can implement computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transfer mechanisms. The term “modulated data signal” may refer to a signal that has been modified in a manner that has one or more of its characteristic sets or encodes information within the signal. For example, but not limited to, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included in the scope of computer-readable media.
[0182] The CPU 806 may be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to execute one or more of the methods and / or processes described herein. The CPU 806 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) each capable of processing a large number of software threads concurrently. The CPU 806 may include any type of processor, and depending on the type of computing device 800 in which it is implemented, it may include different types of processors (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers). For example, depending on the type of computing device 800, the processor may be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 800 may include one or more CPUs 806 in one or more microprocessors or auxiliary coprocessors, such as a computing coprocessor.
[0183] In addition to or instead of the CPU 806, the GPU 808 may be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to execute one or more of the methods and / or processes described herein. One or more of the GPU 808 may be an integrated GPU (for example, with one or more of the CPU 806), and / or one or more of the GPU 808 may be a discrete GPU. In embodiments, one or more of the GPU 808 may be a coprocessor of one or more of the CPU 806. The GPU 808 may be used by the computing device 800 to render graphics (for example, 3D graphics) or to perform general-purpose computing. For example, the GPU 808 may be used for GPU-based general-purpose computing (GPGPU). It can be used for a GPU. The GPU808 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU808 can generate pixel data for an output image in response to rendering commands (for example, rendering commands from CPU806 received via the host interface). The GPU808 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 804. The GPU808 may include two or more GPUs operating in parallel (for example, via a link). The link can connect directly to the GPUs (for example, using NVLINK) or via a switch (for example, using NVSwitch). When coupled together, each GPU808 can generate pixel data or GPGPU data for different parts of an output or for different outputs (for example, the first GPU for the first image and the second GPU for the second image). Each GPU may have its own memory or may share memory with other GPUs.
[0184] In addition to or instead of the CPU 806 and / or GPU 808, the logic unit 820 may be configured to execute at least some computer-readable instructions to control one or more of the computing devices 800 to execute one or more of the methods and / or processes described herein. In embodiments, the CPU 806, GPU 808, and / or logic unit 820 can execute any combination of methods, processes, and / or parts thereof discretely or congruently. One or more of the logic units 820 may be part of and / or integrated with one or more of the CPU 806 and / or GPU 808, and / or one or more of the logic units 820 may be discrete components of the CPU 806 and / or GPU 808 or otherwise external to them. In embodiments, one or more of the logic units 820 may be coprocessors of one or more of the CPU 806 and / or one or more of the GPU 808.
[0185] Examples of the logic unit 820 include one or more processing cores and / or components thereof, such as Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Logical Units (ALUs), Application-Specific Integrated Circuits (ASICs), Floating-Point Units (FPUs), Input / Output (I / O) elements, Peripheral Component Interconnect (PCI) or Peripheral Component Interconnect Express (PCIe) elements, and / or similar.
[0186] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 800 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 810 may include components and functions to enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth®, Bluetooth® LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet® or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0187] I / O port 812 can enable the computing device 800 to be logically connected to other devices, including I / O components 814, presentation components 818, and / or other components, some of which can be built into (e.g., integrated into) the computing device 800. Exemplary I / O components 814 include microphones, mice, keyboards, joysticks, gamepads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. I / O components 814 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input may be transmitted to appropriate network elements for further processing. The NUI may implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, on-screen and beside-screen gesture recognition, air gestures, head and target tracking, and touch recognition related to the display of the computing device 800 (as described in more detail below). The computing device 800 may include depth cameras, such as stereoscope camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 800 may include accelerometers or gyroscopes to enable motion detection (for example, as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by the computing device 800 to render immersive augmented reality or virtual reality.
[0188] The power supply device 816 may include a hardwired power supply device, a battery power supply device, or a combination thereof. The power supply device 816 can provide power to the computing device 800 to enable the components of the computing device 800 to operate.
[0189] The presentation component 818 may include a display (e.g., a monitor, touch screen, television screen, head-up display device (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 818 can receive data from other components (e.g., GPU 808, CPU 806, etc.) and output data (e.g., as images, videos, sounds, etc.).
[0190] Exemplary data center Figure 9 shows an exemplary data center 900 that may be used in at least one embodiment of the present disclosure. The data center 900 may include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and / or an application layer 940.
[0191] As shown in Figure 9, the data center infrastructure layer 910 may include a resource orchestrator 912, grouped computing resources 914, and node computing resources ("node CRs") 916(1) to 916(N), where "N" represents any integer or natural number. In at least one embodiment, the node CRs 916(1) to 916(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and / or cooling modules. In some embodiments, one or more nodes CR916(1) to 916(N) may correspond to a server having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR916(1) to 9161(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or similar, and / or one or more nodes CR916(1) to 916(N) may correspond to a virtual machine (VM).
[0192] In at least one embodiment, the grouped computing resources 914 may include a separate group of nodes CR916 housed in one or more racks (not shown), or a number of racks housed in data centers in various geographical locations (also not shown). The separate group of nodes CR916 within the grouped computing resources 914 may include grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several nodes CR916, including CPUs, GPUs, and / or other processors, may be grouped in one or more racks to provide computing resources to support one or more workloads. The one or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0193] The resource orchestrator 922 can configure or otherwise control one or more nodes CR916(1) to 916(N) and / or grouped computing resources 914. In at least one embodiment, the resource orchestrator 922 may include a software design infrastructure ("SDI") management entity for the data center 900. The resource orchestrator 922 may include hardware, software, or any combination thereof.
[0194] In at least one embodiment, as shown in Figure 9, the framework layer 920 may include a job scheduler 932, a configuration manager 934, a resource manager 936, and / or a distributed file system 938. The framework layer 920 may include a framework to support software 932 of the software layer 930 and / or one or more applications 942 of the application layer 940. The software 932 or application 942 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure, respectively. The framework layer 920 may be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark® ("Spark"), which can use the distributed file system 938 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 932 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 900. The configuration manager 934 may have the ability to configure different layers, for example, a software layer 930 and a framework layer 920 that includes Spark and a distributed file system 938 to support large-scale data processing. The resource manager 936 may have the ability to manage clustered or grouped computing resources that are mapped or allocated for support of the distributed file system 938 and the job scheduler 932. In at least one embodiment, the clustered or grouped computing resources may include computing resources 914 grouped in the data center infrastructure layer 910. The resource manager 1036 can coordinate with the resource orchestrator 912 to manage these mapped or allocated computing resources.
[0195] In at least one embodiment, the software 932 included in the software layer 930 may include software used by at least a portion of nodes CR916(1) to 916(N), grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0196] In at least one embodiment, the application 942 included in the application layer 940 may include one or more types of applications used by at least a portion of the nodes CR916(1) to 916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0197] In at least one embodiment, any of the configuration manager 934, resource manager 936, and resource orchestrator 912 may implement any number and type of self-rewriting actions based on any amount and type of data obtained in any technically possible manner. Self-rewriting actions may free the data center operator of data center 900 from making potentially poor configuration decisions and possibly avoiding underutilized and / or underperforming parts of the data center.
[0198] The data center 900 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters by a neural network architecture using the software and / or computing resources described herein with respect to the data center 900. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described herein with respect to the data center 900 by using weight parameters calculated via one or more training techniques, not limited to those described herein.
[0199] In at least one embodiment, the data center 900 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) for performing training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as services that enable users to train or perform inference of information, such as image recognition, speech recognition, or other artificial intelligence services.
[0200] Exemplary network environment A network environment suitable for use in implementing the embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Each client device, server, and / or other device type (e.g., each device) may be implemented as one or more instances of the computing device 800 in Figure 8, for example, each device may include similar components, features, and / or functionalities of the computing device 800. In addition, if backend devices (e.g., servers, NAS, etc.) are implemented, they may be included as part of the data center 900, examples of which are further detailed herein with respect to Figure 9.
[0201] Components of a network environment may communicate with one another via the network, either wired, wirelessly, or both. A network may include multiple networks, or a network of networks. For example, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. If a network includes a wireless telecommunications network, its components, such as base stations, towers, or access points (and other components), may provide wireless connectivity.
[0202] Compatible network environments may include one or more peer-to-peer network environments (in which case servers may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to the server can be implemented on any number of client devices.
[0203] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, or a combination thereof. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support the software in the software layer and / or one or more applications in the application layer. The software or applications may each include web-based service software or applications. In the embodiment, one or more of the client devices may use the web-based service software or applications (for example, by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software web application framework that may use a distributed file system for, for example, large-scale data processing (e.g., “big data”).
[0204] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more of them) described herein. Any of these various functions may be distributed across multiple locations from a central or core server (e.g., one or more data centers that may be distributed across states, territories, countries, or the world). If the connection to the user (e.g., a client device) is relatively close to the edge server, the core server may delegate at least a portion of its functionality to the edge server. The cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to multiple organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0205] A client device may include at least some of the components, features, and functionalities of the exemplary computing device 800 described herein with respect to Figure 8. As an example, and not limited to, a client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, airship, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, instrument, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0206] This disclosure may be described in general terms with computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by computers or other machines, such as personal digital assistants or other handheld devices. Generally, a program module, including routines, programs, objects, components, and data structures, refers to code that performs a specific task or implements a specific abstract data type. This disclosure may be implemented in a variety of configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked over a communication network.
[0207] In this specification, any “and / or” statement relating to two or more elements should be interpreted as meaning only one element or a combination of elements. For example, “element A, element B, and / or element C” may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. In addition, “at least one of element A or element B” may include at least one element A, at least one element B, or at least one element A and at least one element B. Furthermore, “at least one of element A and element B” may include at least one element A, at least one element B, or at least one element A and at least one element B.
[0208] The subject matter of this disclosure is described in a manner that is specific in order to satisfy statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors intend that the claimed subject matter may be carried out in other ways, including different steps or combinations of steps similar to those described herein, in conjunction with other current or future technologies. Furthermore, the terms “step” and / or “block” may be used herein to imply different elements of the way in which they are used, but these terms should not be construed as implying any particular order among the various steps disclosed herein unless the order of the individual steps is expressly stated and, when so, expressed.
Claims
1. A method performed by one or more processing devices, A step of selecting a first subset of first data, which includes training data or at least one of the components for a first target tracking model, based on at least a set of characteristics of the deployment environment, The steps include: retraining at least a portion of the first target tracking model as a customized target tracking model for the deployment environment using a first subset of the first data; A second step of transmitting data to deploy the aforementioned target tracking model to the deployment environment. Includes, The aforementioned matching target tracking model comprises a first matching target tracking model, and the method is A step of receiving feedback data for the first conforming target tracking model, wherein the feedback data corresponds to the deployment environment, A step of generating a second fitted target tracking model, at least based on retraining the first fitted target tracking model using the aforementioned feedback data, A third step of transmitting data to deploy the second conforming target tracking model in the deployment environment. Methods that further include this.
2. The method according to claim 1, wherein the step of selecting a first subset of the first data is based at least on surround scene data associated with the set of features of the deployment environment, wherein the surround scene data represents a category of data retrieved during surround scene data acquisition using multiple sensors.
3. The method according to claim 1, wherein the first subset of the first data includes training data corresponding to categories of data from each of one or more bench setups and one or more in-vehicle setups.
4. The method according to claim 1, wherein a first subset of the first data includes the training data corresponding to one or more of the set of features of the deployment environment.
5. The first target tracking model comprises a first modified model of a plurality of target tracking modified models, and the plurality of target tracking modified models are Normalized gaze data is generated at least by aligning the axes of gaze data from the camera with faces or eyes depicted within the gaze data. The tagged gaze features of the normalized gaze data, which determine the tagged gaze features corresponding to one or more gaze vectors or gaze points in the corresponding camera space, The first deformation model of the plurality of target tracking deformation models is trained using the first subset of the tagged gaze features. The method according to claim 1, wherein the product is generated based on at least the above.
6. The method according to claim 1, wherein the first target tracking model comprises a subset of the target tracking deformation models of the first data, and the selecting step comprises a step of selecting the first subset from the target tracking deformation models based at least on a face landmark neural network criterion corresponding to the set of features of the deployment environment.
7. The method according to claim 1, wherein the set of features of the deployment environment includes one or more of the following: a set of spatial configuration features, a set of DNN basic configuration features, or a set of gaze type configuration features.
8. The method according to claim 1, wherein a first subset of the first data includes the training data, the set of features includes markings of gaze angle ranges associated with the deployment environment, and the selecting step is at least based on identifying one or more portions of the training data corresponding to the gaze angle ranges.
9. The method described above, The steps include selecting a second subset of the first data corresponding to a second deployment environment, The steps include generating a third adapted target tracking model, adapted to the second deployment environment, using the second subset of the first data, based at least on retraining one or more portions of the second target tracking model, A fourth step of transmitting data to deploy the third conforming target tracking model to the second deployment environment. The method according to claim 1, further comprising:
10. The method according to claim 9, wherein the deployment environment is associated with a first vehicle type, and the second deployment environment is associated with a second vehicle type.
11. Identify a first subset of data that corresponds to a set of features of the deployment environment and includes training data for a target tracking model or at least one of one or more components, As a customized target tracking model for the aforementioned deployment environment, at least a portion of the target tracking model is retrained using a subset of the first data. The second data is transmitted to deploy the aforementioned target tracking model to the deployment environment. A processor comprising one or more circuits, The aforementioned target tracking model comprises a first target tracking model, and the one or more circuits further comprises Feedback data for the first conforming target tracking model, which receives feedback data corresponding to the deployment environment, A second matching target tracking model is generated, at least based on retraining the first matching target tracking model using the aforementioned feedback data. A processor that transmits third data for deploying the second conforming target tracking model to the deployment environment.
12. The aforementioned processor, Control systems for autonomous or semi-autonomous machines Perceptual systems for autonomous or semi-autonomous machines, A system for performing simulation operations. A system for performing deep learning operations. Systems implemented using edge devices, Systems implemented using robots, A system that incorporates one or more virtual machines (VMs). A system that is at least partially implemented in a data center, or A system that is at least partially implemented using cloud computing resources. The processor according to claim 11, which is included in at least one of the following.
13. The processor according to claim 11, wherein the identification of the subset of the first data is based at least on a gaze angle range associated with the deployment environment.
14. The processor according to claim 11, wherein the target tracking model includes a subset of multiple target tracking deformation models of the first data, based on at least the set of features of the deployment environment.
15. The system comprises one or more processing devices and one or more memory devices that are communication-connected to the one or more processing devices and store program instructions that, when executed by the one or more processing devices, cause the one or more processing devices to perform an action, A step of providing first data corresponding to a set of characteristics of the deployment environment, The steps include receiving a fitted target tracking model determined at least on the basis of selecting a first subset of data, which includes training data for a target tracking model or at least one of one or more components, based at least on the set of features of the deployment environment, and retraining at least a portion of the target tracking model as a fitted target tracking model customized for the deployment environment using the first subset of data; A second step of transmitting data to deploy the aforementioned target tracking model to the deployment environment. Includes, The aforementioned adaptive target tracking model comprises a first adaptive target tracking model, and the operation is, A step of providing feedback data to the first conforming target tracking model, wherein the feedback data corresponds to the deployment environment. The steps include receiving a second fitted target tracking model generated at least on the basis of retraining the first fitted target tracking model using the feedback data, A third step of transmitting data to deploy the second conforming target tracking model in the deployment environment. A system that further includes this.
16. The system according to claim 15, wherein the first data is provided from an in-vehicle setup associated with the deployment environment.
17. The system according to claim 15, wherein the first data is provided from a bench setup associated with the deployment environment.