Line of sight detection using one or more neural networks
By creating a virtual camera space within the vehicle and bridging the vehicle coordinate system using reference markers, the problem of insufficient data for training the gaze detection neural network is solved, achieving both accuracy and adaptability in gaze detection under different camera positions.
Patent Information
- Application Number
- CN202010089073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-19
- Filing Date
- 2020-02-12
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-06-21
AI Technical Summary
Existing technologies lack ground-based data when training neural networks for vehicle occupant gaze detection, resulting in insufficient training data and difficulty in adapting to changes in different camera positions, which affects the accuracy and deployment efficiency of the model.
By creating a virtual camera space, bridging the vehicle coordinate system and the camera coordinate system using benchmarks and calibration, ground reality data is generated, and a deep neural network model is trained to enable effective deployment at different camera locations.
It enables gaze detection at any camera position, improving the model's adaptability and accuracy, and ensuring the effective deployment and precision of the gaze detection system in different vehicle environments.
Smart Images

Figure CN112389443B_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment of the present application relates to gaze detection using one or more neural networks. For example, at least one embodiment relates to a processor or computing system for training a neural network according to various new techniques described in this application. Background Art
[0002] Advances in computer technology have led to improvements in object recognition and analysis capabilities. For this purpose, machine learning has been used as a tool to detect objects in image data. However, to train machine learning models to perform object recognition, conventional methods require large amounts of labeled training data, where supervised training data includes ground truth data. Creating this training data can be a lengthy and complex process, which can be prohibitively expensive for various purposes and may result in insufficient training data. Summary of the Invention
[0003] Devices, systems, and techniques are described for determining the location of an object using an image containing a digital representation of the object. In at least one embodiment, the line of sight of one or more occupants of a vehicle is determined independently of the position of one or more sensors used to detect the one or more occupants of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:
[0005] Figure 1A 、 1B , 1C and 1D illustrate views of a vehicle environment according to at least one embodiment;
[0006] Figure 2 illustrates vehicle and camera coordinate systems that may be bridged according to at least one embodiment;
[0007] Figure 3 shows an arrangement of fiducial markers for a region of interest according to at least one embodiment;
[0008] Figure 4A and Figure 4B illustrates portions of an example process for determining a location of an object in accordance with at least one embodiment;
[0009] Figure 5 A process for determining line of sight according to at least one embodiment is shown;
[0010] Figure 6 illustrates a high-level system architecture according to at least one embodiment;
[0011] Figure 7 illustrates a system architecture according to at least one embodiment;
[0012] Figure 8 shows a front portion of a cabin according to at least one embodiment;
[0013] Figure 9 illustrates an MSCM according to at least one embodiment;
[0014] Figure 10 shows a driver user interface and configuration according to at least one embodiment;
[0015] Figure 11 A flow chart illustrating a method for line of sight estimation according to at least one embodiment is shown;
[0016] Figure 12 shows a pipeline of a neural network according to at least one embodiment;
[0017] Figure 13 A gaze detection DNN for classifying a driver's gaze is shown according to at least one embodiment;
[0018] Figure 14 An example environment is shown in accordance with at least one embodiment;
[0019] Figure 15 An example system for training an image synthesis network that can be utilized in accordance with at least one embodiment is shown;
[0020] Figure 16 illustrates some layers of an example statistical model that may be utilized in accordance with at least one embodiment;
[0021] Figure 17A Inference and / or training logic according to at least one embodiment is shown;
[0022] Figure 17B Inference and / or training logic according to at least one embodiment is shown;
[0023] Figure 18 An example data center system is shown in accordance with at least one embodiment;
[0024] Figure 19A An example of an autonomous vehicle according to at least one embodiment is shown;
[0025] Figure 19B According to at least one embodiment, Figure 19A Examples of camera positions and fields of view for autonomous vehicles;
[0026] Figure 19C According to at least one embodiment, Figure 19A Example system architecture for an autonomous vehicle;
[0027] Figure 19D A method for connecting a server to a cloud-based server according to at least one embodiment is shown. Figure 19A Systems for communication between autonomous vehicles;
[0028] Figure 20 A computer system according to at least one embodiment is shown;
[0029] Figure 21 A computer system according to at least one embodiment is shown;
[0030] Figure 22 A computer system according to at least one embodiment is shown;
[0031] Figure 23 A computer system according to at least one embodiment is shown;
[0032] Figure 24A A computer system according to at least one embodiment is shown;
[0033] Figure 24B A computer system according to at least one embodiment is shown;
[0034] Figure 24C A computer system according to at least one embodiment is shown;
[0035] Figure 24D A computer system according to at least one embodiment is shown;
[0036] Figure 24E and 24F illustrates a shared programming model in accordance with at least one embodiment;
[0037] Figure 25 An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment;
[0038] Figures 26A-26B An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment;
[0039] Figures 27A-27B Additional exemplary graphics processor logic is shown in accordance with at least one embodiment;
[0040] Figure 28 A computer system according to at least one embodiment is shown;
[0041] Figure 29A A parallel processor according to at least one embodiment is shown;
[0042] Figure 29B shows a partition unit according to at least one embodiment;
[0043] Figure 29C illustrates a processing cluster according to at least one embodiment;
[0044] Figure 29D A graphics multiprocessor is shown in accordance with at least one embodiment;
[0045] Figure 30 A multi-graphics processing unit (GPU) system is shown in accordance with at least one embodiment;
[0046] Figure 31 A graphics processor according to at least one embodiment is shown;
[0047] Figure 32 shows a microarchitecture of a processor according to at least one embodiment;
[0048] Figure 33 A deep learning application processor according to at least one embodiment is shown;
[0049] Figure 34 An exemplary neuromorphic processor is shown in accordance with at least one embodiment;
[0050] Figure 35 and 36 illustrates at least a portion of a graphics processor according to one or more embodiments;
[0051] Figure 37 illustrates at least a portion of a graphics processor core according to at least one embodiment;
[0052] Figures 38A-38B illustrates at least a portion of a graphics processor core according to at least one embodiment;
[0053] Figure 39 illustrates a parallel processing unit ("PPU") in accordance with at least one embodiment;
[0054] Figure 40 illustrates a general processing cluster ("GPC") in accordance with at least one embodiment;
[0055] Figure 41 illustrates a memory partitioning unit of a parallel processing unit ("PPU") according to at least one embodiment; and
[0056] Figure 42 A streaming multiprocessor in accordance with at least one embodiment is shown. DETAILED DESCRIPTION
[0057] In at least one embodiment, object localization can be performed in a space, such as an interior cabin of a vehicle, as Figure 1A100 . In at least one embodiment, a vehicle can include various components, such as a steering wheel, dashboard, windshield, and rearview mirrors, that a driver 104 or other occupant can look at while in the vehicle. In at least one embodiment, a determination can be made regarding the location of the driver 104 using image or video data captured using at least one camera. In at least one embodiment, cameras can be placed at various locations in the vehicle, such as camera 102 attached to a rearview mirror or camera 106 placed on the dashboard. In at least one embodiment, these cameras can capture image data including a representation of the driver or other occupant of the vehicle, and the image data can be analyzed to determine information about the driver. In at least one embodiment, this information can include the location, posture, head position, and gaze direction of the driver or other occupant.
[0058] In at least one embodiment, gaze direction information can be used to determine what an occupant of a vehicle is looking at at a given point in time. In at least one embodiment, this can include determining whether the driver is looking toward the windshield while driving or looking away from the windshield, such as toward a rearview mirror, the dashboard, or elsewhere in the vehicle. In at least one embodiment, such information can be used to determine whether the driver is distracted or unable to see action occurring around the vehicle, such as a pedestrian walking in front of the vehicle. In at least one embodiment, if the vehicle monitoring system is able to determine that the driver is not looking away from the windshield and may not have had sufficient time to see the pedestrian to stop the vehicle, the vehicle monitoring system can determine to automatically or autonomously stop the vehicle to avoid striking the pedestrian.
[0059] However, in at least one embodiment, the location at which the occupant is looking requires knowledge of the relative positions of multiple objects in the vehicle. In at least one embodiment, the monitoring system will need to have information about the positions of objects (such as windshields or mirrors) in order to be able to determine whether the occupant is looking at a particular object or whether the gaze direction will intersect with that object. In at least one embodiment where the vehicle has a fixed camera and a set of objects, the model can be generated and trained based on these known positions, which can provide ground truth for training. However, in at least one embodiment, the cameras or sensors used to capture such information can be located at various locations within or near the vehicle. In at least one embodiment, these cameras can be movable or optional within the vehicle. In at least one embodiment, the mobility of the device can be problematic because there is no fixed relationship to objects or locations within the vehicle or other such area.
[0060] In at least one embodiment, a neural network or machine learning can be trained to infer information (such as gaze direction or position) based at least in part on captured image or video data from one or more of these cameras or sensors. In at least one embodiment, because the camera can be placed anywhere within the vehicle, it is necessary to train the neural network using training data for various locations within the vehicle. In at least one embodiment, it would be impractical to obtain training data for all possible positions and orientations of the camera. In at least one embodiment, placing cameras at arbitrary locations within the vehicle can result in a lack of ground truth data, which can cause training problems. In at least one embodiment, points or locations within the vehicle may be fixed and known, but the position and orientation of the camera within the vehicle may be unknown, such that the relative positions of these points or objects with respect to the camera will be uncertain and no ground truth is available. In at least one embodiment, as Figure 1B and 1C As shown, image data captured from different cameras will include representations of objects in different locations. In at least one embodiment, Figure 1B The image 120 shown in FIG includes a representation of the driver from the perspective of the first camera 102 on the rearview mirror, and Figure 1C The image 140 shown includes a representation of the driver from the perspective of the second camera 106 on the dashboard. In at least one embodiment, the driver is looking at the dashboard in both images. However, in at least one embodiment, without ground truth data and a known position and orientation of a given camera relative to the vehicle, it is not possible to definitively determine where the driver is looking.
[0061] In at least one embodiment, the coordinate systems of various objects in the vehicle may be considered. Figure 1D As illustrated in view 160 of FIG, there can be a coordinate system 162 associated with a first camera, the origin and orientation of this coordinate system 162 being based on the position and orientation of the camera sensor. In at least one embodiment, there will similarly be a vehicle coordinate system 164 for the vehicle, where points within the vehicle will be fixed relative to the vehicle coordinate system 164. In at least one embodiment, different camera placements will result in different relative distances and orientations between these coordinate systems 164, 164. In at least one embodiment, there can also be additional coordinate systems, such as a coordinate system 166 for the driver or passenger, which can change over time based on factors such as the position, orientation, and posture of that person. In at least one embodiment, another coordinate system 168 corresponding to another camera or sensor in the vehicle can be considered. In at least one embodiment, there can be additional cameras, sensors, or objects, each of which can have one or more corresponding coordinate systems.
[0062] In at least one embodiment, an attempt may be made to relate or bridge at least some of these coordinate systems. Figure 2 In the vehicle orientation shown in view 200 of the vehicle, a camera can be located near a rearview mirror and can be a main cabin camera that continuously captures image data of the vehicle interior, at least during operation or when an occupant is detected. In at least one embodiment, a camera coordinate system 202 determined relative to the camera can be considered, as image data captured using the camera will include representations of objects relative to the camera. In at least one embodiment, representing the positions of these objects relative to the camera can be simplified by determining their positions in the camera coordinate system. However, in at least one embodiment, the positions in the camera coordinate system may not have ground truth data for training, as the camera position can be arbitrary or can be one of multiple possible positions in or near the vehicle.
[0063] In at least one embodiment, a vehicle coordinate system 204 can be considered for the cabin of the vehicle. In at least one embodiment, the coordinate system 204 can be centered at a specific location, such as a location defining an origin in the vehicle. In at least one embodiment, the origin can correspond to a calibration fixture 208 or other such object, which can be placed at a defined point in the vehicle to be modeled. In at least one embodiment, the calibration fixture 208 can include information that can be used to determine the relative position 206 and orientation of the calibration fixture 208 relative to the camera. In at least one embodiment, the calibration fixture 208 can include an asymmetric checkerboard pattern or other asymmetric pattern, which can be used to determine the orientation of the pattern. In at least one embodiment, the planar pattern can also provide information about the relative orientation of the calibration fixture in three dimensions. In at least one embodiment, the image data captured by the camera can include a representation of the calibration fixture 208. In at least one embodiment, the image data can be analyzed to determine the relative position and orientation of the calibration fixture 204 relative to the camera in the camera coordinate system. In at least one embodiment, ground truth data for points in the vehicle is known in the vehicle coordinate system. In at least one embodiment, determining the relative position and orientation of the vehicle coordinate system 204 with respect to the camera coordinate system 202 provides ground truth data that can be used to train a model or neural network for this configuration because points in the vehicle coordinate system 204 can be mapped to corresponding points in the camera coordinate system 202.
[0064] In at least one embodiment, a set of objects or areas of interest in a vehicle may be determined. In at least one embodiment, these objects may include things that a driver or occupant of the vehicle may see, such as areas of the windshield to view objects in front of the vehicle, rearview or side mirrors to view objects to the side or behind the vehicle, the instrument panel to obtain information about vehicle operation, or display screens to obtain other types of content or information. In at least one embodiment, for these and other potential objects of interest, it may be necessary to determine the locations in the vehicle corresponding to these objects.
[0065] In at least one embodiment, markers can be used to designate objects or areas within a vehicle. In at least one embodiment, a fiducial marker such as an AprilTag can be used. In at least one embodiment, the AprilTag provides a visual fiducial element 306 that can be used to determine the position and orientation of a point in an image, which corresponds to a relative position in a camera coordinate system or virtual camera space. In at least one embodiment, a QR code or other fiducial element can be utilized. In at least one embodiment, a set of four fiducial elements 306 can be used to mark an area, where the area can be approximated as a box, trapezoid, or other quadrilateral geometry. In at least one embodiment, two or more fiducial elements can be used to represent the area, depending in part on the shape of the area. In at least one embodiment, image data can be captured by a camera 302 that includes a representation of at least a subset of the fiducial elements 306. In at least one embodiment, the position and orientation of these fiducial elements 306 can be determined from the image data (rather than within virtual camera space). In at least one embodiment, this can include the relative position and orientation of objects such as the windshield area 310, the rearview mirror area 308, or the steering wheel area 312. In at least one embodiment, when an occupant's gaze direction is determined to intersect one of these areas, it can be determined that the occupant is looking at that area, and a decision can be made based on this information as to whether any action should be taken. In at least one embodiment, such action can include selecting information to display at the location the occupant is looking at, taking driving action due to the driver looking away from a certain area, or adjusting the vehicle's operation or configuration, etc.
[0066] In at least one embodiment, there may not be any Figure 3ground truth data for the configuration of the vehicle to provide accurate training of at least one neural network, such as a neural network for a particular type of vehicle corresponding to the configuration. However, in at least one embodiment, ground truth data is available for the positions of the fiducial markers 306 in a vehicle coordinate system or vehicle space relative to the origin of the coordinate system (which may correspond to the calibration frame 304). In at least one embodiment, points in virtual camera space can be mapped to points in vehicle coordinate space, which can provide ground truth data for the fiducial markers 306 represented in the image data captured by the camera 302. In at least one embodiment, ground truth data can be provided for any camera position as long as the vehicle space can be mapped to a corresponding virtual camera space. In at least one embodiment, the calibration frame 304 has an asymmetric pattern, such as a colored checkerboard pattern, which can be used to determine the position, distance, and orientation of the calibration frame based on the appearance of the pattern in the captured image data.
[0067] In at least one embodiment, a virtual camera space is created and used as a bridge to the data collection system and the real automotive environment. In at least one embodiment, this bridging enables the vehicle monitoring system to deploy any trained deep neural network (DNN) model to any vehicle model. In at least one embodiment, a coordinate propagation mechanism is utilized that can transform each physical world point into a point in the virtual camera space. In at least one embodiment, the propagation mechanism can find a relationship between a determined coordinate system (e.g., a camera coordinate system) and one or more adjacent coordinate systems (e.g., a vehicle or occupant coordinate system). In at least one embodiment, these coordinate systems can be linked together or otherwise associated so that the coordinates of any point can be determined in the virtual camera space.
[0068] In at least one embodiment, a virtual camera space and coordinate propagation mechanism can be applied to train multiple vehicle models, generating ground truth for all locations or gaze points that an occupant is required to look at. In at least one embodiment, this can include looking at a light (e.g., a light-emitting diode (LED)) near a fiducial marker on an object in the vehicle. In at least one embodiment, the propagation path begins in LED space, reaches physical LED board space, then fiducial marker space, real-world space, calibration rig space, and finally reaches virtual camera space. In at least one embodiment, data points corresponding to the vehicle, fiducial markers, and other fixed locations can be determined, allowing a single propagation path from vehicle space to virtual camera space. In at least one embodiment, this propagation path corresponds to a vector from that point to the origin in virtual camera space, enabling the determination of the coordinates of that point in a camera coordinate system, which can be used as ground truth data. In one embodiment, this ground truth data can be used to enable user-specific training of a deep neural network model. In at least one embodiment, this ground truth data can be used to empower projection-based DNN gaze modeling. In at least one embodiment, a mirror-based occlusion-invariant camera positioning method is utilized to provide correct camera positioning even when a portion of a geometric pattern plate is obscured (e.g., by a steering wheel). In at least one embodiment, a procedure is utilized that can determine an optimal set of poses for optimizing calibration and positioning by narrowing the pose search space using constraints, such as DMS environment constraints.
[0069] In at least one embodiment, an object detection model can be trained to determine the location of objects in any definable three-dimensional (3D) environment. In at least one embodiment, ground truth data can be determined for a camera or sensor relative to the environment to allow the model to be built and / or trained. In at least one embodiment, objects with definable orientations are placed in the 3D environment to act as a bridge to a virtual camera space having a camera coordinate system. In at least one embodiment, known points in the environment coordinate system can then be converted to ground truth points in the camera coordinate system. In at least one embodiment, this ground truth data is used to train a vision-based deep learning system to be able to infer certain outputs. In at least one embodiment, the output is the direction or position of the vehicle occupant's gaze.
[0070] In at least one embodiment, a gaze estimation network uses a camera in front of the driver or passengers in a vehicle. In at least one embodiment, image data is captured by the camera to attempt to determine where people are looking. In at least one embodiment, during the training portion, a large amount of data is required that includes representations of people looking in various directions at various points of interest (e.g., points in or near the vehicle). In at least one embodiment, once such a network is trained, it can be deployed in a vehicle of that type, which can enable the vehicle or a system communicating with the vehicle to infer where people are looking. In at least one embodiment, applications can be built that take advantage of the availability of this gaze information. In at least one embodiment, the gaze determination information can be used to determine whether the driver is paying attention and paying attention to the road, or if he is distracted and looking elsewhere. In at least one embodiment, the gaze information can be used to modify the operation of the vehicle, such as by modifying car controls, changing displayed information, or modifying the vehicle's driving mode.
[0071] In at least one embodiment, ground truth data is obtained that can be used to describe a three-dimensional point in an arbitrary coordinate system, such as can correspond to a camera or sensor. Then, in at least one embodiment, the ground truth data can be associated with an object represented in an image captured by a camera of at least a partial view of a three-dimensional space (e.g., an interior cabin of a vehicle). In at least one embodiment, position determination can be used for applications that are not related to line of sight or vision. In at least one embodiment, an object determination system can determine the position of an object in an environment (such as a vehicle).
[0072] In at least one embodiment, the pre-trained gaze model can be deployed to any type of vehicle. In at least one embodiment, an LED-based detection system in the vehicle can be utilized, which can be mapped into three-dimensional space. In at least one embodiment, the detection system can be used to determine where the driver is looking, and this determination process can be repeated for multiple locations. In at least one embodiment, a set of LED boards can be used, each including a fiducial marker (such as an AprilTag). In at least one embodiment, the position of each LED board in vehicle space can then be determined. In at least one embodiment, each fiducial marker provides image-based determination of its three-dimensional position, orientation, and identity relative to a camera or sensor. In at least one embodiment, one or more fiducial markers are located near each object of interest, such as one at each corner of the object of interest. Image data including representations of these fiducial markers can be captured, and the corresponding three-dimensional positions reconstructed in vehicle space. In at least one embodiment, these coordinates in vehicle space can be mapped into a camera coordinate system or virtual camera space. In at least one embodiment, a transfer mechanism can be used to propagate the LED board coordinates from the vehicle coordinate system to the camera coordinate system. In at least one embodiment, the difference between the different coordinate systems or virtual spaces includes the origin and six degrees of freedom. In at least one embodiment, a single coordinate system may be useful for environments such as a vehicle, while multiple coordinate systems may be useful for environments with objects that may change orientation or position, such as when multiple occupants are in a vehicle. In at least one embodiment, for environments with multiple coordinate systems, these systems may be linked together to provide ground truth data in any of those associated coordinate systems.
[0073] In at least one embodiment, the detection process includes an initial training phase or data collection phase that can be performed in a controlled environment. In at least one embodiment, the detection process also includes a second inference phase in which the trained model is used, such as in a vehicle or other such environment. In at least one embodiment, the environment for which the model is used can have one or more variable aspects, such as camera position. In at least one embodiment, the variable nature of the camera can introduce uncertainty into the data collection because the camera is not fixed in a consistent position in a common coordinate system, making it difficult for the model to infer information (such as the driver's line of sight position based on image data captured using such a camera). In at least one embodiment, even if the camera is fixed in the environment, a second camera in a second similar environment can be fixed in a different position. In at least one embodiment, the difference in position can make it difficult to deploy a model trained on at most one of these positions for use with a camera placed in a different position.
[0074] In at least one embodiment, a virtual camera space is extracted. In at least one embodiment, this space is referred to as a virtual space because it is not dependent on any physical space. In at least one embodiment, this virtual camera space can be used to help bridge the data collection phase and in-vehicle reasoning. In at least one embodiment, the camera position is fixed, but unknown in a given vehicle or environment. In at least one embodiment, the occupant position can also change from occupant to occupant and over time. In at least one embodiment, the location where the driver is required to look can be fixed, such as the top left corner of the windshield, the top left corner of the left rearview mirror, or the corner of the right rearview mirror. In at least one embodiment, data is collected using cameras or sensors in sub-devices placed at possible locations in the closed environment. In at least one embodiment, a generalized gaze network will be able to determine every possible location where the occupant can look in the vehicle and identify the current gaze location with high accuracy.
[0075] In at least one embodiment, training and ground-truth data are obtained using image data captured by at least one camera or sensor. In at least one embodiment, the image data is processed using a computer vision algorithm to determine, for example, the relative positions of a set of fiducial markers. In at least one embodiment, a bridging mechanism can be used to correlate this position information with known ground-truth data in the environment space (e.g., the vehicle space). Thus, in at least one embodiment, the ground-truth collection process involves at least one localization and reconstruction algorithm. In at least one embodiment, these and other components can be provided as part of a calibration kit. In at least one embodiment, the calibration kit includes a camera calibration component that can determine the imaging characteristics of a camera to account for any specifics of the camera. In at least one embodiment, the calibration kit includes a localization component that can localize a global surveillance camera within any vehicle or enclosed area. In at least one embodiment, the calibration kit also includes a video reconstruction component that can reconstruct the vehicle's geometry. In at least one embodiment, the calibration kit can connect this information through a coordinate propagation process that can be used to obtain ground-truth data for each relevant point in the environment, which can be used to train and deploy a neural network model in a vehicle with such specific geometry.
[0076] In at least one embodiment, the trained model can be used to generate a gaze vector corresponding to where a person is looking. In at least one embodiment, the trained model can also be validated. In at least one embodiment, there is a vector whose origin is between the person's two eyes, and a line of sight with a gaze direction starting from this point can be tracked or determined. In at least one embodiment, this line of sight will intersect with a specific point in the vehicle geometry for a specific vehicle area. In at least one embodiment, this determined area can be compared with ground truth data to determine whether correct inference has been generated. In at least one embodiment, such a process can be used to validate the trained model. In at least one embodiment, the trained model can then be deployed for inference, and inference can be made about where the person is looking. In at least one embodiment, based in part on how the model was trained, this inference can be for any location in the vehicle cabin, or for specific locations. In at least one embodiment, training for specific locations or areas can produce more accurate inferences for those specific locations or areas.
[0077] In at least one embodiment, data capture is performed using a single camera and ambient light. In at least one embodiment, data capture is performed using at least one camera or sensor and at least two illumination sources, such as two infrared (IR) LEDs. In at least one embodiment, the use of two light sources enables line of sight detection to resolve glare or obstructions. In at least one embodiment, data capture is performed using a reference applied to a specific object of interest, such as a windshield, dashboard, left side mirror, right side mirror, and rearview mirror. In at least one embodiment, another data capture process may be used, as it may involve ultrasonic or laser scanning data capture. In at least one embodiment, data captured for a specific type and model of vehicle may be used to train a model, which may then be used with any vehicle of that type and model. In at least one embodiment, different models may be trained for individual variations of the vehicle or environment.
[0078] In at least one embodiment, the gaze data can be used together with other data to determine how to modify or control operational aspects of the vehicle. In at least one embodiment, the driver can say phrases such as "lower that window" or "lower my window." In at least one embodiment, a microphone can capture the voice utterance and speech-to-text analysis can be performed to determine the instruction. In at least one embodiment, the gaze information can be used to help determine the window that the driver is talking about. In at least one embodiment, the vehicle can then automatically lower the determined window based in part on the inferred gaze data. In at least one embodiment, the model can also be personalized for an individual user or person to take individual characteristics or traits into account.
[0079] In at least one embodiment, ground truth data may be generated to train one or more neural networks, such as Figure 4A , as shown in process 400. In at least one embodiment, one or more regions of interest in a vehicle are determined 402. In at least one embodiment, these regions may include areas where an occupant is likely to be gazing, or areas of interest for determining what a user is gazing at (e.g., windows, mirrors, or display panels within or near the vehicle). In at least one embodiment, fiducial markers may be positioned 404 at representative points within those regions, such as at corners of those regions. In at least one embodiment, these fiducial markers may include asymmetric visual aspects that enable determination of information such as position, distance, orientation, and identity of the fiducial markers. In at least one embodiment, these fiducial markers may include April tags or QR codes and may have one or more associated LEDs as discussed herein. In at least one embodiment, position data for these representative points may be determined 406 in vehicle space, for example, by determining absolute positions relative to an origin that remain fixed over time. In at least one embodiment, a camera may be used to capture image data 408, where the image data includes representations of the fiducial markers within the vehicle interior and a calibration frame. In at least one embodiment, the calibration frame is used to specify the origin and orientation of a vehicle coordinate system that defines the virtual vehicle space. In at least one embodiment, the calibration frame may include a checkerboard filter and at least one asymmetric surface, such that the camera is able to determine at least the position, orientation, and scale of the calibration plate represented in the captured image data. In at least one embodiment, 410 analyzes the captured image data to determine the position of a representative point in virtual camera space or according to a camera coordinate system. In at least one embodiment, the point positions are determined by analyzing representations of fiducial markers in the image data. In at least one embodiment, 412 may utilize the calibration frame in the vehicle and represented in the captured image data as a bridging mechanism to associate the vehicle coordinate system and the camera coordinate system, thereby associating points in the camera space with known absolute position data of corresponding points in the vehicle coordinate system. In at least one embodiment, 412 may utilize these known relative positions of these related points as ground truth data to train one or more neural networks.
[0080] In at least one embodiment, executing Figure 4BThe process 450 shown in FIG can use such a trained model to determine a location in an environment (such as a vehicle). In at least one embodiment, an image or video frame to be used for inference is obtained at 452. In at least one embodiment, the image can be provided as input to a trained model or neural network at 454. In at least one embodiment, the trained model can process data, including at least image data, and infer the location of an object in the vehicle at 456, such as an area or object that may correspond to the occupant's gaze.
[0081] In at least one embodiment, Figure 5 The process 500 shown in FIG. 1 can be used to determine gaze data. In at least one embodiment, 502 detects one or more occupants of a vehicle using one or more sensors. In at least one embodiment, these sensors can include imaging, distance, or position sensors, such as cameras, ultrasonic sensors, radar scanning, and LIDAR. In at least one embodiment, 504 can determine the gaze of one or more occupants of the vehicle independently of the position of the one or more sensors.
[0082] Figure 6 A high-level system architecture according to one embodiment of the present invention is shown. System 600 preferably includes a plurality of controllers 602(1)-602(N), including controllers and systems for autonomous or semi-autonomous driving. One or more controllers 602 may include an advanced SoC or platform for executing an intelligent assistant software stack (IX) that performs risk assessment and provides notifications, warnings, and autonomously controls the vehicle in whole or in part to perform the risk assessment and advanced driver assistance functions described herein. Two or more controllers are used to provide autonomous driving functions, executing an autonomous vehicle (AV) software stack to perform autonomous or semi-autonomous driving functions.
[0083] Advanced platforms and SoCs used to execute the present invention preferably have multiple types of processors, thereby providing the "right tool for the job" and processing diversity for functional safety. For example, GPUs are well-suited for high-precision tasks. On the other hand, hardware accelerators can be optimized to perform more specific function sets. By providing a mix of multiple processors, advanced platforms and SoCs include a full set of tools to quickly, reliably, and efficiently execute the complex functions associated with advanced AI-assisted vehicles.
[0084] Figure 7The system architecture according to one embodiment is shown. The system includes a controller and system for autonomous or semi-autonomous driving. The controller (100) receives input from one or more cameras (72, 73, 74, 75) deployed around the vehicle. The controller (100) detects objects and provides information about the presence and trajectory of the objects to the risk assessment module (6000). The system includes a plurality of cameras (77) located inside the vehicle. The cameras (77) can be as follows Figure 8 The arrangement shown, or any other arrangement providing coverage of the driver and other occupants. The camera (77) provides input to a plurality of deep neural networks (5000) to monitor the driver, other occupants and / or conditions in the vehicle. Optionally, a multi-sensor camera module (500), (600(1)-(N)) and / or (700) may be used to view the interior or exterior environment of the vehicle.
[0085] Preferably, the neural network is trained to detect a number of different features and events, including: the presence of a face (5001), the identity of the person in the driver's seat or one or more passenger seats (5002), the driver's head pose (5003), the driver's gaze direction (5004), whether the driver's eyes are open (5005), whether the driver's eyes are closed or occluded (5006), whether the driver is speaking, and if so, what the driver is saying (via audio input or lip reading) (5007), whether a passenger is aggressing or otherwise impairing the driver's ability to control the vehicle (5008), and whether the driver is in distress (5009). In additional embodiments, the network is trained to recognize driver actions, including (but not limited to): detecting cell phones, drinking, smoking, and driver intent based on head and body poses and movements. In one embodiment, both the AV stack and the IX stack can execute on the same platform or SoC (9000).
[0086] Figure 8 An exemplary camera layout for the cabin is shown in . Figure 9 The front of the cabin according to one embodiment is shown. The cabin preferably includes at least two cameras pointed at the driver. In one embodiment, the driver main camera (77(3)) detects infrared light at a wavelength of 940nm, a 60 degree field of view, and captures images at 60fps. The driver main camera (77(3)) is preferably used to determine face ID, and determine the driver's line of sight, head posture, and detect drowsiness. In at least one embodiment, the driver main camera can be replaced by a multi-sensor camera module that provides IR and RGB camera functionality.
[0087] In one embodiment, the driver assist camera (77(4)) is infrared (IR) with a wavelength of 940 nm and a field of view of 60 degrees, capturing images at 60 frames per second. The driver assist camera (77(4)) is preferably used in conjunction with the driver assist camera (77(3)) to determine the driver's line of sight, head posture, and detect drowsiness. Alternatively, the driver assist camera can be replaced with a multi-sensor camera module that provides IR and RGB camera functionality.
[0088] The cabin preferably includes at least one cabin main camera (77(1)) typically mounted overhead. In one embodiment, the cabin main camera (77(1)) is an infrared camera with a wavelength of 940nm, time of flight (ToF) depth, a 90 degree field of view and captures images at 30fps. The cabin main camera (77(1)) is preferably used to determine posture and cabin occupancy. The cabin preferably includes at least one passenger camera 77(5) typically mounted near the passenger storage compartment or passenger side dashboard. In one embodiment, the passenger camera (77(5)) is IR at a wavelength of 940nm, a 60 degree field of view and captures images at 30fps. Alternatively, the driver main camera can be replaced with a multi-sensor camera module (500), (600(1)-(N)) and / or (700) that provides IR and RGB camera functionality.
[0089] The front of the cabin preferably includes a plurality of LED illuminators (78(1)-(2)). The illuminators preferably project 940nm infrared light, are synchronized with the camera, and are eye-safe. The front of the vehicle preferably also includes a low-angle camera to determine when the driver is looking down (compared to when the driver has their eyes closed).
[0090] The cabin also preferably has a "cabin assist" camera (not shown) that provides a view of the entire cabin. The cabin assist camera is preferably mounted in the center of the roof and has a wide-angle lens that provides a view of the entire cabin. This allows the system to determine the number of occupants, estimate the age of the occupants, and perform object detection functions. In other embodiments, the system includes dedicated cameras for front and rear passengers (not shown). Such dedicated cameras allow the system to conduct video conferences with occupants in the front or rear of the vehicle.
[0091] In at least one embodiment, an autonomous vehicle may include one or more multi-sensor camera modules (MSCMs) that provide multiple sensors in a single housing and also allow for interchangeable sensors. MSCMs according to various embodiments may be used in various configurations: (1) IR+IR (IR stereo vision), (2) IR+RGB (stereo vision and paired frames), (3) RGB+RGB (RGB stereo vision). Depending on the desired color and low-light performance, the RGB sensor may be replaced with an RCCB (or other color sensor). The MSCM may be used to cover a camera outside the vehicle, a camera inside the vehicle, or both.
[0092] Figure 9 An embodiment of an MSCM is shown. In this embodiment, the MSCM (900) is coupled to one or more AI supercomputers suitable for controlling an autonomous or semi-autonomous vehicle. In this embodiment, the AI supercomputers (800), (900) include one or more advanced SoCs as described in U.S. Provisional Application No. 62584549, filed on November 10, 2017.
[0093] The multi-sensor camera module 904 includes a serializer (906), an IR image sensor (912), an RGB image sensor (918), a lens and IR filter (914), and a microcontroller (916). Many camera sensors can be used, including the OnSemi AR0144 (1.0 megapixel (1280H x 800V), 60fps, global shutter, CMOS). The AR0144 reduces artifacts in bright and low light conditions and is designed for high shutter efficiency and signal-to-noise ratio to minimize ghosting and noise effects. The AR0144 can be used with both a color sensor (1006) and a monochrome sensor (1007).
[0094] Many different camera lenses (914, 924) can be used. In one embodiment, the camera lens is an LCE-C001 (55HFoV) with a 940nm bandpass. The LED lens is preferably a Ledil Lisa2 FP13026. In one embodiment, each lens is mounted in a molded polycarbonate (PC) housing designed to align with a specific LED, providing precise positioning of the lens at the ideal focus for each qualified brand or style of LED. Other LED lenses can be used.
[0095] exist Figure 9In the embodiment shown, the MSCM controls one or more LEDs (922). These LEDs meet automotive standards and provide infrared illumination for the camera in the form of highly concentrated, invisible infrared light. In one embodiment, the LED is an Osram Opto SFH4725S IR LED (940 nm). The LED (922) is controlled by a switch (920) that flashes the LED. The LED lens is preferably a Ledil Lisa2 FP13026. Other LEDs and lenses may be used.
[0096] The serializer is preferably a MAX9295A GMSL2 SER, although other serializers may be used. Suitable microcontrollers (MCUs) include the Atmel SAMD21. The SAM D21 is a series of low-power microcontrollers using a 32-bit ARM Cortex processor, ranging from 32 to 64 pins, with up to 256KB of flash memory and 32KB of SRAM. The maximum frequency of the SAM D21 device is 48MHz and reaches 2.46CoreMark / MHz. Other MCUs may also be used. The LED driver (922) is preferably the ON-Semi NCV7691-D or equivalent, although other LED drivers may be used.
[0097] Figure 10 An embodiment of the driver UX input / output and configuration is shown. The driver UX includes one or more display screens, including an AV status panel (900), a main display screen (903), an auxiliary display screen (904), a surround display screen (901), and a communication panel (902). The AV status panel (900) is preferably a small (3.5", 4" or 5") display that displays only critical information for a safe driver to operate the vehicle.
[0098] The surround display screen (901) and the auxiliary display screen (904) preferably display information from the cross traffic camera (505), the blind spot camera (506), and the rear cameras (507) and (508). In one embodiment, as Figure 23 As shown, the surround display screen (901) and the auxiliary display screen (904) are arranged to surround the safety driver. In alternative embodiments, the combination or arrangement of the display screens may be different from Figure 23. For example, the AV status panel (900) and the main display screen (903) can be combined in a single forward-facing panel. Alternatively, a portion of the main display screen (903) can be used to display a split-screen view or an overhead view of the advanced AI-assisted vehicle with surrounding objects. Alternatively, the driver UX input / output can include a head-up display ("HUD") (906) of vehicle parameters such as speed, destination, ETA, and number of passengers, or simply the status of the AV system (activated or disabled).
[0099] The driver interface and display may provide information from the autonomous driving stack to assist the driver. For example, the driver interface and display may highlight lanes, cars, signs, pedestrians in the main screen (903) or HUD (906) on the windshield. The driver interface and display may provide a recommended path suggested by the autonomous driving stack, as well as suggestions to stop accelerating or start braking when the vehicle approaches a signal light or traffic sign. The driver interface and display may highlight points of interest, expand the field of view around the car while driving (wide FOV), or assist in parking (e.g., providing a bird's-eye view - if the vehicle has a surround camera).
[0100] The driver interface and display preferably provide alerts including: (1) waiting conditions ahead, including intersections, construction areas, and toll booths; (2) objects in the path of travel, such as pedestrians traveling much slower than the Advanced AI-Assisted Vehicle; (3) vehicles stalled ahead; (4) school zones ahead; (5) children playing on the roadside; (6) animals on the roadside (e.g., deer or dogs); (7) emergency vehicles (e.g., police cars, fire trucks, medical vehicles, or other vehicles with sirens); (8) vehicles likely to cut in front of the travel lane; (9) crossing traffic, especially in situations where traffic lights or signs may be violated; (10) approaching cyclists; (11) unexpected objects in the road (e.g., tires and debris); and (12) poor quality roads ahead (e.g., icy roads and potholes).
[0101] Embodiments may be applicable to any type of vehicle, including but not limited to cars, sedans, buses, taxis, and shuttles. In one embodiment, an advanced AI-assisted vehicle includes a passenger interface for communicating with passengers, including map information, route information, a text-to-speech interface, voice recognition, and external application integration, including integration with calendar applications such as Microsoft Outlook.
[0102] Figure 11A flowchart of a method for gaze estimation according to one embodiment is shown. In step 1, an image of an eye is received. In step 2, a head direction is received. In one embodiment, the head direction data is pre-calculated and may include azimuth and elevation. In another embodiment, the head direction data is an image of the subject's face, and the head direction is determined based on the image. In step 3, a CNN is used to calculate the gaze position based on the image and the head direction data.
[0103] Figure 12 A pipeline of a neural network suitable for determining gaze detection according to one embodiment is shown. FDNet (5001) is trained to detect the presence of a face. HPNet (5003) determines a person's head pose. FPENet (50011) detects reference points. In this embodiment, GazeNet (5004) is a neural network trained using inputs including head position data (x, y, z) and reference points associated with the head. GazeNet uses these inputs to detect the driver's gaze.
[0104] In one embodiment, the risk assessment module determines whether cross traffic is beyond the driver's field of view and provides appropriate warnings. Figure 13 A scenario is shown where the risk assessment module uses information from the DNN for gaze detection (5004) and uses information from the controller to warn the driver.
[0105] The gaze detection DNN classifies the driver’s gaze as falling into an area, such as Figure 13 In one example, the area includes a left intersection (10(1)), a center intersection (10(2)), a right intersection (10(3)), a rearview mirror (10(4)), a left rearview mirror (10(5)), a right rearview mirror (10(5)), an instrument panel (10(7)), and a center console (10(8)).
[0106] While the gaze detection DNN classifies the driver's gaze area, the controller (100(2)) uses the DNN executed on the advanced SoC to detect cross traffic outside the driver's field of view.
[0107] Neural network training and deployment
[0108] In at least one embodiment, an untrained neural network is trained using a training data set. In at least one embodiment, the training framework is the PyTorch framework, while in other embodiments, the training framework is Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework trains the untrained neural network and enables it to be trained using the processing resources described herein to generate a trained neural network. In at least one embodiment, the weights can be randomly selected or selected using pre-training of a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.
[0109] In at least one embodiment, an untrained neural network is trained using supervised learning, where the training dataset includes inputs paired with expected outputs for the inputs, or where the training dataset includes inputs with known outputs and the outputs of the neural network are manually graded. In at least one embodiment, the untrained neural network is trained in a supervised manner to process inputs from the training dataset and compare the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained neural network. In at least one embodiment, a training framework adjusts and controls the weights of the untrained neural network. In at least one embodiment, the training framework includes tools for monitoring the degree to which the untrained neural network converges toward a model (e.g., a trained neural network) adapted to generate correct answers (e.g., results) based on known input data (e.g., new data). In at least one embodiment, the training framework repeatedly trains the untrained neural network while adjusting the weights using a loss function and an adjustment algorithm (such as stochastic gradient descent) to refine the outputs of the untrained neural network. In at least one embodiment, the training framework trains the untrained neural network until the untrained neural network achieves a desired accuracy. In at least one embodiment, the trained neural network can then be deployed to perform any number of machine learning operations.
[0110] In at least one embodiment, the untrained neural network is trained using unsupervised learning, wherein the untrained neural network attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training data set will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 1106 can learn the groupings within the training data set and can determine how the individual inputs relate to the untrained data set. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network that is capable of performing operations useful for reducing the dimensionality of new data. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in a new data set that deviate from the normal pattern of the new data set.
[0111] In at least one embodiment, semi-supervised learning can be used, a technique in which training is performed on a dataset that includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework can be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning enables a trained neural network to adapt to new data without forgetting the knowledge infused into the network during the initial training process.
[0112] As mentioned above, a growing number of industries and applications are leveraging machine learning. For example, deep neural networks (DNNs) developed on processors are being used in a variety of use cases, from self-driving cars to faster drug development, from automated image analysis for security systems to intelligent real-time language translation in video chat applications. Deep learning is a technology that mimics the neural learning process of the human brain, continuously learning, getting smarter, and delivering more accurate results over time. Initially, adults teach children how to correctly identify and classify various shapes, eventually enabling them to recognize shapes without any instruction. Similarly, deep learning or neural learning systems designed to accomplish similar tasks will need to be trained to become smarter and more efficient at recognizing basic objects, occluded objects, and other aspects, while also assigning these objects context.
[0113] At the simplest level, neurons in the human brain look at the various inputs they receive, assign a level of importance to each of these inputs, and then pass outputs to other neurons to act on them. An artificial neuron, or perceptron, is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs representing various features of the object it is being trained to recognize and classify, and assign a specific weight to each of these features based on its importance in defining the object's shape.
[0114] Deep neural network (DNN) models consist of many layers of connected perceptrons (e.g., nodes) that can be trained with large amounts of input data to quickly solve complex problems with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into its components and looks for basic patterns such as lines and angles. The second layer assembles the lines to look for higher-level patterns, such as wheels, windshields, and rearview mirrors. The next layer identifies the type of vehicle, and the final layers generate labels for the input image, identifying the model as a specific car brand. Once a DNN is trained, it can be deployed and used to recognize and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten digits on checks deposited at an ATM, identifying images of friends in photos, providing movie recommendations, identifying and classifying different types of cars, pedestrians, and road hazards in self-driving cars, or translating human speech in near real time.
[0115] During training, data flows through the DNN in a forward propagation phase until a prediction is produced indicating the label corresponding to the input. If the neural network incorrectly labels the input, the difference between the correct and predicted labels is analyzed, and the weights for each feature are adjusted in a backpropagation phase until the DNN correctly labels the input and other inputs in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including support for floating-point multiplication and addition. Inference, which is less computationally intensive than training, is a latency-sensitive process in which a trained neural network is applied to new, previously unseen inputs to classify images, translate speech, and reason about new information.
[0116] Neural networks rely heavily on matrix math operations, and complex, multi-layer networks require significant floating-point performance and bandwidth for efficiency and speed. Computing platforms with thousands of processing cores optimized for matrix math operations and delivering tens to hundreds of TFLOPS of performance can deliver the performance required for deep neural network-based artificial intelligence and machine learning applications.
[0117] Figure 14Components of an example system 1400 that can be used to train and utilize machine learning are shown. As will be discussed, various components can be provided by various combinations of computing devices and resources, which can be under the control of a single entity or multiple entities, or a single computing system. In addition, aspects can be triggered, initiated, or requested by different entities. In at least one embodiment, the training of a neural network can be directed by a vendor associated with a vendor environment 1406, while in at least one embodiment, the training of a neural network can be requested by a customer or other user who can access the vendor environment through a client device 1402 or other such resource. Training data (or data to be analyzed by the trained neural network) can be provided by the vendor, the user, or a third-party content provider 1424. In at least one embodiment, the client device 1402 can be a vehicle or object that can be navigated on behalf of a user, for example, the user can submit requests and / or receive instructions to facilitate device navigation.
[0118] In this example, a request can be submitted via at least one network 1404 to be received by the provider environment 1406. The client device can be any suitable electronic and / or computing device that enables a user to generate and send such a request, such as a desktop computer, a laptop computer, a computer server, a smart phone, a tablet computer, a game console (portable or otherwise), a computer processor, computing logic, and a set-top box. The network 1404 can include any suitable network for sending requests or other such data, such as the Internet, an intranet, an Ethernet network, a cellular network, a local area network (LAN), a network with direct wireless connections between nodes, and the like.
[0119] Requests may be received at interface layer 1408, which, in this example, may forward the data to training and inference manager 1410. The manager may be a system or service comprising hardware and software for managing services and requests consistent with data or content. The manager may receive requests to train a neural network and may provide the requested data to training manager 1412. If the request is unspecified, training manager 1412 may select an appropriate model or network to use and may use the associated training data to train the model. In at least one embodiment, the training data may be a batch of data received from client device 1402 or obtained from a third-party vendor 1424 and stored in training data repository 1414. Training manager 1412 may be responsible for training the data, for example, by using the LARC-based methods discussed herein. The network may be any suitable network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Once the network is trained and successfully evaluated, the trained network may be stored in model repository 1416, which may, for example, store different models or networks for users, applications, services, and the like. As described above, in at least one embodiment, multiple models may exist for a single application or entity, such that multiple models may be utilized based on a number of different factors.
[0120] At a later point in time, a request for content (e.g., a path determination) or data determined or influenced at least in part by a trained neural network may be received from client device 1402 (or another such device). The request may include, for example, input data to be processed using the neural network to obtain one or more inferences or other output values, classifications, or predictions. Although different systems or services may be used, the input data may be received into interface layer 1408 and directed to inference module 1418. If not already stored locally in inference module 1418, inference module 1418 may obtain an appropriate trained network, such as a trained deep neural network (DNN) as described herein, from model repository 1416. Inference module 1418 may provide the data as input to the trained network and may then generate one or more inferences as output. For example, this may include a classification of an instance of the input data. The inferences may then be sent to client device 1402 for display to the user or other communication with the user. The user's context data may also be stored in user context data repository 1422, which may include data about the user that may be used as network input to generate inferences or determine data to be returned to the user after obtaining the instance. Related data, including at least a portion of the input or inference data, may also be stored in a local database 1420 for use in processing future requests. In at least one embodiment, a user may use account or other information to access resources or functionality of the provider environment. If permitted and available, user data may also be collected and used to further train the model to provide more accurate inferences for future requests. Requests for machine learning applications 1426 executed on the client device 1402 may be received through a user interface, and results may be displayed through the same interface. The client device may include resources such as a processor 1428 and memory 1430 for generating requests and processing results or responses, as well as at least one data storage element 1432 for storing data for the machine learning application 1426.
[0121] In at least one embodiment, processor 1428 (or the processor of training manager 1412 or inference module 1418) will be a central processing unit (CPU). However, as described above, resources in such environments can utilize GPUs to process data for at least some types of requests. GPUs, with thousands of cores and designed to handle massively parallel workloads, have become popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building allows for faster training of larger and more complex models, generating predictions offline means that request-time input features cannot be used, or predictions must be generated for all feature permutations and stored in lookup tables to service real-time requests. If the deep learning framework supports CPU mode, and the model is small and simple enough that feedforward execution can be performed on a CPU with reasonable latency, a service on a CPU instance can host the model. In this case, training can be performed offline on the GPU, and inference can be performed in real time on the CPU. If a CPU approach is not feasible, the service can run on a GPU instance. However, because GPUs have different performance and cost characteristics than CPUs, running a service that offloads runtime algorithms to a GPU may require a different design than a CPU-based service.
[0122] Figure 15An example system 1500 is shown that can be used to classify data or generate inferences in at least one embodiment. Based on the teachings and suggestions contained herein, it should be apparent that various types of predictions, labels, or other outputs can also be generated for input data. Furthermore, both supervised and unsupervised training can be used in at least one embodiment discussed herein. In this example, a set of training data 1502 (e.g., classified or labeled data) is provided as input to serve as training data. The training data can include instances of at least one type of object for which a neural network is to be trained, as well as information identifying objects of that type. For example, the training data might include a set of images, each image containing a representation of an object type, wherein each image also contains or is associated with a label, metadata, classification, or other information identifying the type of object represented in the respective image. Various other types of data can also be used as training data, including text data, audio data, video data, and the like. In this example, the training data 1502 is provided as training input to a training manager 1504. The training manager 1504 can be a system or service comprising hardware and software, such as one or more computing devices executing a training application for training a neural network (or other model or algorithm, etc.). In this example, the training manager 1504 receives an instruction or request indicating the type of model to be used for training. The model can be any suitable statistical model, network, or algorithm that can be used for such purposes, and can include, for example, artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, etc. The training manager 1504 can select an initial model or other untrained model from an appropriate repository 1506 and use the training data 1502 to train the model to generate a trained model 1508 (e.g., a trained deep neural network) that can be used to classify similar types of data, or generate other such inferences. In at least one embodiment where training data is not used, the input data can still be trained based on the selection of an appropriate initial model by the training manager 1504.
[0123] Models can be trained in a variety of different ways, which may depend in part on the type of model selected. For example, in one embodiment, a set of training data can be provided to a machine learning algorithm, where the model is a model artifact created by the training process. Each instance of the training data contains a correct answer (e.g., a classification), which can be referred to as a target or target attribute. The learning algorithm finds patterns in the training data that map the input data attributes to the target, the answer to be predicted, and outputs a machine learning model that captures these patterns. The machine learning model can then be used to obtain predictions for new data for which the target is not specified.
[0124] In one example, the training manager 1504 can select from a set of machine learning models, including binary classification, multi-class classification, and regression models. The type of model to be used can depend at least in part on the type of target to be predicted. Machine learning models for binary classification problems can predict a binary outcome, such as one of two possible classes. Learning algorithms (such as logistic regression) can be used to train binary classification models. Machine learning models for multi-class classification problems allow predictions to be generated for multiple classes, such as predicting one of more than two outcomes. Multinomial logistic regression can be useful for training multi-class models. Machine learning models for regression problems can predict numerical values. Linear regression is useful for training regression models.
[0125] In order to train a machine learning model according to one embodiment, the training manager must determine the input training data source and other information, such as the name of the data attribute containing the target to be predicted, the required data transformation instructions, and training parameters to control the learning algorithm. During the training process, the training manager 1504 can automatically select an appropriate learning algorithm based on the target type specified in the training data source. The machine learning algorithm can accept parameters that are used to control certain properties of the training process and the resulting machine learning model. These are referred to as training parameters in this article. If no training parameters are specified, the training manager can utilize known default values that work well for a wide range of machine learning tasks. Examples of training parameters for which values can be specified include the maximum model size, the maximum number of passes through the training data, the type of shuffle, the type of regularization, the learning rate, and the amount of regularization. Default settings can be specified, with options for adjusting the values to fine-tune performance.
[0126] The maximum model size is the total size (in bytes) of the patterns created during model training. By default, a model of the specified size is created, for example, a 100MB model. If the training manager cannot determine that there are enough patterns to fill the model size, a smaller model is created. If the training manager finds that there are more patterns than can be accommodated within the specified size, it enforces a maximum cutoff by pruning the patterns that have the least impact on the quality of the learned model. Choosing the model size allows you to control the tradeoff between the model's predictive quality and cost. Smaller models may cause the training manager to remove many patterns to fit within the maximum size limit, affecting prediction quality. Larger models may be more expensive to query for real-time predictions. Larger input datasets do not necessarily result in larger models because the model stores the patterns, not the input data. If the patterns are few and simple, the resulting model will be smaller. Input data with a large number of raw attributes (input columns) or derived features (output of data transformations) may find and store more patterns during training.
[0127] In at least one embodiment, the training manager 1504 may make multiple passes or iterations through the training data to attempt to discover patterns. There may be a default number of passes, such as ten, and in at least one embodiment, a maximum number of passes may be set, such as up to one hundred passes. In at least one embodiment, there may not be a maximum set, or there may be a set of convergence criteria or other factors that trigger the end of the training process. In at least one embodiment, the training manager 1504 may monitor the quality of the pattern during training (e.g., for model convergence) and may automatically stop training when there are no more data points or patterns to discover. Data sets with only a small number of observations may require more data passes to achieve a sufficiently high model quality. Larger data sets may contain many similar data points, which may reduce the need for a large number of passes. The potential impact of selecting more passes through the data is that model training may take longer and cost more in terms of resources and system utilization.
[0128] In at least one embodiment, the training data is shuffled before training or between training passes. Shuffling is a random or pseudo-random shuffling that produces a truly random ordering, although constraints may exist to prevent certain types of data from grouping. If such grouping occurs, the shuffled data can be reshuffled. Shuffling changes the order or arrangement of the data used for training so that the training algorithm does not encounter groupings of similar data types or too many consecutive observations of a single type of data. For example, a model may be trained to predict objects. Before uploading, the data may be sorted by object type. The algorithm can then process the data alphabetically by object type, initially encountering only data of a specific object type. The model will begin to learn patterns for that object type. Then, the model will only encounter data for the second object type and will attempt to adapt the model to that object type, potentially degrading the patterns that were appropriate for the first object type. This abrupt switch between object types may result in a model that is unable to learn how to accurately predict object types. In at least one embodiment, shuffling can be performed before partitioning the training dataset into training and evaluation subsets, thereby utilizing a relatively even distribution of data types for both phases. In at least one embodiment, training manager 1504 may automatically shuffle the data using, for example, a pseudo-random shuffling technique.
[0129] In at least one embodiment, when creating a machine learning model, training manager 1504 can enable a user to specify settings or apply customization options. For example, a user can specify one or more evaluation settings to indicate a portion of the input data to be retained for evaluating the predictive quality of the machine learning model. A user can specify a policy that indicates which attributes and attribute transformations can be used for model training. The user can also specify various training parameters that control the training process and certain properties of the resulting model.
[0130] Once the training manager determines that the model training is complete, for example by using at least one of the final criteria discussed herein, the trained model 1508 can be provided to the classifier 1514 for use in classifying (or otherwise generating inferences about) validation data 1512. As shown, this involves a logical transition between the model's training mode and the model's inference mode. However, in at least one embodiment, the trained model 1508 will first be passed to an evaluator 1510, which can include an application, process, or service executed on at least one computing resource (e.g., a CPU or GPU of at least one server) for evaluating the quality (or other aspects) of the trained model. The model is evaluated to determine whether it provides at least a minimum acceptable or threshold level of performance when predicting targets for new and future data. If not, the training manager 1504 can continue training the model. Since future data instances will typically have unknown target values, it may be desirable to examine machine learning accuracy metrics on data for which the target answers are known and use this evaluation as a proxy for predictive accuracy for future data.
[0131] In at least one embodiment, a model is evaluated using a subset 1502 of the training data provided for training. This subset can be determined using the shuffling and splitting methods described above. This evaluation data subset is labeled with a target and can therefore serve as a resource for evaluating ground truth. Using the same data used for training to evaluate the predictive accuracy of a machine learning model is ineffective because a model that memorizes the training data rather than generalizing from it may produce a positive evaluation. Once training is complete, the evaluation data subset is processed using the trained model 1508, and an evaluator 1510 can determine the accuracy of the model by comparing the ground truth data with the model's corresponding output (or prediction / observation). The evaluator 1510 in at least one embodiment can provide a summary or performance metric that indicates the degree of match between the predicted and true values. If the trained model does not meet at least a minimum performance criterion or other such accuracy threshold, the training manager 1504 can be instructed to perform further training or, in some cases, attempt to train a new or different model. If the trained model 1508 meets the relevant criteria, the trained model can be provided for use by the classifier 1514.
[0132] When creating and training a machine learning model, in at least one embodiment, it may be desirable to specify model settings or training parameters that will result in a model capable of making accurate predictions. Example parameters include the number of passes (forward and / or backward) to be performed, regularization or refinement, model size, and shuffling type. However, as described above, selecting the model parameter settings that produce the best predictive performance on the evaluation data may result in model overfitting. Overfitting occurs when the model memorizes patterns present in the training and evaluation data sources but fails to generalize to patterns in the data. Overfitting often occurs when the training data includes all the data used in the evaluation. An overfitted model may perform well during evaluation but may not make accurate predictions on new or other validation data. To avoid selecting an overfitted model as the best model, the training manager may retain additional data to verify the model's performance. For example, the training dataset may be divided into 60% for training and 40% for evaluation or validation, which may be divided into two or more stages. After selecting the model parameters that best fit the evaluation data, resulting in convergence on a subset of the validation data (e.g., half of the validation data), a second validation run can be performed using the remaining validation data to ensure the model's performance. If the model meets expectations on the validation data, then the model is not overfitting the data. Optionally, a test or holdout set can be used to test parameters. Using a second validation or testing step can help choose appropriate model parameters to prevent overfitting. However, taking more data out of the training process for validation results in less data available for training. This can be problematic for smaller datasets, as there may not be enough data available for training. One approach in this situation is to perform cross-validation, as described elsewhere in this article.
[0133] There are many metrics or insights that can be used to review and evaluate the predictive accuracy of a given model. A sample evaluation result includes a predictive accuracy metric to report the overall success of the model, as well as visualizations that help explore the model's accuracy beyond the predictive accuracy metric. The results can also provide the ability to view the impact of setting score thresholds (such as for binary classification) and can generate alerts based on the criteria used to check the effectiveness of the evaluation. The choice of metric and visualization may depend at least in part on the type of model being evaluated.
[0134] After satisfactory training and evaluation, the trained machine learning model can be used to build or support a machine learning application. In one embodiment, building a machine learning application is an iterative process involving a series of steps. The core machine learning problem can be constructed based on what is observed and the answer that the model is to predict. Data can then be collected, cleaned, and prepared to make it suitable for use by the machine learning model training algorithm. This data can be visualized and analyzed for integrity checks to verify data quality and understand the data. The original data (e.g., input variables) and answer data (e.g., targets) may not be represented in a way that can be used to train a highly predictive model. Therefore, it may be desirable to build a more predictive input representation or feature from the original variables. The resulting features can be input into a learning algorithm to build a model and evaluate the quality of the model based on the data retained from model construction. The model can then be used to generate predictions of the target answer for new data instances.
[0135] exist Figure 15 In the exemplary system 1500 of FIG. 1 , after providing the evaluation, the trained model 1510 is provided or made available to a classifier 1514, which is capable of processing validation data using the trained model. For example, this may include data received from a user or an unclassified third party, such as a query image for which information regarding the content represented in the image is being sought. The validation data can be processed by the classifier using the trained model, and the resulting results 1516 (e.g., a classification or prediction) can be sent back to the corresponding source or otherwise processed or stored. In at least one embodiment, and where such use is permitted, these currently-classified data instances can be stored in a training data repository and can be used by a training manager for further training of the trained model 1508. In at least one embodiment, the model is trained continuously as new data becomes available, but in at least one embodiment, the model is trained periodically, such as daily or weekly, depending on factors such as the size of the dataset or the complexity of the model.
[0136] The classifier 1514 may include appropriate hardware and software for processing the validation data 1512 using the trained model. In some cases, the classifier will include one or more computer servers, each having one or more graphics processing units (GPUs) capable of processing data. The configuration and design of the GPU may make them more suitable for processing machine learning data than a CPU or other such components. The trained model in at least one embodiment may be loaded into the GPU memory, and the received data instances may be provided to the GPU for processing. The GPU may have many more cores than a CPU, and the GPU core may be less complex. Therefore, a given GPU may be able to process thousands of data instances simultaneously through different hardware threads. The GPU may also be configured to maximize floating point throughput, which can provide significant additional processing advantages for large data sets.
[0137] Even when using GPUs, accelerators, and other such hardware to accelerate tasks such as model training or data classification using such models, such tasks can still require significant time, resource allocation, and cost. For example, if a machine learning model is to be trained using 800 passes, and the dataset includes 1,000,000 data instances to be used for training, each pass will need to process all 1 million instances. Different parts of the architecture can also be supported by different types of devices. For example, training can be performed using a set of servers at a logically centralized location, such as can be provided as a service, while classification of the raw data can be performed by such a service or on client devices, among other such options. These devices can also be owned, operated, or controlled by the same entity or multiple entities.
[0138] Figure 16 An example neural network 1600 that can be trained or otherwise utilized in at least one embodiment is shown. In this example, the statistical model is an artificial neural network (ANN) that includes multiple layers of nodes, including an input layer 1602, an output layer 1606, and multiple layers 1604 of intermediate nodes, typically referred to as "hidden" layers because the internal layers and nodes are typically not visible or accessible in conventional neural networks. Although several intermediate layers are shown for illustrative purposes only, it should be understood that there is no limit to the number of intermediate layers that can be utilized, and any limit on layers will generally be a factor of the resources or time required to process the model. As discussed elsewhere herein, in addition to other such options, other types of models, networks, algorithms, or processes may also be used that may include other numbers or selections of nodes and layers. Validation data may be processed by each layer of the network to generate a set of inference or reasoning scores, which may then be fed into a loss function 1608.
[0139] In this example network 1600, all nodes in a given layer are interconnected to all nodes in adjacent layers. As shown in the figure, the nodes in the middle layer will then be connected to the nodes of the two adjacent layers respectively. In some models, nodes are also called neurons or connected units, and the connections between nodes are called edges. Each node can perform a function for the input received, for example by using a specified function. Nodes and edges can be given different weights during the training process, and each layer of nodes can perform specific types of transformations on the input received, and these transformations can also be learned or adjusted during the training process. Learning can be supervised learning or unsupervised learning, which may depend at least in part on the type of information contained in the training data set. Various types of neural networks can be used, for example, convolutional neural networks (CNNs), which include many convolutional layers and a set of pooling layers and have been shown to be beneficial for applications such as image recognition. CNNs are also easier to train than other networks because the number of parameters to be determined is relatively small.
[0140] In at least one embodiment, various tuning parameters can be used to train such complex machine learning models. Selecting parameters, fitting the model, and evaluating the model are part of the model tuning process, often referred to as hyperparameter optimization. In at least one embodiment, such tuning can include introspecting the underlying model or data. In training or production settings, a robust workflow is important to avoid overfitting of hyperparameters, as described elsewhere herein. Cross-validation and adding Gaussian noise to the training dataset are useful techniques to avoid overfitting to any one dataset. For hyperparameter optimization, it may be desirable to keep the training and validation sets fixed. In at least one embodiment, hyperparameters can be tuned in certain categories, which can include, for example, data preprocessing (e.g., converting words to vectors), CNN architecture definition (e.g., filter size, number of filters), stochastic gradient descent (SGD) parameters (e.g., learning rate), regularization or refinement (e.g., dropout probability), and other such options.
[0141] In an example preprocessing step, instances in a dataset can be embedded into a lower-dimensional space of a specific size. The size of this space is a parameter to be adjusted. The CNN architecture contains many adjustable parameters. The filter size parameter can represent the interpretation of information corresponding to the size of the instance to be analyzed. In computational linguistics, this is called the n-gram size. The example CNN uses three different filter sizes, which represent different possible n-gram sizes. The number of filters in each filter size can correspond to the depth of the filter. Each filter attempts to learn something different from the instance structure, such as the sentence structure of text data. In the convolutional layer, the activation function can be a rectified linear unit, and the pooling type is set to max pooling. The results can then be concatenated into a one-dimensional vector, and the final layer is fully concatenated to the two-dimensional output. This corresponds to binary classification, to which an optimization function can be applied. One such function is an implementation of the root mean square (RMS) propagation method of gradient descent, where example hyperparameters can include the learning rate, batch size, maximum gradient normal, and epochs. For neural networks, regularization can be a very important consideration. In at least one embodiment, the input data can be relatively sparse. In this case, the key hyperparameter might be the dropout of the penultimate layer, which means that a certain percentage of nodes will not "fire" during each training cycle. The example training process can suggest different hyperparameter configurations based on feedback on the performance of previous configurations. The model can be trained using the suggested configurations, evaluated on a specified validation set, and performance reported. This process can be repeated, balancing exploration (learning more about different configurations) and exploitation (leveraging prior knowledge to achieve better results).
[0142] Because CNN training can be parallelized and utilize GPU-powered computing resources, multiple optimization strategies can be tried for different scenarios. Complex scenarios allow for tuning of the model architecture, preprocessing, and stochastic gradient descent parameters. This expands the model configuration space. In the basic scenario, only preprocessing and stochastic gradient descent parameters are tuned. Compared to the basic scenario, complex scenarios allow for more configuration parameters. Tuning of the joint space can be performed using linear or exponential steps and iterated through the model's optimization loop. This tuning process can be significantly less expensive than tuning procedures such as random search and grid search, without any noticeable performance loss.
[0143] In at least one embodiment, back propagation can be used to calculate the gradient for determining the weights of a neural network. Back propagation is a form of differentiation, and as described above, a gradient descent optimization algorithm can be used to adjust the weights applied to various nodes or neurons. The gradient of the relevant loss function can be used to determine the weights. Back propagation can utilize the derivative of the output generated by the loss function to the statistical model. As described above, each node can have an associated activation function that defines the output of each node. Various activation functions can be used appropriately, such as radial basis functions (RBFs) and sigmoid functions, which can be used for data conversion by various support vector machines (SVMs). The activation function of the intermediate layer of the node is referred to as the inner product core in this article. These functions can include, for example, identity functions, step functions, sigmoid functions, ramp functions, etc. The activation function can also be linear or nonlinear, as well as other such options.
[0144] Reasoning and training logic
[0145] Figure 17A Inference and / or training logic 1715 is shown for performing inference and / or training operations associated with one or more embodiments. Figure 17A and / or 17B provide details regarding the inference and / or training logic 1715 .
[0146] In at least one embodiment, inference and / or training logic 1715 may include, but is not limited to, code and / or data memory 1701 to store forward and / or output weights and / or input / output data and / or other parameters to configure neurons or layers of a neural network for training and / or inference in aspects of one or more embodiments. In at least one embodiment, training logic 1715 may include or be coupled to code and / or data memory 1701 to store graphics code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code (such as graphics code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network. In at least one embodiment, code and / or data memory 1701 stores input / output data during training and / or inference using aspects of one or more embodiments and / or weight parameters and / or weight parameters during forward propagation of the weight parameters for each layer of a neural network trained or used in conjunction with one or more embodiments. In at least one embodiment, any portion of code and / or data memory 1701 may be included within other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0147] In at least one embodiment, any portion of code and / or data storage 1701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 1701 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the choice of whether code and / or code and / or data storage 1701 is internal or external to a processor, for example, or composed of DRAM, SRAM, flash memory, or some other memory type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.
[0148] In at least one embodiment, inference and / or training logic 1715 may include, but is not limited to, code and / or data memory 1705 to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in accordance with aspects of one or more embodiments. In at least one embodiment, code and / or data memory 1705 stores input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments, and weight parameters and / or input / output data for each layer of a neural network trained or used with one or more embodiments during backward propagation. In at least one embodiment, training logic 1715 may include or be coupled to code and / or data memory 1705 to store graphics code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code (such as graphics code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data memory 1705 may be included with other on-chip or off-chip data memory, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data memory 1705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, data storage 1705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the choice of whether code and / or data storage 1705 is internal or external to the processor, for example, consisting of DRAM, SRAM, flash memory, or some other type of memory, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in inference and / or training of the neural network, or some combination of these factors.
[0149] In at least one embodiment, code and / or data memory 1701 and code and / or data memory 1705 may be separate memory structures. In at least one embodiment, code and / or data memory 1701 and code and / or data memory 1705 may be the same memory structure. In at least one embodiment, code and / or data memory 1701 and code and / or data memory 1705 may be partially the same memory structure and partially separate memory structures. In at least one embodiment, any portion of code and / or data memory 1701 and code and / or data memory 1705 may be included with other on-chip or off-chip data memory, including the processor's L1, L2, or L3 cache or system memory.
[0150] In at least one embodiment, the inference and / or training logic 1715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1710 , including integer and / or floating point units, to perform logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation memory 1720 , which are functions of input / output and / or weight parameter data stored in code and / or data memory 1701 and / or code and / or data memory 1705 . In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 1710 to generate activations stored in activation memory 1720, where weight values stored in code and / or data store 1705 and / or code and / or data store 1701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data store 1705 or code and / or data store 1701 other on-chip or off-chip memory.
[0151] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 1710, while in another embodiment, one or more ALUs 1710 may be external to the processor or other hardware logic device or circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 1710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution units of a processor, which may be within the same processor or distributed across different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data memory 1701, code and / or data memory 1705, and activation memory 1720 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be on different processors or other hardware logic devices or circuits, or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation memory 1720 may be included with other on-chip or off-chip data memory, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0152] In at least one embodiment, activation memory 1720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, activation memory 1720 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, whether activation memory 1720 is internal or external to the processor, for example, or comprises DRAM, SRAM, flash memory, or other memory types, may be selected based on the memory available on or off chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferring and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 17A The inference and / or training logic 1715 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 17AThe illustrated inference and / or training logic 1715 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).
[0153] Figure 17B Inference and / or training logic 1715 is shown in accordance with at least one or more embodiments. In at least one embodiment, inference and / or training logic 1715 may include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 17B The inference and / or training logic 1715 shown in FIG can be used in conjunction with an application specific integrated circuit (ASIC), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 17B The inference and / or training logic 1715 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 1715 includes, but is not limited to, code and / or data memory 1701 and code and / or data memory 1705, which can be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 17B In at least one embodiment shown in , code and / or data memory 1701 and code and / or data memory 1705 are each associated with dedicated computing resources (e.g., computing hardware 1702 and computing hardware 1706), respectively. In at least one embodiment, computing hardware 1702 and computing hardware 1706 each include one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) solely on information stored in code and / or data memory 1701 and code and / or data memory 1705, respectively, with the results of the functions being stored in activation memory 1720.
[0154] In at least one embodiment, each of code and / or data storage 1701 and 1705 and corresponding computational hardware 1702 and 1706 corresponds to a different layer of a neural network, such that activations from one "storage / computation pair 1701 / 1702" of code and / or data storage 1701 and computational hardware 1702 are provided as inputs to a "storage / computation pair 1705 / 1706" of code and / or data storage 1705 and computational hardware 1706, reflecting the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 1701 / 1702 and 1705 / 1706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) can be included in the inference and / or training logic 1715 after or in parallel with the storage / computation pairs 1701 / 1702 and 1705 / 1706.
[0155] Data Center
[0156] Figure 18 An example data center 1800 is shown in which at least one embodiment may be used. In at least one embodiment, the data center 1800 includes a data center infrastructure layer 1810 , a framework layer 1820 , a software layer 1830 , and an application layer 1840 .
[0157] In at least one embodiment, Figure 18 As shown, the data center infrastructure layer 1810 may include a resource coordinator 1812, grouped computing resources 1814, and node computing resources ("node CRs") 1816(1)-1816(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 1816(1)-1816(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules and cooling modules, etc. In at least one embodiment, one or more of the node CRs 1816(1)-1816(N) may be a server having one or more of the above-mentioned computing resources.
[0158] In at least one embodiment, the grouped computing resources 1814 may include separate groups of node CRs housed in one or more racks (not shown), or many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 1814 may include computing, network, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including CPUs or processors may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0159] In at least one embodiment, resource coordinator 1812 may configure or otherwise control one or more nodes CR 1816(1)-1816(N) and / or grouped computing resources 1814. In at least one embodiment, resource coordinator 1812 may comprise a software design infrastructure ("SDI") management entity for data center 1800. In at least one embodiment, resource coordinator may comprise hardware, software, or some combination thereof.
[0160] In at least one embodiment, Figure 18 As shown, the framework layer 1820 includes a job scheduler 1822, a configuration manager 1824, a resource manager 1826, and a distributed file system 1828. In at least one embodiment, the framework layer 1820 may include a framework that supports software 1832 of the software layer 1830 and / or one or more applications 1842 of the application layer 1840. In at least one embodiment, the software 1832 or the application 1842 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1820 may be, but is not limited to, a free and open source software web application framework, such as Apache Spark, which may utilize the distributed file system 1828 for large-scale data processing (e.g., "big data"). TM(hereinafter referred to as "Spark"). In at least one embodiment, job scheduler 1822 may include a Spark driver to facilitate scheduling of workloads supported by various layers of data center 1800. In at least one embodiment, configuration manager 1824 may be capable of configuring different layers, such as software layer 1830 and framework layer 1820 including Spark and a distributed file system 1828 for supporting large-scale data processing. In at least one embodiment, resource manager 1826 may be capable of managing clustered or grouped computing resources mapped to or allocated to support distributed file system 1828 and job scheduler 1822. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1814 on data center infrastructure layer 1810. In at least one embodiment, resource manager 1826 may coordinate with resource coordinator 1812 to manage these mapped or allocated computing resources.
[0161] In at least one embodiment, software 1832 included in software layer 1830 may include software used by at least a portion of node CRs 1816(1)-1816(N), grouped computing resources 1814, and / or distributed file system 1828 of framework layer 1820. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0162] In at least one embodiment, the applications 1842 included in the application layer 1840 may include one or more types of applications used by at least a portion of the node CRs 1816(1)-1816(N), the grouped computing resources 1814, and / or the distributed file system 1828 of the framework layer 1820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0163] In at least one embodiment, any of configuration manager 1824, resource manager 1826, and resource coordinator 1812 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 1800 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0164] In at least one embodiment, data center 1800 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 1800. In at least one embodiment, using the weight parameters calculated using one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using the resources described above with respect to data center 1800.
[0165] In at least one embodiment, a data center can use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.
[0166] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 18 In some embodiments, the present invention provides a method for performing inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0167] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0168] autonomous vehicles
[0169] Figure 19A An example of an autonomous vehicle 1900 according to at least one embodiment is shown. In at least one embodiment, autonomous vehicle 1900 (alternatively referred to herein as "vehicle 1900") can be, but is not limited to, a passenger vehicle, such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 1a00 can be a semi-tractor-trailer for hauling cargo. In at least one embodiment, vehicle 1a00 can be an aircraft, a robotic vehicle, or another type of vehicle.
[0170] Automated driving vehicles may be described according to the levels of automation defined by the National Highway Traffic Safety Administration ("NHTSA") and the Society of Automotive Engineers ("SAE") under the U.S. Department of Transportation, "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, dated June 15, 2018, Standard No. J3016-201609, dated September 30, 2016, and previous and future versions of this standard). In one or more embodiments, the vehicle 1900 may be capable of functioning according to one or more of the levels 1 to 5 of automated driving. For example, in at least one embodiment, the vehicle 1900 may be capable of conditional automation (level 3), high automation (level 4), and / or full automation (level 5), depending on the embodiment.
[0171] In at least one embodiment, vehicle 1900 may include, but is not limited to, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, vehicle 1900 may include, but is not limited to, a propulsion system 1950, such as an internal combustion engine, a hybrid power plant, an all-electric engine, and / or another type of propulsion system. In at least one embodiment, propulsion system 1950 may be connected to a drive train of vehicle 1900, which may include, but is not limited to, a transmission, to enable propulsion of vehicle 1900. In at least one embodiment, propulsion system 1950 may be controlled in response to receiving a signal from throttle / accelerator 1952.
[0172] In at least one embodiment, when propulsion system 1950 is operating (e.g., when the vehicle is traveling), a steering system 1954 (which may include, but is not limited to, a steering wheel) is used to steer vehicle 1900 (e.g., along a desired path or route). In at least one embodiment, steering system 1954 may receive signals from steering actuator 1956. A steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, brake sensor system 1946 may be used to operate the vehicle brakes in response to signals received from brake actuator 1948 and / or brake sensors.
[0173] In at least one embodiment, the controller 1936 may include, but is not limited to, one or more system-on-chips ("SoCs") ( Figure 19A1900 ). The controller 1936 may include a graphics processing unit (GPU) (not shown) and / or a graphics processing unit ("GPU") to provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1900. For example, in at least one embodiment, the controller 1936 may send signals to operate the vehicle brakes via the brake actuator 1948, to operate the steering system 1954 via the steering actuator 1956, and / or to operate the propulsion system 1950 via the throttle / accelerator 1952. The controller 1936 may include one or more onboard (e.g., integrated) computing devices (e.g., a supercomputer) that processes sensor signals and outputs operational commands (e.g., signals representing commands) to implement autonomous driving and / or assist the driver in driving the vehicle 1900. In at least one embodiment, the controller 1936 may include a first controller 1936 for autonomous driving functionality, a second controller 1936 for functional safety functionality, a third controller 1936 for artificial intelligence functionality (e.g., computer vision), a fourth controller 1936 for infotainment functionality, a fifth controller 1936 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 1936 may handle two or more of the above functions, two or more controllers 1936 may handle a single function, and / or any combination thereof.
[0174] In at least one embodiment, the controller 1936 provides signals for controlling one or more components and / or systems of the vehicle 1900 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, the sensor data can be received from sensors such as, but not limited to, a global navigation satellite system ("GNSS") sensor 1958 (e.g., a global positioning system sensor), a RADAR sensor 1960, an ultrasonic sensor 1962, a LIDAR sensor 1964, an inertial measurement unit (IMU) sensor 1966 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1996, a stereo camera 1968, a wide angle camera 1970 (e.g., a fisheye camera), an infrared camera 1972, a surround camera 1974 (e.g., a 360-degree camera), a telephoto camera (e.g., a 360-degree camera), or a 360-degree camera. Figure 19A Not shown), mid-range camera ( Figure 19A ), speed sensor 1944 (e.g., for measuring the speed of vehicle 1900), vibration sensor 1942, steering sensor 1940, brake sensor (e.g., as part of brake sensor system 1946), and / or other sensor types are received.
[0175] In at least one embodiment, one or more controllers 1936 may receive input (e.g., represented by input data) from a dashboard 1932 of the vehicle 1900 and provide output (e.g., represented by output data, display data, etc.) via a human machine interface ("HMI") display 1934, an audible annunciator, a speaker, and / or other components of the vehicle 1900. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high definition map ( Figure 19A ), location data (e.g., the location of the vehicle 1900, such as on a map), directions, the locations of other vehicles (e.g., occupancy barriers), information about objects and the states of objects sensed by the controller 1936, etc. For example, in at least one embodiment, the HMI display 1934 can display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about the driving maneuvers the vehicle has made, is making, or will make (e.g., changing lanes now, taking exit 34B in two miles, etc.).
[0176] In at least one embodiment, the vehicle 1900 further includes a network interface 1924 that can communicate via one or more networks using a wireless antenna 1926 and / or a modem. For example, in at least one embodiment, the network interface 1924 may be capable of communicating via Long Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. In at least one embodiment, the wireless antenna 1926 can also use a local area network (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or a low power wide area network (hereinafter referred to as "LPWAN") (e.g., LoRaWAN, SigFox, etc.) to enable communication between objects in the environment (e.g., vehicles, mobile devices).
[0177] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 19A for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0178] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0179] Figure 19B According to at least one embodiment, Figure 19A Examples of camera locations and fields of view for autonomous vehicle 1900 are shown. In at least one embodiment, the cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located in different locations on vehicle 1900.
[0180] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera that may be suitable for use with components and / or systems of the vehicle 1900. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level ("ASIL") B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 120fps, 240fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-clear-clear ("RCCC") filter array, a red-clear-clear-blue ("RCCB") filter array, a red-blue-green-clear ("RBGC") filter array, a Foveon X3 filter array, a Bayer sensor ("RGGB") filter array, a monochrome sensor filter array, and / or other types of filter arrays. In at least one embodiment, a clear pixel camera, such as one having an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase photosensitivity.
[0181] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system ("ADAS") functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0182] In at least one embodiment, one or more cameras can be mounted in a mounting assembly, such as a custom designed (three-dimensional ("3D") printed) assembly, so as to cut out stray light and reflections from within the car (e.g., reflections from the dashboard reflecting in the windshield mirror) that might interfere with the camera's ability to capture image data. With respect to the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly can be 3D printed custom so that the camera mounting plate matches the shape of the rearview mirror. In at least one embodiment, the camera can be integrated into the rearview mirror. For side-view cameras, in at least one embodiment, the camera can also be integrated into the four pillars at each corner of the cabin.
[0183] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view that includes a portion of the environment in front of the vehicle 1900 can be used for surround vision, as well as to help identify the forward path and obstacles with the assistance of one or more controllers 1936 and / or control SoCs, thereby providing information that is critical for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning ("LDW"), automatic cruise control ("ACC"), and / or other functions (e.g., traffic sign recognition).
[0184] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS ("Complementary Metal Oxide Semiconductor") color imager. In at least one embodiment, a wide-angle camera 1970 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the road, or bicycles). Although Figure 19B Only one wide-angle camera 1970 is shown in FIG. 1 , however, in other embodiments, any number (including zero) of wide-angle cameras 1970 may be present on the vehicle 1900. In at least one embodiment, any number of remote cameras 1998 (e.g., a remote stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the remote cameras 1998 may also be used for object detection and classification, as well as basic object tracking.
[0185] In at least one embodiment, any number of stereo cameras 1968 may also be included in a forward-facing configuration. In at least one embodiment, one or more stereo cameras 1968 may include an integrated control unit including a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the vehicle 1900's environment, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1968 may include, but are not limited to, a compact stereo vision sensor, which may include, but are not limited to, two camera lenses (one on each side) and an image processing chip that may measure the distance from the vehicle 1900 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1968 may be used in addition to those described herein.
[0186] In at least one embodiment, a camera having a field of view of a portion of the environment including the sides of the vehicle 1900 (e.g., a side-view camera) can be used for surround viewing to provide information for creating and updating occupancy grids and generating side collision warnings. For example, in at least one embodiment, the surround camera 1974 (e.g., Figure 19B Four surround cameras 1974 (shown) can be positioned on the vehicle 1900. In at least one embodiment, the surround cameras 1974 can include, but are not limited to, any number and combination of wide-angle cameras 1970, fisheye cameras, 360-degree cameras, and / or the like. For example, in at least one embodiment, four fisheye cameras can be located on the front, rear, and sides of the vehicle 1900. In at least one embodiment, the vehicle 1900 can use three surround cameras 1974 (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0187] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 1900 (e.g., a rearview camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. In at least one embodiment, a variety of cameras can be used, including but not limited to cameras that are also suitable as forward-facing cameras (e.g., long-range camera 1998 and / or mid-range camera 1976, stereo camera 1968, infrared camera 1972, etc.), as described herein.
[0188] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be used to Figure 19B In a system of the present invention, an inference or prediction operation is performed based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0189] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0190] Figure 19C According to at least one embodiment, Figure 19A A block diagram of an example system architecture for an autonomous vehicle 1900 is provided. In at least one embodiment, Figure 19C Each of one or more components, one or more features, and one or more systems of vehicle 1900 is shown as being connected via bus 1902. In at least one embodiment, bus 1902 may include, but is not limited to, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, the CAN bus can be a network internal to vehicle 1900 that helps control various features and functions of vehicle 1900, such as brake actuation, acceleration, braking, steering, wipers, etc. In one embodiment, bus 1902 can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1902 can be read to find steering wheel angle, ground speed, engine revolutions per minute ("RPM"), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1902 can be an ASIL B compliant CAN bus.
[0191] In at least one embodiment, FlexRay and / or Ethernet may be used in addition to or in addition to CAN. In at least one embodiment, there may be any number of buses 1902, which may include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 1902 may be used to perform different functions and / or for redundancy. For example, a first bus 1902 may be used for collision avoidance functionality, and a second bus 1902 may be used for actuation control. In at least one embodiment, each bus 1902 may communicate with any component of the vehicle 1900, and two or more buses 1902 may communicate with the same component. In at least one embodiment, each of any number of system-on-chips ("SoCs") 1904, each of the controllers 1936, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 1900) and may be connected to a common bus, such as a CAN bus.
[0192] In at least one embodiment, the vehicle 1900 may include one or more controllers 1936, such as those described herein with respect to Figure 19A Those described. Controller 1936 can be used for a variety of functions. In at least one embodiment, controller 1936 can be coupled to any of the various other components and systems of vehicle 1900 and can be used to control vehicle 1900, vehicle 1900's artificial intelligence, vehicle 1900's infotainment, etc.
[0193] In at least one embodiment, the vehicle 1900 may include any number of SoCs 1904. Each of the SoCs 1904 may include, but is not limited to, a central processing unit ("CPU") 1906, a graphics processing unit ("GPU") 1908, a processor 1910, a cache 1912, an accelerator 1914, a data store 1916, and / or other components and features not shown. In at least one embodiment, the SoCs 1904 may be used to control the vehicle 1900 in a variety of platforms and systems. For example, in at least one embodiment, the SoCs 1904 may be combined in a system (e.g., a system of the vehicle 1900) with a high-definition ("HD") map 1922 that may be downloaded from one or more servers (e.g., a system of the vehicle 1900) via a network interface 1924. Figure 19C (not shown) to obtain map refreshes and / or updates.
[0194] In at least one embodiment, CPU 1906 may include a CPU cluster or CPU complex (alternatively referred to herein as "CCPLEX"). In at least one embodiment, CPU 1906 may include multiple cores and / or a level 2 ("L2") cache. For example, in at least one embodiment, CPU 1906 may include eight cores in a mutually coupled multiprocessor configuration. In at least one embodiment, CPU 1906 may include four dual-core clusters, each with a dedicated L2 cache (e.g., a 2MB L2 cache). In at least one embodiment, CPU 1906 (e.g., CCPLEX) may be configured to support simultaneous cluster operations such that any combination of CPU 1906's clusters may be active at any given time.
[0195] In at least one embodiment, one or more CPUs 1906 can implement power management functionality including, but not limited to, one or more of the following features: automatic clock gating of various hardware modules when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to executing a wait for interrupt ("WFI") / wait for event ("WFE") instruction; each core can be independently powered; each core cluster can be independently clock gated when all cores are clock gated or power gated; and / or each core cluster can be independently power gated when all cores are power gated. In at least one embodiment, CPU 1906 can further implement an enhanced algorithm for managing power states, where allowed power states and expected wakeup times are specified, and hardware / microcode determines the optimal power state for the core, cluster, and CCPLEX input. In at least one embodiment, the processing core can support a simplified power state entry sequence in software, where the work is offloaded to the microcode.
[0196] In at least one embodiment, GPU 1908 may include an integrated GPU (or alternatively referred to herein as an "iGPU"). In at least one embodiment, GPU 1908 may be programmable and may be efficient for parallel workloads. In at least one embodiment, GPU 1908, in at least one embodiment, may utilize an enhanced tensor instruction set. In at least one embodiment, GPU 1908 may include one or more streaming microprocessors, wherein each streaming microprocessor may include a level 1 ("L1") cache (e.g., an L1 cache having at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache having 512KB of storage capacity). In at least one embodiment, GPU 1908 may include at least eight streaming microprocessors. In at least one embodiment, GPU 1908 may utilize a computing application programming interface (API). In at least one embodiment, GPU 1908 may utilize one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0197] In at least one embodiment, one or more GPUs 1908 can be power-optimized for optimal performance in automotive and embedded use cases. For example, in one embodiment, GPU 1908 can be fabricated on fin field-effect transistors (“FinFETs”). In at least one embodiment, each streaming microprocessor can include multiple mixed-precision processing cores divided into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, a level 0 (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads that mix compute and addressing operations. In at least one embodiment, the streaming microprocessor can include independent thread scheduling capabilities to enable finer-grained synchronization and collaboration between parallel threads. In at least one embodiment, a streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0198] In at least one embodiment, one or more GPUs 1908 may include high bandwidth memory ("HBM") and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / s in some examples. In at least one embodiment, synchronous graphics random access memory ("SGRAM"), such as graphics double data rate type five synchronous random access memory ("GDDR5"), may be used in addition to or in lieu of HBM memory.
[0199] In at least one embodiment, the GPU 1908 may include unified memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU 1908 to directly access the CPU 1906 page tables. In at least one embodiment, when the GPU 1908 memory management unit ("MMU") experiences a miss, an address translation request may be sent to the CPU 1906. In response, in at least one embodiment, the CPU 1906 may look up the virtual-to-physical mapping of the address in its page tables and transmit the translation back to the GPU 1908. In at least one embodiment, unified memory technology may allow a single unified virtual address space to be used for both the CPU 1906 and GPU 1908 memory, thereby simplifying programming the GPU 1908 and porting applications to the GPU 1908.
[0200] In at least one embodiment, GPU 1908 may include any number of access counters that can track how frequently GPU 1908 accesses other processors' memory. In at least one embodiment, the access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the page most frequently, thereby improving the efficiency of memory ranges shared between processors.
[0201] In at least one embodiment, one or more SoCs 1904 may include any number of caches 1912, including those described herein. For example, in at least one embodiment, cache 1912 may include a level 3 ("L3") cache available to both CPU 1906 and GPU 1908 (e.g., connecting two CPUs 1906 and GPU 1908). In at least one embodiment, cache 1912 may include a write-back cache that can track the state of a line, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4MB or more, depending on the embodiment, although smaller cache sizes may be used.
[0202] In at least one embodiment, one or more SoCs 1904 may include one or more accelerators 1914 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC 1904 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement the GPU 1908 and offload some tasks of the GPU 1908 (e.g., freeing up more cycles of the GPU 1908 to perform other tasks). In at least one embodiment, the accelerator 1914 may be used for target workloads that are sufficiently stable to withstand acceleration (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, the CNN may include a region-based or region-based convolutional neural network (“RCNN”) and a fast RCNN (e.g., as used for object detection) or other types of CNNs.
[0203] In at least one embodiment, the accelerator 1914 (e.g., a hardware acceleration cluster) may include a deep learning accelerator ("DLA"). The DLA may include, but is not limited to, one or more Tensor Processing Units ("TPUs"), which may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be an accelerator configured and optimized to perform image processing functions (e.g., for CNN, RCNN, etc.). The DLA may be further optimized for a specific set of neural network types and floating point operations and inference. In at least one embodiment, the design of the DLA may provide higher performance per millimeter than a typical general-purpose GPU, and generally significantly exceeds the performance of a CPU. In at least one embodiment, the TPU may perform several functions, including single-instance convolution functions that support, for example, INT8, INT16, and FP16 data types for features and weights, as well as post-processor functions. In at least one embodiment, the DLA can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, for example, but not limited to: a CNN for object recognition and detection using data from a camera sensor; a CNN for distance estimation using data from a camera sensor; a CNN for emergency vehicle detection and recognition and detection using data from a microphone 1996; a CNN for face recognition and vehicle owner recognition using data from a camera sensor; and / or a CNN for safety and / or security-related events.
[0204] In at least one embodiment, the DLA can perform any function of the GPU 1908, and by using an inference accelerator, for example, designers can target either the DLA or the GPU 1908 for any function. For example, in at least one embodiment, designers can focus CNN processing and floating-point operations on the DLA and leave other functions to the GPU 1908 and / or other accelerators 1914.
[0205] In at least one embodiment, the accelerators 1914 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator ("PVA"), which may be referred to herein alternatively as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems ("ADAS") 1938, autonomous driving, augmented reality ("AR") applications, and / or virtual reality ("VR") applications. The PVA may strike a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, for example, but not limited to, any number of reduced instruction set computer ("RISC") cores, direct memory access ("DMA"), and / or any number of vector processors.
[0206] In at least one embodiment, the RISC core can interact with an image sensor (e.g., an image sensor of any camera described herein), an image signal processor, and the like. In at least one embodiment, each RISC core can include any amount of memory. In at least one embodiment, the RISC core can use any of a variety of protocols, depending on the embodiment. In at least one embodiment, the RISC core can execute a real-time operating system ("RTOS"). In at least one embodiment, the RISC core can be implemented using one or more integrated circuit devices, application specific integrated circuits ("ASICs"), and / or memory devices. For example, in at least one embodiment, the RISC core can include an instruction cache and / or tightly coupled RAM.
[0207] In at least one embodiment, the DMA can enable components of the PVA to access system memory independently of the CPU 1906. In at least one embodiment, the DMA can support any number of features for providing optimizations to the PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, the DMA can support up to six or more dimensions of addressing, which can include, but are not limited to, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0208] In at least one embodiment, the vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core can include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem can serve as the main processing engine of the PVA and can include a vector processing unit ("VPU"), an instruction cache, and / or a vector memory (e.g., "VMEM"). In at least one embodiment, the VPU can include a digital signal processor, such as a single instruction multiple data ("SIMD"), a very long instruction word ("VLIW") digital signal processor. In at least one embodiment, the combination of SIMD and VLIW can increase throughput and speed.
[0209] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to execute independently of the other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to exploit data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of images. In at least one embodiment, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA, among other things. In at least one embodiment, the PVAs may include additional error correction code ("ECC") memory to enhance overall system security.
[0210] In at least one embodiment, the accelerator 1914 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory ("SRAM") to provide high bandwidth, low latency SRAM to the accelerator 1914. In at least one embodiment, the on-chip memory may include at least 4MB of SRAM, including, for example, but not limited to, eight field-configurable memory blocks that can be accessed by both the PVA and the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and the DLA may access the memory via a backbone network that provides high-speed access to the memory to the PVA and the DLA. In at least one embodiment, the backbone network may include an on-chip computer vision network that interconnects the PVA and the DLA to the memory (e.g., using APB).
[0211] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for sending control signals / addresses / data, as well as burst-type communication for continuous data transmission. In at least one embodiment, the interface may comply with the International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("IEC") 61508 standards, although other standards and protocols may be used.
[0212] In at least one embodiment, one or more SoCs 1904 may include a real-time gaze tracking hardware accelerator. In at least one embodiment, the real-time gaze tracking hardware accelerator may be used to quickly and efficiently determine the position and range of objects (e.g., within a world model) to generate real-time visual simulations for use in RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.
[0213] In at least one embodiment, the accelerator 1914 (e.g., a cluster of hardware accelerators) has a wide range of uses for autonomous driving. In at least one embodiment, the PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of the PVA at low power and low latency are well matched to the algorithmic domain that requires predictable processing. In other words, the PVA performs well in semi-intensive or intensive conventional computations, even on small data sets, that require predictable runtimes with low latency and low power. In at least one embodiment, such as for an autonomous vehicle (vehicle 1900), the PVA is designed to run classic computer vision algorithms because they are efficient at object detection and integer math operations.
[0214] For example, according to at least one embodiment of the technology, PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for Level 3-5 autonomous driving use dynamic estimation / stereo matching on the fly (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, PVA can perform computer stereo vision functions on input from two monocular cameras.
[0215] In at least one embodiment, the PVA can be used to perform dense optical flow. For example, in at least one embodiment, the PVA can process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to provide processed time-of-flight data.
[0216] In at least one embodiment, the DLA can be used to run any type of network to enhance control and driving safety, including, for example, but not limited to, a neural network that outputs a confidence score for each object detection. In at least one embodiment, the confidence score can be expressed or interpreted as a probability, or as providing a relative "weight" of each detection relative to other detections. In at least one embodiment, the confidence score enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, in at least one embodiment, the system can set a threshold for the confidence score and only consider detections that exceed the threshold as true positive detections. In an embodiment using an automatic emergency braking ("AEB") system, a false positive detection will cause the vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, a highly confident detection can be considered a trigger for AEB. In at least one embodiment, the DLA can run a neural network for regressing the confidence score. In at least one embodiment, the neural network may take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), outputs of the IMU sensor 1966 associated with vehicle 1900 heading, distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LIDAR sensor 1964 or RADAR sensor 1960), and the like.
[0217] In at least one embodiment, one or more SoCs 1904 may include a data storage device 1916 (e.g., memory). In at least one embodiment, the data storage device 1916 may be on-chip memory of the SoC 1904 that may store the neural network to be executed on the GPU 1908 and / or DLA. In at least one embodiment, the data storage device 1916 may have a capacity large enough to store multiple instances of the neural network for redundancy and safety. In at least one embodiment, the data storage device 1916 may include an L2 or L3 cache.
[0218] In at least one embodiment, one or more SoCs 1904 may include any number of processors 1910 (e.g., embedded processors). In at least one embodiment, the processors 1910 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and related security implementations. In at least one embodiment, the boot and power management processor may be part of the SoC 1904 boot sequence and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assist in system low power state transitions, SoC 1904 thermal and temperature sensor management, and / or SoC 1904 power state management. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1904 may use the ring oscillator to detect the temperature of the CPU 1906, GPU 1908, and / or accelerator 1914. In at least one embodiment, if the temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine and place the SoC 1904 into a lower power consumption state and / or place the vehicle 1900 into a driver's safe parking pattern (e.g., bringing the vehicle 1900 to a safe stop).
[0219] In at least one embodiment, one or more processors 1910 may further include a set of embedded processors that can be used as an audio processing engine. In at least one embodiment, the audio processing engine can be an audio subsystem that can provide hardware with full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core that has a digital signal processor with dedicated RAM.
[0220] In at least one embodiment, processor 1910 may further include an always-on processor engine that may provide the necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the processor on the always-on processor engine may include, but is not limited to, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0221] In at least one embodiment, the processor 1910 may further include a safety cluster engine, which may include but is not limited to a dedicated processor subsystem for handling safety management of automotive applications. In at least one embodiment, the safety cluster engine may include but is not limited to two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.) and / or routing logic. In safety mode, in at least one embodiment, the two or more cores may operate in lockstep mode and may function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, the processor 1910 may further include a real-time camera engine, which may include but is not limited to a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, the processor 1910 may further include a high dynamic range signal processor, which may include but is not limited to an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0222] In at least one embodiment, processor 1910 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final video to produce the final image for the player window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1970, the surround camera 1974, and / or the in-cabin monitoring camera sensor. In at least one embodiment, the in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of SoC 1904, the neural network being configured to identify cabin events and respond accordingly. In at least one embodiment, the in-cabin system may perform, but is not limited to, lip reading to activate cellular service and place calls, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain features are available to the driver when the vehicle is operating in autonomous mode, but are otherwise disabled.
[0223] In at least one embodiment, the video image compositor may include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in the video, the noise reduction appropriately weights the spatial information, thereby reducing the weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, the temporal noise reduction performed by the video image compositor may use information from previous images to reduce noise in the current image.
[0224] In at least one embodiment, the video image compositor can also be configured to perform stereo rectification on the input stereo lens frames. In at least one embodiment, the video image compositor can also be used for user interface composition when using the operating system desktop, without requiring the GPU 1908 to continuously render new surfaces. In at least one embodiment, when the GPU 1908 is powered and actively performing 3D rendering, the video image compositor can be used to offload the GPU 1908 to improve performance and responsiveness.
[0225] In at least one embodiment, one or more SoCs 1904 may further include a Mobile Industry Processor Interface ("MIPI") camera serial interface for receiving video and input from a camera, a high-speed interface, and / or a video input block that may be used for a camera and associated pixel input functionality. In at least one embodiment, one or more SoCs 1904 may further include an input / output controller that may be controlled by software and may be used to receive I / O signals that are not assigned to a specific role.
[0226] In at least one embodiment, the one or more SoCs 1904 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. The SoC 1904 may be used to process data from cameras (e.g., connected via Gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensor 1964, RADAR sensor 1960, etc., which may be connected via Ethernet), data from the bus 1902 (e.g., vehicle 1900 speed, steering wheel position, etc.), data from the GNSS sensor 1958 (e.g., connected via Ethernet or CAN bus), etc. In at least one embodiment, the one or more SoCs 1904 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and may be used to offload the CPU 1906 from routine data management tasks.
[0227] In at least one embodiment, SoC 1904 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, thereby providing a comprehensive functional safety architecture that leverages and effectively uses computer vision and ADAS technologies to achieve diversity and redundancy, providing a platform that can provide a flexible and reliable driving software stack and deep learning tools. In at least one embodiment, SoC 1904 can be faster and more reliable than conventional systems, and even more energy efficient and space efficient. For example, in at least one embodiment, accelerator 1914, when combined with CPU 1906, GPU 1908, and data storage device 1916, can provide a fast and effective platform for level 3-5 autonomous vehicles.
[0228] In at least one embodiment, computer vision algorithms can be executed on a CPU, which can be configured using a high-level programming language (e.g., the C programming language) to execute a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs generally cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, which are used in in-vehicle ADAS applications and actual Level 3-5 autonomous vehicles.
[0229] The embodiments described herein allow for the execution of multiple neural networks simultaneously and / or sequentially, and for the results to be combined together to achieve Level 3-5 autonomous driving functionality. For example, in at least one embodiment, a CNN executed on a DLA or a discrete GPU (e.g., GPU 1920) may include text and word recognition, thereby allowing the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. In at least one embodiment, the DLA may also include a neural network that can recognize, interpret, and provide semantic understanding of the symbols and pass this semantic understanding to a path planning module running on the CPU Complex.
[0230] In at least one embodiment, for Level 3, 4, or 5 driving, multiple neural networks can be run simultaneously. For example, in at least one embodiment, a warning sign consisting of "Caution: flashing lights indicate icy conditions" along with connected lights can be interpreted by multiple neural networks independently or collectively. In at least one embodiment, the sign itself can be identified as a traffic sign by a first deployed neural network (e.g., an already trained neural network), and the text "flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which notifies the vehicle's path planning software (preferably executing on the CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be identified by operating a third deployed neural network over multiple frames, notifying the vehicle's path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks can run simultaneously, for example within the DLA and / or on GPU 1908.
[0231] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 1900. In at least one embodiment, the always-on sensor processor engine can be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and, in security mode, can be used to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 1904 provides protection against theft and / or carjacking.
[0232] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphone 1996 to detect and identify emergency vehicle sirens. In at least one embodiment, the SoC 1904 uses the CNN to classify environmental and urban sounds, as well as classify visual data. In at least one embodiment, the CNN running on the DLA is trained to identify the relative approaching speed of the emergency vehicle (e.g., by using the Doppler effect). In at least one embodiment, the CNN can also be trained to identify emergency vehicles for the area in which the vehicle is operating, as identified by the GNSS sensor 1958. In at least one embodiment, when operating in Europe, the CNN will seek to detect European sirens, while when operating in the United States, the CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used with the assistance of the ultrasonic sensor 1962 to execute emergency vehicle safety routines, slow the vehicle, pull over, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.
[0233] In at least one embodiment, the vehicle 1900 may include a CPU 1918 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 1904 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU 1918 may include an X86 processor, such as one of the CPUs 1918. The CPU 1918 may be used to perform any of a variety of functions, including, for example, arbitrating potential inconsistent results between ADAS sensors and the SoC 1904, and / or monitoring the status and health of the controller 1936 and / or the information system on chip ("information SoC") 1930.
[0234] In at least one embodiment, the vehicle 1900 may include a GPU 1920 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 1904 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, the GPU 1920 may provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and may be used to train and / or update the neural networks based at least in part on input from sensors of the vehicle 1900 (e.g., sensor data).
[0235] In at least one embodiment, vehicle 1900 may further include a network interface 1924, which may include, but is not limited to, a wireless antenna 1926 (e.g., one or more wireless antennas 1926 for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, network interface 1924 may be used to wirelessly connect to a cloud (e.g., a server and / or other network device), other vehicles, and / or computing devices (e.g., a passenger's client device) via the internet. In at least one embodiment, to communicate with other vehicles, a direct link may be established between vehicle 190 and the other vehicle, and / or an indirect link may be established (e.g., via a network and the internet). In at least one embodiment, a vehicle-to-vehicle communication link may be used to provide the direct link. The vehicle-to-vehicle communication link may provide vehicle 1900 with information about vehicles in its vicinity (e.g., vehicles in front of, to the sides of, and / or behind vehicle 1900). In at least one embodiment, the aforementioned functionality may be part of the cooperative adaptive cruise control functionality of vehicle 1900.
[0236] In at least one embodiment, the network interface 1924 may include a SoC that provides modulation and demodulation functionality and enables the controller 1936 to communicate over a wireless network. In at least one embodiment, the network interface 1924 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion may be performed by a well-known process and / or using a superheterodyne process. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functionality for communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0237] In at least one embodiment, the vehicle 1900 may further include data storage 1928, which may include, but is not limited to, off-chip (e.g., SoC 10904) memory. In at least one embodiment, the data storage 1928 may include, but is not limited to, one or more storage elements including RAM, SRAM, dynamic random access memory ("DRAM"), video random access memory ("VRAM"), flash memory, a hard disk, and / or other components and / or devices that can store at least one bit of data.
[0238] In at least one embodiment, the vehicle 1900 may further include a GNSS sensor 1958 (e.g., GPS and / or assisted GPS sensor) to assist with mapping, perception, occupancy raster generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1958 may be used, including, for example, but not limited to, a GPS connected to a serial interface (e.g., RS-232) bridge using a USB connector with Ethernet.
[0239] In at least one embodiment, the vehicle 1900 may further include one or more RADAR sensors 1960. The RADAR sensors 1960 may be used by the vehicle 1900 for remote vehicle detection, even in darkness and / or in adverse weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. The RADAR sensor 1960 may use CAN and / or bus 1902 (e.g., to transmit data generated by the RADAR sensor 1960) for control and access to object tracking data, and in some examples may access Ethernet to access raw data. In at least one embodiment, a variety of RADAR sensor types may be used. For example, but not limited to, the RADAR sensor 1960 may be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 1960 are pulse Doppler RADAR sensors.
[0240] In at least one embodiment, the RADAR sensor 1960 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range with side coverage, and the like. In at least one embodiment, the long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, the long-range RADAR system can provide a wide field of view achieved through two or more independent scans (e.g., within a range of 250 meters). In at least one embodiment, the RADAR sensor 1960 can help distinguish between static and moving objects and can be used by the ADAS system 1938 for emergency brake assistance and forward collision warning. The sensor 1960 included in the long-range RADAR system can include, but is not limited to, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In at least one embodiment, the six antennas, with the central four antennas, can create a focused beam pattern designed to record the vehicle 1900's surroundings at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, the additional two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 1900's lane.
[0241] In at least one embodiment, as an example, a medium-range RADAR system may include a range of up to 160m (front) or 80m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system may include, but is not limited to, any number of RADAR sensors 1960 designed to be mounted on both ends of the rear bumper. When mounted on both ends of the rear bumper, in at least one embodiment, the RADAR sensor system can generate two light beams that continuously monitor the rear of the vehicle and nearby blind spots. In at least one embodiment, the short-range RADAR system can be used in an ADAS system 1938 for blind spot detection and / or lane change assistance.
[0242] In at least one embodiment, the vehicle 1900 may further include one or more ultrasonic sensors 1962. The ultrasonic sensors 1962, which may be located on the front, rear, and / or sides of the vehicle 1900, may be used for parking assistance and / or for creating and updating occupancy barriers. In at least one embodiment, a variety of ultrasonic sensors 1962 may be used, and different ultrasonic sensors 1962 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensors 1962 may operate at an ASIL B functional safety level.
[0243] In at least one embodiment, the vehicle 1900 may include one or more LIDAR sensors 1964. The LIDAR sensors 1964 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensors 1964 may be functionally safe up to ASIL B. In at least one embodiment, the vehicle 1900 may include multiple (e.g., two, four, six, etc.) LIDAR sensors 1964 that may utilize Ethernet (e.g., providing data to a Gigabit Ethernet switch).
[0244] In at least one embodiment, LIDAR sensor 1964 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, commercially available LIDAR sensors 1964 may, for example, have an advertised range of approximately 100 meters, an accuracy of 2-3 cm, and support 100 Mbps Ethernet connectivity. In at least one embodiment, one or more non-obtrusive LIDAR sensors 1964 may be used. In such an embodiment, LIDAR sensor 1964 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 1900. In at least one embodiment, LIDAR sensor 1964 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, and have a range of 200 meters. In at least one embodiment, forward-facing LIDAR sensor 1964 may be configured for a horizontal field of view between 45 and 135 degrees.
[0245] In at least one embodiment, LIDAR technology (such as 3D flash LIDAR) may also be used. 3D flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200 meters around vehicle 1900. In at least one embodiment, the flash LIDAR unit includes, but is not limited to, a receiver that records the laser pulse propagation time and the reflected light at each pixel, which in turn corresponds to the range from vehicle 1900 to the object. In at least one embodiment, flash LIDAR can allow for the generation of a highly accurate and distortion-free image of the surrounding environment using each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 1900. In at least one embodiment, the 3D flash LIDAR system includes, but is not limited to, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, the flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture the reflected laser light in the form of a 3D range point cloud and co-registered intensity data.
[0246] In at least one embodiment, the vehicle may further include an IMU sensor 1966. In at least one embodiment, the IMU sensor 1966 may be located at the center of the rear axle of the vehicle 1900. In at least one embodiment, the IMU sensor 1966 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In at least one embodiment, for example, in a six-axis application, the IMU sensor 1966 may include, but not limited to, an accelerometer and a gyroscope. In at least one embodiment, for example, in a nine-axis application, the IMU sensor 1966 may include, but not limited to, an accelerometer, a gyroscope, and a magnetometer.
[0247] In at least one embodiment, IMU sensor 1966 can be implemented as a miniature, high-performance GPS-aided inertial navigation system ("GPS / INS") that combines microelectromechanical system ("MEMS") inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude; in at least one embodiment, IMU sensor 1966 can enable vehicle 1900 to estimate heading without input from a magnetic sensor by directly observing and correlating velocity changes from GPS to IMU sensor 1966. In at least one embodiment, IMU sensor 1966 and GNSS sensor 1958 can be combined in a single integrated unit.
[0248] In at least one embodiment, the vehicle 1900 can include microphones 1996 positioned within and / or around the vehicle 1900. In at least one embodiment, the microphones 1996 can be used for emergency vehicle detection and identification, among other things.
[0249] In at least one embodiment, the vehicle 1900 may further include any number of camera types, including stereo cameras 1968, wide angle cameras 1970, infrared cameras 1972, surround cameras 1974, long range cameras 1998, mid range cameras 1976, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire periphery of the vehicle 1900. In at least one embodiment, the type of camera used depends on the vehicle 1900. In at least one embodiment, any combination of camera types may be used to provide the necessary coverage around the vehicle 1900. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, the vehicle 1900 may include six cameras, seven cameras, ten cameras, twelve cameras, or another number of cameras. The cameras may, by way of example but not limitation, support Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet. In at least one embodiment, the present disclosure previously referred to herein may provide a plurality of cameras. Figure 19A and Figure 19B Each camera is described in more detail.
[0250] In at least one embodiment, vehicle 1900 may further include a vibration sensor 1942. In at least one embodiment, vibration sensor 1942 may measure vibration of a component (e.g., an axle) of vehicle 1900. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 1942 are used, the difference between the vibrations may be used to determine friction or slippage in the road surface (e.g., when there is a vibration difference between a powered drive shaft and a freely rotating shaft).
[0251] In at least one embodiment, the vehicle 1900 may include an ADAS system 1938. In some examples, the ADAS system 1938 may include, but is not limited to, an SoC. In at least one embodiment, the ADAS system 1938 may include, but is not limited to, any number of autonomous / adaptive / automatic cruise control ("ACC") systems, cooperative adaptive cruise control ("CACC") systems, forward collision warning ("FCW") systems, automatic emergency braking ("AEB") systems, lane departure warning ("LDW") systems, lane keeping assist ("LKA") systems, blind spot alert ("BSW") systems, rear cross traffic alert ("RCTW") systems, collision warning ("CW") systems, lane centering ("LC") systems, and / or other systems, features, and / or functions, and combinations thereof.
[0252] In at least one embodiment, the ACC system may utilize RADAR sensors 1960, LIDAR sensors 1964, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to vehicles immediately adjacent to vehicle 1900 and automatically adjusts the speed of vehicle 1900 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance keeping and recommends that vehicle 1900 change lanes when necessary. In at least one embodiment, lateral ACC is associated with other ADAS applications, such as LC and CW.
[0253] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from the other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via the network interface 1924 and / or wireless antenna 1926. In at least one embodiment, the direct link may be provided by a vehicle-to-vehicle ("V2V") communication link, while the indirect link may be provided by an infrastructure-to-vehicle ("I2V") communication link. Typically, the V2V communication concept provides information about the vehicle immediately ahead (e.g., the vehicle immediately ahead of vehicle 1900 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. In at least one embodiment, the CACC system may include one or both of the I2V and V2V information sources. In at least one embodiment, given information about vehicles ahead of vehicle 1900, the CACC system may be more reliable and have the potential to improve the smoothness of traffic flow and reduce road congestion.
[0254] In at least one embodiment, the FCW system is designed to warn the driver of hazards so that the driver can take corrective action. In at least one embodiment, the FCW system uses a forward-facing camera and / or RADAR sensor 1960, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system can provide warnings, such as in the form of audible, visual warnings, vibrations, and / or rapid brake pulses.
[0255] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. In at least one embodiment, the AEB system can utilize a forward-facing camera and / or RADAR sensor 1960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first warns the driver to take corrective action to avoid the collision, and, if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. In at least one embodiment, the AEB system can include technologies such as dynamic brake support and / or braking on impending collisions.
[0256] In at least one embodiment, the LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when vehicle 1900 crosses a lane marking. In at least one embodiment, the LDW system is inactive when the driver indicates an intentional lane departure by activating a turn signal. In at least one embodiment, the LDW system may utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to components such as a display, speaker, and / or vibration. In at least one embodiment, a LKA system is a variation of the LDW system. If vehicle 1900 begins to leave its lane, the LKA system provides steering input or braking to correct vehicle 1900.
[0257] In at least one embodiment, the BSW system detects and warns the driver of vehicles in the car's blind spot. In at least one embodiment, the BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system can provide additional warnings when the driver uses a turn signal. In at least one embodiment, the BSW system can use a rear-facing camera and / or RADAR sensor 1960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0258] In at least one embodiment, the RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 1900 is in reverse. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system can use one or more rear-facing RADAR sensors 1960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.
[0259] In at least one embodiment, conventional ADAS systems can be prone to generating false positive results, which can be annoying and distracting to the driver, but are generally not catastrophic because conventional ADAS systems alert the driver and allow the driver to decide whether a safe condition truly exists and act accordingly. In at least one embodiment, in the event of conflicting results, the vehicle 1900 itself decides whether to follow the results of the primary or secondary computer (e.g., the first controller 1936 or the second controller 1936). For example, in at least one embodiment, the ADAS system 1938 can be a backup and / or secondary computer that provides perception information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor can run redundant software on hardware components to detect failures in perception and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1938 can be provided to a supervisory MCU. In at least one embodiment, if the outputs from the primary and secondary computers conflict, the supervisory MCU determines how to reconcile the conflict to ensure safe operation.
[0260] In at least one embodiment, the primary computer can be configured to provide a confidence score to the supervisory MCU to indicate the primary computer's confidence in the selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the primary computer's instructions regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, if the confidence score does not meet the threshold, and if the primary computer and the secondary computer indicate different results (e.g., a conflict), the supervisory MCU can arbitrate between the computers to determine the appropriate result.
[0261] In at least one embodiment, the supervisory MCU can be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides a false alarm based, at least in part, on outputs from the primary and secondary computers. In at least one embodiment, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system identifies a metal object that is not actually a danger, such as a drain grate or manhole cover, that would trigger an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to override the LDW when a cyclist or pedestrian is present and lane departure is actually the safest action. In at least one embodiment, the supervisory MCU can include at least one of a DLA or a GPU suitable for running a neural network with associated memory. In at least one embodiment, the supervisory MCU can include and / or be included as a component of SoC 1904.
[0262] In at least one embodiment, the ADAS system 1938 may include an auxiliary computer that uses traditional computer vision rules to perform ADAS functions. In at least one embodiment, the auxiliary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU may improve reliability, safety, and performance. For example, in at least one embodiment, the diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially with respect to failures caused by software (or software-hardware interface) functions. For example, in at least one embodiment, if there is a software vulnerability or bug in the software running on the main computer, and a different software code running on the auxiliary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that the vulnerability in the software or hardware on the main computer did not cause a significant error.
[0263] In at least one embodiment, the output of the ADAS system 1938 can be input into the primary computer's perception module and / or the primary computer's dynamic driving task module. For example, in at least one embodiment, if the ADAS system 1938 indicates a forward collision warning due to an object directly ahead, the perception module can use this information when identifying the object. In at least one embodiment, as described herein, the secondary computer can have its own neural network that has been trained to reduce the risk of false positives.
[0264] In at least one embodiment, the vehicle 1900 may further include an infotainment SoC 1930 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, in at least one embodiment, the infotainment system 1930 may not be a SoC and may include, but is not limited to, two or more discrete components. In at least one embodiment, the infotainment SoC 1930 may include, but is not limited to, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., a navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / closed, air filter information, etc.) to the vehicle 1900. For example, the infotainment SoC 1930 may include a radio, a disk player, a navigation system, a video player, USB and Bluetooth connectivity, a car, an in-vehicle entertainment system, WiFi, steering wheel audio controls, hands-free voice control, a heads-up display ("HUD"), an HMI display 1934, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 1930 may be further configured to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from an ADAS system 1938, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0265] In at least one embodiment, the infotainment SoC 1930 can include any number and type of GPU functionality. In at least one embodiment, the infotainment SoC 1930 can communicate with other devices, systems, and / or components of the vehicle 1900 via a bus 1902 (e.g., a CAN bus, Ethernet, etc.). In at least one embodiment, the infotainment SoC 1930 can be coupled to a supervisory MCU so that the infotainment system's GPU can perform some autonomous driving functions in the event of a failure of the main controller 1936 (e.g., the vehicle's 1900 main computer and / or backup computer). In at least one embodiment, the infotainment SoC 1930 can cause the vehicle 1900 to enter a driver-to-safety stop mode, as described herein.
[0266] In at least one embodiment, the vehicle 1900 may further include an instrument panel 1932 (e.g., a digital instrument panel, an electronic instrument panel, a digital instrument panel, etc.). In at least one embodiment, the instrument panel 1932 may include, but is not limited to, a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). In at least one embodiment, the instrument panel 1932 may include, but is not limited to, any number and combination of a set of instruments, such as a speedometer, fuel level, oil pressure, a tachometer, an odometer, a turn indicator, a gear position indicator, a seat belt warning light, a parking brake warning light, an engine check light, information about a supplemental restraint system (e.g., an airbag), lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1930 and the instrument panel 1932. In at least one embodiment, the instrument panel 1932 may be included as part of the infotainment SoC 1930, or vice versa.
[0267] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 19C to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0268] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0269] Figure 19D In accordance with at least one embodiment, a cloud-based server and Figure 19A19. A diagram of a system 1976 for communicating between autonomous vehicles 1900. In at least one embodiment, system 1976 may include, but is not limited to, a server 1978, a network 1990, and any number and type of vehicles, including vehicle 1900. In at least one embodiment, server 1978 may include, but is not limited to, multiple GPUs 1984(A)-1984(H) (collectively referred to herein as GPUs 1984), PCIe switches 1982(A)-1982(D) (collectively referred to herein as PCIe switches 1982), and / or CPUs 1980(A)-1980(B) (collectively referred to herein as CPUs 1980). GPUs 1984, CPUs 1980, and PCIe switches 1982 may be interconnected with high-speed connections, such as, but not limited to, NVLink interface 1988 developed by NVIDIA and / or PCIe connections 1986. In at least one embodiment, the GPUs 1984 are connected via NVLink and / or NVSwitch SoC, and the GPUs 1984 and PCIe switches 1982 are connected via a PCIe interconnect. In at least one embodiment, although eight GPUs 1984, two CPUs 1980, and four PCIe switches 1982 are shown, this is not intended to be limiting. In at least one embodiment, each of the servers 1978 may include, but is not limited to, any number of GPUs 1984, CPUs 1980, and / or PCIe switches 1982 in any combination. For example, in at least one embodiment, the servers 1978 may each include eight, sixteen, thirty-two, and / or more GPUs 1984.
[0270] In at least one embodiment, server 1978 can receive image data representing an image from a vehicle via network 1990 that depicts unexpected or altered road conditions, such as recently begun road construction. In at least one embodiment, server 1978 can transmit neural network 1992, updated neural network 1992, and / or map information 1994, including, but not limited to, information regarding traffic and road conditions, to the vehicle via network 1990. In at least one embodiment, updates to map information 1994 can include, but not limited to, updates to HD map 1922, such as information regarding construction sites, potholes, access roads, flooding, and / or other obstacles. In at least one embodiment, neural network 1992, updated neural network 1992, and / or map information 1994 can be generated by new training and / or experience represented by data received from any number of vehicles in the environment, and / or based at least on training performed at a data center (e.g., using server 1978 and / or other servers).
[0271] In at least one embodiment, server 1978 can be used to train a machine learning model (e.g., a neural network) based at least in part on the training data. In at least one embodiment, the training data can be generated by the vehicle and / or can be generated in simulation (e.g., using a game engine). In at least one embodiment, any amount of the training data is labeled (e.g., where the associated neural network benefits from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, no amount of the training data is labeled and / or pre-processed (e.g., where the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1990, and / or the machine learning model can be used by server 1978 to remotely monitor the vehicle.
[0272] In at least one embodiment, server 1978 can receive data from vehicles and apply the data to the latest real-time neural networks for real-time intelligent reasoning. In at least one embodiment, server 1978 can include a deep learning supercomputer powered by GPU 1984 and / or a dedicated AI computer, such as the DGX and DGXStation machines developed by NVIDIA. However, in at least one embodiment, server 1978 can include the deep learning infrastructure of a data center using CPU power.
[0273] In at least one embodiment, the deep learning infrastructure of server 1978 may be capable of fast, real-time inference and may use this capability to assess and verify the health of the processor, software, and / or related hardware in vehicle 1900. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1900, such as an image sequence and / or objects located within that image sequence by vehicle 1900 (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to those identified by vehicle 1900, and if the results do not match and the deep learning infrastructure concludes that the AI in vehicle 1900 is malfunctioning, server 1978 may send a signal to vehicle 1900 instructing the vehicle's 1900 fail-safe computer to take control, notify passengers, and complete a safe parking maneuver.
[0274] In at least one embodiment, the server 1978 may include a GPU 1984 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). In at least one embodiment, the combination of GPU-driven servers and inference acceleration can enable real-time responses. In at least one embodiment, for example, in situations where performance is less critical, servers driven by CPUs, FPGAs, and other processors can be used for inference. In at least one embodiment, inference and / or training logic 1715 is used to execute one or more embodiments. Figure 17A and / or 17B provide details regarding the inference and / or training logic 1715 .
[0275] Computer system
[0276] Figure 20 2 is a block diagram illustrating an exemplary computer system according to at least one embodiment, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof, formed into a processor 2000, which may include execution units to execute instructions. In at least one embodiment, in accordance with the present disclosure, such as the embodiments described herein, the computer system 2000 may include, but is not limited to, components such as a processor 2002, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, the computer system 2000 may include a processor such as the Intel Corporation of Santa Clara, California. Processor family, Xeon TM 、 XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 2000 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0277] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include microcontrollers, digital signal processors ("DSPs"), system-on-chips, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system that can execute one or more instructions according to at least one embodiment.
[0278] In at least one embodiment, the computer system 2000 may include, but is not limited to, a processor 2002, which may include, but is not limited to, one or more execution units 2008 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 2000 is a single-processor desktop or server system, but in another embodiment, the computer system 2000 may be a multi-processor system. In at least one embodiment, the processor 2002 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 2002 may be coupled to a processor bus 2010, which may transmit data signals between the processor 2002 and other components in the computer system 2000.
[0279] In at least one embodiment, processor 2002 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 2004. In at least one embodiment, processor 2002 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 2002. Other embodiments may include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 2006 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.
[0280] In at least one embodiment, an execution unit 2008, including but not limited to logic for performing integer and floating-point operations, is also located in the processor 2002. In at least one embodiment, the processor 2002 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 2008 may include logic for processing a packed instruction set 2009. In at least one embodiment, by including the packed instruction set 2009 in the instruction set of the general-purpose processor 2002, along with associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 2002. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations one data element at a time.
[0281] In at least one embodiment, execution unit 2008 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 2000 may include, but is not limited to, memory 2020. In at least one embodiment, memory 2020 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 2020 may store instructions 2019 and / or data 2021 represented by data signals that may be executed by processor 2002.
[0282] In at least one embodiment, the system logic chip can be coupled to the processor bus 2010 and the memory 2020. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 2016, and the processor 2002 can communicate with the MCH 2016 via the processor bus 2010. In at least one embodiment, the MCH 2016 can provide a high-bandwidth memory path 2018 to the memory 2020 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 2016 can initiate data signals between the processor 2002, the memory 2020, and other components in the computer system 2000, and bridge data signals between the processor bus 2010, the memory 2020, and the system I / O 2022. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 2016 may be coupled to the memory 2020 via a high-bandwidth memory path 2018 , and the graphics / video card 2012 may be coupled to the MCH 2016 via an Accelerated Graphics Port (“AGP”) interconnect 2014 .
[0283] In at least one embodiment, the computer system 2000 can use the system I / O 2022 as a proprietary hub interface bus to couple the MCH 2016 to the I / O controller hub ("ICH") 2030. In at least one embodiment, the ICH 2030 can provide direct connections to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 2020, the chipset, and the processor 2002. Examples can include, but are not limited to, an audio controller 2029, a firmware hub ("Flash BIOS") 2028, a wireless transceiver 2026, data storage 2024, a traditional I / O controller 2023 including user input and a keyboard interface 2025, a serial expansion port 2027 (e.g., a universal serial bus (USB)), and a network controller 2034. The data storage 2024 can include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0284] In at least one embodiment, Figure 20 The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Figure 20An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, the devices may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 2000 are interconnected using a Compute Express Link (CXL) interconnect.
[0285] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 20 for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0286] In at least one embodiment, such components may be used to determine the position of an object relative to a vehicle.
[0287] Figure 21 2 is a block diagram illustrating an electronic device 2100 for utilizing a processor 2110 in accordance with at least one embodiment. In at least one embodiment, the electronic device 2100 may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0288] In at least one embodiment, the system 2100 may include, but is not limited to, a processor 2110 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, the processor 2110 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Figure 21 shows a system comprising interconnected hardware devices or "chips", while in other embodiments, Figure 21 An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 21 The devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 21 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.
[0289] In at least one embodiment, Figure 21 It may include a display 2124, a touch screen 2125, a touchpad 2130, a near field communication unit ("NFC") 2145, a sensor hub 2140, a thermal sensor 2146, a fast chipset ("EC") 2135, a trusted platform module ("TPM") 2138, a BIOS / firmware / flash memory ("BIOS, FW Flash") 2122, a DSP 2160, a drive 2120 (e.g., a solid state disk ("SSD") or a hard disk drive ("HDD")), a wireless local area network unit ("WLAN") 2150, a Bluetooth unit 2152, a wireless wide area network unit ("WWAN") 2156, a global positioning system (GPS) 2155, a camera ("USB 3.0 camera") 2154 (e.g., a USB 3.0 camera) and / or a low power double data rate ("LPDDR") memory unit ("LPDDR3") 2115 implemented in, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0290] In at least one embodiment, other components may be communicatively coupled to the processor 2110 via the components discussed above. In at least one embodiment, an accelerometer 2141, an ambient light sensor (“ALS”) 2142, a compass 2143, and a gyroscope 2144 may be communicatively coupled to the sensor hub 2140. In at least one embodiment, a thermal sensor 2139, a fan 2137, a keyboard 2146, and a touchpad 2130 may be communicatively coupled to the EC 2135. In at least one embodiment, a speaker 2163, an earpiece 2164, and a microphone (“mic”) 2165 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 2162, which in turn may be communicatively coupled to the DSP 2160. In at least one embodiment, the audio unit 2164 may include, for example, but not limited to, an audio codec / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a SIM card (“SIM”) 2157 may be communicatively coupled to the WWAN unit 2156. In at least one embodiment, components such as the WLAN unit 2150 and the Bluetooth unit 2152 and the WWAN unit 2156 may be implemented as a next generation form factor (NGFF).
[0291] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 21for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0292] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0293] Figure 22 A computer system 2200 is shown in accordance with at least one embodiment. In at least one embodiment, the computer system 2200 is configured to implement the various processes and methods described throughout this disclosure.
[0294] In at least one embodiment, the computer system 2200 includes, but is not limited to, at least one central processing unit ("CPU") 2202 connected to a communication bus 2210 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 2200 includes, but is not limited to, a main memory 2204 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in the main memory 2204 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 2222 provides an interface to other computing devices and networks for receiving data from the computer system 2200 and transmitting data to other systems.
[0295] In at least one embodiment, computer system 2200 includes, but is not limited to, input device 2208, parallel processing system 2212, and display device 2206, which can be implemented using conventional cathode ray tubes ("CRTs"), liquid crystal displays ("LCDs"), light emitting diodes ("LEDs"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 2208 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the aforementioned modules can be located on a single semiconductor platform to form a processing system.
[0296] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 22to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0297] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0298] Figure 23 A computer system 2300 is shown according to at least one embodiment. In at least one embodiment, computer system 2300 includes, but is not limited to, a computer 2310 and a USB drive 2320. In at least one embodiment, computer 2310 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, computer 2310 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.
[0299] In at least one embodiment, the USB disk 2320 includes, but is not limited to, a processing unit 2330, a USB interface 2340, and USB interface logic 2350. In at least one embodiment, the processing unit 2330 can be any instruction execution system, device, or device capable of executing instructions. In at least one embodiment, the processing unit 2330 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 2330 includes an application-specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 2330 is a tensor processing unit ("TPC") that is optimized to perform machine learning reasoning operations. In at least one embodiment, the processing core 2330 is a vision processing unit ("VPU") that is optimized to perform machine vision and machine learning reasoning operations.
[0300] In at least one embodiment, USB interface 2340 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 2340 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 2340 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 2350 can include any number and type of logic that enables processing unit 2330 to connect to a device (e.g., computer 2310) via USB connector 2340.
[0301] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A17B provides details about the reasoning and / or training logic 1715. In at least one embodiment, the reasoning and / or training logic 1715 may be configured to perform the following operations: Figure 23 In some embodiments, the present invention provides a method for performing inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0302] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0303] Figure 24A An exemplary architecture is shown in which multiple GPUs 2410-2413 are communicatively coupled to multiple multi-core processors 2405-2406 via high-speed links 2440-2443 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 2440-2443 support 4 GB / s, 30 GB / s, 80 GB / s, or higher communication throughput. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.
[0304] Additionally, in one embodiment, two or more of the GPUs 2410-2413 are interconnected via high-speed links 2429-2430, which may be implemented using the same or different protocols / links as used for high-speed links 2440-2443. Similarly, two or more multi-core processors 2405-2406 may be connected via high-speed link 2428, which may be a symmetric multiprocessor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, Figure 24A All communications between the various system components shown in FIG. 5 can be accomplished using the same protocols / links (eg, through a common interconnect fabric).
[0305] In one embodiment, each multi-core processor 2405-2406 is communicatively coupled to processor memory 2401-2402 via memory interconnects 2426-2427, respectively, and each GPU 2410-2413 is communicatively coupled to GPU memory 2420-2423 via GPU memory interconnects 2450-2453, respectively. Memory interconnects 2426-2427 and 2450-2453 can utilize the same or different memory access technologies. By way of example and not limitation, processor memory 2401-2402 and GPU memory 2420-2423 can be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portion of the processor memory 2401-2402 may be volatile memory, while another portion may be non-volatile memory (eg, using a two-level memory (2LM) hierarchy).
[0306] As described below, although the various processors 2405-2406 and GPUs 2410-2413 may each be physically coupled to a specific memory 2401-2402, 2420-2423, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as an "effective address" space) is distributed among the various physical memories. For example, the processor memories 2401-2402 may each include 64GB of system memory address space, while the GPU memories 2420-2423 may each include 32GB of system memory address space (for a total of 256GB of addressable memory in this example).
[0307] Figure 24B 2407 and a graphics acceleration module 2446. The graphics acceleration module 2446 may include one or more GPU chips integrated on a line card coupled to the processor 2407 via a high-speed link 2440. Alternatively, the graphics acceleration module 2446 may be integrated with the processor 2407 on the same package or chip.
[0308] In at least one embodiment, the processor 2407 shown includes multiple cores 2460A-2460D, each core having a translation lookaside buffer 2461A-2461D and one or more caches 2462A-2462D. In at least one embodiment, the cores 2460A-2460D may include various other components for executing instructions not shown and processing data. The caches 2462A-2462D may include level 1 (L1) and level 2 (L2) caches. In addition, one or more shared caches 2456 may be included in the caches 2462A-2462D and shared by a group of cores 2460A-2460D. For example, one embodiment of the processor 2407 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 2407 and the graphics acceleration module 2446 are connected to the system memory 2414, which may include Figure 24A Processor memory 2401-2402.
[0309] Coherence is maintained for data and instructions stored in the various caches 2462A-2462D, 2456, and system memory 2414 via inter-core communication on the coherent bus 2464. For example, each cache may have cache coherence logic / circuitry associated therewith to communicate over the coherent bus 2464 in response to a detected read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented on the coherent bus 2464 to snoop cache accesses.
[0310] In one embodiment, the proxy circuit 2425 communicatively couples the graphics acceleration module 2446 to the coherent bus 2464, thereby allowing the graphics acceleration module 2446 to participate in a cache coherence protocol as a peer of the cores 2460A-2460D. Specifically, the interface 2435 provides a connection to the proxy circuit 2425 via a high-speed link 2440 (e.g., a PCIe bus, NVLink, etc.), and the interface 2437 connects the graphics acceleration module 2446 to the link 2440.
[0311] In one implementation, the accelerator integrated circuit 2436 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 2431, 2432, N of the graphics acceleration module 2446. The graphics processing engines 2431, 2432, N can each include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 2431, 2432, N can include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 2446 can be a GPU having multiple graphics processing engines 2431-2432, N, or the graphics processing engines 2431-2432 can be individual GPUs integrated into a common package, line card, or chip.
[0312] In one embodiment, the accelerator integrated circuit 2436 includes a memory management unit (MMU) 2439 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation) and memory access protocols for accessing system memory 2414. The MMU 2439 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective addresses into physical / real address translations. In one implementation, a cache 2438 stores commands and data for efficient access by the graphics processing engines 2431-2432, N. In one embodiment, data stored in the cache 2438 and graphics memory 2433-2434, M is kept consistent with the core caches 2462A-2462D, 2456 and system memory 2414. As described above, this can be accomplished via proxy circuitry 2425 acting on behalf of cache 2438 and memory 2433-2434, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 2462A-2462D, 2456 to cache 2438 and receiving updates from cache 2438).
[0313] A set of registers 2445 stores context data for threads executed by graphics processing engines 2431-2432, N, and context management circuitry 2448 manages thread contexts. For example, context management circuitry 2448 can perform save and restore operations to save and restore the contexts of various threads during context switching (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). For example, upon a context switch, context management circuitry 2448 can store current register values to a designated area in memory (e.g., identified by a context pointer). Then, upon returning to a context, it can restore the register values. In one embodiment, interrupt management circuitry 2447 receives and processes interrupts received from system devices.
[0314] In one embodiment, the MMU 2439 converts virtual / effective addresses from the graphics processing engine 2431 into real / physical addresses in the system memory 2414. One embodiment of the accelerator integrated circuit 2436 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 2446 and / or other accelerator devices. The graphics accelerator module 2446 can be dedicated to a single application executing on the processor 2407, or can be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is proposed in which the resources of the graphics processing engines 2431-2432, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into "slices" based on the processing requirements and priorities associated with the virtual machines and / or applications and allocated to different virtual machines and / or applications.
[0315] In at least one embodiment, the accelerator integrated circuit 2436 acts as a bridge to the system for the graphics acceleration module 2446 and provides address translation and system memory cache services. In addition, the accelerator integrated circuit 2436 can provide virtualization facilities for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 2431-2432, N.
[0316] Because the hardware resources of the graphics processing engines 2431-2432, N are explicitly mapped to the actual address space seen by the host processor 2407, any host processor can directly address these resources using effective address values. In one embodiment, one function of the accelerator integrated circuit 2436 is to physically separate the graphics processing engines 2431-2432, N so that they appear as independent units to the system.
[0317] In at least one embodiment, one or more graphics memories 2433-2434, M are respectively coupled to each of the graphics processing engines 2431-2432, N. The graphics memories 2433-2434, M store instructions and data processed by each of the graphics processing engines 2431-2432, N. The graphics memories 2433-2434, M may be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memory, such as 3D XPoint or Nano-Ram.
[0318] In one embodiment, to reduce data traffic on link 2440, a biasing technique is used to ensure that the data stored in graphics memory 2433-2434, M will be the data most frequently used by graphics processing engines 2431-2432, N, and preferably not used (at least not frequently) by cores 2460A-2460D. Similarly, the biasing mechanism attempts to keep data needed by cores (preferably not graphics processing engines 2431-2432, N) in caches 2462A-2462D, 2456 of the core and system memory 2414.
[0319] Figure 24C Another exemplary embodiment is shown in which an accelerator integrated circuit 2436 is integrated within the processor 2407. In at least this embodiment, the graphics processing engines 2431-2432, N communicate directly to the accelerator integrated circuit 2436 via interfaces 2437 and 2435 over a high-speed link 2440 (where again any form of bus or interface protocol may be used). The accelerator integrated circuit 2436 may perform operations related to Figure 24B The same operations described above may be performed with higher throughput given their close proximity to the coherent bus 2464 and caches 2462A-2462D, 2456. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integrated circuit 2436 and a programming model controlled by the graphics acceleration module 2446.
[0320] In at least one embodiment, the graphics processing engines 2431-2432, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to the graphics processing engines 2431-2432, N, thereby providing virtualization within a VM / partition.
[0321] In at least one embodiment, graphics processing engines 2431-2432, N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a hypervisor to virtualize graphics processing engines 2431-2432, N to allow access by each operating system. For a single-partition system without a hypervisor, the operating system owns graphics processing engines 2431-2432, N. In at least one embodiment, the operating system can virtualize graphics processing engines 2431-2432, N to provide access to each process or application.
[0322] In at least one embodiment, the graphics acceleration module 2446 or the separate graphics processing engines 2431-2432, N use a process handle to select a process element. In at least one embodiment, the processing element is stored in the system memory 2414 and can be addressed using the effective address to real address conversion techniques described herein. In at least one embodiment, the process handle can be an implementation-specific value provided to the host process when registering its environment with the graphics processing engine 2431-2432, N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.
[0323] Figure 24D An exemplary accelerator integrated slice 2490 is shown. As used herein, a "slice" includes a specified portion of the processing resources of the accelerator integrated circuit 2436. The application effective address space 2482 within the system memory 2414 stores process elements 2483. In one embodiment, the process element 2483 is stored in response to a GPU call 2481 from an application 2480 executing on the processor 2407. The process element 2483 contains the processing state of the corresponding application 2480. The work descriptor (WD) 2484 contained in the processing element 2483 can be a single job requested by the application or may contain a pointer to a job queue. In at least one embodiment, the WD 2484 is a pointer to a job request queue in the application address space 2482.
[0324] The graphics acceleration module 2446 and / or the various graphics processing engines 2431-2432, N can be shared by all or part of the processes in the system. In at least one embodiment, an infrastructure for establishing a processing state and sending a WD 2484 to the graphics acceleration module 2446 to start a job in a virtualized environment can be included.
[0325] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns a graphics acceleration module 2446 or a single graphics processing engine 2431. Because the graphics acceleration module 2446 is owned by a single process, the hypervisor initializes the accelerator integrated circuit 2436 for the owning partition, and the operating system initializes the accelerator integrated circuit 2436 for the owning partition when the graphics acceleration module 2446 is allocated.
[0326] In operation, the WD fetch unit 2491 in the accelerator integrated slice 2490 fetches the next WD 2484, which includes an indication of work to be performed by one or more graphics processing engines of the graphics acceleration module 2446. Data from the WD 2484 can be stored in registers 2445 for use by the MMU 2439, the interrupt management circuit 2447, and / or the context management circuit 2448, as shown. For example, one embodiment of the MMU 2439 includes segment / page roaming circuitry for accessing the segment / page tables 2486 within the OS virtual address space 2485. The interrupt management circuit 2447 can process interrupt events 2492 received from the graphics acceleration module 2446. When executing graphics operations, the effective addresses 2493 generated by the graphics processing engines 2431-2432, N are converted into real addresses by the MMU 2439.
[0327] In one embodiment, the same register set 2445 is replicated for each graphics processing engine 2431-2432, N, and / or graphics acceleration module 2446 and can be initialized by the hypervisor or operating system. Each of these replicated registers can be included in the accelerator integration slice 2490. Table 1 shows exemplary registers that can be initialized by the hypervisor.
[0328] Table 1 - Registers initialized by the hypervisor
[0329] 1 Chip Control Register 2 Processing area pointer for real address (RA) plans 3 Authorization Mask Override Register 4 Interrupt vector table input offset 5 Interrupt vector table entry restriction 6 Status Register 7 Logical partition ID 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Description Register
[0330] Example registers that may be initialized by the operating system are shown in Table 2.
[0331] Table 2 - Operating System Initialization Registers
[0332] 1 Process and thread identification 2 Effective Address (EA) environment save / restore pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) stores the segment table pointer 5 Mask of Authority 6 Job Descriptor
[0333] In one embodiment, each WD 2484 is specific to a particular graphics acceleration module 2446 and / or graphics processing engine 2431-2432, N. It contains all the information needed for the graphics processing engine 2431-2432, N to do the work or work, or it can be a pointer to a memory location where the application has set up a command queue for the work to be done.
[0334] Figure 24E 24. Additional details of an exemplary embodiment of a sharing model are shown. This embodiment includes a hypervisor real address space 2498 in which a process element list 2499 is stored. The hypervisor real address space 2498 can be accessed by a hypervisor 2496, which virtualizes the graphics acceleration module engine for an operating system 2495.
[0335] In at least one embodiment, the shared programming model allows all or some processes in all or some partitions of the system to use the graphics acceleration module 2446. There are two programming models for sharing the graphics acceleration module 2446 by multiple processes and partitions: time slice sharing and graphics direction sharing.
[0336] In this model, the hypervisor 2496 owns the graphics acceleration module 2446, and its functionality is available to all operating systems 2495. For the graphics acceleration module 2446 to support virtualization within the hypervisor 2496, the graphics acceleration module 2446 must adhere to the following requirements: 1) Application job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 2446 must provide a context save and restore mechanism. 2) Application job requests must be guaranteed by the graphics acceleration module 2446 to complete within a specified timeframe, including any translation errors, or the graphics acceleration module 2446 must provide the ability to preempt job processing. 3) When operating within a directed sharing programming model, fairness must be ensured between processes.
[0337] In at least one embodiment, an application 2480 is required to make an operating system 2495 system call with a graphics acceleration module 2446 type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore region pointer (CSRP). In at least one embodiment, the graphics acceleration module 2446 type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module 2446 type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 2446 and can take the form of a graphics acceleration module 2446 command, an effective address pointer to a user-defined structure, an effective address pointer to a command queue, or any other data structure describing the work to be performed by the graphics acceleration module 2446. In one embodiment, the AMR value is the AMR state to be used for the current process. In at least one embodiment, the value passed to the operating system is similar to the application setting for the AMR. If the accelerator integrated circuit 2436 and graphics acceleration module 2446 implementation do not support the User Authority Mask Override Register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. The hypervisor 2496 can optionally apply the current authorization mask override register (AMOR) value before placing the AMR into the process element 2483. In at least one embodiment, the CSRP is one of the registers 2445 that contains the effective address of an area in the application's effective address space 2482 for the graphics acceleration module 2446 to save and restore context state. This pointer is optional if state does not need to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be fixed system memory.
[0338] After receiving the system call, the operating system 2495 can verify that the application 2480 has been registered and granted permission to use the graphics acceleration module 2446. The operating system 2495 then calls the hypervisor 2496 using the information shown in Table 3.
[0339] Table 3 - OS to Hypervisor call parameters
[0340] 1 Work Descriptor (WD) 2 Authorization Mask Register (AMR) value (may be masked) 3 Effective Address (EA) Context Save / Restore Area Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Utilization Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN)
[0341] After receiving the hypervisor call, the hypervisor 2496 verifies that the operating system 2495 has registered and been granted permission to use the graphics acceleration module 2446. The hypervisor 2496 then places the process element 2483 into a linked list of process elements of the corresponding type of graphics acceleration module 2446. The process element may contain the information shown in Table 4.
[0342] Table 4 - Process element information
[0343]
[0344]
[0345] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 2490 registers 2445 .
[0346] like Figure 24F As shown, in at least one embodiment, a unified memory is used that can be addressed by a common virtual memory address space used to access physical processor memories 2401-2402 and GPU memories 2420-2423. In this implementation, operations executed on GPUs 2410-2413 utilize the same virtual / effective memory address space to access processor memories 2401-2402, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 2401, a second portion is allocated to second processor memory 2402, a third portion is allocated to GPU memory 2420, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memories 2401-2402 and GPU memories 2420-2423, thereby allowing any processor or GPU to access any physical memory having a virtual address mapped to that memory.
[0347] In one embodiment, bias / coherence management circuitry 2494A-2494E within one or more MMUs 2439A-2439E ensures cache coherence between the caches of one or more host processors (e.g., 2405) and GPUs 2410-2413 and implements biasing techniques that indicate physical memory where certain types of data should be stored. Figure 24F , multiple instances of bias / coherence management circuits 2494A- 2494E are shown, which may be implemented within an MMU of one or more host processors 2405 and / or within an accelerator integrated circuit 2436 .
[0348] One embodiment allows GPU-attached memory 2420-2423 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, but without the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability to access GPU-attached memory 2420-2423 as system memory without heavy cache coherence overhead provides a beneficial operating environment for GPU offloading. This arrangement allows host processor 2405 software to set operands and access computation results without incurring the overhead of traditional I / O DMA data copies. Such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU-attached memory 2420-2423 without cache coherence overhead is critical to the execution time of offloaded computations. For example, in situations with heavy streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 2410-2413. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation may play a role in determining the efficiency of GPU offloading.
[0349] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granular structure (i.e., controlled at the granularity of a memory page), with each GPU-connected memory page comprising 1 or 2 bits. In at least one embodiment, the bias table can be implemented in the stolen memory range of one or more GPU-attached memories 2420-2423, with or without a bias cache in the GPUs 2410-2413 (e.g., to cache frequently / recently used bias table entries). Alternatively, the entire bias table can be maintained within the GPU.
[0350] In at least one embodiment, the bias table entry associated with each access to the GPU-attached memory 2420-2423 is accessed before the GPU memory is actually accessed, resulting in the following operations. First, local requests from GPUs 2410-2413 whose pages are found in the GPU bias are forwarded directly to the corresponding GPU memory 2420-2423. Local requests from the GPUs that find their pages in the host bias are forwarded to processor 2405 (e.g., via a high-speed link as described above). In one embodiment, requests from processor 2405 to find the requested page in the host processor bias complete a request similar to a normal memory read. Alternatively, requests for pages in the GPU bias can be forwarded to GPUs 2410-2413. In at least one embodiment, if the GPU is not currently using the page, the GPU can convert the page to the host processor bias. In at least one embodiment, the bias state of a page can be changed by a software-based mechanism, a hardware-assisted software-based mechanism, or in limited cases, a purely hardware-based mechanism.
[0351] One mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn calls the GPU's device driver, which in turn sends a message (or queues a command descriptor) to the GPU, directing it to change the bias state and, in some transitions, perform a cache flush operation in the host. In at least one embodiment, a cache flush operation is used for transitions from host processor 2405 bias to GPU bias, but not for the reverse transition.
[0352] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by host processor 2405. To access these pages, processor 2405 may request access from GPU 2410, which may or may not immediately grant access. Therefore, to reduce communication between processor 2405 and GPU 2410, it is advantageous to ensure that GPU-biased pages are pages required by the GPU, not by host processor 2405, and vice versa.
[0353] Reasoning and / or training logic 1715 is used to execute one or more embodiments. Figure 17A and / or 17B provide details regarding the inference and / or training logic 1715 .
[0354] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0355] Figure 25An exemplary integrated circuit and associated graphics processor according to various embodiments described herein are shown, which can be manufactured using one or more IP cores. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0356] Figure 25 25 is a block diagram illustrating an exemplary system on a chip integrated circuit 2500 that can be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 2500 includes one or more application processors 2505 (e.g., CPUs), at least one graphics processor 2510, and may additionally include an image processor 2515 and / or a video processor 2520, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 2500 includes peripheral or bus logic that includes a USB controller 2525, a UART controller 2530, an SPI / SDIO controller 2535, and an I / O controller. 2 S / I 2 C controller 2540. In at least one embodiment, the integrated circuit 2500 may include a display device 2545 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 2550 and a Mobile Industry Processor Interface (MIPI) display interface 2555. In at least one embodiment, memory may be provided by a flash memory subsystem 2560, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 2565 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 2570.
[0357] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 may be used within integrated circuit 2500 to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0358] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0359] Figures 26A-26BAn exemplary integrated circuit and associated graphics processor according to various embodiments described herein are shown, which can be manufactured using one or more IP cores. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0360] Figures 26A-26B is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. Figure 26A An exemplary graphics processor 2610 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 26B Another exemplary graphics processor 2640 of a system on a chip integrated circuit is shown, which can be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, Figure 26A The graphics processor 2610 is a low power graphics processor core. In at least one embodiment, Figure 26B The graphics processor 2640 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 2610, 2640 can be Figure 25 A variant of the graphics processor 2510.
[0361] In at least one embodiment, the graphics processor 2610 includes a vertex processor 2605 and one or more fragment processors 2615A-2615N (e.g., 2615A, 2615B, 2615C, 2615D through 2615N-1 and 2615N). In at least one embodiment, the graphics processor 2610 can execute different shader programs via separate logic, such that the vertex processor 2605 is optimized to perform operations for the vertex shader program, while one or more fragment processors 2615A-2615N perform fragment (e.g., pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 2605 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 2615A-2615N use the primitives and vertex data generated by the vertex processor 2605 to generate a frame buffer for display on a display device. In at least one embodiment, fragment processors 2615A-2615N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.
[0362] In at least one embodiment, graphics processor 2610 additionally includes one or more memory management units (MMUs) 2620A-2620B, caches 2625A-2625B, and circuit interconnects 2630A-2630B. In at least one embodiment, one or more MMUs 2620A-2620B provide virtual to physical address mapping for graphics processor 2610, including for vertex processor 2605 and / or fragment processors 2615A-2615N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 2625A-2625B. In at least one embodiment, one or more MMUs
[0363] 2620A-2620B can synchronize with other MMUs within the system, including one or more MMUs associated with one or more application processors 2505, image processor 2515, and / or video processor 2520 of the graphics 25, so that each processor 2505-2520 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 2630A-2630B enable the graphics processor 2610 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.
[0364] In at least one embodiment, graphics processor 2640 includes Figure 26A One or more MMUs 2620A-2620B, caches 2625A-2625B, and circuit interconnects 2630A-2630B of the graphics processor 2610. In at least one embodiment, the graphics processor 2640 includes one or more shader cores 2655A-2655N (e.g., 2655A, 2655B, 2655C, 2655D, 2655E, 2655F, up to 2655N-1 and 2655N), which provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 2640 includes an inter-core task manager 2645 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 2655A-2655N and a tiling unit 2658 to accelerate tile-based rendering operations in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within the scene or to optimize use of internal caches.
[0365] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provide details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 may be used within integrated circuits 26A and / or 26B to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions or architectures, or neural network use cases described herein.
[0366] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0367] Figures 27A-27B Additional exemplary graphics processor logic according to embodiments described herein is shown. In at least one embodiment, Figure 27A Shows that can be included in Figure 25 The graphics core 2700 within the graphics processor 2510 may, in at least one embodiment, be Figure 26B Unified shader cores 2655A-2655N. Figure 27B A highly parallel, general-purpose graphics processing unit 2730 suitable for deployment on a multi-chip module in at least one embodiment is shown.
[0368] In at least one embodiment, graphics core 2700 includes a shared instruction cache 2702, texture units 2718, and cache / shared memory 2720, which are common to execution resources within graphics core 2700. In at least one embodiment, graphics core 2700 may include multiple slices 2701A-2701N, or partitions of each core, and a graphics processor may include multiple instances of graphics core 2700. Slices 2701A-2701N may include support logic including local instruction caches 2704A-2704N, thread schedulers 2706A-2706N, thread dispatchers 2708A-2708N, and a set of registers 2710A-2710N. In at least one embodiment, slices 2701A-2701N may include a set of additional function units (AFUs 2712A-2712N), floating point units (FPUs 2714A-2714N), integer arithmetic logic units (ALUs 2716-2716N), address calculation units (ACUs 2713A-2713N), double-precision floating point units (DPFPUs 2715A-2715N), and matrix processing units (MPUs 2717A-2717N).
[0369] In at least one embodiment, the FPUs 2714A-2714N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 2715A-2715N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 2716A-2716N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 2717A-2717N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPUs 2717A-2717N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFUs 2712A-2712N can perform additional logical operations not supported by the floating-point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).
[0370] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provide details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 may be used in graphics core 2700 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0371] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0372] Figure 27BA general purpose processing unit (GPGPU) 2730 is shown in at least one embodiment, which can be configured to enable highly parallel computing operations to be performed by an array of graphics processing units. In at least one embodiment, GPGPU 2730 can be directly linked to other instances of GPGPU 2730 to create a multi-GPU cluster to increase the training speed for deep neural networks. In at least one embodiment, GPGPU 2730 includes a host interface 2732 to enable connection to a host processor. In at least one embodiment, host interface 2732 is a PCI Express interface. In at least one embodiment, host interface 2732 can be a vendor-specific communication interface or communication structure. In at least one embodiment, GPGPU 2730 receives commands from the host processor and dispatches the execution threads associated with those commands to a set of compute clusters 2736A-2736H using a global scheduler 2734. In at least one embodiment, compute clusters 2736A-2736H share a cache memory 2738. In at least one embodiment, cache memory 2738 may serve as a higher level cache for cache memories within compute clusters 2736A-2736H.
[0373] In at least one embodiment, GPGPU 2730 includes memory 2744A-2744B coupled to compute clusters 2736A-2736H via a set of memory controllers 2742A-2742B. In at least one embodiment, memory 2744A-2744B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.
[0374] In at least one embodiment, computing clusters 2736A-2736H each include a set of graphics cores, such as Figure 27A The graphics core 2700 may include multiple types of integer and floating-point logic units, including those for performing computational operations within a range of precision suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each compute cluster 2736A-2736H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units may be configured to perform 64-bit floating-point operations.
[0375] In at least one embodiment, multiple instances of GPGPU 2730 can be configured to operate as a compute cluster. In at least one embodiment, the communications used by compute clusters 2736A-2736H for synchronization and data exchange vary between embodiments. In at least one embodiment, multiple instances of GPGPU 2730 communicate via a host interface 2732. In at least one embodiment, GPGPU 2730 includes an I / O hub 2739 that couples GPGPU 2730 to a GPU link 2740, enabling direct connections to other instances of GPGPU 2730. In at least one embodiment, GPU link 2740 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 2730. In at least one embodiment, GPU link 2740 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 2730 reside in separate data processing systems and communicate via a network device accessible via host interface 2732. In at least one embodiment, GPU link 2740 may be configured to enable connection to a host processor, in addition to or in place of host interface 2732 .
[0376] In at least one embodiment, the GPGPU 2730 can be configured to train a neural network. In at least one embodiment, the GPGPU 2730 can be used within an inference platform. In at least one embodiment where the GPGPU 2730 is used for inference, the GPGPU can include fewer compute clusters 2736A-2736H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with the memories 2744A-2744B can differ between the inference and training configurations, with higher bandwidth memory technology being dedicated to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 2730 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during inference operations of a deployed neural network.
[0377] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 may be used in GPGPU 2730 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.
[0378] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0379] Figure 28 2 is a block diagram illustrating a computing system 2800 according to at least one embodiment. In at least one embodiment, computing system 2800 includes a processing subsystem 2801 having one or more processors 2802 and a system memory 2804 communicating via an interconnect path that may include a memory hub 2805. In at least one embodiment, memory hub 2805 may be a separate component within a chipset assembly or integrated within one or more processors 2802. In at least one embodiment, memory hub 2805 is coupled to an I / O subsystem 2811 via a communication link 2806. In one embodiment, I / O subsystem 2811 includes an I / O hub 2807, which enables computing system 2800 to receive input from one or more input devices 2808. In at least one embodiment, I / O hub 2807 may enable a display controller, included in one or more processors 2802, to provide output to one or more display devices 2810A. In at least one embodiment, the one or more display devices 2810A coupled to the I / O hub 2807 may include local, internal, or embedded display devices.
[0380] In at least one embodiment, the processing subsystem 2801 includes one or more parallel processors 2812 coupled to the memory hub 2805 via a bus or other communication link 2813. In at least one embodiment, the communication link 2813 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 2812 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 2812 form a graphics processing subsystem that can output pixels to one of one or more display devices 2810A coupled via the I / O hub 2807. In at least one embodiment, the one or more parallel processors 2812 can also include a display controller and display interface (not shown) to enable direct connection to the one or more display devices 2810B.
[0381] In at least one embodiment, a system storage unit 2814 can be connected to the I / O hub 2807 to provide a storage mechanism for the computing system 2800. In at least one embodiment, an I / O switch 2816 can be used to provide an interface mechanism to enable connections between the I / O hub 2807 and other components, such as a network adapter 2818 and / or a wireless network adapter 2819 that can be integrated into the platform, as well as various other devices that can be added via one or more add-on devices 2820. In at least one embodiment, the network adapter 2818 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2819 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radios.
[0382] In at least one embodiment, computing system 2800 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to I / O hub 2807. Figure 28 The communication paths interconnecting the various components in the system may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols (e.g., NV-Link high-speed interconnect or interconnect protocol).
[0383] In at least one embodiment, one or more parallel processors 2812 include circuits optimized for graphics and video processing (including, for example, video output circuitry) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2812 include circuits optimized for general-purpose processing. In at least one embodiment, the components of the computing system 2800 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2812, memory hub 2805, processor 2802, and I / O hub 2807 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of the computing system 2800 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 2800 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules into a modular computing system.
[0384] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A17B provides details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 can be used in system diagram 2800 to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0385] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0386] processor
[0387] Figure 29A 2900 in accordance with at least one embodiment. In at least one embodiment, the various components of the parallel processor 2900 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 2900 is shown as a processor according to an exemplary embodiment. Figure 28 A variation of the one or more parallel processors 2812 shown.
[0388] In at least one embodiment, parallel processor 2900 includes parallel processing unit 2902. In at least one embodiment, parallel processing unit 2902 includes an I / O unit 2904 that enables communication with other devices, including other instances of parallel processing unit 2902. In at least one embodiment, I / O unit 2904 can be directly connected to other devices. In at least one embodiment, I / O unit 2904 connects to other devices using a hub or switch interface (e.g., memory hub 2805). In at least one embodiment, the connection between memory hub 2805 and I / O unit 2904 forms a communication link 2813. In at least one embodiment, I / O unit 2904 is connected to a host interface 2906 and a memory crossbar switch 2916, where host interface 2906 receives commands for performing processing operations and memory crossbar switch 2916 receives commands for performing memory operations.
[0389] In at least one embodiment, when host interface 2906 receives command buffers via I / O unit 2904, host interface 2906 can direct work operations to execute those commands to front end 2908. In at least one embodiment, front end 2908 is coupled to scheduler 2910, which is configured to distribute commands or other work items to processing cluster array 2912. In at least one embodiment, scheduler 2910 ensures that processing cluster array 2912 is properly configured and in a valid state before distributing tasks to processing cluster array 2912. In at least one embodiment, scheduler 2910 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, a microcontroller-implemented scheduler 2910 can be configured to perform complex scheduling and work distribution operations at both coarse and fine granularity, thereby enabling rapid preemption and context switching of threads executing on processing array 2912. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 2912 through one of multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 2912 by scheduler 2910 logic within a microcontroller that includes scheduler 2910 .
[0390] In at least one embodiment, processing cluster array 2912 may include up to "N" processing clusters (e.g., cluster 2914A, cluster 2914B, through cluster 2914N). In at least one embodiment, each cluster 2914A-2914N of processing cluster array 2912 may execute a large number of concurrent threads. In at least one embodiment, scheduler 2910 may allocate work to clusters 2914A-2914N of processing cluster array 2912 using various scheduling and / or work distribution algorithms, which may vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling may be handled dynamically by scheduler 2910 or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by processing cluster array 2912. In at least one embodiment, different clusters 2914A-2914N of processing cluster array 2912 may be assigned to process different types of programs or to perform different types of computations.
[0391] In at least one embodiment, processing cluster array 2912 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 2912 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing cluster array 2912 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.
[0392] In at least one embodiment, processing cluster array 2912 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 2912 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 2912 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 2902 may transfer data from system memory via I / O units 2904 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 2922) during processing and then written back to system memory.
[0393] In at least one embodiment, when parallel processing unit 2902 is used to perform graphics processing, scheduler 2910 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 2914A-2914N of processing cluster array 2912. In at least one embodiment, portions of processing cluster array 2912 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of clusters 2914A-2914N can be stored in a buffer to allow the intermediate data to be transferred between clusters 2914A-2914N for further processing.
[0394] In at least one embodiment, the processing cluster array 2912 can receive processing tasks to be executed via the scheduler 2910, which receives commands defining the processing tasks from the front end 2908. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 2910 can be configured to obtain an index corresponding to a task, or can receive the index from the front end 2908. In at least one embodiment, the front end 2908 can be configured to ensure that the processing cluster array 2912 is configured in a valid state before starting a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.).
[0395] In at least one embodiment, each of one or more instances of parallel processing unit 2902 can be coupled to parallel processor memory 2922. In at least one embodiment, parallel processor memory 2922 can be accessed via memory crossbar 2916, which can receive memory requests from processing cluster array 2912 and I / O unit 2904. In at least one embodiment, memory crossbar 2916 can access parallel processor memory 2922 via memory interface 2918. In at least one embodiment, memory interface 2918 can include multiple partition units (e.g., partition unit 2920A, partition unit 2920B, through partition unit 2920N), each of which can be coupled to a portion of parallel processor memory 2922 (e.g., a memory unit). In at least one embodiment, the plurality of partition units 2920A-2920N are configured to be equal to the number of storage units, such that the first partition unit 2920A has a corresponding first storage unit 2924A, the second partition unit 2920B has a corresponding storage unit 2924B, and the Nth partition unit 2920N has a corresponding Nth storage unit 2924N. In at least one embodiment, the number of partition units 2920A-2920N may not be equal to the number of storage devices.
[0396] In at least one embodiment, memory units 2924A-2924N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2924A-2924N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across memory units 2924A-2924N, allowing partition units 2920A-2920N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 2922. In at least one embodiment, local instances of parallel processor memory 2922 may be eliminated in favor of a unified memory design utilizing system memory in combination with local cache memory.
[0397] In at least one embodiment, any of the clusters 2914A-2914N in the processing cluster array 2912 can process data to be written to any memory unit 2924A-2924N within the parallel processor memory 2922. In at least one embodiment, the memory crossbar 2916 can be configured to transmit the output of each cluster 2914A-2914N to any partition unit 2920A-2920N or another cluster 2914A-2914N, which can perform other processing operations on the output. In at least one embodiment, each cluster 2914A-2914N can communicate with a memory interface 2918 via the memory crossbar 2916 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 2916 has connections to memory interface 2918 for communicating with I / O unit 2904, and connections to local instances of parallel processor memory 2922, thereby enabling processing units within different processing clusters 2914A-2914N to communicate with system memory or other memory that is not local to parallel processing unit 2902. In at least one embodiment, memory crossbar switch 2916 can use virtual channels to separate traffic flows between clusters 2914A-2914N and partition units 2920A-2920N.
[0398] In at least one embodiment, multiple instances of parallel processing unit 2902 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 2902 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 2902 can include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 2902 or parallel processor 2900 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0399] Figure 29B is a block diagram of a partition unit 2920 according to at least one embodiment. In at least one embodiment, the partition unit 2920 is Figure 29A29 . In at least one embodiment, the partition unit 2920 includes an L2 cache 2921, a frame buffer interface 2925, and a raster operations unit ("ROP") 2926. The L2 cache 2921 is a read / write cache that is configured to perform load and store operations received from the memory crossbar 2916 and the ROP 2926. In at least one embodiment, the L2 cache 2921 outputs read misses and urgent writeback requests to the frame buffer interface 2925 for processing. In at least one embodiment, updates may also be sent to the frame buffer via the frame buffer interface 2925 for processing. In at least one embodiment, the frame buffer interface 2925 interacts with one of the memory units in the parallel processor memory, such as the memory units 2924A-2924N of FIG. 29 (e.g., within the parallel processor memory 2922).
[0400] In at least one embodiment, ROP 2926 is a processing unit that performs raster operations such as stenciling, z-testing, blending, etc. In at least one embodiment, ROP 2926 then outputs processed graphics data that is stored in graphics memory. In at least one embodiment, ROP 2926 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic implemented by ROP 2926 can vary based on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, incremental color compression is performed based on depth and color data on a per-tile basis.
[0401] In at least one embodiment, ROP 2926 is included within each processing cluster (e.g., Figure 29A In at least one embodiment, read and write requests for pixel data are transmitted through the memory crossbar 2916 rather than through the pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device such as a Figure 28 2802 for further processing, or by Figure 29A One of the processing entities within parallel processor 2900 is routed for further processing.
[0402] Figure 29C is a block diagram of a processing cluster 2914 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Figure 29AIn at least one embodiment, one or more of the one or more processing clusters 2914 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of generally synchronized threads, which uses a common instruction unit that is configured to issue instructions to a group of processing engines within each processing cluster.
[0403] In at least one embodiment, the operation of the processing cluster 2914 can be controlled by a pipeline manager 2932 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2932 Figure 29A The scheduler 2910 receives instructions and manages the execution of these instructions by the graphics multiprocessor 2934 and / or the texture unit 2936. In at least one embodiment, the graphics multiprocessor 2934 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2914. In at least one embodiment, one or more instances of the graphics multiprocessor 2934 may be included within the processing cluster 2914. In at least one embodiment, the graphics multiprocessor 2934 may process data, and the data crossbar 2940 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, the pipeline manager 2932 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed to the data crossbar 2940.
[0404] In at least one embodiment, each graphics multiprocessor 2934 within a processing cluster 2914 may include the same set of function execution logic (e.g., arithmetic logic units, load-store units, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions have completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.
[0405] In at least one embodiment, instructions transmitted to processing cluster 2914 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 2934. In at least one embodiment, a thread group can include fewer threads than the number of processing engines within graphics multiprocessor 2934. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the processing of a loop within the thread group. In at least one embodiment, a thread group can also include more threads than the number of processing engines within graphics multiprocessor 2934. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 2934, processing can be performed within consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 2934.
[0406] In at least one embodiment, the graphics multiprocessor 2934 includes an internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 2934 can abandon the internal cache and use cache memory within the processing cluster 2914 (e.g., L1 cache 2948). In at least one embodiment, each graphics multiprocessor 2934 can also access a partition unit (e.g., Figure 29A L2 cache within partition units 2920A-2920N) of the graphics multiprocessor 2934 is shared across all processing clusters 2914 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 2934 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 2902 can be used as global memory. In at least one embodiment, processing cluster 2914 includes multiple instances of graphics multiprocessor 2934, which can share common instructions and data, which can be stored in L1 cache 2948.
[0407] In at least one embodiment, each processing cluster 2914 may include a memory management unit ("MMU") 2945 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2945 may reside in Figure 29A2918. In at least one embodiment, the MMU 2945 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and optionally to cache memory lines. In at least one embodiment, the MMU 2945 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 2934 or L1 cache or processing cluster 2914. In at least one embodiment, the physical address is processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, a cache line index may be used to determine whether a request for a cache line is a hit or a miss.
[0408] In at least one embodiment, the processing clusters 2914 can be configured such that each graphics multiprocessor 2934 is coupled to a texture unit 2936 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2934, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 2934 outputs processed tasks to a data crossbar 2940 to provide the processed tasks to another processing cluster 2914 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 2916. In at least one embodiment, a preROP 2942 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 2934 and direct the data to a ROP unit, which can communicate with a partitioning unit (e.g., a partitioning unit) as described herein. Figure 29A In at least one embodiment, the PreROP 2942 unit can perform optimizations for color mixing, organize pixel color data, and perform address translation.
[0409] Reasoning and / or training logic 1715 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 17A 17B provides details regarding inference and / or training logic 1715. In at least one embodiment, inference and / or training logic 1715 can be used in graphics processing cluster 2914 to perform inference or prediction operations based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0410] In at least one embodiment, such a component may be used to determine the position of an object relative to a vehicle.
[0411] Figure 29D A graphics multiprocessor 2934 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 2934 is coupled to a pipeline manager 2932 of a processing cluster 2914. In at least one embodiment, the graphics multiprocessor 2934 has an execution pipeline that includes, but is not limited to, an instruction cache 2952, an instruction unit 2954, an address mapping unit 2956, a register file 2958, one or more general purpose graphics processing unit (GPGPU) cores 2962, and one or more load / store units 2966. The GPGPU cores 2962 and the load / store units 2966 are coupled to a cache memory 2972 and a shared memory 2970 via a memory and cache interconnect 2968.
[0412] In at least one embodiment, the instruction cache 2952 receives a stream of instructions to be executed from the pipeline manager 2932. In at least one embodiment, the instructions are cached in the instruction cache 2952 and dispatched for execution by the instruction unit 2954. In one embodiment, the instruction unit 2954 can dispatch instructions as thread groups (e.g., warps), assigning each thread group to a different execution unit within the GPGPU core 2962. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 2956 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by the load / store unit 2966.
[0413] In at least one embodiment, register file 2958 provides a set of registers for the functional units of graphics multiprocessor 2934. In at least one embodiment, register file 2958 provides temporary storage for operands for the data paths of the functional units (e.g., GPGPU core 2962, load / store unit 2966) connected to graphics multiprocessor 2934. In at least one embodiment, register file 2958 is divided between each functional unit such that a dedicated portion of register file 2958 is allocated to each functional unit. In at least one embodiment, register file 2958 is divided between the different warps being executed by graphics multiprocessor 2934.
[0414] In at least one embodiment, the GPGPU cores 2962 may each include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions for the graphics multiprocessor 2934. The GPGPU cores 2962 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 2962 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 2934 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed-function or special-function logic.
[0415] In at least one embodiment, the GPGPU core 2962 includes SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, the GPGPU core 2962 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.
[0416] In at least one embodiment, the memory and cache interconnect 2968 is an interconnect network that connects each functional unit of the graphics multiprocessor 2934 to the register file 2958 and the shared memory 2970. In at least one embodiment, the memory and cache interconnect 2968 is a crossbar interconnect that allows the load / store unit 2966 to perform load and store operations between the shared memory 2970 and the register file 2958. In at least one embodiment, the register file 2958 can operate at the same frequency as the GPGPU core 2962, resulting in very low latency for data transfers between the GPGPU core 2962 and the register file 2958. In at least one embodiment, the shared memory 2970 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 2934. In at least one embodiment, the cache memory 2972 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2936. In at least one embodiment, the shared memory 2970 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 2972, threads executing on GPGPU core 2962 may also programmatically store data in shared memory.
[0417] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate g...
Claims
1. A processor, comprising: one or more circuits to help determine the line of sight of an occupant of the vehicle independent of the angle at which the occupant is detected from one or more sensors in the vehicle, wherein the line of sight is determined based at least in part on one or more neural networks, The one or more neural networks are trained using position data known for a vehicle coordinate system and mapped to at least one virtual coordinate system corresponding to the one or more sensors. 2 . The processor of claim 1 , wherein the position data is obtained in part by identifying reference features located in the vehicle and represented in data detected during a training data collection process. 3 . The processor of claim 1 , wherein the vehicle coordinate system is mapped to the at least one virtual coordinate system using a calibration stand positioned at a fixed location of the vehicle coordinate system in the vehicle. 4 . The processor of claim 1 , wherein the one or more circuits further assist in determining the line of sight by determining an intersection of an occupant's line of sight vector with an area of the vehicle.
5. A vehicle comprising: one or more sensors; as well as one or more processors to facilitate determining a line of sight of an occupant of the vehicle independent of an occupant's perspective detected from the one or more sensors, wherein the line of sight is determined based at least in part on one or more neural networks, The one or more neural networks are trained using position data known for a vehicle coordinate system and mapped to at least one virtual coordinate system corresponding to the one or more sensors.
6. The vehicle of claim 5, wherein the position data is obtained in part by identifying fiducial features located in the vehicle and represented in data detected during a training data collection process. 7 . The vehicle of claim 5 , wherein the vehicle coordinate system is mapped to the at least one virtual coordinate system using a calibration stand positioned at a fixed location of the vehicle coordinate system in the vehicle.
8. The vehicle of claim 5, wherein the one or more processors are further configured to assist in determining the line of sight by determining an intersection of an occupant's line of sight vector with an area of the vehicle.
9. A sight line detection method, comprising: detecting one or more occupants of the vehicle via one or more sensors; as well as determining a line of sight of the one or more occupants independently of a position of the one or more sensors, wherein the line of sight is determined based at least in part on one or more neural networks, The one or more neural networks are trained using position data known for a vehicle coordinate system and mapped to at least one virtual coordinate system corresponding to the one or more sensors.
10. The method of claim 9, wherein the position data is obtained in part by identifying fiducial features represented in data detected during a training data collection process. 11 . The method of claim 9 , wherein the vehicle coordinate system is mapped to the at least one virtual coordinate system using a calibration stand positioned at a fixed location of the vehicle coordinate system in the vehicle. 12 . The method of claim 9 , wherein the line of sight is determined in part by identifying an intersection of an occupant's line of sight vector and an identification area of the vehicle.
13. A processor comprising: one or more circuits to facilitate training one or more neural networks to determine the gaze of one or more individuals independently of the angle at which the gaze of the one or more individuals is detected, The one or more neural networks are trained using position data known for a reference coordinate system and mapped to at least one virtual coordinate system corresponding to one or more sensors used to detect the one or more persons.
14. The processor of claim 13, wherein the position data is obtained in part by identifying fiducial features represented in data detected during a training data collection process.
15. The processor of claim 13, wherein the reference coordinate system is mapped to the at least one virtual coordinate system using a calibration stand located at a fixed position of the reference coordinate system in physical space.
16. The processor of claim 13, wherein the one or more circuits further assist in determining the line of sight by determining an intersection of a line of sight vector with an identification area.
17. A sight line detection system comprising: one or more processors to facilitate training one or more neural networks to determine the gaze of one or more individuals independent of the angle at which the gaze of the one or more individuals is detected, The one or more neural networks are trained using position data known for a reference coordinate system and mapped to at least one virtual coordinate system corresponding to one or more sensors used to detect the one or more persons.
18. The system of claim 17, wherein the position data is obtained in part by identifying fiducial features represented in data detected during a training data collection process.
19. The system of claim 17, wherein the reference coordinate system is mapped to the at least one virtual coordinate system using a calibration stand located at a fixed position of the reference coordinate system in physical space.
20. The system of claim 17, wherein the one or more processors further assist in determining the line of sight by determining an intersection of a line of sight vector and an identification area.
Citation Information
Patent Citations
A method and system for superimposing a human eye line of sight and a foreground image
CN109493305A
Facial information measuring system
JP2005267258A