System and method for generating images using dithered motion vectors
By using dithered motion vectors and a self-sharpening cubic interpolation filter in the image rendering channel, the problem of jagged edges on the display is solved, improving the resolution and realism of the image.
Patent Information
- Application Number
- CN202110997742.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-28
- Filing Date
- 2021-08-27
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-01-16
AI Technical Summary
Due to the limited pixel size of existing displays, images often exhibit jagged edges. Current anti-aliasing technologies struggle to effectively smooth out sharp boundaries, thus affecting the realism of the image.
By employing a dithering motion vector technique, the amplitude and direction of the motion vector are modified by adding dithering values to the image rendering channel. Combined with a self-sharpening cubic interpolation filter, a more realistic image is generated.
It effectively reduces jagged edges in images, improves image resolution and realism, and enhances image quality.
Smart Images

Figure CN114119381B_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to image generation. For example, at least one embodiment relates to image generation using jittered motion vectors. Background Technology
[0002] Current displays often produce a jagged effect, typically perceived by users as jagged edges in images. This is due to the limited size of display pixels, making it inherently unable to reproduce truly curved or non-linear surfaces. Anti-aliasing techniques have been developed to attempt to compensate for this limitation, thereby generating more realistic-looking images on modern displays. One such technique, for example, is temporal anti-aliasing, which samples subpixel values from one or more previous frames to determine the subpixel values for the current frame, attempting to effectively improve the resolution of the current image. This and other techniques can interpolate between selected subpixel values from previous frames based on motion vectors, introducing time-varying elements designed to effectively smooth sharp boundaries in the image. Ongoing efforts exist to improve these and other techniques involving the temporal accumulation of image values. Attached Figure Description
[0003] The above and other objects and advantages of this disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings, wherein like reference numerals always refer to like parts, wherein:
[0004] FIG. 1 The determination of the motion vector of jitter according to at least one embodiment is conceptually illustrated;
[0005] FIG. 2 This is a generalized embodiment of an illustrative processing system constructed for use according to at least one embodiment;
[0006] FIG. 3A The inference and / or training logic according to at least one embodiment is illustrated;
[0007] FIG. 3B The inference and / or training logic according to at least one embodiment is illustrated;
[0008] FIG. 4 The training and deployment of a neural network according to at least one embodiment are illustrated;
[0009] FIG. 5 An example data center system according to at least one embodiment is shown;
[0010] FIG. 6A An example of an autonomous vehicle according to at least one embodiment is shown;
[0011] FIG. 6B The illustration shows an embodiment according to at least one of the embodiments. FIG. 6Aan example of a camera position and field of view of an autonomous vehicle;
[0012] FIG. 6C is a block diagram illustrating an example system architecture of an autonomous vehicle in accordance with at least one embodiment; FIG. 6A
[0013] FIG. 6D is a diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle in accordance with at least one embodiment; FIG. 6A
[0014] FIG. 6B is a block diagram illustrating a computer system in accordance with at least one embodiment;
[0015] FIG. 6B is a block diagram illustrating a computer system in accordance with at least one embodiment;
[0016] FIG. 3A illustrates a computer system in accordance with at least one embodiment;
[0017] FIG. 3B illustrates a computer system in accordance with at least one embodiment;
[0018] FIG. 6B illustrates a computer system in accordance with at least one embodiment;
[0019] FIG. 6C illustrates a computer system in accordance with at least one embodiment;
[0020] FIG. 6A illustrates a computer system in accordance with at least one embodiment;
[0021] FIG. 6C illustrates a computer system in accordance with at least one embodiment;
[0022] FIG. 6A and FIG. 6C illustrates a shared programming model in accordance with at least one embodiment;
[0023] FIG. 6A illustrates an exemplary integrated circuit and associated graphics processor in accordance with at least one embodiment.
[0024] FIG. 6B illustrates an exemplary integrated circuit and associated graphics processor in accordance with at least one embodiment.
[0025] FIG. 3A and FIG. 3B illustrates an additional exemplary graphics processor logic in accordance with at least one embodiment;
[0026] FIG. 6C A computer system, in accordance with at least one embodiment, is shown;
[0027] FIG. 6D A parallel processor, in accordance with at least one embodiment, is shown;
[0028] FIG. 6A A partition unit, in accordance with at least one embodiment, is shown;
[0029] FIG. 3A A processing cluster, in accordance with at least one embodiment, is shown;
[0030] FIG. 3B A graphics multiprocessor, in accordance with at least one embodiment, is shown;
[0031] FIG. 7 A multi-GPU system, in accordance with at least one embodiment, is shown;
[0032] FIG. 7 A graphics processor, in accordance with at least one embodiment, is shown;
[0033] FIG. 7 A block diagram illustrating a processor micro-architecture for a processor, in accordance with at least one embodiment, is shown;
[0034] FIG. 7 A deep learning application processor, in accordance with at least one embodiment, is shown;
[0035] FIG. 3A A block diagram illustrating an example neuromorphic processor, in accordance with at least one embodiment, is shown;
[0036] FIG. 3B At least a portion of a graphics processor, in accordance with one or more embodiments, is shown;
[0037] FIG. 7 At least a portion of a graphics processor, in accordance with one or more embodiments, is shown;
[0038] FIG. 8 At least a portion of a graphics processor, in accordance with one or more embodiments, is shown;
[0039] FIG. 8 A block diagram illustrating a graphics processing engine of a graphics processor, in accordance with at least one embodiment, is shown;
[0040] FIG. 8 A block diagram illustrating at least a portion of a graphics processor core, in accordance with at least one embodiment, is shown;
[0041] FIG. 8 Thread execution logic, in accordance with at least one embodiment, is shown including an array of processing elements of a graphics processor core.
[0042] FIG. 8 A parallel processing unit (“PPU”) is shown in accordance with at least one embodiment;
[0043] FIG. 8 A general processing cluster (“GPC”) is shown in accordance with at least one embodiment;
[0044] FIG. 3A A memory partition unit of a parallel processing unit (“PPU”) is shown in accordance with at least one embodiment;
[0045] FIG. 3B A streaming multiprocessor is shown in accordance with at least one embodiment;
[0046] FIG. 8 An example dataflow graph of an advanced compute pipeline in accordance with at least one embodiment;
[0047] FIG. 9 A system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced compute pipeline in accordance with at least one embodiment;
[0048] FIG. 3A An example illustration of an advanced compute pipeline for processing imaging data in accordance with at least one embodiment;
[0049] FIG. 3B An example dataflow graph of a virtual instrument supporting an ultrasound device in accordance with at least one embodiment;
[0050] FIG. 9 An example dataflow graph of a virtual instrument supporting a CT scanner in accordance with at least one embodiment;
[0051] FIG. 10 A dataflow graph of a process for training a machine learning model is shown in accordance with at least one embodiment; and
[0052] FIG. 3A An example illustration of a client-server architecture to enhance annotation tools with pre-trained annotation models in accordance with at least one embodiment;
[0053] FIG. 3B A flow diagram showing processing steps for generating image values in accordance with at least one embodiment;
[0054] FIG. 10 And FIG. 11A The disadvantages of using motion vector generation without de-jittering of images are illustrated;
[0055] FIG. 11A And FIG. 11Billustrates the disadvantages of images generated using motion vectors without jitter and Catmull-ROM filtering; and
[0056] FIG. 11A and FIG. 11C illustrates the advantages of images generated using motion vectors with jitter determined according to at least one embodiment (Catmull-ROM filtering applied). DETAILED DESCRIPTION
[0057] In at least one embodiment, systems and methods relate to image generation using a temporal accumulation process that employs jittered motion vectors for any image portion, such as an intermediate pass image portion. As one example, image values for an intermediate pass (such as a shadow pass) can be determined according to jittered motion vectors. In at least one embodiment, the jitter values used in those shadow passes are the same as those used in conjunction with other passes (such as a default pass). To compensate for blurring that can be introduced using motion vectors in image rendering passes (such as intermediate image rendering passes), a filter can be applied, such as a sharpening cubic interpolation filter, one example being a Catmull-ROM filter.
[0058] FIG. 11B Determination of jittered motion vectors according to at least one embodiment is conceptually illustrated. In at least one embodiment, motion vectors can be employed in temporal accumulation of previous image values for determining image values for a current frame. In at least one embodiment, such accumulation methods can be employed in any image pass, such as a default image pass (e.g., a process for determining image values for objects in an image), and / or in an intermediate image pass (e.g., a process for modifying or generating additional visual effects, modifying the appearance of objects generated in a default image pass). In at least one embodiment, the jitter values applied to motion vectors for individual image rendering passes can be substantially the same as those employed in any aspect of other image rendering passes. In at least one embodiment, “jitter” refers to vectors that shift the magnitude and / or direction of motion vectors when added to the motion vectors. In at least one embodiment, jitter vectors can be generated in any manner, such as by randomly generating direction and magnitude values. In at least one embodiment, jitter can be added to motion vectors during intermediate passes such as shadow passes, occlusion passes, reflection passes, specular reflection passes, and the like.
[0059] In FIG. 11DIn particular, one motion vector 100 describes the motion of the object 10 from one frame (n-1) to the next frame (n) and has a dither value applied to slightly modify its magnitude and / or direction. In at least one embodiment, the undithered motion vector 90 can measure the distance a point on the object 10 moves from the n-1 frame to the n frame. In particular, the point 20 is the center of a pixel, shown as one of the square grid overlaying the object 10. In at least one embodiment, the undithered motion vector 90 can thus extend from the point 20 to the same point 30 in the object 10 in frame n-1, representing the distance and direction that the point 30 travels from the n-1 frame to the n frame. In at least one embodiment, dither can be applied to each point 20, 30 in each frame n-1, n. That is, a dither value or vector can be applied to each frame to slightly move the motion vector and generate some amount of blur to compensate for the jagged edges in the appearance of the object 10 when represented by a discrete set of pixels. In particular, a dither value 70 can be applied to the object 10 in the n-1 frame and a dither value 60 can be applied to the object 10 in the n frame. In at least one embodiment, the dithered motion vector 100 can be used by any image rendering pass to determine image values for any portion of an image, such as any portion of the object 10. In at least one embodiment, the same dither values 60, 70 used in any selected intermediate pass can be the dither values used in any aspect of another or more image rendering passes.
[0060] In at least one embodiment, the image value for each pixel can be determined in part by mapping the pixel center to a corresponding point in the previous frame and determining the color value for that mapped point by interpolation of neighboring pixels. That is, the pixel value for a current image frame can be determined at least in part as an estimate of the neighboring pixel values of a previous image frame. For example, the color value for a pixel having a center point 20 can be computed by determining the location of the corresponding point 30, receiving the current dither vector 60, retrieving the previous dither vector 70 from memory (such as a buffer), and computing the dithered motion vector 100, as the vector sum of the vectors drawn between points 20 and 30 (and vectors 60 and 70).
[0061] In at least one embodiment, the position of point 30 or point 40 can then be determined from motion vector 90 or 100, as appropriate. In at least one embodiment, the color value of the point can then be estimated from the color values of the neighboring pixel centers 110, 120, 130, 140. Estimation can be performed in any manner. In at least one embodiment, the color value of point 30, 40 can be estimated by bilinear interpolation of the color values of points 110, 120, 130, 140. However, any other method of estimating the color value of point 30, 40 from the color values of neighboring points 110, 120, 130, 140 is contemplated. In at least one embodiment, and as further described below, interpolation can be performed by a process such as the Catmull-Rom interpolation method.
[0062] In at least one embodiment, the color value of point 20 or one pixel of frame n is determined at least in part from the estimated color value of point 30 / 40. That is, the color value of the pixel of frame n is determined using the estimate of the color of those corresponding points in the previous frame n-1. In at least one embodiment, other values can also contribute to the color value of the pixel of frame n. In at least one embodiment, color samples can also optionally be taken at new jitter points 150 (different from the jitter values 60, 70 applied to the motion vectors). That is, the image can be sampled at new jitter points 150 such that the final color value at point 20 also includes information of the current image that it represents. In at least one embodiment, one or more samples of object 10 can be taken at frame n (at one or more jitter points 150), and the color values of these samples can be combined with the interpolated color value of point 30 / 40 as determined above. The combination or blending of these samples and interpolated values can be performed in any manner, such as a simple average, any weighted average using any one or more fixed or adaptive blending weights, or the like. In this manner, the color value of a particular pixel in the current image frame n can be determined from an interpolation of neighboring color values of the previous image frame n-1 and a combination of one or more samples taken at jitter points of the current image frame n. In at least one embodiment, this sampling using new jitter points 150 is optional and can or can not be used as desired.
[0063] FIG. 11E is a generalized embodiment of an illustrative electronic computing device configured for use in accordance with at least one embodiment. In at least one embodiment, computing device 200 can be any device capable of performing the operations of the embodiments. For example, computing device 200 can perform any of the above-described processes to generate motion vectors and determine pixel color values accordingly.
[0064] As a non-limiting example, the computing device 200 can be any electronic computing device, such as a system on a chip (SoC), an embedded processor or microprocessor, and the like, as well as any associated devices or hardware. In at least one embodiment, the computing device 200 can send and receive data via input / output (hereafter “I / O”) paths 202 and 214, which can be in electronic communication with any other device, for example, through an electronic communication medium, such as through the public Internet. In at least one embodiment, the I / O paths 202 can provide data and other input to control circuitry 204, which includes processing circuitry 206 and storage 208. In at least one embodiment, the control circuitry 204 can be used to send and receive commands, requests, and other suitable data using the I / O paths 202. In at least one embodiment, the I / O paths 202 can connect the control circuitry 204, and in particular the processing circuitry 206, to one or more communication paths. In at least one embodiment, I / O functionality can be provided by one or more of these communication paths, but in FIG. 11F The user input interface 210 can be any suitable user interface, such as a remote control, a mouse, a trackball, a keypad, a keyboard, a touchscreen, a touchpad, a stylus input, a joystick, a voice recognition interface, or other user input interface. The display 212 can be provided as a standalone device or integrated with other elements of the computing device 200. For example, the display 212 can be a touchscreen or a touch-sensitive display. In such a case, the user input interface 210 can be integrated or combined with the display 212. The display 212 can be one or more of a monitor, a television, a liquid crystal display (LCD) for a mobile device, amorphous silicon display, low-temperature polysilicon display, electronic ink display, electrophoretic display, active matrix display, electro wetting display, electrofluidic display, cathode ray tube display, light-emitting diode display, electroluminescent display, plasma display panel, high-performance addressing display, thin-film transistor display, organic light-emitting diode display, surface-conduction electron-emitter display (SED), laser television, carbon nanotube, quantum dot display, interferometric modulator display, or any other suitable device for displaying visual images.
[0065] The control circuit 204 can be based on any suitable processing circuitry, such as the processing circuitry 206. As referred to herein, processing circuitry can be understood to mean circuitry based on one or more microprocessors, microcontrollers, digital signal processors, programmable logic devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. and can include a plurality of cores and / or a plurality of processing units (e.g., CPUs, GPUs, etc.). In at least one embodiment, the processing circuitry can be distributed across a plurality of separate processors or processing units, such as a plurality of or Volta TM processors, processors, etc.) or a plurality of different processors (e.g., processors and processors, etc.). Any type and form of processing circuitry can be employed. For example, the processing circuitry 206 can include a multi-core processor, a multi-core processor configured as a graphics or compute pipeline, a neuromorphic processor, any other parallel processor or graphics processor, etc. In at least one embodiment, the processing circuitry 206 can include, without limitation, a complex instruction set computer (“CISC”) microprocessor, reduced instruction set computing (“RISC”) microprocessor, very long instruction word (“VLIW”) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor or a graphics processor.
[0066] In at least one embodiment, the control circuit 204 executes instructions for secure authentication, where the instructions can be embedded instructions or can be part of an application running on an operating system. In at least one embodiment, the computing device 100 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.
[0067] In at least one embodiment, memory can be an electronic storage device provided as storage 208 that is part of control circuit 204. As referred to herein, an “electronic storage device” or “storage device” can be understood to mean any device for storing electronic data, computer software or firmware, e.g., random access memory, read-only memory, hard drives, solid state devices, quantum storage devices, or any other suitable fixed or removable storage device, and / or any combination of these. In at least one embodiment, storage 208 can be used to store code modules as described below. In at least one embodiment, non-volatile memory can also be used (e.g., to hold the boot-up routines and other instructions). In at least one embodiment, cloud-based storage can be used in addition to or in place of storage 208.
[0068] In at least one embodiment, storage 208 can also store instructions or code for anti-aliasing processing described above to perform operations of at least one embodiment. In operation, processing circuit 206 can retrieve and execute instructions stored in storage 208 to perform processes herein.
[0069] In at least one embodiment, storage 208 can be a memory that stores a plurality of program or instruction modules for execution by processing circuit 206. For example, storage 208 can store a rendering engine 216, an anti-aliasing module 218, and storage 220, which can include buffers and other data structures and storage for performing anti-aliasing processing of at least one embodiment. In at least one embodiment, rendering engine 216 can be a set of instructions for rendering image frames or generating color values for pixels of image frames. In at least one embodiment, rendering engine 216 can retrieve dither values buffered in storage 220 and generate motion vectors as described herein, passing determined motion vectors to anti-aliasing module 218. In at least one embodiment, anti-aliasing module 218 can be a set of instructions for performing temporal accumulation processes described above, including color value interpolation and image sampling or resampling, correction and accumulation or blending of sample values to generate pixel color values. In at least one embodiment, storage 220 can be any storage for storing pixel color values of previous image frames as well as storage for dither values used to determine motion vectors of at least one embodiment. In at least one embodiment, storage 220 can include one or more buffers that store these image frames and associated dither values for retrieval by rendering engine 216 and anti-aliasing module 218. In at least one embodiment, storage 220 can be local storage (e.g., a partition or other portion of storage 208), or remote storage implemented in a remote device such as a remote database or a remote computing device such as a secure server or the like.
[0070] In at least one embodiment, computing device 200 can be a standalone computing device such as a desktop or laptop computer, a server computer, etc. However, embodiments are not limited to this configuration and other implementations of computing device 200 are contemplated. For example, computing device 200 can be a remote computing device in wired or wireless communication with another electronic computing device via an electronic communications network such as the public Internet. In such latter embodiments, a user can remotely instruct computing device 200 to implement processes described herein to select a version of a program for execution on device 200.
[0071] In at least one embodiment, computing device 200 can be any electronic computing device capable of performing pixel color value determination and anti-aliasing processing. For example, computing device 200 can be an embedded processor, a microcontroller, a locally or remotely located desktop computer, a tablet computer, or a server in electronic communication with camera 90 and actuator 70, etc. In at least one embodiment, computing device 200 can have any configuration or architecture that allows it to select and execute a version of a program in accordance with any embodiment, such as any of the configurations or architectures described below.
[0072] Inference and Training Logic
[0073] FIG. 11F Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments described herein. FIG. 3A And / or FIG. 3B Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6.
[0074] In at least one embodiment, inference and / or training logic 315 can include, without limitation, code and / or data storage 301 for storing forward and / or output weights and / or input / output data, and / or other parameters of neurons or layers of a neural network configured in aspects of one or more embodiments that are trained and / or used for inferencing. In at least one embodiment, training logic 315 can include or be coupled to code and / or data storage 301 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 301 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 301 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.
[0075] In at least one embodiment, any portion of code and / or data storage 301 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 301 can be cache memory, dynamic random addressable memory (“DRAM”), static random addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 301 is internal or external to a processor, e.g., or comprised of DRAM, SRAM, flash or some other storage type, can depend on available storage space to store on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.
[0076] In at least one embodiment, inference and / or training logic 315 can include, without limitation, code and / or data storage 305 to store backward and / or output weights and / or input / output data for neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 305 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 315 can include or be coupled to code and / or data storage 305 to store graph code or other software to control timing and / or order, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).
[0077] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 305 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 305 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 305 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash) or other storage. In at least one embodiment, whether code and / or data storage 305 is internal or external to a processor, e.g., whether made up of DRAM, SRAM, Flash, or some other storage type, is a choice dictated by design parameters including whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data being used in inferencing and / or training of a neural network, or some combination of these factors.
[0078] In at least one embodiment, code and / or data storage 301 and code and / or data storage 305 can be separate storage structures. In at least one embodiment, code and / or data storage 301 and code and / or data storage 305 can be the same storage structure. In at least one embodiment, code and / or data storage 301 and code and / or data storage 305 can be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 301 and code and / or data storage 305 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.
[0079] In at least one embodiment, inference and / or training logic 315 can include, without limitation, one or more arithmetic logic units (“ALUs”) 310 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. Results of such operations can result in activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 320, which are functions of input / output and / or weight parameter data stored in code and / or data storage 301 and / or code and / or data storage 305. In at least one embodiment, activations stored in activation storage 320 are generated by ALUs 310 in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALUs 310, where weight values stored in code and / or data storage 305 and / or code and / or data storage 301 are used as operands having other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 305 or code and / or data storage 301 or other on-chip or off-chip storage.
[0080] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 310, while in another embodiment, one or more ALUs 310 may be located outside the processor or other hardware logic device or the circuitry using them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 310 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 301, code and / or data storage 305, and activation storage 320 may share a processor or other hardware logic device or circuitry, while in another embodiment, they may be located in different processors or other hardware logic devices or circuitry, or in some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation storage 320 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0081] In at least one embodiment, the active memory 320 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 320 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 320 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types.
[0082] In at least one embodiment, FIG. 12 The inference and / or training logic 315 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel (e.g., "Lake Crest") processor. In at least one embodiment, FIG. 12The illustrated inference and / or training logic 315 can be used in combination with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (e.g., field programmable gate arrays (“FPGAs”)).
[0083] FIG. 3A Inference and / or training logic 315 are used in conjunction with the central processing unit (CPU), graphics processing unit (GPU), tensor processing unit (TPU), or other hardware that can be used to process neural network model instructions in accordance with at least one embodiment. Inference and / or training logic 315 can be used for processing of neural network model instructions in accordance with the techniques of this disclosure. FIG. 3B The inference and / or training logic 315 illustrated in FIG. 3 can be used in combination with an application-specific integrated circuit (ASIC), such as Google’s Tensor Processing Unit (TPU), an inference processing unit (IPU) from Graphcore®AI, processing units from NVIDIA®, an inference processing unit (IPU) from Graphcore®AI, TM or an Intel®“Lake Crest” processor from Intel®Corporation. In at least one embodiment, the inference and / or training logic 315 illustrated in FIG. 3 can be used in combination with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate arrays (FPGAs)). In at least one embodiment, the inference and / or training logic 315 includes, without limitation, code and / or data storage 301 and code and / or data storage 305, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In FIG. 13A-13B In at least one embodiment, each of code and / or data storage 301 and code and / or data storage 305 is respectively associated with a dedicated computing resource, such as computing hardware 302 and computing hardware 306. In at least one embodiment, each of computing hardware 302 and computing hardware 306 includes one or more ALUs that only perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data storage 301 and code and / or data storage 305, respectively, and results of performing the functions are stored in activation storage 320. FIG. 13A-13B
[0084] In at least one embodiment, each of code and / or data stores 301 and 305 and corresponding compute hardware 302 and 306, respectively, correspond to different layers of a neural network, such that activations resulting from one “store / compute pair 301 / 302” of code and / or data store 301 and compute hardware 302 provide input to next “store / compute pair 305 / 306” of code and / or data store 305 and compute hardware 306, in order to reflect a conceptual organization of a neural network. In at least one embodiment, each store / compute pair 301 / 302 and 305 / 306 can correspond to more than one neural network layer. In at least one embodiment, additional store / compute pairs (not shown) can be included in inference and / or training logic 315, either subsequent to or in parallel with store / compute pairs 301 / 302 and 305 / 306.
[0085] Neural network training and deployment
[0086] FIG. 13A Training and deployment of a deep neural network is shown, in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 406 is trained using a training dataset 402. In at least one embodiment, training framework 404 is a PyTorch framework, while in other embodiments, training framework 404 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, training framework 404 trains untrained neural network 406 and enables it to train using processing resources described herein, to generate a trained neural network 408. In at least one embodiment, weights can be chosen at random or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.
[0087] In at least one embodiment, an untrained neural network 406 is trained using supervised learning, where a training dataset 402 includes inputs paired with desired outputs for inputs, or where a training dataset 402 includes inputs with known outputs and the output is manually graded by a human. In at least one embodiment, an untrained neural network 406 is trained in a supervised manner and processes an input from a training dataset 402 and compares a resulting output to a set of expected or desired outputs. In at least one embodiment, an error is then propagated back through the untrained neural network 406. In at least one embodiment, a training framework 404 adjusts weights that control the untrained neural network 406. In at least one embodiment, a training framework 404 includes tools to monitor how well an untrained neural network 406 is converging towards a model (e.g., a trained neural network 408) that is suitable for generating correct answers (e.g., results 414) based on input data (e.g., new datasets 412). In at least one embodiment, a training framework 404 trains an untrained neural network 406 repeatedly while adjusting weights to improve outputs of the untrained neural network 406 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, a training framework 404 trains an untrained neural network 406 until the untrained neural network 406 reaches a desired accuracy. In at least one embodiment, a trained neural network 408 can then be deployed to perform any number of machine learning operations.
[0088] In at least one embodiment, an untrained neural network 406 is trained using unsupervised learning, where an untrained neural network 406 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 402 will include input data without any associated output data or “ground truth” data. In at least one embodiment, an untrained neural network 406 can learn groupings within a training dataset 402 and can determine how individual inputs relate to the untrained dataset 402. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 408 that is capable of performing operations useful for reducing a dimensionality of new datasets 412. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in new datasets 412 that deviate from a normal pattern of new datasets 412.
[0089] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mix of labeled and unlabeled data is included in training dataset 402. In at least one embodiment, training framework 404 can be used to perform incremental learning, for example, through a transferred learning technique. In at least one embodiment, incremental learning enables trained neural network 408 to adapt to new dataset 412 without forgetting knowledge that was imprinted into trained neural network 408 during initial training.
[0090] data center
[0091] FIG. 13B An example data center 500 that can use at least one embodiment is shown. In at least one embodiment, data center 500 includes a data center infrastructure layer 510, a framework layer 520, a software layer 530, and an application layer 540.
[0092] In at least one embodiment, as shown in FIG. 13A data center infrastructure layer 510 can include a resource orchestrator 512, grouped computing resources 514, and node computing resources (“node C.R.s”) 516(1)-516(N), where “N” represents a positive integer (which can be a different integer “N” than used in other figures). In at least one embodiment, node C.R.s 516(1)-516(N) can include, but are not limited to, any number of central processing units (“CPUs” or other processors including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 518(1)-518(N) (such as dynamic random access memory, solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of node C.R.s 516(1)-516(N) can be a server having one or more of above-described computing resources.
[0093] In at least one embodiment, the grouped computing resource 514 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographic locations. In at least one embodiment, the individual groups of node CRs within the grouped computing resource 514 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0094] In at least one embodiment, resource coordinator 512 may configure or otherwise control one or more nodes CR516(1)-516(N) and / or grouped computing resources 514. In at least one embodiment, resource coordinator 512 may include a software design infrastructure (“SDI”) management entity for data center 500. In at least one embodiment, resource coordinator 512 may include hardware, software, or some combination thereof.
[0095] In at least one embodiment, such as FIG. 13B As shown, framework layer 520 includes a job scheduler 522, a configuration manager 524, a resource manager 526, and a distributed file system 528. In at least one embodiment, framework layer 520 may include a framework of software 532 supporting software layer 530 and / or one or more applications 542 supporting application layer 540. In at least one embodiment, software 532 or application 542 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 520 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 528 for large-scale data processing (e.g., "big data"). TM(hereinafter “Spark”). In at least one embodiment, job scheduler 532 can include a Spark driver to facilitate scheduling of workloads supported by various layers of data center 500. In at least one embodiment, configuration manager 524 can be capable of configuring different layers, such as software layer 530 and framework layer 520 including Spark and a distributed file system 528 for supporting large-scale data processing. In at least one embodiment, resource manager 526 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting distributed file system 528 and job scheduler 522. In at least one embodiment, clustered or grouped computing resources can include grouped computing resources 514 on data center infrastructure layer 510. In at least one embodiment, resource manager 526 can coordinate with resource orchestrator 512 to manage these mapped or allocated computing resources.
[0096] In at least one embodiment, software 532 included in software layer 530 can include software used by at least a portion of node C.R.s 516(1)-516(N), grouped computing resources 514, and / or distributed file system 528 of framework layer 520. In at least one embodiment, one or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0097] In at least one embodiment, one or more applications 542 included in application layer 540 can include one or more types of applications used by at least a portion of node C.R.s 516(1)-516(N), grouped computing resources 514, and / or distributed file system 528 of framework layer 520. In at least one embodiment, one or more types of applications can include, but are not limited to, any number of genomics applications, cognitive computing, applications, and machine learning applications, including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0098] In at least one embodiment, any of configuration manager 524, resource manager 526, and resource orchestrator 512 can implement any number and type of self-modifying actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, self-modifying actions can relieve data center operators of data center 500 from making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a data center.
[0099] In at least one embodiment, data center 500 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained according to a neural network architecture by computing weight parameters using software and computing resources described above with respect to data center 500. In at least one embodiment, using weight parameters computed by one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 500.
[0100] In at least one embodiment, data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0101] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A and / or 3B. In at least one embodiment, inference and / or training logic 315 can be used in system 300 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. FIG. 12 and / or FIG. 12 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A and / or 3B. In at least one embodiment, inference and / or training logic 315 can be used in system 300 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. FIG. 13B
[0102] Autonomous vehicle
[0103] FIG. 3A An example of an autonomous vehicle 600 is shown in accordance with at least one embodiment. In at least one embodiment, autonomous vehicle 600 (alternatively referred to herein as “vehicle 600”) can be, without limitation, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 600 can be a semi-truck tractor-trailer used to haul cargo. In at least one embodiment, vehicle 600 can be an airplane, a robotic vehicle, or another type of vehicle.
[0104] Autonomous vehicles can be described in terms of automation levels defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). In at least one embodiment, vehicle 600 can be capable of functioning according to one or more of Levels 1 through 5 of the automated driving levels. For example, in at least one embodiment, vehicle 600 can be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on embodiment.
[0105] In at least one embodiment, vehicle 600 can include, without limitation, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 600 can include, without limitation, a propulsion system 650 such as a combustion engine, a hybrid electric device, a fully electric engine, and / or another propulsion system type. In at least one embodiment, propulsion system 650 can be connected to a drivetrain of vehicle 600, which can include, without limitation, a transmission to enable propulsion of vehicle 600. In at least one embodiment, propulsion system 650 can be controlled in response to receiving signals from throttle / accelerator 652.
[0106] In at least one embodiment, steering system 654 (which can include, without limitation, a steering wheel) is used to steer vehicle 600 (e.g., along a desired path or route) when propulsion system 650 is operating (e.g., when vehicle 600 is in motion). In at least one embodiment, steering system 654 can receive signals from steering actuator 656. In at least one embodiment, a steering wheel can be optional for full automation (Level 7) functionality. In at least one embodiment, brake sensor system 646 can be used to operate vehicle brakes in response to signals received from brake actuator 648 and / or brake sensors.
[0107] In at least one embodiment, controller 636 can include, without limitation, one or more system on a chip (“SoC”) (e.g., one or more processors, one or more microprocessors, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more ASICs, and / or one or more other hardware components that control functionality of vehicle 600). FIG. 3BA controller 636 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 600. For example, in at least one embodiment, controller 636 may send signals to operate vehicle braking via brake actuator 648, to operate steering system 654 via one or more steering actuators 656, and to operate propulsion system 650 via one or more throttles / accelerators 652. In at least one embodiment, one or more controllers 636 may include one or more onboard (e.g., integrated) computing devices that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 600. In at least one embodiment, one or more controllers 636 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the functions described above, and two or more controllers may handle a single function and / or any combination thereof.
[0108] In at least one embodiment, one or more controllers 636 provide signals for controlling one or more components and / or systems of vehicle 600 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors, including but not limited to one or more Global Navigation Satellite System (“GNSS”) sensors 658 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 660, one or more ultrasonic sensors 662, one or more LIDAR sensors 664, one or more inertial measurement unit (IMU) sensors 666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 696, one or more stereo cameras 668, one or more wide-angle cameras 670 (e.g., fisheye cameras), one or more infrared cameras 672, one or more surround cameras 674 (e.g., 360-degree cameras), and remote cameras (…). FIG. 13A (not shown in the image), medium-range camera ( FIG. 13B (Not shown in the diagram) One or more speed sensors 644 (e.g., for measuring the speed of vehicle 600), one or more vibration sensors 642, one or more steering sensors 640, one or more brake sensors (e.g., as part of brake sensor system 646) and / or other sensor types are received.
[0109] In at least one embodiment, one or more controllers 636 can receive inputs (e.g., represented by input data) from an instrument cluster 632 of vehicle 600 and provide outputs (e.g., represented by output data, display data, etc.) through a human-machine interface (“HMI”) display 634, audible annunciators, speakers, and / or other components of vehicle 600. In at least one embodiment, outputs can include information such as vehicle speed, velocity, time, map data (e.g., high-definition map FIG. 14A-14B data (e.g., location of vehicle 600, for example, on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects, and state of objects perceived by one or more controllers 636, etc. For example, in at least one embodiment, HMI display 634 can display information about presence of one or more objects (e.g., a street sign, a warning sign, a traffic signal changing, etc.) and / or information about driving operations vehicle has made, is making, or will make (e.g., changing lanes now, taking exit 36B in two miles, etc.).
[0110] In at least one embodiment, vehicle 600 further includes a network interface 624 that can communicate over one or more networks using one or more wireless antennas 626 and / or one or more modems. For example, in at least one embodiment, network interface 624 can be capable of communicating over a Long-Term Evolution (“LTE”), Wideband Code-Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”) network, etc. In at least one embodiment, one or more wireless antennas 626 can also use one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low power wide area networks (hereinafter “LPWANs”) (e.g., LoRaWAN, SigFox, etc. protocols) to enable communication between objects (e.g., vehicles, mobile devices) in an environment.
[0111] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 14A and / or FIG. 12 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, inference and / or training logic 315 can be used in system FIG. 13B for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0112] FIG. 14B An example of camera locations and fields of view of an autonomous vehicle 600 of FIG. 6A is shown, in accordance with at least one embodiment. FIG. 3A In at least one embodiment, the camera and respective field of view is one example embodiment and is not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras can be included and / or cameras can be located at different locations on vehicle 600.
[0113] In at least one embodiment, camera types for cameras can include, but are not limited to, digital cameras that can be suitable for use with components and / or systems of vehicle 600. In at least one embodiment, one or more cameras can operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, camera types can have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, cameras can be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, a color filter array can include a red clear clear (“RCCC”) color filter array, a red clear clear blue (“RCCB”) color filter array, a red blue green clear (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a clear pixel camera, such as a camera with an RCCC, RCCB, and / or RBGC color filter array, can be used in an effort to improve photosensitivity.
[0114] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) can record and provide image data (e.g., video) simultaneously.
[0115] In at least one embodiment, one or more cameras can be mounted in mounting assemblies, such as custom designed (three-dimensional (“3D”) printed) assemblies, in order to cut out stray light and reflections from within vehicle 600 (e.g., reflections of instrument panel in windshield mirror), which can interfere with camera’s image data capture capabilities. With regard to rearview mirror mounting assemblies, in at least one embodiment, rearview mirror assemblies can be 3D printed custom designed such that camera mounting plates match shape of rearview mirror. In at least one embodiment, one or more cameras can be integrated into rearview mirror. In at least one embodiment, for side view cameras, one or more cameras can also be integrated within four pillars at each corner of a cabin.
[0116] In at least one embodiment, cameras with a field of view that includes portions of environment in front of vehicle 600 (e.g., front-facing cameras) can be used for surround view, as well as to help identify forward path and obstacles with the help of one or more controllers 636 and / or control SoCs, thereby providing information that is critical to generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, front-facing cameras can be used to perform many ADAS functions similar to LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, front-facing cameras can also be used for ADAS functions and systems, including but not limited to Lane Departure Warning (“LDW”), Adaptive Cruise Control (“ACC”), and / or other functions (such as traffic sign recognition).
[0117] In at least one embodiment, various cameras can be used in a front-facing configuration, including, for example, monocular camera platforms that include a CMOS (“complementary metal-oxide semiconductor”) color imager. In at least one embodiment, wide-view cameras 670 can be used to perceive objects (e.g., pedestrians, crossing or bicycles) entering from the periphery. Although only one wide-view camera 670 is shown in FIG. 6B, in other embodiments, there can be any number (including zero) of wide-view cameras on vehicle 600. In at least one embodiment, any number of long-range cameras 698 (e.g., long-range stereo camera pairs) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. In at least one embodiment, long-range cameras 698 can also be used for object detection and classification, as well as basic object tracking. FIG. 3B
[0118] In at least one embodiment, any number of stereo cameras 668 can also be included in a forward-facing configuration. In at least one embodiment, one or more stereo cameras 668 can include an integrated control unit that includes a scalable processing unit that can provide programmable logic (“FPGA”) and multi-core microprocessors with integrated controller area network (“CAN”) or Ethernet interfaces on a single chip. In at least one embodiment, such a unit can be used to generate a 3D map of an environment of vehicle 600, including distance estimates for all points in an image. In at least one embodiment, one or more stereo cameras 668 can include, without limitation, a compact stereo-vision sensor that can include, without limitation, two camera lenses (one each on left and right) and an image processing chip that can measure distances from vehicle 600 to target objects and use generated information (e.g., metadata) to activate autonomous emergency braking and lane-departure warning functions. In at least one embodiment, other types of stereo cameras 668 can be used in addition to those described herein.
[0119] In at least one embodiment, cameras with a field of view that includes a portion of an environment to the sides of vehicle 600 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating an occupancy grid, as well as generating side collision warnings. For example, in at least one embodiment, surround cameras 674 (e.g., four surround cameras as shown) can be positioned on vehicle 600. In at least one embodiment, one or more surround cameras 674 can include, without limitation, any number and combination of wide-view cameras, fisheye lenses, 360-degree cameras, and / or the like. For example, in at least one embodiment, four fisheye lens cameras can be located on front, back, and sides of vehicle 600. In at least one embodiment, vehicle 600 can use three surround cameras 674 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround view camera. FIG. 14B
[0120] In at least one embodiment, cameras with a field of view that includes a portion of an environment to the rear of vehicle 600 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear collision warnings, and creating and updating an occupancy grid. In at least one embodiment, a wide variety of cameras can be used, including without limitation cameras that are also suitable as one or more forward-facing cameras (e.g., long-range cameras 698 and / or one or more mid-range cameras 676, one or more stereo cameras 668, one or more infrared cameras 672, etc.), as described herein.
[0121] Inference and / or training logic 315 is used to perform inference and / or training operations associated with one or more embodiments. FIG. 14A and / or FIG. 3A This document provides details regarding inference and / or training logic 315. In at least one embodiment, inference and / or training logic 315 may be... FIG. 3B Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0122] FIG. 15 The illustration shows an embodiment according to at least one of the embodiments. FIG. 15 A block diagram of an example system architecture for an autonomous vehicle 600. In at least one embodiment, FIG. 3A Each of one or more components, one or more features, and one or more systems of vehicle 600 is shown as connected via bus 602. In at least one embodiment, bus 602 may include, but is not limited to, a CAN data interface (which may alternatively be referred to herein as “CAN bus”). In at least one embodiment, CAN may be a network within vehicle 600 used to help control various features and functions of vehicle 600, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. In one embodiment, bus 602 may be configured to have dozens or even hundreds of nodes, each node having its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 602 can be read to find steering wheel angle, ground speed, engine rotation speed (“RPM”), button position, and / or other vehicle status indicators. In at least one embodiment, bus 602 may be an ASIL B compliant CAN bus.
[0123] In at least one embodiment, FlexRay and / or Ethernet protocols may be used in addition to or from CAN. In at least one embodiment, there may be any number of molded buses 602, which may include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or may be used for redundancy. For example, a first bus may be used for a collision avoidance function, and a second bus may be used for actuation control. In at least one embodiment, each of the buses 602 may communicate with any component of the vehicle 600, and two or more buses 602 may communicate with corresponding components. In at least one embodiment, each of any number of system-on-chip (“SoC”) 604 (e.g., SoC 604(A) and SoC 604(B)), each of one or more controllers 636, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of the vehicle 600) and may be connected to a common bus, such as a CAN bus.
[0124] In at least one embodiment, vehicle 600 may include one or more controllers 636, such as those described herein. FIG. 3B As described above. In at least one embodiment, controller 636 can be used for a variety of functions. In at least one embodiment, controller 636 can be coupled to any of various other components and systems of vehicle 600, and can be used to control vehicle 600, artificial intelligence of vehicle 600, infotainment and / or other functions of vehicle 600.
[0125] In at least one embodiment, vehicle 600 may include any number of SoCs 604. In at least one embodiment, each of the SoCs 604 may include, but is not limited to, a central processing unit (“one or more CPUs”) 606, a graphics processing unit (“one or more GPUs”) 608, one or more processors 610, one or more caches 612, one or more accelerators 614, one or more data storage 616, and / or other components and features not shown. In at least one embodiment, one or more SoCs 604 may be used to control vehicle 600 on various platforms and systems. For example, in at least one embodiment, one or more SoCs 604 may be combined with a high-definition (“HD”) map 622 in a system (e.g., the system of vehicle 600), the high-definition map 622 being accessible from one or more servers via a network interface 624. FIG. 15 (Not shown in the image) Get map refresh and / or update.
[0126] In at least one embodiment, one or more CPU(s) 606 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, one or more CPU(s) 606 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, one or more CPU(s) 606 can include eight cores in a multi-processor configuration coupled to one another. In at least one embodiment, one or more CPU(s) 606 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., 2 MB L2 cache). In at least one embodiment, one or more CPU(s) 606 (e.g., CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of one or more CPU(s) 606 can be active at any given time.
[0127] In at least one embodiment, one or more CPU(s) 606 can implement power management functionality including, without limitation, one or more of the following features: individual hardware modules can be automatically clock-gated at idle to conserve dynamic power; each core clock can be gated when the core is not actively executing instructions due to execution of a wait for interrupt (“WFI”) / wait for event (“WFE”) instruction; each core can be independently powered; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. In at least one embodiment, one or more CPU(s) 606 can further implement enhanced algorithms for managing power states with allowed power states and expected wake-up times specified and hardware / microcode determining optimal power states for core, cluster, and CCPLEX inputs. In at least one embodiment, processing cores can support a simplified power state input sequence in software with work offloaded to microcode.
[0128] In at least one embodiment, GPU(s) 608 can include an integrated GPU (also referred to herein as an “iGPU”). In at least one embodiment, GPU(s) 608 can be programmable and efficient for parallel workloads. In at least one embodiment, GPU(s) 608 can use an enhanced tensor instruction set. In one embodiment, GPU(s) 608 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“LI”) cache (e.g., an LI cache with at least 116 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with 712 KB storage capacity). In at least one embodiment, GPU(s) 608 can include at least eight streaming microprocessors. In at least one embodiment, GPU(s) 608 can use a compute application programming interface (“API”). In at least one embodiment, GPU(s) 608 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA model).
[0129] In at least one embodiment, GPU(s) 608 can be power-optimized to achieve best performance in automotive and embedded use cases. For example, in at least one embodiment, GPU(s) 608 can be fabricated on a finned field effect transistor (“FinFET”) circuit. In at least one embodiment, each streaming microprocessor can contain a plurality of mixed-precision processing cores divided into a plurality of blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 18 FP32 cores, 10 FP64 cores, 18 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a level zero (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or an 84 KB register file. In at least one embodiment, a streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads that mix compute and address operations. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to enable both simplified programming and increased performance.
[0130] In at least one embodiment, one or more GPU(s) 608 can include high bandwidth memory (“HBM”) and / or 18 GB HBM2 memory subsystems to provide, in some examples, a peak memory bandwidth of about 1100 GB / s. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random access memory (“SGRAM”) can be used, for example, graphics double data rate type five synchronous random access memory (“GDDR5”).
[0131] In at least one embodiment, one or more GPU(s) 608 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow one or more GPU(s) 608 to directly access one or more CPU(s) 606 page tables. In at least one embodiment, when a memory management unit (“MMU”) of a GPU of one or more GPU(s) 608 experiences a miss, an address translation request can be sent to one or more CPU(s) 606. In response, a CPU of one or more CPU(s) 606 can look up a virtual-to-physical mapping for an address in its page tables and transmit a translation back to one or more GPU(s) 608, in at least one embodiment. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both one or more CPU(s) 606 and one or more GPU(s) 608, simplifying programming of one or more GPU(s) 608 and porting of applications to one or more GPU(s) 608.
[0132] In at least one embodiment, one or more GPU(s) 608 can include any number of access counters that can track how frequently one or more GPU(s) 608 is accessing memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved into physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.
[0133] In at least one embodiment, one or more SoC(s) 604 can include any number of caches 612, including those described herein. For example, in at least one embodiment, one or more cache(s) 612 can include a level three (“L3”) cache that can be available to and / or connected to one or more CPU(s) 606 and one or more GPU(s) 608. In at least one embodiment, one or more cache(s) 612 can include a write-back cache that can track state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, although smaller cache sizes can be used, L3 cache can include 6 MB of memory or more, according to embodiments.
[0134] In at least one embodiment, one or more SoC(s) 604 can include one or more accelerator(s) 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 604 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 6 MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement and offload some tasks of one or more GPU(s) 608 (e.g., freeing up more cycles of one or more GPU(s) 608 to perform other tasks). In at least one embodiment, one or more accelerator(s) 614 can be used for target workloads that are stable enough to pass the acceleration test (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, CNNs can include region-based or region with convolutional neural networks (“RCNNs”) and fast RCNNs (e.g., as used for object detection) or other types of CNNs.
[0135] In at least one embodiment, one or more accelerators 614 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). In at least one embodiment, one or more DLAs can include, without limitation, one or more Tensor Processing Units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can be further optimized for a particular set of neural network types and floating point operations and inferencing. In at least one embodiment, design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often significantly outperform CPUs. In at least one embodiment, one or more TPUs can perform several functions including support for INT8, INT16, and FP16 data types for features and weights, single instance convolution functionality, and post-processor functionality, for example. In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, as well as recognition and detection, using data from microphones; CNNs for facial recognition and vehicle owner identification using data from camera sensors; and / or CNNs for safety and / or safety related events.
[0136] In at least one embodiment, a DLA can perform any of functions of one or more GPU(s) 608, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or one or more GPU(s) 608 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations for CNNs on one or more DLAs, and leave other functions to one or more GPU(s) 608 and / or one or more accelerators 614.
[0137] In at least one embodiment, one or more accelerators 614 can include a programmable vision accelerator (“PVA”), which can alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, one or more PVAs can be designed and configured to accelerate computer vision algorithms used for advanced driver assistance systems (“ADAS”) 638, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, one or more PVAs can strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs can include, for example and without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.
[0138] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any camera described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.
[0139] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 606. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, but are not limited to, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0140] In at least one embodiment, a vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem can function as a primary processing engine for a PVA and can include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction multiple data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.
[0141] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to a dedicated memory. As a result, in at least one embodiment, each vector processor can be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA can be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA can execute a general purpose computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA can execute different computer vision algorithms on one image at a time, or even different algorithms on a sequence of images or portions of images. In at least one embodiment, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA, among other things. In at least one embodiment, a PVA can include additional error correcting code (“ECC”) memory to enhance overall system security.
[0142] In at least one embodiment, one or more accelerators 614 can include on-chip computer vision networks and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 614. In at least one embodiment, on-chip memory can include at least 6 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include on-chip computer vision networks that interconnect PVA and DLA to memory (e.g., using APB).
[0143] In at least one embodiment, on-chip computer vision networks can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 28262 or International Electrotechnical Commission (“IEC”) 81508 standards.
[0144] In at least one embodiment, one or more SoC 604 can include real-time line-of-sight tracking hardware accelerators. In at least one embodiment, real-time line-of-sight tracking hardware accelerators can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.
[0145] In at least one embodiment, one or more accelerators 614 have broad use for autonomous driving. In at least one embodiment, PVAs can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, capabilities of PVAs at low power and low latency are well matched to algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that can require predictable runtimes with low latency and low power. In at least one embodiment, PVAs can be designed to run classic computer vision algorithms, such as in vehicle 600, as they can be efficient at object detection and integer math operations.
[0146] For example, in accordance with at least one embodiment of technology, PVAs are used to perform computer stereo vision. In at least one embodiment, semi-global matching based algorithms can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use dynamic estimation / stereo matching in run-time (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0147] In at least one embodiment, PVAs can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 6D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, PVAs are used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.
[0148] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a confidence level for each object detection. In at least one embodiment, confidence level can be represented or interpreted as a probability, or as providing a relative “weight” of each detection relative to other detections. In at least one embodiment, a confidence measurement enables system to make further decisions as to which detections should be considered as true positive detections and not false positive detections. In at least one embodiment, system can set a threshold for confidence level, and only consider detections that exceed threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in vehicle automatically performing emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, neural network can take as its input at least some subset of parameters, such as bounding box size, ground plane estimates obtained (e.g., from another subsystem), output of one or more IMU sensors 666 related to object’s vehicle 600 direction, distance, 3D position estimates obtained from neural network and / or other sensors (e.g., one or more LIDAR sensors 664 or one or more RADAR sensors 660), etc.
[0149] In at least one embodiment, one or more SoC(s) 604 can include one or more data storage(s) 616 (e.g., memory). In at least one embodiment, one or more data storage(s) 616 can be on-chip memory of one or more SoC(s) 604 that can store neural networks to be executed on one or more GPU(s) 608 and / or DLA. In at least one embodiment, one or more data storage(s) 616 can have sufficient capacity to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data storage(s) 616 can include L2 or L3 cache.
[0150] In at least one embodiment, one or more SoC(s) 604 can include any number of processor(s) 610 (e.g., embedded processors). In at least one embodiment, one or more processor(s) 610 can include a boot and power management processor that can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. In at least one embodiment, a boot and power management processor can be part of a one or more SoC(s) 604 boot sequence and can provide run-time power management services. In at least one embodiment, a boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 604 thermal and temperature sensor management, and / or one or more SoC(s) 604 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 604 can use ring oscillators to detect temperature of one or more CPU(s) 606, one or more GPU(s) 608, and / or one or more accelerator(s) 614. In at least one embodiment, if a temperature is determined to exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoC(s) 604 into a lower power state and / or put vehicle 600 into a safe park pattern for the driver (e.g., cause vehicle 600 to safely park).
[0151] In at least one embodiment, one or more processor(s) 610 can further include a set of embedded processors that can function as an audio processing engine that can be an audio subsystem that is capable of providing full hardware support for multi-channel audio to hardware through a number of interfaces as well as a broad and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0152] In at least one embodiment, one or more processor(s) 610 can further include an always-on processor engine that can provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, a processor on an always-on processor engine can include, but is not limited to, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0153] In at least one embodiment, one or more processors 610 can further include a safety cluster engine including, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 610 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 610 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0154] In at least one embodiment, one or more processors 610 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final video needed to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-view cameras 670, one or more surround cameras 674, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, preferably, in-cabin monitoring camera sensors are monitored by a neural network running on another instance of SoC 604 that is configured to identify cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change an infotainment system and settings of a vehicle, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.
[0155] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.
[0156] In at least one embodiment, video image compositor can also be configured to perform stereo correction on input stereoscopic lens frames. In at least one embodiment, when using an operating system desktop, video image compositor can also be used for user interface composition and one or more GPUs 608 are not required to continuously render new surfaces. In at least one embodiment, when one or more GPUs 608 are powered and active for 3D rendering, video image compositor can be used to offload one or more GPUs 608 to improve performance and responsiveness.
[0157] In at least one embodiment, one or more SoCs in SoC 604 can further include mobile industry processor interface (“MIPI”) camera serial interfaces for receiving video and input from cameras, high speed interfaces, and / or video input blocks that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoCs 604 can further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0158] In at least one embodiment, one or more SoCs in SoC 604 can further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. In at least one embodiment, one or more SoCs 604 can be used to process data from cameras (e.g., connected over gigabit multimedia serial link and Ethernet channels), sensors (e.g., one or more LIDAR sensors 664, one or more RADAR sensors 660, etc., which can be connected over Ethernet channels), data from bus 602 (e.g., speed of vehicle 600, steering wheel position, etc.), data from one or more GNSS sensors 658 (e.g., connected over Ethernet bus or CAN bus), etc. In at least one embodiment, one or more SoCs in SoC 604 can further include dedicated high performance mass storage controllers that can include their own DMA engines and can be used to free one or more CPUs 606 from regular data management tasks.
[0159] In at least one embodiment, SoC(s) 604 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technology for diversity and redundancy, which provides a platform that can provide a flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 604 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 614, when combined with CPU(s) 606, GPU(s) 608, and data storage(s) 616, can provide a fast, efficient platform for level 3-5 autonomous vehicles.
[0160] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C) to perform a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs typically cannot meet performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.
[0161] Embodiments described herein allow for simultaneous and / or sequential execution of multiple neural networks, and allow for results to be combined together to enable level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 620) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that a neural network has not been specifically trained for. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing that semantic understanding to a path planning module running on a CPU Complex.
[0162] In at least one embodiment, multiple neural networks can be run simultaneously for a level 3, 4, or 5 drive. For example, in at least one embodiment, a warning sign consisting of a “Caution: flashing lights indicate icy conditions” sign with electric lights flashing can be interpreted by multiple neural networks independently or collectively. In at least one embodiment, the warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has already been trained), the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network, which informs a vehicle’s path planning software (preferably executing on a CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be recognized by a third deployed neural network operating over multiple frames, informing a vehicle’s path planning software of the existence (or non-existence) of flashing lights. In at least one embodiment, all three neural networks can be run simultaneously, e.g., within a DLA and / or on one or more GPU(s) 608.
[0163] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from a camera sensor to identify presence of an authorized driver and / or owner of vehicle 600. In at least one embodiment, when an owner approaches a driver door and turns on a light, a normally open sensor processor engine can be used to unlock the vehicle, and, in a safe mode, when the owner leaves the vehicle, can be used to disable the vehicle. In this way, one or more SoC(s) 604 provide a safeguard against theft and / or carjacking.
[0164] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphones 696 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 604 use a CNN to classify ambient and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify relative proximity of emergency vehicles (e.g., by using Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles for regions in which a vehicle is operating, as identified by one or more GNSS sensors 658. In at least one embodiment, when operating in Europe, a CNN will seek to detect European sirens, while in North America, a CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, slow vehicle down, pull vehicle to side of road, stop, and / or idle vehicle until emergency vehicle passes, with assistance of one or more ultrasonic sensors 662.
[0165] In at least one embodiment, vehicle 600 can include one or more CPUs 618 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 604 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 618 can include an X86 processor, such as one or more CPUs 618 can be used to perform any of a variety of functions, such as including potentially arbitrating inconsistent results between ADAS sensors and one or more SoCs 604, and / or one or more monitoring controllers 636 for status and health and / or an information system on a chip (“Info SoC”) 630.
[0166] In at least one embodiment, vehicle 600 can include one or more GPUs 620 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 604 via a high-speed interconnect (e.g., NVIDIA’s NVLINK channel). In at least one embodiment, one or more GPUs 620 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 600.
[0167] In at least one embodiment, vehicle 600 can further include network interface 624, which can include, without limitation, one or more wireless antennas 626 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 624 can be used to enable wireless connectivity through an Internet cloud service (e.g., with servers and / or other network equipment) with other vehicles and / or computing devices (e.g., client devices of passengers). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 600 and another vehicle and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a direct link can be provided using a vehicle-to-vehicle communication link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 600 with information about vehicles in a vicinity of vehicle 600 (e.g., vehicles in front of, to the side of, and / or behind vehicle 600). In at least one embodiment, this aforementioned functionality can be part of a cooperative adaptive cruise control functionality of vehicle 600.
[0168] In at least one embodiment, network interface 624 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 636 to communicate over wireless networks. In at least one embodiment, network interface 624 can include radio frequency front ends for up-conversion from baseband to radio frequency and for down-conversion from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0169] In at least one embodiment, vehicle 600 can further include one or more data stores 628, which can include, without limitation, off-chip (e.g., of SoC(s) 604) storage. In at least one embodiment, one or more data stores 628 can include, without limitation, one or more storage elements including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.
[0170] In at least one embodiment, vehicle 600 can further include one or more GNSS sensors 658 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 658 can be used including, for example and without limitation, a GPS using a USB connector with Ethernet connected to a serial interface (e.g., RS-232) bridge.
[0171] In at least one embodiment, vehicle 600 can further include one or more RADAR sensors 660. In at least one embodiment, one or more RADAR sensors 660 can be used by vehicle 600 for long range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. In at least one embodiment, one or more RADAR sensors 660 can use CAN bus and / or bus 602 (e.g., to transmit data generated by one or more RADAR sensors 660) for control and access to object tracking data, in certain examples can access an Ethernet channel for access to raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more of RADAR sensors 660 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 660 are pulse Doppler RADAR sensors.
[0172] In at least one embodiment, RADAR sensor(s) 660 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, long-range RADAR systems can provide a wide field of view implemented through two or more independent scans (e.g., up to 270 m range). In at least one embodiment, RADAR sensor(s) 660 can help distinguish between static and moving objects, and can be used by ADAS system 638 for emergency brake assist and forward collision warning. In at least one embodiment, sensor(s) 660 included in a long-range RADAR system can include, without limitation, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 600 at higher speeds with minimal traffic interference from adjacent lanes. In at least one embodiment, other two antennas can expand the field of view, such that vehicles 600 entering or leaving a lane can be quickly detected.
[0173] In at least one embodiment, as an example, a mid-range RADAR system can include, for example, a range of up to 180 m (front) or 100 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 660 designed to be mounted at either end of a rear bumper. When mounted at either end of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor the vehicle’s rearward direction and a blind spot close by. In at least one embodiment, a short-range RADAR system can be used in ADAS system 638 for blind spot detection and / or lane change assist.
[0174] In at least one embodiment, vehicle 600 can further include ultrasonic sensor(s) 662. In at least one embodiment, ultrasonic sensor(s) 662, which can be positioned in front, rear, and / or side locations of vehicle 600, can be used for parking assist and / or to create and update occupancy grids. In at least one embodiment, a wide variety of ultrasonic sensors 662 can be used, and different ultrasonic sensors 662 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 662 can operate at a functional safety level of ASIL B.
[0175] In at least one embodiment, vehicle 600 can include one or more LIDAR sensors 664. In at least one embodiment, one or more LIDAR sensors 664 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 664 can operate at a functional safety level of ASIL B. In at least one embodiment, vehicle 600 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 664 that can use Ethernet channels (e.g., provide data to a Gigabit Ethernet switch).
[0176] In at least one embodiment, one or more LIDAR sensors 664 can be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, one or more LIDAR sensors 664 that are commercially available can have an advertised range of approximately 120 m, have an accuracy of 2 cm - 3 cm, and support an Ethernet connection of 120 Mbps, for example. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such embodiments, one or more LIDAR sensors 664 can include small devices that can be embedded into front, rear, side, and / or corner locations of vehicle 600. In at least one embodiment, one or more LIDAR sensors 664, in such embodiments, can provide a horizontal field of view of up to 140 degrees and a vertical field of view of 35 degrees, with a range of 220 m, even for low reflectivity objects.
[0177] In at least one embodiment, a forward-facing LIDAR sensor 664 can be configured for a horizontal field of view between 65 degrees and 155 degrees.
[0178] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. In at least one embodiment, 3D Flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200 m around vehicle 600. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 600 to an object. In at least one embodiment, flash LIDAR can allow for highly accurate and distortion-free images of a surrounding environment to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 600. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond class I (eye-safe) laser pulses per frame, and can capture reflected laser light as a 3D ranging point cloud and co-registered intensity data.
[0179] In at least one embodiment, vehicle 600 can also include one or more IMU sensors 666. In at least one embodiment, one or more IMU sensors 666 can be located at a center of a rear axle of vehicle 600. In at least one embodiment, one or more IMU sensors 666 can include, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one magnetic compass, multiple magnetic compasses, and / or other sensor types. In at least one embodiment, such as in a six-axis application, one or more IMU sensors 666 can include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as in a nine-axis application, one or more IMU sensors 666 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0180] In at least one embodiment, one or more IMU sensors 666 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receivers, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude; in at least one embodiment, one or more IMU sensors 666 can enable vehicle 600 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from GPS to one or more IMU sensors 666. In at least one embodiment, one or more IMU sensors 666 and one or more GNSS sensors 658 can be combined in a single integrated unit.
[0181] In at least one embodiment, vehicle 600 can include one or more microphones 696 placed within and / or around vehicle 600. In at least one embodiment, additionally, one or more microphones 696 can be used for emergency vehicle detection and identification.
[0182] In at least one embodiment, vehicle 600 can further include any number of camera types, including one or more stereo cameras 668, one or more wide-view cameras 670, one or more infrared cameras 672, one or more surround cameras 674, one or more long-range cameras 698, one or more mid-range cameras 676, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around an entire periphery of vehicle 600. In at least one embodiment, a type of camera used depends on vehicle 600. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 600. In at least one embodiment, a number of cameras deployed can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 600 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communications, by way of example and without limitation. In at least one embodiment, cameras can be described in greater detail herein previously with reference to FIG. 6A. FIG. 16A and FIG. 15 Each camera can be described in greater detail.
[0183] In at least one embodiment, vehicle 600 can further include one or more vibration sensors 642. In at least one embodiment, one or more vibration sensors 642 can measure vibrations of components of vehicle 600 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in road surface. In at least one embodiment, when two or more vibration sensors 642 are used, differences between vibrations can be used to determine road surface friction or slippage (e.g., when there is a difference in vibration between a power driven axle and a free spinning axle).
[0184] In at least one embodiment, vehicle 600 can include ADAS system 638. In at least one embodiment, ADAS system 638 can include, without limitation, an SoC. In at least one embodiment, ADAS system 638 can include, without limitation, any number of adaptive / autonomous / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane-keep assist (“LKA”) systems, blind-spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality, and combinations thereof.
[0185] In at least one embodiment, an ACC system can use one or more RADAR sensors 660, one or more LIDAR sensors 664, and / or any number of cameras. In at least one embodiment, an ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls distance to another vehicle immediately in front of vehicle 600 and automatically adjusts speed of vehicle 600 to maintain a safe distance from the vehicle in front. In at least one embodiment, a lateral ACC system performs distance keeping and advises vehicle 600 to change lanes when necessary. In at least one embodiment, lateral ACC is relevant to other ADAS applications, such as LC and CW.
[0186] In at least one embodiment, a CACC system uses information from other vehicles that can be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via network interface 624 and / or one or more wireless antennas 626. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, V2V communications provide information about the immediately preceding vehicle (e.g., a vehicle immediately ahead of and in the same lane as vehicle 600), while I2V communications provide information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V sources of information. In at least one embodiment, a CACC system can be more reliable with information about vehicles in front of vehicle 600, and has potential to improve smoothness of traffic flow and reduce road congestion.
[0187] In at least one embodiment, an FCW system is designed to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or one or more RADAR sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an FCW system can provide a warning, for example, in the form of a sound, a visual warning, a vibration, and / or a quick brake pulse.
[0188] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, an AEB system can use one or more forward-facing cameras and / or one or more RADAR sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, it typically first warns a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, an AEB system can include techniques such as dynamic brake support and / or crash imminent braking.
[0189] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, e.g., steering wheel or seat vibration, to warn a driver when vehicle 600 crosses lane markers. In at least one embodiment, an LDW system is not active when a driver indicates an intentional lane departure, such as by activating turn signals. In at least one embodiment, an LDW system can use a forward-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system. In at least one embodiment, if vehicle 600 begins to deviate from a lane, an LKA system provides a steering input or brake to correct vehicle 600.
[0190] In at least one embodiment, a BSW system detects and warns vehicle drivers of vehicles in a car’s blind spot. In at least one embodiment, a BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system can provide additional warnings when a driver uses turn signals. In at least one embodiment, a BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0191] In at least one embodiment, a RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside of a rear camera range while vehicle 600 is backing up. In at least one embodiment, a RCTW system includes an AEB system to ensure application of vehicle brakes to avoid a collision. In at least one embodiment, a RCTW system can use one or more rear-facing RADAR sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide driver feedback such as displays, speakers, and / or vibrating components.
[0192] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems warn drivers and allow that driver to decide whether a safety situation is truly present and take appropriate action. In at least one embodiment, in the event of a result conflict, vehicle 600 itself decides whether to heed the results of a primary computer or a secondary computer (e.g., a first controller or a second controller of controllers 636). For example, in at least one embodiment, ADAS system 638 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run redundant varieties of software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, outputs from ADAS system 638 can be provided to a supervisory MCU. In at least one embodiment, if outputs from a primary computer and outputs from a secondary computer conflict, a supervisory MCU decides how to reconcile the conflict to ensure safe operation.
[0193] In at least one embodiment, a host computer can be configured to provide a confidence score to a supervisory MCU to indicate a host computer’s confidence in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the host computer’s indication regardless of whether the secondary computer provides a conflicting or inconsistent result. In at least one embodiment, in cases where a confidence score does not satisfy a threshold, and in cases where the host computer and secondary computer indicate different results (e.g., a conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.
[0194] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which a secondary computer provides false alarms based at least in part on output from a host computer and output from a secondary computer. In at least one embodiment, a neural network in a supervisory MCU can learn when to trust output of a secondary computer, and when not to. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, a neural network in a supervisory MCU can learn when the FCW system identifies metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alert. In at least one embodiment, when the secondary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when there is a bicyclist or pedestrian present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 604.
[0195] In at least one embodiment, ADAS system 638 can include a secondary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, the secondary computer can use classic computer vision rules (if-then), and presence of a neural network in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, a diversified implementation and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and non-identical software code running on a secondary computer provides consistent overall results, a supervisory MCU can be more confident that overall results are correct, and that a bug in software or hardware on the host computer did not cause a significant error.
[0196] In at least one embodiment, output of ADAS system 638 can be input into a perception module of host computer and / or a dynamic driving task module of host computer. For example, in at least one embodiment, if ADAS system 638 indicates a forward collision warning due to an object directly in front of vehicle 600, perception block can use this information when identifying the object. In at least one embodiment, as described herein, a secondary computer can have its own neural network that is trained to reduce risk of false positives.
[0197] In at least one embodiment, vehicle 600 can further include infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, in at least one embodiment, infotainment SoC 630 can not be a SoC and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 630 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear park assist, radio data system, vehicle related information such as fuel level, total range, brake fuel level, oil level, doors open / close, air filter information, etc.) to vehicle 600. For example, infotainment SoC 630 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car, in-car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 634, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 630 can be further used to provide information (e.g., visual and / or audible) to a user of vehicle 600, such as information from ADAS system 638, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0198] In at least one embodiment, infotainment SoC 630 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 630 can communicate with other devices, systems, and / or components of vehicle 600 over bus 602. In at least one embodiment, infotainment SoC 630 can be coupled to a supervisory MCU such that a GPU of an infotainment system can perform some autonomous driving functions in the event of a failure of a host controller 636 (e.g., a primary computer and / or a backup computer of vehicle 600). In at least one embodiment, infotainment SoC 630 can cause vehicle 600 to enter a driver-to-safe-stop mode, as described herein.
[0199] In at least one embodiment, vehicle 600 can further include an instrument cluster 632 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument cluster, etc.). In at least one embodiment, instrument cluster 632 can include, without limitation, a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 632 can include, without limitation, any number and combination of gauges such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 630 and instrument cluster 632. In at least one embodiment, instrument cluster 632 can be included as part of infotainment SoC 630, and vice versa.
[0200] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 16B and / or FIG. 16A Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 16A In at least one embodiment, inference and / or training logic 315 can be used in system
[0201] FIG. 16A is a cloud-based server in communication with a vehicle in accordance with at least one embodiment. FIG. 15FIG. 6 is a diagram of a system 676 in communication between autonomous vehicles 600. In at least one embodiment, system 676 can include, without limitation, one or more servers 678, one or more networks 690, and any number and type of vehicles, including vehicles 600. In at least one embodiment, one or more servers 678 can include, without limitation, a plurality of GPUs 684(A)-684(H) (collectively referred to herein as GPUs 684), PCIe switches 682(A)-682(D) (collectively referred to herein as PCIe switches 682), and / or CPUs 680(A)-680(B) (collectively referred to herein as CPUs 680). GPUs 684, CPUs 680, and PCIe switches 682 can be interconnected with high-speed connection lines such as, but not limited to, NVLink interfaces 688 developed by NVIDIA and / or PCIe connections 686. In at least one embodiment, GPUs 684 are connected by NVLink and / or NVSwitch SoC connections, and GPUs 684 and PCIe switches 682 are connected by PCIe interconnects. Although eight GPUs 684, two CPUs 680, and four PCIe switches 682 are illustrated, this is not intended to be limiting. In at least one embodiment, each of one or more servers 678 can include, without limitation, any number of GPUs 684, CPUs 680, and / or PCIe switches 682 in any combination. For example, in at least one embodiment, one or more servers 678 can each include eight, sixteen, thirty-two, and / or more GPUs 684.
[0202] In at least one embodiment, one or more servers 678 can receive, through one or more networks 690 and from vehicles, image data representative of images showing unexpected or changed road conditions, such as road work that recently started. In at least one embodiment, one or more servers 678 can transmit, through one or more networks 690 and to vehicles, updated ego neural networks 692, and / or map information 694 including, without limitation, information about traffic and road conditions. In at least one embodiment, updates to map information 694 can include, without limitation, updates to HD map 622, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural networks 692 and / or map information 694 can be the result of new training and / or experience represented in data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 678 and / or other servers).
[0203] In at least one embodiment, one or more servers 678 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using game engines). In at least one embodiment, any amount of training data is labeled (e.g., where associated neural networks benefit from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, no training data is labeled and / or pre-processed (e.g., where associated neural networks do not require supervised learning). In at least one embodiment, once a machine learning model is trained, machine learning model can be used by vehicles (e.g., transferred to vehicles by one or more networks 690, and / or machine learning model can be used by one or more servers 678 to monitor vehicles remotely.
[0204] In at least one embodiment, one or more servers 678 can receive data from vehicles and apply data to up-to-date, real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 678 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 684, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 678 can include deep learning infrastructure of data centers powered using CPUs.
[0205] In at least one embodiment, deep learning infrastructure of one or more servers 678 can be capable of fast, real-time inference, and can use this capability to assess and validate health of processors, software, and / or related hardware in vehicles 600. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicles 600, such as sequences of images and / or objects located by vehicles 600 in that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural networks to identify objects and compare them to objects identified by vehicles 600, and, if results do not match and deep learning infrastructure concludes that Al in vehicles 600 is malfunctioning, one or more servers 678 can send a signal to vehicles 600 instructing failsafe computers of vehicles 600 to take control, notify passengers, and complete safe parking operations.
[0206] In at least one embodiment, one or more servers 678 can include one or more GPUs 684 and one or more programmable inference accelerators (such as NVIDIA’s TensorRT 3 devices). In at least one embodiment, a combination of GPU-driven servers and inference-accelerated servers can make real-time responses possible. In at least one embodiment, CPU-, FPGA-, and other processor-driven servers can be used for inference, for example in cases where performance is less critical. In at least one embodiment, hardware structure 315 is used to perform one or more embodiments. Details regarding hardware structure 315 are provided herein. FIG. 16A and / or FIG. 16C Details regarding hardware structure 315 are provided herein.
[0207] Computer system
[0208] FIG. 16A is a block diagram illustrating an exemplary computer system that can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units to execute an instruction, according to at least one embodiment. In at least one embodiment, according to the present disclosure, such as embodiments described herein, computer system 700 can include, without limitation, components such as processor 702 that has execution units including logic to perform algorithms for process data. In at least one embodiment, computer system 700 can include a processor, such as a Pentium TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessors, although other systems (including PCs, workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, computer system 700 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux, for example), embedded software, and / or graphical user interfaces can also be used.
[0209] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area
[0210] In at least one embodiment, computer system 700 can include, but is not limited to, processor 702 that can include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or inferencing according to techniques described herein. In at least one embodiment, computer system 700 is a single processor desktop or server system, but in another embodiment, computer system 700 can be a multiprocessor system. In at least one embodiment, processor 702 can include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combo of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 702 can be coupled to a processor bus 710 that can transmit data signals between processor 702 and other components in computer system 700.
[0211] In at least one embodiment, processor 702 can include, but is not limited to, level 1 ("Ll") internal cache memory ("cache") 704. In at least one embodiment, processor 702 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 702's external. Other embodiments can include a combination of internal and external caches based on specific implementation and requirements. In at least one embodiment, register file 706 can store different types of data within various registers including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.
[0212] In at least one embodiment, execution unit 708, including but not limited to logic to perform integer and floating point operations, also resides in processor 702. In at least one embodiment, processor 702 can also include microcode ("ucode") read-only memory ("ROM"), which stores microcode for certain macroinstructions. In at least one embodiment, execution unit 708 can also include logic to handle a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in a general-purpose processor, operations used by many multimedia applications can be performed using the packed data of processor 702 without direct instruction of a separate multimedia processor. In at least one embodiment, many multimedia applications can be accelerated using the full width of processor's data bus for performing operations on packed data, which can not need to be transmitted over the processor's data bus in smaller units for performing one or more operations on one data element at a time.
[0213] In at least one embodiment, execution unit 708 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 700 can include, but not be limited to, memory 720. In at least one embodiment, memory 720 can be a Dynamic Random Access Memory ("DRAM") device, a Static Random Access Memory ("SRAM") device, a flash memory device, or another memory device. In at least one embodiment, memory 720 can store data 721 and / or instructions 719 that can be executed by processor 702 as signaled by data signals.
[0214] In at least one embodiment, a system logic chip can be coupled to processor bus 710 and memory 720. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 716 and processor 702 can communicate with MCH 716 via processor bus 710. In at least one embodiment, MCH 716 can provide a high bandwidth memory path 718 to memory 720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 716 can direct data signals between processor 702, memory 720, and other components in computer system 700 and can translate virtual addresses into physical addresses, perform memory protection, and
[0215] In at least one embodiment, computer system 700 can use system I / O interface 722 as a proprietary hub interface bus to couple MCH 716 to I / O controller hub (“ICH”) 730. In at least one embodiment, ICH 730 can provide a direct connection to some I / O devices and can indirect connections to other I / O devices through a local I / O bus. In at least one embodiment, local I / O bus can include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 720, chipset, and processor 702. Examples can include, without limitation, audio controller 729, firmware hub (“Flash BIOS”) 728, wireless transceiver 726, data storage 724, legacy I / O controller 723 containing user input and keyboard interfaces 725, serial expansion port 727 (e.g., Universal Serial Bus (USB) port), and network controller 734. In at least one embodiment, data storage 724 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0216] In at least one embodiment, FIG. 16A A system including interconnected hardware devices or “chips” is shown, while in other embodiments, FIG. 16A A SoC can be shown. In at least one embodiment, FIG. 16AThe devices illustrated in FIG. 8 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 800 are interconnected using a compute express link (CXL) interconnect.
[0217] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A and / or 3B. In examples in which FIG. 16A and / or FIG. 3A Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A and / or 3B. In examples in which inference and / or training logic 315 are used for predicting or forecasting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein, inference and / or training logic 315 can be used in place of, or in conjunction with, inference and / or training logic 215 described in conjunction with FIG. 2. FIG. 3B Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A and / or 3B. In examples in which inference and / or training logic 315 are used for predicting or forecasting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein, inference and / or training logic 315 can be used in place of, or in conjunction with, inference and / or training logic 215 described in conjunction with FIG. 2.
[0218] FIG. 16D is a block diagram illustrating an electronic device 800 for utilizing processor 810, in accordance with at least one embodiment. In at least one embodiment, electronic device 800 can be, for example and without limitation, a laptop computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0219] In at least one embodiment, electronic device 800 can include, without limitation, processor 810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 810 is coupled using a bus or interface, such as an Industry Standard 2 In at least one embodiment, processor 810 is coupled using a bus or interface, such as an Industry Standard Architecture (ISA), Extended ISA (EISA), Advanced Technology Attachment (AT A) bus, or a bus or interface that implements a version of the Peripheral Component Interconnect (PCI) such as PCI, PCI Extended (PCI-X), PCI Express (PCIe), or a bus or interface that implements a version of the Accelerated Graphics Port (AGP) such as AGP 1.X, AGP 4X, or AGP 8X. In at least one embodiment, processor 810 is coupled to other components of electronic device 800 using a bus or interface, such as a Serial Peripheral Interface (SPI) bus, a High-Definition Audio (HDA) bus, a Serial Advanced Technology Attachment (SATA) bus, a Universal Serial Bus (USB) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (UART) bus. FIG. 3A In at least one embodiment, a system is illustrated that includes hardware devices or “chips” interconnected, while in other embodiments, FIG. 3B In at least one embodiment, an exemplary SoC can be illustrated. In at least one embodiment, FIG. 17 The devices illustrated in FIG. 8 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, FIG. 3A one or more components of computer system 800 are interconnected using a compute express link (CXL) interconnect.
[0220] In at least one embodiment, FIG. 3BThe display 824, touch screen 825, touch pad 830, near field communication unit ("NFC") 845, sensor hub 840, thermal sensor 846, express chip set ("EC") 835, trusted platform module ("TPM") 838, BIOS / firmware / flash ("BIOS, FW Flash") 822, DSP 860, drive 820 (e.g., solid state disk ("SSD") or hard disk drive ("HDD")), wireless local area network unit ("WLAN") 850, Bluetooth unit 852, wireless wide area network unit ("WWAN") 856, global positioning system ("GPS") unit 855, camera ("USB 3.0 camera") 854 (e.g., USB 3.0 camera), and / or low power double data rate ("LPDDR") memory unit ("LPDDR3") 815 implemented in, for example, LPDDR3 standard can be included. These components can each be implemented in any suitable manner.
[0221] In at least one embodiment, other components can be communicatively coupled to processor 810 by components described herein. In at least one embodiment, accelerometer 841, ambient light sensor ("ALS") 842, compass 843, and gyroscope 844 can be communicatively coupled to sensor hub 840. In at least one embodiment, thermal sensor 839, fan 837, keyboard 836, and touch pad 830 can be communicatively coupled to EC 835. In at least one embodiment, speaker 863, earpiece 864, and microphone ("mic") 865 can be communicatively coupled to audio unit ("audio codec and class D amplifier") 862, which in turn can be communicatively coupled to DSP 860. In at least one embodiment, audio unit 862 can include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, SIM card ("SIM") 857 can be communicatively coupled to WWAN unit 856. In at least one embodiment, components such as WLAN unit 850 and Bluetooth unit 852, as well as WWAN unit 856, can be implemented as a next generation form factor ("NGFF").
[0222] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 18 and / or FIG. 3A Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, inference and / or training logic 315 can be used in system FIG. 3B for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0223] FIG. 19 A computer system 900 according to at least one embodiment is shown. In at least one embodiment, the computer system 900 is configured to implement various processes and methods described throughout this disclosure.
[0224] In at least one embodiment, the computer system 900 includes, but is not limited to, at least one central processing unit (“CPU”) 902 connected to a communication bus 910 implemented using any suitable protocol, such as PCI (“Peripheral Device Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 900 includes, but is not limited to, main memory 904 and control logic (e.g., implemented in hardware, software, or a combination thereof), and data may be stored in main memory 904 in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“Network Interface”) 922 provides an interface to other computing devices and networks for receiving data using the computer system 900 and transferring data to other systems.
[0225] In at least one embodiment, the computer system 900 includes, but is not limited to, an input device 908, a parallel processing system 912, and a display device 906, which may be implemented using conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light-emitting diode (“LED”) display, plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 908 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the modules described herein may reside on a single semiconductor platform to form the processing system.
[0226] Inference and / or training logic 315 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... FIG. 3A and / or FIG. 3B Details are provided regarding the inference and / or training logic 315. In at least one embodiment, the inference and / or training logic 315 can be in the system. FIG. 20 It is used in the context of reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0227] FIG. 3AA computer system 1000 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 1000 includes, without limitation, a computer 1010 and a USB stick 1020. In at least one embodiment, computer 1010 can include, without limitation, any number and type of processor (not shown) and memory (not shown). In at least one embodiment, computer 1010 includes, without limitation, a server, a cloud instance, a laptop computer, and a desktop computer.
[0228] In at least one embodiment, USB stick 1020 includes, without limitation, a processing unit 1030, a USB interface 1040, and USB interface logic 1050. In at least one embodiment, processing unit 1030 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1030 can include, without limitation, any number and type of processing core (not shown). In at least one embodiment, processing unit 1030 includes an application specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, processing unit 1030 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, processing unit 1030 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.
[0229] In at least one embodiment, USB interface 1040 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1040 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1040 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1050 can include any number and type of logic that enables processing unit 1030 to interface with a device (e.g., computer 1010) via USB connector 1040.
[0230] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided herein in conjunction with FIGS. 3A and / or 3B. In at least one embodiment, inference and / or training logic 315 can be used in system 300 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. FIG. 3B and / or FIG. 21 Details regarding inference and / or training logic 315 are provided. In at least one embodiment, inference and / or training logic 315 can be used in system 300 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. FIG. 22
[0231] FIG. 3A An exemplary architecture is shown in which a plurality of GPUs 1110(1)- 1110(N) are communicatively coupled to a plurality of multi-core processors 1105(1)- 1105(M) over high-speed links 1140(1)-1140(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, high-speed links 1140(1)-1140(N) support communication
[0232] Further, in at least one embodiment, two or more GPUs 1110 are interconnected by high-speed links 1129(1)-1129(2), which can be implemented using similar or different protocols / links than those used for high-speed links 1140(1)-1140(N). Similarly, two or more multi-core processors 1105 can be connected by a high-speed link 1128, which can be an SMP bus that runs at 20 GB / s, 30 GB / s, 120 GB / s or higher. Alternatively, similar protocols / links (e.g., over common interconnect fabric) can be used to FIG. 3B all communication between various system components shown in FIG. 11B.
[0233] In at least one embodiment, each multi-core processor 1105 is communicatively coupled to processor memory 1101(1)-1101(M) via memory interconnects 1126(1)- 1126(M), respectively, and each GPU 1110(1)-1110(N) is communicatively coupled to GPU memory 1120(1)-1120(N) through GPU memory interconnects 1150(1)-1150(N), respectively. In at least one embodiment, memory interconnects 1126 and 1150 can utilize similar or different memory access technologies. By way of non-limiting example, processor memory 1101(1)-1101(M) and GPU memory 1120 can be volatile memory such as dynamic random access memory (DRAM) (including stack DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or can be non-volatile memory such as 3D XPoint or Nano-Ram. In at least one embodiment, certain portions of processor memory 1101 can be volatile memory while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0234] As described herein, although various multi-core processors 1105 and GPUs 1110 can be physically coupled to particular memories 1101, 1120, respectively, and / or can implement a unified memory architecture in which a virtual system address space (also referred to as an “effective address” space) is distributed among various physical memories. For example, processor memories 1101(1)-1101(M) can each contain 64 GB of system memory address space, and GPU memories 1120(1)-1120(N) can each contain 32 GB of system memory address space, resulting in a total of 276 GB of addressable memory size when M = 2 and N = 4. N and M can also be other values.
[0235] FIG. 3A Additional details are shown for an interconnect between multi-core processor 1107 and graphics acceleration module 1146, in accordance with one exemplary embodiment. In at least one embodiment, graphics acceleration module 1146 can include one or more GPU chips integrated on a line card that is coupled via a high-speed link 1140 (e.g., a PCIe bus, NVLink, etc.) to processor 1107. In at least one embodiment, graphics acceleration module 1146 can alternatively be integrated on a package or chip with processor 1107.
[0236] In at least one embodiment, processor 1107 includes a number of cores 1160A-1160D, each having a translation lookaside buffer (“TLB”) 1161A-1161D and one or more caches 1162A-1162D. In at least one embodiment, cores 1160A-1160D can include various other components not shown for purposes of performing instructions and processing data. In at least one embodiment, caches 1162A-1162D can include Level 1 (LI) and Level 2 (L2) caches. Additionally, one or more shared caches 1156 can be included in caches 1162A-1162D and shared by sets of cores 1160A-1160D. For example, one embodiment of processor 1107 includes 24 cores, each with its own LI cache, twelve shared L2 caches, and twelve shared L3 caches. In that embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, processor 1107 and graphics acceleration module 1146 are connected with system memory 1114, which can include processor memories 1101(1)-1101(M) in FIG. 3B
[0237] In at least one embodiment, consistency for data and instructions stored in respective caches 1162A-1162D, 1156, and system memory 1114 is maintained through inter-core communications over coherence bus 1164. In at least one embodiment, each cache can have cache coherence logic / circuitry associated with it to communicate over coherence bus 1164 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snoop protocol is implemented over coherence bus 1164 to snoop cache accesses.
[0238] In at least one embodiment, agent circuitry 1125 communicatively couples graphics acceleration module 1146 to coherence bus 1164, allowing graphics acceleration module 1146 to participate in a cache coherence protocol as a peer to cores 1160A-1160D. In particular, in at least one embodiment, interface 1135 provides connectivity from graphics acceleration module 1146 to agent circuitry 1125 over high-speed link 1140, and interface 1137 connects graphics acceleration module 1146 to high-speed link 1140.
[0239] In at least one embodiment, accelerator integration circuit 1136 provides graphics acceleration module 1146 with cache management, memory access, context management, and interrupt management services on behalf of a plurality of graphics processing engines 1131(1)-1131(N) represented by graphics acceleration module 1146. In at least one embodiment, graphics processing engines 1131(1)-1131(N) can each comprise a separate graphics processing unit (GPU). In at least one embodiment, graphics processing engines 1131(1)-1131(N) alternatively can comprise different types of graphics processing engines within a GPU such as, for example, a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, graphics acceleration module 1146 can be a GPU with a plurality of graphics processing engines 1131(1)-1131(N) or graphics processing engines 1131(1)-1131(N) can be individual GPUs integrated on a common package, line card, or chip.
[0240] In at least one embodiment, accelerator integration circuit 1136 includes a memory management unit (MMU) 1139 to process memory read and write requests between graphics processing engines 1131(1)-1131(N) and system memory 1114, in at least one embodiment, MMU 1139 includes address translation lookaside buffer (TLB) to improve translation of virtual addresses to real addresses, in at least one embodiment, MMU 1139 also includes a memory access unit to perform memory access operations.
[0241] In at least one embodiment, a set of registers 1145 store program context data for threads executed by graphics processing engines 1131(1)-1131(N), and context management circuit 1148 manages thread contexts. For example, context management circuit 1148 can perform save and restore operations to save and restore context for individual threads during context switches (e.g., where a first thread is saved and a second thread is stored so that second thread can be executed by graphics processing engines). For example, context management circuit 1148 can store current register values to a designated area in memory (e.g., identified by a context pointer) upon a context switch. Register values can then be restored when returning to a context. In at least one embodiment, interrupt management circuit 1147 receives and processes interrupts received from system devices.
[0242] In at least one embodiment, MMU 1139 translates virtual / effective addresses from graphics processing engines 1131 into real / physical addresses in system memory 1114. In at least one embodiment, accelerator integration circuit 1136 supports a number of graphics accelerator modules 1146 and / or other accelerator devices (e.g., 6, 10, 18). In at least one embodiment, graphics accelerator modules 1146 can be dedicated to a single application executing on processor 1107 or can be shared between multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources of graphics processing engines 1131(1)-1131(N) are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into“slices” that are assigned to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.
[0243] In at least one embodiment, accelerator integration circuit 1136 performs as a bridge to system for a system of graphics acceleration modules 1146 and provides address translation and system memory cache services. Further, in at least one embodiment, accelerator integration circuit 1136 can provide virtualization facilities for a host processor to manage virtualization of graphics processing engines 1131(1)-1131(N), interrupts, and memory management.
[0244] In at least one embodiment, because hardware resources of graphics processing engines 1131(1)-1131(N) are explicitly mapped to real address space seen by host processor 1107, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integration circuit 1136 is to physically separate graphics processing engines 1131(1)-1131(N) so that they appear as independent units to a system.
[0245] In at least one embodiment, one or more graphics memories 1133(1)-1133(M) are respectively coupled to each graphics processing engine 1131(1)-1131(N) and N=M. In at least one embodiment, graphics memories 1133(1)-1133(M) store instructions and data for processing by each of graphics processing engines 1131(1)-1131(N). In at least one embodiment, graphics memories 1133(1)-1133(M) can be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory such as 3D XPoint or Nano-Ram.
[0246] In at least one embodiment, to reduce data traffic on high-speed link 1140, a biasing technique can be used to ensure that data stored in graphics memory 1133(1)-1133(M) is that which is most frequently used by graphics processing engines 1131(1)-1131(N) and is preferably not used (at least not frequently) by cores 1160A-1160D. Similarly, in at least one embodiment, a biasing mechanism attempts to keep data needed by cores (and preferably not graphics processing engines 1131(1)-1131(N)) in caches 1162A-1162D, 1156, and system memory 1114.
[0247] FIG. 23 Another exemplary embodiment is shown in which accelerator integration circuit 1136 is integrated within processor 1107. In this embodiment, graphics processing engines 1131(1)-1131(N) communicate directly over high-speed link 1140 to accelerator integration circuit 1136 via interface 1137 and interface 1135 (which can be any form of bus or interface protocol as desired). In at least one embodiment, accelerator integration circuit 1136 can perform similar operations to those described with respect to FIG. 3A accelerator integration circuit 1136 can perform similar operations to those described with respect to
[0248] In at least one embodiment, graphics processing engines 1131(1)-1131(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1131(1)-1131(N), providing virtualization within a VM / partition.
[0249] In at least one embodiment, graphics processing engines 1131(1)-1131(N) can be shared by multiple VM / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize graphics processing engines 1131(1)-1131(N) to allow access by each operating system. In at least one embodiment, for a single-partition system without a hypervisor, an operating system owns graphics processing engines 1131(1)-1131(N). In at least one embodiment, an operating system can virtualize graphics processing engines 1131(1)-1131(N) to provide access to each process or application.
[0250] In at least one embodiment, graphics acceleration module 1146 or individual graphics processing engines 1131(1)-1131(N) use a process handle to select a process element. In at least one embodiment, a process element is stored in system memory 1114 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation-specific value provided to a host process when registering its context with a graphics processing engine 1131(1)-1131(N) (i.e., calling system software to add a process element to a process element linked list). In at least one embodiment, a lower 16 bits of a process handle can be an offset of a process element in a process element linked list.
[0251] FIG. 3B An exemplary accelerator integration slice 1190 is shown. In at least one embodiment, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 1136. In at least one embodiment, an application is an effective address space 1182 in system memory 1114 that stores a process element 1183. In at least one embodiment, process element 1183 is stored in response to a GPU invocation 1181 from an application 1180 executing on processor 1107. In at least one embodiment, process element 1183 contains process state for a respective application 1180. In one embodiment, a work descriptor (WD) 1184 contained in process element 1183 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 1184 is a pointer to a job request queue in an application’s effective address space 1182.
[0252] In at least one embodiment, graphics acceleration module 1146 and / or individual graphics processing engines 1131(1)-1131(N) can be shared by all or a subset of processes in a system. In at least one embodiment, can include infrastructure for setting up process state and sending a WD 1184 to graphics acceleration module 1146 to start a job in a virtualized environment.
[0253] In at least one embodiment, a dedicated process programming model is implementation specific. In at least one embodiment, in this model, a single process owns a graphics acceleration module 1146 or individual graphics processing engines 1131. In at least one embodiment, when a graphics acceleration module 1146 is owned by a single process, a hypervisor initializes an accelerator integration circuit for the owned partition, and an operating system initializes an accelerator integration circuit 1136 for the owned process when a graphics acceleration module 1146 is assigned.
[0254] In at least one embodiment, in operation, a WD fetch unit 1191 in an accelerator integration slice 1190 fetches a next WD 1184 that includes an indication of work to be completed by one or more graphics processing engines of a graphics acceleration module 1146. In at least one embodiment, data from WD 1184 can be stored in registers 1145 and used by MMU 1139, interrupt management circuit 1147, and / or context management circuit 1148, as shown. For example, one embodiment of MMU 1139 includes segment / page walk circuitry for accessing segment / page tables 1186 within an OS virtual address space 1185. In at least one embodiment, interrupt management circuit 1147 can handle interrupt events 1192 received from a graphics acceleration module 1146. In at least one embodiment, effective addresses 1193 generated by graphics processing engines 1131(1)-1131(N) are translated to real addresses by MMU 1139 when performing graphics operations.
[0255] In at least one embodiment, registers 1145 are replicated for each graphics processing engine 1131(1)-1131(N) and / or graphics acceleration module 1146, and can be initialized by a hypervisor or operating system. In at least one embodiment, each of these replicated registers can be included in an accelerator integration slice 1190. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.
[0256] Table 1 - Hypervisor-Initialized Registers
[0257]
[0258] Exemplary registers that can be initialized by an operating system are shown in Table 2.
[0259] Table 2 - Operating System-Initialized Registers
[0260]
[0261] In at least one embodiment, each WD 1184 is specific to a particular graphics acceleration module 1146 and / or graphics processing engine 1131(1)-1131(N). In at least one embodiment, it contains all information that graphics processing engine 1131(1)-1131(N) needs to complete the work assigned to WD 1184, or it can be a pointer to a memory location where application has set up a command queue of work to be completed.
[0262] FIG. 23 Additional details of one exemplary embodiment of a shared model are shown. This embodiment includes a hypervisor real address space 1198 in which a list of process elements 1199 is stored. In at least one embodiment, hypervisor real address space 1198 is accessible via hypervisor 1196, which virtualizes graphics acceleration module engines for operating system 1195.
[0263] In at least one embodiment, a shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in a system to use graphics acceleration module 1146. In at least one embodiment, there are two programming models in which graphics acceleration module 1146 is shared by multiple processes and partitions, i.e., time-sliced sharing and graphics-directed sharing.
[0264] In at least one embodiment, in this model, system hypervisor 1196 owns graphics acceleration module 1146 and makes its functionality available to all operating systems 1195. In at least one embodiment, for graphics acceleration module 1146 to support virtualization by system hypervisor 1196, graphics acceleration module 1146 can adhere to certain requirements, such as (1) application’s job requests must be autonomous (i.e., no state needs to be kept between jobs), or graphics acceleration module 1146 must provide a context save and restore mechanism, (2) graphics acceleration module 1146 guarantees that an application’s job request completes within a specified amount of time, including any translation faults, or graphics acceleration module 1146 provides the ability to preempt processing of a job, and (3) fairness between graphics acceleration module 1146 processes must be ensured when operating in a directed shared programming model.
[0265] In at least one embodiment, application 1180 is required to use a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP) for an operating system 1195 system call. In at least one embodiment, the graphics acceleration module type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system specific value. In at least one embodiment, the WD is formatted specifically for a graphics acceleration module 1146 and can take the form of a graphics acceleration module 1146 command, a valid address pointer to a user defined structure, a valid address pointer to a queue of commands, or any other data structure describing work to be done by a graphics acceleration module 1146.
[0266] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to how an application program sets the AMR. In at least one embodiment, if an accelerator integration circuit 1136 (not shown) and graphics acceleration module 1146 implementation does not support a user authority mask override register (UAMOR), then the operating system can apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. In at least one embodiment, the hypervisor 1196 can selectively apply the current authority mask override register (AMOR) value before placing the AMR in the process element 1183. In at least one embodiment, the CSRP is one of registers 1145 that contains a valid address of an area in application’s effective address space 1182 for a graphics acceleration module 1146 to save and restore context state. In at least one embodiment, this pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be a fixed system memory.
[0267] Upon receiving the system call, operating system 1195 can verify that application 1180 has registered and is granted authority to use graphics acceleration module 1146. Operating system 1195 then, in at least one embodiment, uses the information shown in Table 3 to call hypervisor 1196.
[0268] Table 3 - Operating system to hypervisor call parameters
[0269]
[0270] In at least one embodiment, upon receiving a hypervisor call, hypervisor 1196 verifies that operating system 1195 has been registered and granted permission to use graphics acceleration module 1146. Then, in at least one embodiment, hypervisor 1196 adds process element 1183 to a linked list of process elements of the corresponding graphics acceleration module 1146 type. In at least one embodiment, the process element may include the information shown in Table 4.
[0271] Table 4 – Process Element Information
[0272]
[0273]
[0274] In at least one embodiment, the hypervisor initializes registers 1145 of multiple accelerator integration slices 1190.
[0275] like FIG. 3A As shown, in at least one embodiment, a unified memory is used, which is addressable via a common virtual memory address space for accessing physical processor memories 1101(1)-1101(N) and GPU memories 1120(1)-1120(N). In this implementation, operations performed on GPUs 1110(1)-1110(N) utilize the same virtual / effective memory address space to access processor memories 1101(1)-1101(M) and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1101(1), a second portion to second processor memory 1101(N), a third portion to GPU memory 1120(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memory 1101 and GPU memory 1120, thereby allowing any processor or GPU to access that memory using a virtual address mapped to any physical memory.
[0276] In at least one embodiment, the bias / coherence management circuitry 1194A-1194E within one or more MMUs 1139A-1139E ensures cache coherence between one or more host processors (e.g., 1105) and the cache of GPU 1110, and implements biasing techniques to indicate the physical memory in which certain types of data should be stored. In at least one embodiment, although in FIG. 3BMultiple instances of bias / coherency management circuitry 1194A-1194E are shown in the example of FIG. 11, but can be implemented within the MMU(s) of the host processor(s) 1105 and / or within the accelerator integration circuit 1136.
[0277] One embodiment allows GPU memory 1120 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability to access GPU memory 1120 as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. In at least one embodiment, this arrangement allows software of the host processor(s) 1105 to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies include driver calls, interrupts, and memory mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 1120 without cache coherency overhead can be critical to the execution time of offloaded computations. In at least one embodiment, for example, in cases with a large amount of streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 1110. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offload.
[0278] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granularity structure (e.g., controlled at the granularity of a memory page) that includes 1 or 2 bits per GPU-attached memory page. In at least one embodiment, with or without a bias cache in GPU 1110 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in the stolen memory range of the GPU memory(s) 1120. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.
[0279] In at least one embodiment, prior to actually accessing GPU memory, a bias table entry associated with each access to GPU-attached memory 1120 is accessed, resulting in the following operations. In at least one embodiment, local requests from GPU 1110 that find their pages in GPU bias are forwarded directly to corresponding GPU memory 1120. In at least one embodiment, local requests from GPU that find their pages in host bias are forwarded to processor 1105 (e.g., over a high-speed link as described herein). In at least one embodiment, requests from processor 1105 that find requested pages in host processor bias complete requests similar to normal memory reads. Alternatively, requests that point to GPU bias pages can be forwarded to GPU 1110. In at least one embodiment, if GPU is not currently using a page, GPU can then migrate the page to host processor bias. In at least one embodiment, bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in limited cases purely hardware-based mechanisms.
[0280] In at least one embodiment, a mechanism for changing bias state employs an API call (e.g., OpenCL) that in turn invokes a device driver of a GPU that in turn sends a message (or causes a command descriptor to be enqueued) to a GPU, directing the GPU to change bias state and, in certain migrations, to perform a cache flush operation in a host. In at least one embodiment, a cache flush operation is used for a migration from host processor 1105 bias to GPU bias, but not for the reverse migration.
[0281] In one embodiment, cache coherency is maintained by temporarily rendering GPU bias pages that host processor 1105 cannot cache. In at least one embodiment, to access these pages, processor 1105 can request access from GPU 1110, which can or can not grant access immediately. Thus, in at least one embodiment, to reduce communication between processor 1105 and GPU 1110, it is beneficial to ensure that GPU bias pages are pages that are needed by GPU and not by host processor 1105, and vice versa.
[0282] One or more hardware structures 135 are used to perform one or more embodiments. Details regarding one or more hardware structures 135 can be found in this document in connection with FIG. 24 and / or FIG. 3A Details regarding one or more hardware structures 135 are provided.
[0283] FIG. 3BExemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown that can be fabricated using one or more IP cores. In addition to the illustrated, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.
[0284] FIG. 3A is a block diagram illustrating an exemplary system on a chip integrated circuit 1200 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1200 includes one or more application processor(s) 1205 (e.g., CPUs), at least one graphics processor 1210, and can additionally include an image processor 1215 and / or a video processor 1220, any of which can be a modular IP core. In at least one embodiment, integrated circuit 1200 includes peripheral or bus logic including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I2S / I2C controller 1240. In at least one embodiment, integrated circuit 1200 can include a display device 1245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display controller 1255. In at least one embodiment, storage can be provided by a flash memory subsystem 1260 including flash memory and a flash memory controller. In at least one embodiment, memory interface can be provided via a memory controller 1265 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1270. 2 2S / I 2 2C controller 1240. In at least one embodiment, integrated circuit 1200 can include a display device 1245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display controller 1255. In at least one embodiment, storage can be provided by a flash memory subsystem 1260 including flash memory and a flash memory controller. In at least one embodiment, memory interface can be provided via a memory controller 1265 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1270.
[0285] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 3B and / or FIG. 25 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, inference and / or training logic 315 can be used in integrated circuit 1200 to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0286] FIG. 24 Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown that can be fabricated using one or more IP cores. In addition to the illustrated, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.
[0287] FIG. 3A is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. FIG. 3B An exemplary graphics processor 1310 of a system on a chip integrated circuit is shown according to at least one embodiment, which can be fabricated using one or more IP cores. FIG. 3A Another exemplary graphics processor 1340 of a system on a chip integrated circuit is shown according to at least one embodiment, which can be fabricated using one or more IP cores. In at least one embodiment, FIG. 25 The graphics processor 1310 of FIG. 13A is a low power graphics processor core. In at least one embodiment, FIG. 3A The graphics processor 1340 of FIG. 13B is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1310, 1340 can be a FIG. 3B a variant of the graphics processor 1210 of FIG. 12.
[0288] In at least one embodiment, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processor(s) 1315A-1315N (e.g., 1315A, 1315B, 1315C, 1315D, through 1315N-1, and 1315N). In at least one embodiment, the graphics processor 1310 can execute different shader programs via separate logical
[0289] In at least one embodiment, graphics processor 1310 additionally includes one or more memory management units (MMUs) 1320A-1320B, cache memory 1325A-1325B, and one or more circuit interconnects 1330A-1330B. In at least one embodiment, one or more MMU(s) 1320A-1320B provide for virtual to physical address mapping for graphics processor 1310, including for vertex processor 1305 and / or fragment processor 1315A-1315N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more cache memories 1325A-1325B. In at least one embodiment, one or more MMU(s) 1320A-1320B can be synchronized with one or more MMUs within FIG. 26 application processor(s) 1205, image processors 1215, and / or video processors 1220, such that each processor 1205-1220 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1330A-1330B enable graphics processor 1310 to interface with other IP cores within a SoC, either via an internal bus, or via a direct connection.
[0290] In at least one embodiment, graphics processor 1340 includes one or more shader cores 1355A-1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, through 1355N-1, and 1355N), as shown in FIG. 13C, which provides a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code to implement vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, a plurality of shader cores can vary. In at least one embodiment, graphics processor 1340 includes an inter-core task manager 1345, which acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1355A-1355N and tiling unit 1358 to accelerate tiling operations for tile-based rendering in which rendering operations for a scene are subdivided in image space, e.g., to exploit local spatial coherence or to optimize use of internal caches. FIG. 3A
[0291] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGs. 3A-3C. FIG. 3B and / or FIG. 26 Details regarding the inference and / or training logic 315 are provided. In at least one embodiment, inference and / or training logic 315 can be used in graphics processing cluster 1114 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases as described herein. FIG. 3A and / or FIG. 3B
[0292] FIG. 27A-27B Additional exemplary graphics processor logic in accordance with embodiments described herein is shown. In at least one embodiment, graphics processor 1210 can include graphics core(s) 1400 that can be a unified shader core 1355A-1355N as shown in FIG. 13B, in at least one embodiment. FIG. 27A FIG. 27B FIG. 27A FIG. 27B A highly parallel general purpose graphics processing unit (“GPGPU”) 1430 suitable for deployment on a multi-chip module is shown in at least one embodiment.
[0293] In at least one embodiment, graphics core 1400 includes shared instruction cache 1402, texture unit(s) 1418, and cache / shared memory 1420, which are common to the execution resources within graphics core 1400. In at least one embodiment, graphics core 1400 can include a number of slices 1401A-1401N or partitions of each core and a graphics processor can include multiple instances of graphics core 1400. In at least one embodiment, slices 1401A-1401N can include support logic including a local instruction cache 1404A-1404N, a thread scheduler 1406A-1406N, a thread dispatcher 1408A-1408N, and a set of registers 1410A-1410N. In at least one embodiment, slices 1401A-1401N can include a set of additional functional units (AFUs) 1412A-1412N, floating point units (FPUs) 1414A-1414N, integer arithmetic logic units (ALUs) 1416A-1416N, address computation units (ACUs) 1413A-1413N, double precision floating point units (DPFPUs) 1415A-1415N, and matrix processing units (MPUs) 1417A-1417N.
[0294] In at least one embodiment, FPUs 1414A-1414N can perform single-precision (32-bit) and half-precision (16-bit) floating point operations, while DPFPUs 1415A-1415N perform double-precision (64-bit) floating point operations. In at least one embodiment, ALUs 1416A-1416N can perform variable precision integer operations at 10, 18, and 34-bit precision, and can be configured to operate in mixed precision formats. In at least one embodiment, MPUs 1417A-1417N can also be configured for mixed precision matrix operations including half-precision floating point and 10-bit integer operations. In at least one embodiment, MPUs 1417-1417N can perform a variety of matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix to matrix multiplication (GEMM). In at least one embodiment, AFUs 1412A-1412N can perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).
[0295] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, 3C, 3D, 3E, 3F, 3G, 3H, 3I, 3J, 3K, and 3L. FIG. 3A and / or FIG. 3B Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, 3C, 3D, 3E, 3F, 3G, 3H, 3I, 3J, 3K, and 3L. In at least one embodiment, inference and / or training logic 315 can be used in graphics core 1400 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.
[0296] FIG. 3AA general-purpose processing unit (GPGPU) 1430 is illustrated in at least one embodiment, which can be configured to enable highly parallel computational operations to be performed by a set of graphics processing units. In at least one embodiment, the GPGPU 1430 can be directly linked to other instances of the GPGPU 1430 to create a multi-GPU cluster to improve the training speed for deep neural networks. In at least one embodiment, the GPGPU 1430 includes a host interface 1432 for connection to a host processor. In at least one embodiment, the host interface 1432 is a PCI Express interface. In at least one embodiment, the host interface 1432 may be a vendor-specific communication interface or communication structure. In at least one embodiment, the GPGPU 1430 receives commands from the host processor and uses a global scheduler 1434 to allocate execution threads associated with those commands to a set of compute clusters 1436A-1436H. In at least one embodiment, compute clusters 1436A-1436H share a cache memory 1438. In at least one embodiment, cache memory 1438 can be used as a higher-level cache within the cache memory of computing clusters 1436A-1436H.
[0297] In at least one embodiment, the GPGPU 1430 includes memories 1444A-1444B, which are coupled to the computing cluster 1436A-1436H via a set of memory controllers 1442A-1442B. In at least one embodiment, memories 1444A-1444B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), which includes graphics double data rate (GDDR) memory.
[0298] In at least one embodiment, each of the computing clusters 1436A-1436H includes a set of graphics cores, for example... FIG. 3B The graphics core 1400 may include various types of integer and floating-point logic units that can perform computational operations across a range of precisions, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each computing cluster 1436A-1436H may be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units may be configured to perform 64-bit floating-point operations.
[0299] In at least one embodiment, multiple instances of GPGPU 1430 can be configured to function as a compute cluster. In at least one embodiment, communication for synchronization and data exchange for compute clusters 1436A-1436H varies between embodiments. In at least one embodiment, multiple instances of GPGPU 1430 communicate via host interface 1432. In at least one embodiment, GPGPU 1430 includes an I / O hub 1439 that couples the GPGPU 1430 with a GPU link 1440 that enables a direct connection to other instances of GPGPU 1430. In at least one embodiment, GPU link 1440 couples to a specialized GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGP 1430. In at least one embodiment, GPU link 1440 couples with a high-speed interconnect to transmit and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1430 are located in separate data processing systems and communicate via a network device accessible via host interface 1432. In at least one embodiment, GPU link 1440 can be configured to enable connection to a host processor in addition to or as an alternative to host interface 1432.
[0300] In at least one embodiment, GPGPU 1430 can be configured to train neural networks. In at least one embodiment, GPGPU 1430 can be used within an inferencing platform. In at least one embodiment, where GPGPU 1430 is used for inferencing, GPGPU 1430 can include fewer compute clusters 1436A-1436H relative to when GPGPU 1430 is used to train neural networks. In at least one embodiment, memory technology associated with memory 1444A-1444B can vary between inferencing and training configurations, with higher bandwidth memory technology dedicated to training configurations. In at least one embodiment, an inferencing configuration of GPGPU 1430 can support inferencing specific instructions. For example, in at least one embodiment, an inferencing configuration can provide support for one or more 10 -bit integer dot product instructions that can be used during inferencing operations for deployed neural networks.
[0301] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 315 are used in systems to facilitate generation of trees, such as decision trees, for use in performing inferencing operations associated with one or more embodiments. FIG. 28 and / or FIG. 28Details regarding the inference and / or training logic 315 are provided. In at least one embodiment, the inference and / or training logic 315 can be used in GPGPU 1430 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0302] FIG. 28 A block diagram of a computer system 1500 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 1500 includes a processing subsystem 1501 with one or more processors 1502 and a system memory 1504 communicating via an interconnection path 1505, which can include a memory hub 1505. In at least one embodiment, memory hub 1505 can be a separate component coupled with one or more processors 1502 via communication links 1506 to memory hub 1505. In at least one embodiment, memory hub 1505 can be integrated into one or more processors 1502 or communication links 1506 can be provided internal to one or more processors 1502.
[0303] In at least one embodiment, processing subsystem 1501 includes one or more parallel processor(s) 1512 coupled to memory hub 1505 via a bus or other communication link 1513. In at least one embodiment, communication link 1513 can be implemented with any interconnection mechanism under which communication medium 1514 couples a storage location to a storage access unit. In at least one embodiment, communication medium 1514 is one in which communication medium properties allow efficient, favorable and strong data transfer between computing devices.
[0304] In at least one embodiment, system storage 1514 can connect to I / O hub 1507 to provide storage mechanisms for computing system 1500. In at least one embodiment, I / O switches 1516 can be used to provide an interface mechanism to enable connections between I / O hub 1507 and other components, such as network adapter 1518 and / or wireless network adapter 1519 that can be integrated into a platform, as well as various other devices that can be added via one or more add-in devices 1520. In at least one embodiment, network adapter 1518 can be an Ethernet adapter or another wired
[0305] In at least one embodiment, computing system 1500 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to I / O hub 1507. In at least one embodiment, communication paths between various components shown in FIG. 15 can be implemented using any suitable protocols, such as PCI-based protocols (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols, including those that can be employed at a system-on-a-chip. FIG. 28 In at least one embodiment, communication paths between various components in FIG. 15 can be implemented using any suitable protocols, such as NV-Link high-speed interconnect or interconnect protocols.
[0306] In at least one embodiment, parallel processor 1512 includes circuitry optimized for graphics and video processing, including for example video output circuitry, and is configured for use in a gaming console, a personal computer, or other system. In at least one embodiment, parallel processor 1512 includes circuitry optimized for general use applications, including for example high-precision floating point, integer and / or Boolean logic. In at least one embodiment, system 1500 of FIG. 15 can be used in a gaming console, a personal computer, or other system.
[0307] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Inferences can be made using one or more examples of a trained model.FIG. 28 and / or FIG. 30 Details regarding inference and / or training logic 315 are provided. In at least one embodiment, inference and / or training logic 315 can be used in system 1500 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein. FIG. 30
[0308] Processor
[0309] FIG. 3A A parallel processor 1600, according to at least one embodiment, is shown. In at least one embodiment, various components of parallel processor 1600 can be implemented using one or more integrated circuits, which can be programmable integrated circuits, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, parallel processor 1600 is a graphics processor, according to an exemplary embodiment. In at least one embodiment, parallel processor 1600 is a general purpose processor for use in a GPGPU implementation. FIG. 3B Variations of one or more parallel processors 1512 are shown.
[0310] In at least one embodiment, parallel processor 1600 includes a parallel processing unit 1602. In at least one embodiment, parallel processing unit 1602 includes an I / O unit 1604 that enables communication with other devices, including other instances of parallel processing unit 1602. In at least one embodiment, I / O unit 1604 can be directly connected to other devices. In at least one embodiment, I / O unit 1604 connects with other devices using a hub or switch interface, such as memory hub 2105. In at least one embodiment, connections between memory hub 1605 and I / O unit 1604 form a communication link 1613. In at least one embodiment, I / O unit 1604 connects with a host interface 1606 and a memory crossbar 1616, where host interface 1606 receives commands directed to processing operations and memory crossbar 1616 receives commands directed to memory operations.
[0311] In at least one embodiment, when host interface 1606 receives a command buffer via I / O unit 1604, host interface 1606 can direct a work operation to execute those commands to front end 1608. In at least one embodiment, front end 1608 is coupled with scheduler 1610, which is configured to assign commands or other work items to processing cluster array 1612. In at least one embodiment, scheduler 1610 ensures that processing cluster array 1612 is properly configured and in an active state before assigning tasks to processing cluster array 1612. In at least one embodiment, scheduler 1610 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, microcontroller- implemented scheduler 1610 is configurable to perform complex scheduling and work distribution operations with both coarse and fine grain, enabling fast preemption and context switching for threads executing on processing array 1612. In at least one embodiment, host software can prove a workload for scheduling on processing array 1612 through one of multiple graphics processing paths. In at least one embodiment, workload can then be automatically distributed by scheduler 1610 logic within a microcontroller including scheduler 1610 on processing array 1612.
[0312] In at least one embodiment, processing cluster array 1612 can include up to “N” processing clusters (e.g., cluster 1614A, cluster 1614B, through cluster 1614N), where “N” represents a positive integer (which can be a different integer than integer “N” used in other Figures). In at least one embodiment, each cluster 1614A-1614N of processing cluster array 1612 can execute a large number of concurrent threads. In at least one embodiment, scheduler 1610 can use various scheduling and / or work distribution algorithms to assign work to clusters 1614A-1614N of processing cluster array 1612, which can vary depending on workload produced by each program or type of computation. In at least one embodiment, scheduling can be handled by scheduler 1610 dynamically, or can be assisted in part by compiler logic during compilation of program logic configured for execution by processing cluster array 1612. In at least one embodiment, different clusters 1614A-1614N of processing cluster array 1612 can be allocated for processing different types of programs or for performing different types of computations.
[0313] In at least one embodiment, processing cluster array 1612 can be configured to perform a wide variety of types of parallel processing operations. In at least one embodiment, processing cluster array 1612 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 1612 can include logic to perform processing tasks including filtering of video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.
[0314] In at least one embodiment, processing cluster array 1612 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 1612 can include additional logic to support the performance of such graphics processing operations including, but not limited to, texture sampling logic to perform texture operations for 3D workflows, and tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 1612 can be configured to execute graphics processing related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel
[0315] In at least one embodiment, when parallel processing unit 1602 is used to perform graphics processing, scheduler 1610 can be configured to divide the processing workload into approximately equal sized tasks, to better enable distribution of the graphics processing operations across multiple clusters 1614A-1614N of processing cluster array 1612. In at least one embodiment, portions of processing cluster array 1612 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation,
[0316] In at least one embodiment, processing cluster array 1612 can receive processing tasks to be executed via a scheduler 1610, which can receive commands defining processing tasks from front end 1608. In at least one embodiment, processing tasks can include indices of data to be processed, e.g., surface (patch) data, raw data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., what programs to execute for various processing tasks) that define how the data should be processed. In at least one embodiment, scheduler 1610 can be configured to fetch the indices corresponding to the tasks, or can receive the indices from front end 1608. In at least one embodiment, front end 1608 can be configured to ensure that processing cluster array 1612 is configured in an effective state before launching a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).
[0317] In at least one embodiment, each of one or more instances of parallel processing unit 1602 can be coupled to parallel processor memory 1622. In at least one embodiment, parallel processor memory 1622 can be accessed by the processing clusters 1612, as well as the I / O units 1604, via memory crossbar 1616. In at least one embodiment, memory crossbar 1616 can be configured to include a set of programmable logic. In at least one embodiment, memory crossbar 1616 can be configured to implement any combination of memory operations commonly used by general-purpose processors, tensor processing units, graphics processors, neural processors, image signal processors, and / or the like. In at least one embodiment, memory crossbar 1616 can be configured to implement any combination of load, store, add, subtract, multiply, divide, bit shift, bit mask, compare, logical AND, logical OR, logical NOT, exclusive OR, exclusive NOR, NAND, NOR, XOR, XNOR, and / or the like.
[0318] In at least one embodiment, memory units 1624A-1624N can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 1624A-1624N can also include 3D stacked memory including, but not limited to, high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across memory units 1624A-1624N, allowing partition units 1620A-1620N to write portions of each rendering target in parallel to effectively use available bandwidth of parallel processor memory 1622. In at least one embodiment, local instances of parallel processor memory 1622 can be excluded in favor of a unified memory design that utilizes system memory in combination with local cache memory.
[0319] In at least one embodiment, any of clusters 1614A-1614N of processing cluster array 1612 can process data that is to be written to any of memory units 1624A-1624N within parallel processor memory 1622. In at least one embodiment, memory crossbar 1616 can be configured to transmit outputs of each cluster 1614A-1614N to any partition unit 1620A-1620N or another cluster 1614A-1614N, which can perform further processing operations on the outputs. In at least one embodiment, each cluster 1614A-1614N can communicate with memory interface 1618 through memory crossbar 1616 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 1616 has a connection to memory interface 1618 to communicate with I / O unit 1604, and a connection to a local instance of parallel processor memory 1622, enabling processing clusters 1614A-1614N within different processing clusters 1614A-1614N to communicate with system memory or other memories not local to parallel processor 1602. In at least one embodiment, memory crossbar 1616 can separate transactions according to software address or according to hardware address.
[0320] In at least one embodiment, multiple instances of parallel processing unit 1602 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1602 can be configured to operate in coordination with one another (e.g., in a clustered computing environment). In at least one embodiment, parallel processing unit 1602 of different instances can be configured to operate in a lockstep fashion to provide fault tolerance and redundancy. In at least one embodiment, parallel processing units 1602 of different instances can be configured to operate to provide parallel processing capabilities. For example, in at least one embodiment, some instances of parallel processing unit 1602 can be configured to perform general processing tasks, while other instances of parallel processing unit 1602 can be configured to perform compute tasks.
[0321] FIG. 29 is a block diagram of a partition unit 1620 in accordance with at least one embodiment. In at least one embodiment, partition unit 1620 is an instance of one of partition units 1620A-1620N of FIG. 28 In at least one embodiment, partition unit 1620 includes an L2 cache 1621, a frame buffer interface 1625, and a ROP 1626 (raster operations unit). In at least one embodiment, L2 cache 1621 is a read / write cache that is configured to perform loads and stores FIG. 28 In at least one embodiment, frame buffer interface 1625 interacts with one of memory units 1624A-1624N (e.g., within parallel processor memory 1622) in parallel processor memory in order to perform reads and writes requested by partition unit 1620.
[0322] In at least one embodiment, ROP 1626 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, ROP 1626 then outputs processed graphics data into graphics memory. In at least one embodiment, ROP 1626 includes compression logic to compress depth or color data that is written into memory and decompress depth or color data read from memory. In at least one embodiment, compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. In at least one embodiment, a type of compression performed by ROP 1626 can vary based on statistical characteristics of data to be compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis on depth and color data.
[0323] In at least one embodiment, ROP 1626 is included within each processing cluster (e.g., processing clusters 1614A-1614N) rather than in a segment unit 1620. In at least one embodiment, read and write requests to pixel data are transmitted by memory crossbar 1616 rather than pixel fragment data. In at least one embodiment, processed graphics data can be displayed on one of display device(s) 1510, routed by processor 1502 for further processing, or routed by one of processing entities within parallel processor 1600 for further processing. FIG. 28 FIG. 3A In at least one embodiment, processed graphics data can be displayed on one of display device(s) 1510, routed by processor 1502 for further processing, or routed by one of processing entities within parallel processor 1600 for further processing. FIG. 3B
[0324] FIG. 30 is a block diagram of a processing cluster 1614 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of processing clusters 1614A-1614N. In at least one embodiment, processing cluster 1614 can be configured to execute a large number of threads concurrently, where a thread FIG. 28 is an instance of one of processing clusters 1614A-1614N. In at least one embodiment, processing cluster 1614 can be configured to execute a large number of threads concurrently, where a thread refers to an instance of a particular program executed by a particular set of input data. In at least one embodiment, Single-Instruction, Multiple-Data (SIMD) instruction issue techniques are used to support execution of a large number of threads concurrently without providing multiple independent instruction units to apply instructions to multiple sets of data. In at least one embodiment, Single-Instruction, Multiple Thread (SIMT) techniques are used to support execution of a large number of generally synchronous threads, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0325] In at least one embodiment, operation of processing cluster 1614 can be controlled via a pipeline manager 1632 that distributes processing tasks to SIMT parallel processor. In at least one embodiment, pipeline manager 1632 receives instructions from FIG. 29 The scheduler 1610 of the processor 1600 receives instructions, and manages execution of those instructions, by the graphics multiprocessor 1634 and / or the texture unit(s) 1636. In at least one embodiment, graphics multiprocessor 1634 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of differing architectures can be included within processing cluster 1614. In at least one embodiment, one or more instances of graphics multiprocessor 1634 can be included within a processing cluster 1614. In at least one embodiment, graphics multiprocessor 1634 can process data, and a data crossbar 1640 can be used to distribute the processed data to one of multiple possible destinations, including other shader units. In at least one embodiment, pipeline manager 1632 can facilitate distribution by specifying destinations for processed data as a function of that processed data.
[0326] In at least one embodiment, each graphics multiprocessor 1634 within processing cluster 1614 can include an identical set of functional execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner, where new instructions can be issued before previous instructions are complete. In at least one embodiment, functional execution logic supports a variety of operations including integer and floating point arithmetic, comparison operations, Boolean operations, shift operations, and the like. In at least one embodiment, same functional- unit hardware can be leveraged to perform a number of different operations and any combination of hardware
[0327] In at least one embodiment, instructions transmitted to processing cluster 1614 form a thread for execution. In at least one embodiment, a set of threads executing across a set of parallel processing engines forms a warp. In at least one embodiment, threads within a warp are grouped into sub-sets of threads, known as a slice, that execute identical instructions as one another in parallel. In at least one embodiment, a thread can be an instantiation of a thread block, which is a grouping of threads that collectively instantiate a portion of a program to be executed on processing cluster 1614. In at least one embodiment, a thread block is bounded in that it is limited in the number of threads per thread block, where that number can be power of two. In at least one embodiment, a thread block is an atomic unit of execution that contains sufficient state to be managed independently by processing cluster 1614. In at least one embodiment, a thread block contains the state that is required for private variables that are maintained across execution of a program by processing cluster 1614. In at least one embodiment, a thread block can be bound to a number of threads, where the number of threads is the number of threads executed in parallel by processing cluster 1614. In at least one embodiment, a thread block can be bound to a number of warps, where the number of warps is the number of warps that can be executed by processing cluster 1614 in parallel. In at least one embodiment, a thread block can be bound to a number of registers, where the number of registers is the number of registers available to a thread block for use in storing register data.
[0328] In at least one embodiment, the graphics multiprocessor 1634 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 1634 may forgo the internal cache and use a cache memory within the processing cluster 1614 (e.g., L1 cache 1648). In at least one embodiment, each graphics multiprocessor 1634 may also access partition units (e.g., FIG. 28 The L2 cache is located within partition units 1620A-1620N, which are shared among all processing clusters 1614 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 1634 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of the parallel processing unit 1602 can be used as global memory. In at least one embodiment, the processing cluster 1614 includes multiple instances of the graphics multiprocessor 1634, which can share common instructions and data that can be stored in the L1 cache 1648.
[0329] In at least one embodiment, each processing cluster 1614 may include a memory management unit (“MMU”) 1645 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1645 may reside in FIG. 31 The memory interface 1618 is located within the MMU 1645. In at least one embodiment, the MMU 1645 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles and optionally to cache line indices. In at least one embodiment, the MMU 1645 may include an address translation back buffer (TLB) or a cache that may reside within the graphics multiprocessor 1634, the L1 cache 1648, or the processing cluster 1614. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, cache line indices may be used to determine whether a request for a cache line is a hit or a miss.
[0330] In at least one embodiment, processing cluster 1614 can be configured such that each graphics multiprocessor 1634 is coupled to a texture unit 1636 for performing texture mapping operations in accordance with texture coordinate values. In at least one embodiment, texture data is read from an internal texture Ll cache (not shown) or from an L2 cache (not shown) as needed. In at least one embodiment, texture data is also fetched from a graphics processor memory, such as a shared memory 1670, an L2 cache, or a system memory, as needed. In at least one embodiment, each graphics multiprocessor 1634 outputs processed tasks to data crossbar 1640 in order to provide processed task data to another processing cluster 1614 for further processing or to store processed task data in an L2 cache, a shared memory 1670, or a system memory via memory crossbar 1616. FIG. 29 In at least one embodiment, preROP 1642 is configured to receive data from graphics multiprocessors 1634, direct data to a ROP unit that can be located within or outside of processing cluster 1614, and perform optimizations related to color blending, pixel ordering, and address translations.
[0331] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3L and / or 3M. FIG. 3A and / or FIG. 3B Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3L and / or 3M. In at least one embodiment, inference and / or training logic 315 can be used in graphics processing cluster 1614 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0332] FIG. 32 A graphics multiprocessor 1634 according to at least one embodiment is shown. In at least one embodiment, graphics multiprocessor 1634 is coupled to a pipeline manager 1632 of processing cluster 1614. In at least one embodiment, graphics multiprocessor 1634 has an execution pipeline with an instruction cache 1652, an instruction unit 1654, an address mapping unit 1656, a register file 1658, one or more general-purpose GPU cores (GPGPU cores) 1662, and one or more load / store units 1666. In at least one embodiment, GPGPU cores 1662 and load / store units 1666 are coupled with cache memory 1672 and shared memory 1670 via a memory and cache interconnect 1668.
[0333] In at least one embodiment, instruction cache 1652 receives a stream of instructions 1650 to be executed by graphics processing engine 1630 from pipeline manager 1632. In at least one embodiment, instructions 1650 are cached in instruction cache 1652 and dispatched for execution by instruction unit 1654. In at least one embodiment, instruction unit 1654 can dispatch instructions to threads assigned to different ones of GPGPU cores 1662 as thread groups, for example, thread warps. In at least one embodiment, instructions can be accessed from an unified address space within a single program by specifying an address within the unified address space. In at least one embodiment, address mapping unit 1656 is used to convert an address in the unified address space into addresses used by loading / store unit 1666 to access the different memory addresses.
[0334] In at least one embodiment, register file 1658 provides a set of registers for functional units of graphics processing engine 1634. In at least one embodiment, register file 1658 provides temporary storage for operands of the data
[0335] In at least one embodiment, GPGPU cores 1662 can each include floating point
[0336] In at least one embodiment, GPGPU cores 1662 include SIMD logic capable of performing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 1662 can physically execute SIMD 4, SIMD 8, and SIMD 16 instructions and logically execute a SIMD 1, SIMD 2, and SIMD 32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated by a shader compiler during compilation of code written by a programmer. In at least one embodiment, a programmer writing code for the programmable processing unit 1600 can write SIMD code in a high level programming language, which is then compiled into multiple instruction packets that can include one or more SIMD instructions and / or one or more SIMD control instructions. In at least one embodiment, a single
[0337] In at least one embodiment, memory and cache interconnect 1668 is an interconnect network that connects each functional unit of graphics multiprocessor 1634 to register file 1658 and shared memory 1670. In at least one embodiment, memory and cache interconnect 1668 is a crossbar interconnect that allows load / store units 1666 to implement load and store operations between shared memory 1670 and register file 1658. In at least one embodiment, register file 1658 can operate at the same frequency as GPGPU cores 1662, so that there is very little latency in transferring data between GPGPU cores 1662 and register file 1658. In at least one embodiment, shared memory 1670 can be used to enable communication between threads executing on functional units within graphics multiprocessor 1634. In at least one embodiment, cache memory 1672 can be used to cache texture data that is accessed by texture unit 1636 in conjunction with normal, lighting, coordinate, and / or texture operations performed on the texture data. In at least one embodiment, shared memory 1670 can also be used for program management.
[0338] In at least one embodiment, parallel processor(s) or GPGPUs as described herein are communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, GPU(s) can be communicatively coupled to host processor / cores over a bus or other interconnect (e.g., a high-speed
[0339] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 32 and / or FIG. 33 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, inference and / or training logic 315 can be used in graphics processing unit 1634 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0340] FIG. 33A multi-GPU computing system 1700 is shown in accordance with at least one embodiment. In at least one embodiment, multi-GPU computing system 1700 can include a processor 1702 coupled to a plurality of general purpose graphics processing units (GPGPUs) 1706A-D via a host interface switch 1704. In at least one embodiment, host interface switch 1704 is a PCI Express switch device that couples processor 1702 to a PCI Express bus over which processor 1702 can communicate with GPGPUs 1706A-D. In at least one embodiment, GPGPUs 1706A-D can be interconnected via a set of high-speed P2P GPU-to-GPU links 1716. In at least one embodiment, GPU-to-GPU links 1716 connect to each of GPGPUs 1706A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 1716 enable direct communication between each GPGPU 1706A-D without having to communicate through host interface bus 1704 to which processor 1702 is connected. In at least one embodiment, where GPU-to-GPU traffic is directed to P2P GPU links 1716, host interface bus 1704 remains available for system memory access or communication with other instances of multi-GPU computing system 1700, e.g., via one or more network devices. While in at least one embodiment GPGPUs 1706A-D are connected to processor 1702 via host interface switch 1704, in at least one embodiment processor 1702 includes direct support for P2P GPU links 1716 and can be directly connected to GPGPUs 1706A-D.
[0341] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 33 and / or FIG. 33 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, inference and / or training logic 315 can be used in multi-GPU computing system 1700 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0342] FIG. 33is a block diagram of a graphics processor 1800 according to at least one embodiment. In at least one embodiment, graphics processor 1800 includes ring interconnect 1802, front-end pipeline 1804, media engine 1837, and graphics cores 1880A-1880N. In at least one embodiment, ring interconnect 1802 couples graphics processor 1800 to other processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 1800 is one of a number of processors integrated within a multi-core processing system.
[0343] In at least one embodiment, graphics processor 1800 receives batches of commands via ring interconnect 1802. In at least one embodiment, incoming commands are interpreted by a command streamer 1803 in pipeline front-end 1804. In at least one embodiment, graphics processor 1800 includes scalable execution logic to perform 3D geometry processing and media processing via the graphics cores 1880A-1880N. In at least one embodiment, for 3D geometry processing commands, command streamer 1803 supplies commands to geometry pipeline 1836. In at least one embodiment, for at least some media processing commands, command streamer 1803 supplies commands to a video front end 1834, which couples with a media engine 1837. In at least one embodiment, media engine 1837 includes a Video Quality Engine (VQE) 1830 for video and image post-processing, and a multi-format encode / decode (MFX) 1833 engine to provide hardware-accelerated
[0344] In at least one embodiment, graphics processor 1800 includes a scalable thread execution resource having (featureing) graphics cores 1880A-1880N (which can be modular and sometimes referred to as core slices), each including multiple sub-cores 1850A-1850N, 1860A-1860N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 1800 can have any number of graphics cores 1880A. In at least one embodiment, graphics processor 1800 includes graphics core(s) 1880A having at least a first sub-core 1850A and a second sub-core 1860A. In at least one embodiment, graphics processor 1800 is a low power processor with a single sub-core (e.g., 1850A). In at least one embodiment, graphics processor 1800 includes multiple graphics cores 1880A, each including a set of first sub-cores 1850A-1850N and a set of second sub-cores 1860A-1860N. In at least one embodiment, each sub-core in first sub-cores 1850A-1850N includes at least a first set of execution units 1852A-1852N and media / texture samplers 1854A-1854N. In at least one embodiment, each sub-core in second sub-cores 1860A-1860N includes at least a second set of execution units 1862A-1862N and samplers 1864A-1864N. In at least one embodiment, each sub-core 1850A-1850N, 1860A-1860N shares a set of shared resources 1870A-1870N. In at least one embodiment, shared resources include shared cache memory and pixel operation logic.
[0345] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided herein in conjunction with FIGS. 1, 6, 7, 8, 9, and 10. FIG. 33 and / or FIG. 33 Details regarding inference and / or training logic 315 are provided herein in conjunction with FIGS. 1, 6, 7, 8, 9, and 10. In at least one embodiment, inference and / or training logic 315 can be used in graphics processor 1800 to perform inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0346] FIG. 33is a block diagram illustrating micro-architecture of a processor 1900 according to at least one embodiment that can include logic circuits to execute instructions. In at least one embodiment, processor 1900 can execute instructions including x86 instructions, ARM instructions, special purpose instructions for application specific integrated circuits (ASICs), and so forth. In at least one embodiment, processor 1900 can include registers to store packed data, such as 84-bit wide MMX® registers enabled by Intel® Corporation of Santa Clara, California in microprocessors employing MMX® technology, Streaming SIMD (Single Instruction, Multiple Data) extensions (SSE), SSE2, SSE3, SSE4, AVX or higher (generically referred to as SSEx), and so forth. In at least one embodiment, processor 1900 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference. TM In at least one embodiment, MMX registers available in integer and floating point form can operate with packed data elements along with single instruction multiple data (“SIMD”) and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, 148-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX or higher (generically referred to as “SSEx”) technology can hold such packed data operands. In at least one embodiment, processor 1900 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0347] In at least one embodiment, processor 1900 includes an in-order front-end (“front-end”) 1901 to fetch instructions to be executed and to prepare instructions for execution by an out-of-order execution logic of processor pipeline. In at least one embodiment, front-end 1901 can include several units. In at least one embodiment, instruction prefetcher 1926 fetches instructions from memory and provides instructions to instruction decoder 1928 which, in turn, decodes or interprets instructions. For example, in at least one embodiment, instruction decoder 1928 decodes a received instruction into one or more operations called “micro-instructions” or “micro-operations” (also called “micro-op” or “uops”) that the machine can execute. In at least one embodiment, instruction decoder 1928 parses the instruction into an operation code that tells what is to be performed, as well as data for the operation. In at least one embodiment, tracking cache 1930 can assemble decoded micro-instructions into program ordered sequences or traces in micro-instruction queue 1934 for execution. In at least one embodiment, when tracking cache 1930 encounters a complex instruction, microcode ROM 1932 provides the micro-instructions needed to complete the operation.
[0348] In at least one embodiment, some instructions can be converted into a single micro- operation, while others can require several micro-operations to complete. In at least one embodiment, if more than four micro-instructions are needed to complete a single instruction, then instruction decoder 1928 can access microcode ROM 1932 to perform that instruction. In at least one embodiment, instructions can be decoded into a small number of micro-instructions to handle at instruction decoder 1928. In at least one embodiment, if multiple micro-instructions are needed to complete an operation, then an instruction can be stored in microcode ROM 1932. In at least one embodiment, a trace cache 1930 references an entry point programmable logic array (“PLA”) to determine a correct micro-instruction pointer for reading a microcode sequence from microcode ROM 1932 to complete one or more instructions, in accordance with at least one embodiment. In at least one embodiment, after microcode ROM 1932 completes sequencing of micro-operations for an instruction, a front end 1901 of a machine can resume fetching micro-operations from trace cache 1930.
[0349] In at least one embodiment, out-of-order execution engine (“out-of-order engine”) 1903 can prepare instructions for execution. In at least one embodiment, out-of-order execution logic has multiple buffers to smooth and reorder instruction flow to optimize performance as instructions are pipelined down and dispatched for execution. In at least one embodiment, out-of-order execution engine 1903 includes, without limitation, an allocator / register renamer 1940, a memory micro instruction queue 1942, an integer / float micro instruction queue 1944, a memory scheduler 1946, a fast scheduler 1902, a slow / general floating point scheduler (“slow / general FP scheduler”) 1904, and a simple floating point scheduler (“simple FP scheduler”) 1906. In at least one embodiment, fast scheduler 1902, slow / general floating point scheduler 1904, and simple floating point scheduler 1906 are also collectively referred to as “micro instruction schedulers 1902, 1904, 1906.” In at least one embodiment, allocator / register renamer 1940 allocates machine buffers and resources needed by each micro instruction to execute in sequence. In at least one embodiment, allocator / register renamer 1940 renames logical registers to entries in a register file. In at least one embodiment, allocator / register renamer 1940 also allocates entries for each micro instruction in one of two micro instruction queues, memory micro instruction queue 1942 for memory operations and integer / float micro instruction queue 1944 for non-memory operations, in front of memory scheduler 1946 and micro instruction schedulers 1902, 1904, 1906. In at least one embodiment, micro instruction schedulers 1902, 1904, 1906 determine when micro instructions are ready to execute based on readiness of their dependent input register operand sources and availability of execution resource micro instructions needed to complete. In at least one embodiment, fast scheduler 1902 can schedule on every half of a main clock cycle, while slow / general floating point scheduler 1904 and simple floating point scheduler 1906 can schedule once per main processor clock cycle. In at least one embodiment, micro instruction schedulers 1902, 1904, 1906 arbitrate for a dispatch port to dispatch micro instructions for execution.
[0350] In at least one embodiment, execution block 1911 includes, without limitation, integer register file / bypass network 1908, floating point register file / bypass network (“FP register file / bypass network”) 1910, address generation units (“AGUs”) 1912 and 1914, fast arithmetic logic units (“fast ALUs”) 1916 and 1918, slow arithmetic logic unit (“slow ALU”) 1920, floating point ALU (“FP”) 1922, and floating point move unit (“FP move”) 1924. In at least one embodiment, integer register file / bypass network 1908 and floating point register file / bypass network 1910 are also referred to herein as “register files 1908, 1910.” In at least one embodiment, AGUs 1912 and 1914, fast ALUs 1916 and 1918, slow ALU 1920, floating point ALU 1922, and floating point move unit 1924 are also referred to herein as “execution units 1912, 1914, 1916, 1918, 1920, 1922, and 1924.” In at least one embodiment, execution block 1911 can include, without limitation, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).
[0351] In at least one embodiment, register networks 1908, 1910 can be arranged between micro-instruction schedulers 1902, 1904, 1906 and execution units 1912, 1914, 1916, 1918, 1920, 1922, and 1924. In at least one embodiment, integer register file / bypass network 1908 performs integer operations. In at least one embodiment, floating point register file / bypass network 1910 performs floating point operations. In at least one embodiment, each of register networks 1908, 1910 can include, without limitation, a bypass network that can bypass or forward a just-completed result that has not yet been written into a register file to a new dependee. In at least one embodiment, register networks 1908, 1910 can communicate data with each other. In at least one embodiment, integer register file / bypass network 1908 can include, without limitation, two separate register files, one for low order 32 bits data, a second for high order 32 bits data. In at least one embodiment, floating point register file / bypass network 1910 can include, without limitation, 148 bit wide entries, as floating point instructions typically have operands that are 84 to 148 bits wide.
[0352] In at least one embodiment, execution units 1912, 1914, 1916, 1918, 1920, 1922, 1924 can execute instructions. In at least one embodiment, register files 1908, 1910 store integer and floating point data operand values that micro-instructions need to execute. In at least one embodiment, processor 1900 can include, without limitation, any number of execution units 1912, 1914, 1916, 1918, 1920, 1922, 1924 and combinations thereof. In at least one embodiment, floating point ALU 1922 and floating point move unit 1924 can execute floating point, MMX, SIMD, AVX and SSE, or other operations, including specialized machine learning instructions. In at least one embodiment, floating point ALU 1922 can include, without limitation, an 84-bit by 84-bit floating point divider to execute divide, square root, and remainder micro-ops. In at least one embodiment, instructions for dealing with floating point values can be handled with floating point hardware. In at least one embodiment, ALU operations can be passed to fast ALUs 1916, 1918. In at least one embodiment, fast ALUs 1916, 1918 can execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 1920 as slow ALU 1920 can include, without limitation, integer execution hardware for long latency type of operations such as multiplies, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be executed by AGUs 1912, 1914. In at least one embodiment, fast ALU 1916, fast ALU 1918, and slow ALU 1920 can execute integer operations on 84-bit data operands. In at least one embodiment, fast ALU 1916, fast ALU 1918, and slow ALU 1920 can be implemented to support a range of operand data bit sizes including sixteen, thirty-two, 148, 276, etc. In at least one embodiment, floating point ALU 1922 and floating point move unit 1924 can be implemented to support a range of operands with bit widths including 148-bit packed data operands, which can be operated on in connection with SIMD and multimedia instructions.
[0353] In at least one embodiment, micro-instruction schedulers 1902, 1904, 1906 schedule dependent operations prior to completion of parent load execution. In at least one embodiment, because micro-instructions can be speculatively scheduled and executed in processor 1900, processor 1900 can also include logic to handle memory misses. In at least one embodiment, if a data load in a data cache misses, there can be dependent operations running in a pipeline that cause a scheduler to temporarily have incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, dependent operations can need to be replayed and independent operations can be allowed to complete. In at least one embodiment, a scheduler and replay mechanism of at least one embodiment of a processor can also be designed to capture instruction sequences for text string compare operations.
[0354] In at least one embodiment, a “register” can refer to an on-board processor storage location that can be used as part of an instruction that identifies an operand. In at least one embodiment, a register can be one that can be used from outside of a processor (from a programmer’s perspective). In at least one embodiment, a register can not be limited to a particular type of circuit. Rather, in at least one embodiment, a register can store data, provide data, and perform functions described herein. In at least one embodiment, registers described herein can be implemented by circuitry within a processor using a variety of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 34-bit integer data. A register file of at least one embodiment also contains eight multimedia SIMD registers for packing data.
[0355] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 32 and / or FIG. 32 Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. In at least one embodiment, portions of inference and / or training logic 315 can be incorporated into execution block 1911 and other memory or registers shown or not shown. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more ALUs shown in execution block 1911. Further, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure ALUs of execution block 1911 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0356] FIG. 32A deep learning application processor 2000 is shown, in accordance with at least one embodiment. In at least one embodiment, deep learning application processor 2000 uses instructions that, if executed by deep learning application processor 2000, cause deep learning application processor 2000 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, deep learning application processor 2000 is an application specific integrated circuit (ASIC). In at least one embodiment, application processor 2000 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions or both. In at least one embodiment, deep learning application processor 2000 includes, without limitation, processing clusters 2010(1)-2010(12), inter-chip links ("ICLs") 2020(1)-2020(12), inter-chip controllers ("ICCs") 2030(1)-2030(2), second generation high bandwidth memory ("HBM2") 2040(1)-2040(4), memory controllers ("Mem Ctrlrs") 2042(1)-2042(4), high bandwidth memory physical layers ("HBM PHYs") 2044(1)-2044(4), management controller central processing units ("management controller CPUs") 2050, serial peripheral interfaces, internal integrated circuits, and general purpose input / output blocks ("SPI, I2C, GPIO") 2060, peripheral component interconnect express controllers and direct memory access blocks ("PCIe controllers and DMA") 2070, and sixteen lane peripheral component interconnect express ports ("PCI Express x 18") 2080.
[0357] In at least one embodiment, processing clusters 2010 can perform deep learning operations, including inferencing or prediction operations based on weight parameters calculated based on one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 2010 can include, without limitation, any number and type of processors. In at least one embodiment, deep learning application processor 2000 can include any number and type of processing clusters 2000. In at least one embodiment, inter-chip links 2020 are bidirectional. In at least one embodiment, inter-chip links 2020 and inter-chip controllers 2030 enable multiple deep learning application processors 2000 to exchange information, including activation information resulting from execution of one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 2000 can include any number (including zero) and type of ICLs 2020 and ICCs 2030.
[0358] In at least one embodiment, HBM2 2040 provides a total of 34 GB of memory. In at least one embodiment, HBM2 2040(i) is associated with both a memory controller 2042(i) and an HBM PHY 2044(i), where “i” is any integer. In at least one embodiment, any number of HBM2s 2040 can provide any type and total amount of high bandwidth memory and can be associated with any number (including zero) and type of memory controllers 2042 and HBM PHYs 2044. In at least one embodiment, SPI, I2C, GPIO 3360, PCIe controller 2060, and DMA 2070 and / or PCIe 2080 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.
[0359] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, 3C, 3D, 3E, 3F, 3G, 3H, 3I, 3J, 3K, and 3L. FIG. 32 and / or FIG. 32 Details regarding inference and / or training logic 315 are provided. In at least one embodiment, deep learning application processor is used to train a machine learning model (e.g., neural network) to predict or infer information provided to deep learning application processor 2000. In at least one embodiment, deep learning application processor 2000 is used to infer or predict information based on a trained machine learning model (e.g., neural network) that has been trained by another processor or system or by deep learning application processor 2000. In at least one embodiment, processor 2000 can be used to perform one or more neural network use cases described herein.
[0360] FIG. 36Bis a block diagram of a neuromorphic processor 2100, in accordance with at least one embodiment. In at least one embodiment, neuromorphic processor 2100 can receive one or more inputs from a source external to neuromorphic processor 2100. In at least one embodiment, these inputs can be transmitted to one or more neurons 2102 within neuromorphic processor 2100. In at least one embodiment, neurons 2102 and components thereof can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, neuromorphic processor 2100 can include, without limitation, thousands or millions of instances of neurons 2102, although any suitable number of neurons 2102 can be used. In at least one embodiment, each instance of neurons 2102 can include neuron input 2104 and neuron output 2106. In at least one embodiment, neurons 2102 can generate outputs that can be transmitted to inputs of other instances of neurons 2102. In at least one embodiment, neuron input 2104 and neuron output 2106 can be interconnected via synapses 2108.
[0361] In at least one embodiment, neurons 2102 and synapses 2108 can be interconnected such that neuromorphic processor 2100 operates to process or analyze information received by neuromorphic processor 2100. In at least one embodiment, a neuron 2102 can send out an output spike (or “spike” or “peak”) when input received through neuron input 2104 exceeds a threshold value. In at least one embodiment, neuron 2102 can sum or integrate signals received at neuron input 2104. For example, in at least one embodiment, neuron 2102 can be implemented as a leaky integrate-and-fire neuron, where neuron 2102 can produce an output (or “spike”) using a transfer function such as a sigmoid or threshold function if a sum (referred to as “membrane potential”) exceeds a threshold value. In at least one embodiment, a leaky integrate-and-fire neuron can sum signals received at neuron input 2104 into a membrane potential, and can apply a program decay factor (or leak) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-fire neuron can spike if multiple input signals are received at neuron input 2104 fast enough to exceed a threshold value (i.e., before the membrane potential decays too low to spike). In at least one embodiment, neuron 2102 can be implemented using circuitry or logic that receives input, integrates input into a membrane potential, and decays the membrane potential. In at least one embodiment, input can be averaged, or any other suitable transfer function can be used. Furthermore, in at least one embodiment, neuron 2102 can include, without limitation, comparator circuitry or logic that produces an output spike at neuron output 2106 when a result of applying a transfer function to neuron input 2104 exceeds a threshold value. In at least one embodiment, once neuron 2102 spikes, it can ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 2102 can resume normal operation after a suitable period of time (or refractory period).
[0362] In at least one embodiment, neurons 2102 can be interconnected by synapses 2108. In at least one embodiment, synapses 2108 can operate to transmit a signal from an output of a first neuron 2102 to an input of a second neuron 2102. In at least one embodiment, a neuron 2102 can transmit information over more than one instance of a synapse 2108. In at least one embodiment, one or more instances of neuron output 2106 can be connected through an instance of synapse 2108 to an instance of neuron input 2104 in the same neuron 2102. In at least one embodiment, an instance of a neuron 2102 that produces an output to be transmitted over an instance of a synapse 2108 can be referred to as a “presynaptic neuron” with respect to that instance of synapse 2108. In at least one embodiment, an instance of a neuron 2102 that receives input transmitted over an instance of a synapse 2108 can be referred to as a “postsynaptic neuron” with respect to that instance of synapse 2108. In at least one embodiment, with respect to various instances of synapse 2108, because an instance of a neuron 2102 can receive input from one or more instances of synapse 2108 and can also transmit output through one or more instances of synapse 2108, a single instance of a neuron 2102 can be both a “presynaptic neuron” and a “postsynaptic neuron”.
[0363] In at least one embodiment, neurons 2102 can be organized into one or more layers. In at least one embodiment, each instance of neuron 2102 can have one neuron output 2106 that can fan out to one or more neuron inputs 2104 through one or more synapses 2108. In at least one embodiment, neuron outputs 2106 of neurons 2102 in a first layer 2110 can be connected to neuron inputs 2104 of neurons 2102 in a second layer 2112. In at least one embodiment, layer 2110 can be referred to as a “feedforward layer”. In at least one embodiment, each instance of neuron 2102 in an instance of first layer 2110 can fan out to each instance of neuron 2102 in a second layer 2112. In at least one embodiment, first layer 2110 can be referred to as a “fully connected feedforward layer”. In at least one embodiment, each instance of neuron 2102 in each instance of second layer 2112 fans out to fewer than all instances of neuron 2102 in a third layer 2114. In at least one embodiment, second layer 2112 can be referred to as a “sparsely connected feedforward layer”. In at least one embodiment, neurons 2102 in second layer 2112 can fan out to neurons 2102 in multiple other layers, including to neurons 2102 in second layer 2112. In at least one embodiment, second layer 2112 can be referred to as a “recurrent layer”. In at least one embodiment, neuromorphic processor 2100 can include any suitable combination of recurrent and feedforward layers, including but not limited to sparsely connected feedforward layers and fully connected feedforward layers.
[0364] In at least one embodiment, neuromorphic processor 2100 can include, without limitation, a reconfigurable interconnect architecture or a dedicated hardwired interconnect to connect synapses 2108 to neurons 2102. In at least one embodiment, neuromorphic processor 2100 can include, without limitation, circuitry or logic that, depending on a neural network topology and neuron fan-in / fan-out, allows synapses to be allocated to different neurons 2102 as needed. For example, in at least one embodiment, synapses 2108 can be connected to neurons 2102 using an interconnect structure such as a network-on-chip or through dedicated connections. In at least one embodiment, synapse interconnects and components thereof can be implemented using circuitry or logic.
[0365] FIG. 34A processing system is shown in accordance with at least one embodiment. In at least one embodiment, system 2200 includes one or more processor(s) 2202 and one or more graphics processor(s) 2208, and can be a single processor desktop system, a multiprocessor workstation system, or a server system having many processor(s) 2202 or processor core(s) 2207. In at least one embodiment, system 2200 is a processing platform incorporated within a system on a chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0366] In at least one embodiment, system 2200 can include or be coupled to a graphics processing unit(s) 2208 (GPU(s) 2208) or graphics processing unit(s) on a processor(s) 2202 (processor(s) 2202 GPUs). In at least one embodiment, system 2200 is a server, mobile, handheld, or embedded device that includes one or more instances of processor(s) 2202 and graphics processor(s) 2208. In at least one embodiment, system 2200 is a processor(s) 2202 or graphics processor(s) 2208 that is integrated within a system on a chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0367] In at least one embodiment, processor(s) 2202 each include one or more processor cores 2207 to process instructions which, when executed, implement the operations for system and user software. In at least one embodiment, each of the one or more processor cores 2207 is configured to process a specific instruction sequence 2209. In at least one embodiment, instruction sequence 2209 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via a very long instruction word (VLIW). In at least one embodiment, processor cores 2207 can each process different instruction sequences 2209 that can include instructions to facilitate emulation of other instruction sequences. In at least one embodiment, processor cores 2207 can also include other processing devices, such as a digital signal processor (DSP).
[0368] In at least one embodiment, processor 2202 includes cache memory 2204. In at least one embodiment, processor 2202 can have a single level of internal cache or multiple levels of internal caches. In at least one embodiment, cache memory is shared among multiple components of processor 2202. In at least one embodiment, processor 2202 also uses an external cache (e.g., a level three (L3) cache, or last level cache (LLC)) (not shown), which can be shared between processor cores 2207 using known cache coherency techniques. In at least one embodiment, additionally included in processor 2202 are register files 2206, which processor can include different types of registers to store different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, register files 2206 can include general registers or other registers.
[0369] In at least one embodiment, one or more processors 2202 are coupled with one or more interface buses 2210 to transmit communications signals between processor 2202 and other components in system 2200, for example, address, data, or control signals. In at least one embodiment, interface bus 2210 can be a processor bus, such as a version of direct media interface (DMI) bus. In at least one embodiment, interface bus 2210 is not limited to DMI bus, and can include one or more peripheral component interconnect buses (e.g., a PCI, a PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, processor 2202 includes an integrated memory controller 2216 and platform controller hub 2230. In at least one embodiment, memory controller 2216 facilitates communication between memory devices and other components of processing system 2200, while platform controller hub 2230 provides connections to input / output (I / O) devices via local I / O bus.
[0370] In at least one embodiment, memory device 2220 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, or a phase change memory device, among others. In at least one embodiment, memory device 2220 can be a system memory of processing system 2200, to store data 2222 and instructions 2221 for use when one or more processors 2202 executes an application or process. In at least one embodiment, memory controller 2216 also couples with an optional external graphics processor 2212, which can communicate with one or more graphics processors 2208 in processors 2202 to perform graphics and media operations.
[0371] In at least one embodiment, platform controller hub 2230 enables peripherals coupled to bridge 2250 to interact with a processor and / or each other over high-speed I / O buses 2220 and 2210. In at least one embodiment, I / O peripherals include, but are not limited to, an audio controller 2246, a network controller 2234, a firmware interface 2228, a wireless transceiver 2226, a touch sensor 2225, a data storage device 2224 (e.g., solid-state drive (SSD), floppy drive, optical drive, etc.), a graphics processor 2214, and a high-definition multimedia
[0372] In at least one embodiment, memory controller 2216 and instances of platform controller hub 2230 can be integrated into a discrete external graphics processor, such as external graphics processor 2212. In at least one embodiment, platform controller hub 2230 and / or memory controller 2216 can be external to one or more processor(s) 2202. For example, in at least one embodiment, system 2200 can include an external memory controller 2216 and platform controller hub 2230, which can be configured as a memory controller hub and a peripheral controller hub within a system chipset that is distinct from processor(s) 2202.
[0373] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 and 6. FIG. 34 and / or FIG. 34Details regarding the inference and / or training logic 315 are provided. In at least one embodiment, some or all of inference and / or training logic 315 can be incorporated with graphics processor 2200. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs embodied in a 3D pipeline. Further, in at least one embodiment, the inference and / or training operations described herein can be accomplished with logic other than that shown. FIG. 34 or FIG. 35A-35B In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of graphics processor 2200 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0374] FIG. 35A is a block diagram of a processor 2300 having one or more processor cores 2302A-2302N, an integrated memory controller 2314, and an integrated graphics processor 2308, according to at least one embodiment. In at least one embodiment, processor 2300 can include additional cores, up to and including an additional core 2302N denoted in phantom in FIG. 23. In at least one embodiment, each processor core 2302A-2302N includes one or more internal cache units 2304A-2304N. In at least one embodiment, each processor core can also access one or more shared cache units 2306.
[0375] In at least one embodiment, internal cache units 2304A-2304N and shared cache unit 2306 represent a cache memory hierarchy within processor 2300. In at least one embodiment, cache memory units 2304A-2304N can include at least one level of instruction and data caches per processor core and one or more shared level caches, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other level cache, where the highest level of cache prior to main memory is classified as an LLC. In at least one embodiment, cache coherence logic maintains coherence between various cache units 2306 and 2304A-2304N.
[0376] In at least one embodiment, processor 2300 also includes a set of one or more bus controller units 2316 and a system agent core 2310. In at least one embodiment, one or more bus controller units 2316 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, system agent core 2310 provides management functionality for various processor components. In at least one embodiment, system agent core 2310 includes one or more integrated memory controllers 2314 to manage access to various external memory devices (not shown), including support for data bus protocols such as DDR DRAM.
[0377] In at least one embodiment, one or more processor cores 2302A-2302N include support to run in multiple threads simultaneously. In at least one embodiment, system agent core 2310 includes components for coordination and operation of cores 2302A-2302N during multi-threaded processing. In at least one embodiment, system agent core 2310 can additionally include a power control unit (PCU), including logic and components to govern one or more power states of processor cores 2302A-2302N and graphics processor 2308.
[0378] In at least one embodiment, processor 2300 also includes graphics processor 2308, which can be configured to perform a graphics processing applications or a graphics processing- intensive applications. In at least one embodiment, graphics processor 2308 couples with shared cache unit 2306 and system agent core 2310, including one or more integrated memory controllers 2314. In at least one embodiment, system agent core 2310 further includes a display controller 2311 for driving one or more coupled displays to present graphics processor output to a
[0379] In at least one embodiment, ring based interconnect unit 2312 is used to couple the internal components of the processor 2300. In at least one embodiment, an alternative interconnect unit can be used, such as a point-to-point interconnect, a switched interconnect, or other technology. In at least one embodiment, graphics processor 2308 couples with ring interconnect 2312 through I / O link 2313.
[0380] In at least one embodiment, I / O link 2313 represents at least one of a variety of I / O interconnects, including a package I / O interconnect that facilitates communication between various processor components and a high performance embedded memory module 2318 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 2302A-2302N and graphics processor 2308 uses embedded memory module 2318 as a shared last level cache.
[0381] In at least one embodiment, processor cores 2302A-2302N are homogenous cores executing a common instruction set architecture. In at least one embodiment, processor cores 2302A-2302N are heterogeneous with respect to instruction set architecture (ISA) in that one or more processor cores 2302A-2302N execute a common instruction set while one or more other processor cores 2302A-2302N execute a subset or a different instruction set. In at least one embodiment, processor cores 2302A-2302N are heterogeneous with respect to microarchitecture in that one or more cores have a relatively high power consumption coupled with one or more power cores having a lower power consumption. In at least one embodiment, processor 2300 can be implemented on or as a SoC integrated circuit.
[0382] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, 3C, 3D, 3E, 3F, 3G, 3H, 3I, 3J, 3K, and 3L. In various embodiments, inference and / or training logic 315 can be used in place of or in conjunction with inference and / or training logic 310 described above. FIG. 35B and / or FIG. 36A Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, 3C, 3D, 3E, 3F, 3G, 3H, 3I, 3J, 3K, and 3L. In various embodiments, inference and / or training logic 315 can be used in place of or in conjunction with inference and / or training logic 310 described above. FIG. 33 In at least one embodiment, inference and / or training logic 315 can be used in place of or in conjunction with graphics processor 2308 described above. In at least one embodiment, inference and / or training logic 315 uses one or more of 3D pipeline, graphics core 2302, shared function logic, or other logic to perform one or more of the training and / or inferencing techniques described herein. Further, in at least one embodiment, inference and / or training operations described herein can be accomplished using logic other than that shown in FIG. 32 or FIG. 32 In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of processor 2300 to perform one or more of machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0383] FIG. 32is a block diagram for a graphics processor 2400 that can be a discrete graphics processing unit, or can be graphics processor integrated with a multiple core processor. In at least one embodiment, the graphics processor 2400 communicates with registers on the graphics processor 2400 using a memory mapped I / O interface 2408. In at least one embodiment, graphics processor 2400 includes a memory interface 2414 to access a memory such as a shared L2 cache, shared L3 cache, system memory, or a combination thereof.
[0384] In at least one embodiment, graphics processor 2400 also includes a display controller 2402 to drive a display 2420 coupled to the graphics processor 2400. In at least one embodiment, display controller 2402 includes hardware for one or more overlay planes for the display 2420, as well as hardware for composition of multiple layers of video or user interface elements. In at least one embodiment, display 2420 can be an internal display 2420 or an external display coupled to the graphics processor 2400. In at least one embodiment, display 2420 is a head mounted display, such as a virtual reality (VR) display or an augmented reality (AR) display. In at least one embodiment, graphics processor 2400 includes a video codec engine 2406 to encode, decode, or transcode media to, from, or between one or more media encoding formats, including, but not limited to Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4 AVC, and Society of Motion Picture
[0385] In at least one embodiment, graphics processor 2400 includes a block image transfer (BLIT) engine 2404 to perform two-dimensional (2D) rasterizer operations including, for example, bit-boundary block transfers. However, in at least one embodiment, 2D graphics operations are performed using one or more components of graphics processing engine (GPE) 2410. In at least one embodiment, GPE 2410 is a compute engine to perform graphics processing operations including three-dimensional (3D) graphics processing operations and media processing operations.
[0386] In at least one embodiment, GPE 2410 includes a 3D pipeline 2412 for performing 3D operations, such as rendering three-dimensional graphics and images. In at least one embodiment, 3D pipeline 2412 includes programmable and fixed function elements that perform various tasks necessary to render images and graphics from 3D models. In at least one embodiment, while 3D pipeline 2412 can be used to perform media operations, in at least one embodiment GPE 2410 also includes a media pipeline 2416 to perform media operations such as video post-processing and image enhancements.
[0387] In at least one embodiment, media pipeline 2416 includes fixed function or programmable logic for performing one or more specialized media operations, such as video decode acceleration, video de-interlacing, and video codec acceleration, in place of or on behalf of video codec engine 2406. In at least one embodiment, media pipeline 2416 also includes a thread spawning unit to spawn threads to perform computations for the media operations in 3D / Media subsystem 2415. In at least one embodiment, spawned threads run computations for media operations in place of or in addition to the computations performed by fixed function and programmable logic elements in media pipeline 2416.
[0388] In at least one embodiment, 3D / Media subsystem 2415 includes logic to execute threads spawned by 3D pipeline 2412 and media pipeline 2416. In at least one embodiment, 3D pipeline 2412 and media pipeline 2416 send thread execution requests to 3D / Media subsystem 2415, which includes thread dispatch logic to arbitrate various requests and dispatch them to available thread execution resources. In at least one embodiment, execution resources include an array of graphics execution units to process 3D and media threads. In at least one embodiment, 3D / Media subsystem 2415 includes one or more internal caches to cache it thread instructions and data. In at least one embodiment, subsystem 2415 also includes shared memory, which is shared among various processing clusters and includes register space for each logic element.
[0389] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, and 6. FIG. 36B and / or FIG. 36B Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3 A, 3B, and 6. In at least one embodiment, portions or all of inference and / or training logic 315 can be incorporated into processor 2400. For example, in at least one embodiment, training and / or inference techniques described herein can use one or more of ALUs included in 3D pipeline 2412. Moreover, in at least one embodiment, inference and / or training operations described herein can use one or more of ALUs included in media pipeline 2416.FIG. 3A or FIG. 3B logic other than that shown in FIG. 9. In at least one embodiment, the weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure ALUs of graphics processor 2400 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0390] FIG. 9 FIG. 9 is a block diagram of a graphics processing engine 2510 of a graphics processor of FIG. 8 according to at least one embodiment. In at least one embodiment, a graphics processing engine (GPE) 2510 is a version of GPE 2410 shown in FIG. 8. In at least one embodiment, a media pipeline 2516 is optional and can not be explicitly included in GPE 2510. In at least one embodiment, a separate media and / or image processor is coupled to GPE 2510. FIG. 37
[0391] In at least one embodiment, GPE 2510 is coupled with or includes a command streamer 2503 that provides a command stream to 3D pipeline 2512 and / or media pipeline 2516. In at least one embodiment, command streamer 2503 is coupled to memory, which can be system memory, internal cache memory, and shared cache memory, among others. In at least one embodiment, command streamer 2503 receives commands from memory and sends commands to 3D pipeline 2512 and / or media pipeline 2516. In at least one embodiment, commands are instructions, primitives, or micro-operations fetched from a ring buffer, which stores commands for 3D pipeline 2512 and media pipeline 2516. In at least one embodiment, ring buffer can also include batch command buffers that store batches of commands. In at least one embodiment, commands for 3D pipeline 2512 can also include references to data stored in memory, such as, but not limited to, vertex and geometry data for 3D pipeline 2512 and image data and memory objects for media pipeline 2516. In at least one embodiment, 3D pipeline 2512 and media pipeline 2516 process commands and data by executing them or by dispatching them for execution on graphics core array 2514. In at least one embodiment, graphics core array 2514 includes one or more graphics core blocks (e.g., one or more graphics cores 2515A, one or more graphics cores 2515B), each including one or more graphics cores. In at least one embodiment, each graphics core includes a set of graphics execution resources, including general and graphics specific execution logic to perform graphics and compute operations, and fixed function texture processing and / or machine learning and artificial intelligence acceleration logic, includingFIG. 37 and FIG. 37 Inference and / or training logic 315 may, individually or collectively, be used to implement inference and / or training logic 315 of FIG. 10.
[0392] In at least one embodiment, 3D pipeline 2512 includes fixed function and programmable logic to process one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to graphics core array 2514. In at least one embodiment, graphics core array 2514 provides unified execution resources for processing shader programs. In at least one embodiment, multipurpose execution logic (e.g., execution units) within graphics core array 2514 includes support for variable data size operations and multiple thread control blocks with context-specific thread control
[0393] In at least one embodiment, graphics core array 2514 also includes execution logic to perform media functions such as video and / or image processing. In at least one embodiment, execution units include fixed function or programmable logic that is optimized to perform
[0394] In at least one embodiment, output data can be output to memory in a unified return buffer (URB) 2518, the output data generated by threads executing on graphics core array 2514. In at least one embodiment, URB 2518 can store data for multiple threads. In at least one embodiment, URB 2518 can be used to send data between different threads executing on graphics core array 2514. In at least one embodiment, URB 2518 can also be used for synchronization between threads on graphics core array 2514 and fixed function logic within shared function logic 2520.
[0395] In at least one embodiment, graphics core array 2514 is scalable, such that graphics core array 2514 includes variable numbers of graphics cores, each of which include variable numbers of execution units based on target performance and power levels of GPE 2510. In at least one embodiment, execution resources are dynamically scalable such that execution resources can be enabled or disabled as needed.
[0396] In at least one embodiment, graphics core array 2514 is coupled to shared function logic 2520, which includes a number of resources shared among graphics cores in graphics core array 2514. In at least one embodiment, shared functions performed by shared function logic 2520 are embodied in hardware logic units that provide specialized supplemental functionality to graphics core array 2514. In at least one embodiment, shared function logic 2520 includes, without limitation, sampler units 2521, math units 2522, and inter-thread communication (ITC) logic 2523. In at least one embodiment, one or more caches 2525 are included in or coupled to shared function logic 2520.
[0397] In at least one embodiment, shared functions are used if demand for a specialized function is insufficient to include in graphics core array 2514. In at least one embodiment, a single instance of a specialized function is used in shared function logic 2520 and shared among other execution resources within graphics core array 2514. In at least one embodiment, a particular shared function can be included within shared function logic 2816 within graphics core array 2514 that is within shared function logic 2520 that is used extensively by graphics core array 2514. In at least one embodiment, shared function logic 2816 within graphics core array 2514 can include some or all of the logic within shared function logic 2520. In at least one embodiment, all logical elements within shared function logic 2520 can be replicated within shared function logic 2526 of graphics core array 2514. In at least one embodiment, shared function logic 2520 is excluded in favor of shared function logic 2526 within graphics core array 2514.
[0398] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 315 are provided below in conjunction with FIGS. 3A and / or 3B. In at least one embodiment, portions or all of inference and / or training logic 315 can be incorporated in graphics processor 2510. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more of ALUs embodied in 3D pipeline 2512, graphics core 2515, shared function logic 2526, shared function logic 2520, or other logic included in graphics processor 2510. In at least one embodiment, portions or all of inference and / or training logic 315 can be incorporated within graphics processor 2510. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more of ALUs embodied in 3D pipeline 2512, graphics core 2515, shared function logic 2526, shared function logic 2520, or other logic included in graphics processor 2510. In at least one embodiment, portions or all of inference and / or training logic 315 can be incorporated within graphics processor 2510. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more of ALUs embodied in 3D pipeline 2512, graphics core 2515, shared function logic 2526, shared function logic 2520, or other logic included in graphics processor 2510. or Logic outside of the illustrated logic to accomplish. In at least one embodiment, the weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure ALUs of graphics processor 2510 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0399] is a block diagram of hardware logic of a graphics processor core 2600 in accordance with at least one embodiment described herein. In at least one embodiment, graphics processor core 2600 is included within a graphics core array. In at least one embodiment, graphics processor core 2600 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, graphics processor core 2600 is an example of one graphics core slice and graphics processors described herein can include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 2600 can include fixed function blocks 2630, also known as sub-slices, coupled with multiple sub-cores 2601A-2601F that include modules of general-purpose and fixed function logic.
[0400] In at least one embodiment, fixed function blocks 2630 include geometry and fixed function pipeline 2636, for example, that can be shared by all sub-cores in graphics processor 2600 in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry and fixed function pipeline 2636 includes a 3D fixed function pipeline, a video front-end unit, thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.
[0401] In at least one embodiment, fixed function blocks 2630 also include a graphics SoC interface 2637, a graphics microcontroller 2638, and a media pipeline 2639. In at least one embodiment, graphics SoC interface 2637 provides an interface between graphics core 2600 and other processor cores within a System on a Chip (SoC) integrated circuit. In at least one embodiment, graphics microcontroller 2638 is a programmable sub-processor that is configurable to manage various functions of graphics processor 2600, including thread dispatch, scheduling, and pre-emption.
[0402] In at least one embodiment, SoC interface 2637 enables graphics core 2600 to communicate with general application processor cores (e.g., CPUs) and / or other components within the SoC, including memory hierarchy elements such as a shared last level cache, system RAM, and / or embedded on-chip or package DRAM. In at least one embodiment, SoC interface 2637 can also enable communication with fixed function devices within the SoC, such as a camera image signal processor (ISP) and / or a video codec, and to use and / or implement global memory atoms that can be shared between graphics core 2600 and a CPU within the SoC. In at least one embodiment, graphics SoC interface 2637 also implements power management controls for graphics processor core 2600 and enables an interface between a clock domain of a graphics processor core 2600 and other clock domains within the SoC. In at least one embodiment, SoC interface 2637 enables receipt of command buffers from a command streamer and global thread dispatcher that are configured to provide commands and instructions to each of one or more graphics cores within graphics processor.
[0403] In at least one embodiment, graphics microcontroller 2638 can be configured to perform various scheduling and management tasks for graphics core 2600. In at least one embodiment, graphics microcontroller 2638 can perform graphics and / or compute workload scheduling on various graphics processing engines within execution unit (EU) arrays 2602A-2602F, 2604A-2604F in sub-cores 2601A-2601F. In at least one embodiment, host software executing on CPU cores of an SoC including graphics core 2600 can submit workloads for one of multiple graphics processor paths, which invoke scheduling operations on appropriate graphics engines. In at least one embodiment, scheduling operations include determining which workload to run next, submitting a workload to a command streamer, pre-empting existing workloads running on an engine, monitoring progress of a workload, and informing host software when a workload is complete. In at least one embodiment, graphics microcontroller 2638 can also facilitate low power or idle states for graphics core 2600, providing the ability to save and restore registers across low power states independently of operating systems and / or graphics driver software on the system.
[0404] In at least one embodiment, graphics core 2600 can have up to N more or less modular cores than shown in FIG. 26A. For each set of N cores, graphics core 2600 can also include shared function logic 2610, shared and / or cache memory 2612, geometry / fixed function pipeline 2614, and additional fixed function logic 2616 to accelerate various graphics and compute operations. In at least one embodiment, shared function logic 2610 can include logic units (e.g., samplers, math, and / or inter-thread communication logic) that are shared by each N core within graphics core 2600. In at least one embodiment, shared and / or cache memory 2612 can be a last level cache memory for N cores 2601A-2601F within graphics core 2600, and can also be used as shared memory that can be accessed by multiple cores. In at least one embodiment, geometry / fixed function pipeline 2614 can be included instead of geometry / fixed function pipeline 2636 within fixed function block 2630, and can include similar logic units.
[0405] In at least one embodiment, graphics core 2600 includes additional fixed function logic 2616 that can include various fixed function acceleration logic used by graphics core 2600. In at least one embodiment, additional fixed function logic 2616 includes an additional geometry pipeline used for position only shading. In position only shading, there are at least two geometry pipelines, while in full geometry and fixed function pipeline 2614, 2636, and cull pipeline, which is a trimmed down version of full geometry pipeline that can be included in additional fixed function logic 2616. In at least one embodiment, cull pipeline is a trimmed down version of full geometry pipeline. In at least one embodiment, full and cull pipelines can execute different instances of an application, each with separate state. In at least one embodiment, position only shading can hide long cull runs of triangles that are discarded, which can complete shading earlier in some cases. For example, in at least one embodiment, cull pipeline logic in additional fixed function logic 2616 can execute position shaders in parallel with a main application, and often generate critical results faster than full pipeline because cull pipeline takes and shades position attributes of vertices without performing rasterization and rendering pixels to a frame buffer. In at least one embodiment, cull pipeline can use generated critical results to compute visibility information for all triangles, regardless of whether those triangles are culled or not. In at least one embodiment, full pipeline, which can be referred to as replay pipeline in this case, can consume visibility information to skip culled triangles to only shade visible triangles that are ultimately passed to rasterization stage.
[0406] In at least one embodiment, additional fixed function logic 2616 can also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementing optimizations including for machine learning training or inferencing.
[0407] In at least one embodiment, within each graphics sub-core 2601A-2601F includes a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests by graphics pipeline, media pipeline, or shader program. In at least one embodiment, graphics sub-cores 2601A-2601F include multiple arrays of execution units 2602A-2602F, 2604A-2604F, thread dispatch and inter-thread communication (TD / IC) logic 2603A-2603F, 3D (e.g., texture) samplers 2605A-2605F, media samplers 2606A-2606F, shader processors 2607A-2607F, and shared local memory (SLM) 2608A-2608F. In at least one embodiment, arrays of execution units 2602A-2602F, 2604A-2604F each include multiple execution units that are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logic operations, including graphics, media, and compute operations in response to graphic, media, and compute shader programs stored in memory. In at least one embodiment, TD / IC logic 2603A-2603F performs local thread dispatch and thread control operations for execution units within a sub-core and facilitate
[0408] Inference and / or training logic 315 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inferencing and / or training operations include machine learning training operations or machine learning inferencing operations, or both. In at least one embodiment, inference and / or training logic 315 include, without limitation, hardware and / or software. In at least one embodiment, inference and / or training logic 315 include, without limitation, one or more of processors 302, storage 304, and / or memory 306. and / or Details regarding the inference and / or training logic 315 are provided. In at least one embodiment, portions of inference and / or training logic 315 can be incorporated into graphics processor 2610. For example, in at least one embodiment, training and / or inference techniques described herein can be performed using one or more ALUs embodied in 3D pipeline, graphics microcontroller 2638, geometry and fixed function pipeline 2614 and 2636, or other logic in Additionally, in at least one embodiment, inference and / or training operations described herein can be done using logic other than that shown in or In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of graphics processor 2600 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0409] Thread execution logic 2700 including an array of processing elements that each include a graphics processor core is shown in accordance with at least one embodiment. At least one embodiment is shown in which thread execution logic 2700 is used. Exemplary internal details of graphics processing unit 2708 are shown in accordance with at least one embodiment.
[0410] As As shown in FIG. 27, in at least one embodiment, thread execution logic 2700 includes shader processor 2702, thread dispatcher 2704, instruction cache 2706, scalable execution unit array including a plurality of execution units 2707A-2707N and 2708A-2708N, sampler 2710, data cache 2712, and data port 2714. In at least one embodiment, the scalable execution unit array can be dynamically scalable, to permit varying numbers of execution units to be used as found appropriate for the particular load 5 requiring execution. In at least one embodiment, the execution units can include floating point units, arithmetic logic units, and the like. In at least one embodiment, the execution units can include single program multiple data instruction execution units and / or multiple program multiple data instruction execution units. In at least one embodiment, execution units can be configured to execute multiple threads concurrently.
[0411] In at least one embodiment, execution units 2707 and / or 2708 are primarily used for executing shader programs. In at least one embodiment, shader processor 2702 can process various shader programs and dispatch execution threads associated with shader programs via thread dispatcher 2704. In at least one embodiment, thread dispatcher 2704 includes logic to arbitrate thread initiation requests from graphics and media pipelines and instantiate requested threads on one or more execution units 2707 and / or 2708. For example, in at least one embodiment, a geometry pipeline can dispatch a vertex, tessellation or geometry shader to thread execution logic for processing. In at least one embodiment, thread dispatcher 2704 can also handle runtime thread spawning requests from running shader programs.
[0412] In at least one embodiment, execution units 2707 and / or 2708 support a instruction set that includes native support for many standard 3D graphics shader ins...
Claims
1. A processor comprising: one or more circuits to cause dithering to be applied to one or more motion vectors during rendering of an image; wherein the processor is further to determine at least one image value of a current image frame at least in part from an interpolation of image values of one or more previous image frames, the interpolation performed from the one or more motion vectors to which values of the dithering are applied, the one or more motion vectors describing motion of an object from one frame to a next, the dithering referring to vectors that shift a magnitude and / or a direction of the one or more motion vectors when added to the one or more motion vectors.
2. The processor of claim 1, wherein the motion vectors are motion vectors of one or more intermediate image rendering passes.
3. The processor of claim 2, wherein the intermediate image rendering passes comprise one or more of a shadow pass, an occlusion pass, a reflection pass, or a specular reflection pass.
4. The processor of claim 1, further comprising: one or more memories to store values of the dithering to be applied to the motion vectors.
5. The processor of claim 4, wherein the one or more circuits are further to retrieve the values of the dithering from a buffer and apply the values of the dithering to the one or more motion vectors during the rendering.
6. The processor of claim 1, wherein the interpolation is a Catmull-Rom interpolation.
7. A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: cause dithering to be applied to one or more motion vectors during rendering of an image; wherein, the one or more processors are further to determine at least one image value of a current image frame at least in part from an interpolation of image values of one or more previous image frames, the interpolation performed from the one or more motion vectors to which values of the dithering are applied, the one or more motion vectors describing motion of an object from one frame to a next, the dithering referring to vectors that shift a magnitude and / or a direction of the one or more motion vectors when added to the one or more motion vectors.
8. The machine-readable medium of claim 7, wherein the motion vectors are motion vectors of one or more intermediate image rendering passes.
9. The machine-readable medium of claim 8, wherein the intermediate image rendering passes comprise one or more of a shadow pass, an occlusion pass, a reflection pass, or a specular reflection pass.
10. The machine-readable medium of claim 7, wherein the instructions, if performed by one or more processors, further cause the one or more processors to: store values of the dithering to be applied to the motion vectors.
11. The machine-readable medium of claim 10, wherein the instructions, if performed by one or more processors, further cause the one or more processors to: retrieve the values of the dithering from a buffer; and applying values of the dither to the one or more motion vectors during rendering of the image.
12. The machine-readable medium of claim 7, wherein the interpolation is a Catmull-Rom interpolation.
13. A display system comprising: one or more processors to generate an image for display on a display device at least in part by causing a dither to be applied to one or more motion vectors during rendering of the image; wherein the one or more processors are further to determine at least one image value of a current image frame at least in part from an interpolation of image values of one or more previous image frames, the interpolation performed from the one or more motion vectors to which values of the dither have been applied, the one or more motion vectors describing motion of an object from one frame to a next frame, the dither referring to a vector that shifts a magnitude and / or a direction of the one or more motion vectors when added to the one or more motion vectors.
14. The system of claim 13, wherein the motion vectors are motion vectors of one or more intermediate image rendering passes.
15. The system of claim 14, wherein the intermediate image rendering passes comprise one or more of a shadow pass, an occlusion pass, a reflection pass, or a specular reflection pass.
16. The system of claim 13, further comprising: one or more memories to store values of the dither to be applied to the motion vectors.
17. The system of claim 16, wherein the one or more processors are further to retrieve the values of the dither from a buffer and to apply the values of the dither to the one or more motion vectors during rendering of the image.
18. The system of claim 13, wherein the interpolation is a Catmull-Rom interpolation.
Citation Information
Patent Citations
Adding greater realism to a computer-generated image by smoothing jagged edges
CN110390644A
Rendering antialiased geometry to an image buffer using jittering
US8063914B1