Video upsampling using one or more neural networks
A neural network-based upsampling system enhances video frames by integrating historical data and applying kernel factors, addressing the issue of inadequate video quality for display devices and reducing artifacts for improved image clarity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2023-11-02
- Publication Date
- 2026-05-27
AI Technical Summary
Existing video content quality is often inadequate for the display device it is consumed on, leading to artifacts and lower quality than desired, especially in live video scenarios.
A neural network-based upsampling system that utilizes deep learning to enhance video frames by integrating historical frame data and applying kernel factors for improved resolution and reduced artifacts.
The system effectively upsamples video content to higher resolutions with reduced artifacts, providing superior image quality and smooth transitions, suitable for real-time display on various devices.
Smart Images

Figure 0007866536000005 
Figure 0007866536000006 
Figure 0007866536000007
Abstract
Description
Technical Field
[0001] This application is a PCT application, claiming priority to U.S. Patent Application No. 16 / 565,088, filed on September 9, 2019, and entitled "VIDEO UPSAMPLING USING ONE OR MORE NEURAL NETWORKS", the entire disclosure of which is incorporated herein by reference for all purposes.
[0002] At least one embodiment relates to processing resources used to facilitate the execution of artificial intelligence. For example, at least one embodiment relates to a processor or computing system used to train a neural network with various novel techniques described herein.
Background Art
[0003] Video content is being consumed in an ever-increasing variety of ways from a variety of resources on a variety of devices, so there are situations where the quality of video content is less than appropriate for the type of device used to display that content. Solutions for improving content quality often experience artifacts or are of lower quality than desired and may be difficult to obtain live video.
Summary of the Invention
Means for Solving the Problems
[0004] Various embodiments according to the present disclosure are described with reference to the drawings.
Brief Description of the Drawings
[0005] [Figure 1A] A diagram showing image data that can be processed or generated according to at least one embodiment. [Figure 1B] A diagram showing image data that can be processed or generated according to at least one embodiment. [Figure 2A]This figure shows a solution for upsampling video content, according to at least one embodiment. [Figure 2B] This figure shows a solution for upsampling video content, according to at least one embodiment. [Figure 3] This figure shows the components of a system for temporary anti-aliasing and upscaling of video content, according to at least one embodiment. [Figure 4] This figure shows a process for upsampling video content, according to at least one embodiment. [Figure 5] This figure shows part of the process for inferring upsampled video frames, according to at least one embodiment. [Figure 6] This figure shows a system for training and inference using one or more neural networks, according to at least one embodiment. [Figure 7] This figure shows a system for training one or more neural networks, according to at least one embodiment. [Figure 8] This figure shows the structure of a neural network according to at least one embodiment. [Figure 9A] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 9B] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 10] This figure shows an exemplary data center system according to at least one embodiment. [Figure 11] This figure shows a computer system according to at least one embodiment. [Figure 12] This figure shows a computer system according to at least one embodiment. [Figure 13] This figure shows a computer system according to at least one embodiment. [Figure 14] This figure shows a computer system according to at least one embodiment. [Figure 15A] A diagram showing a computer system according to at least one embodiment. [Figure 15B] A diagram showing a computer system according to at least one embodiment. [Figure 15C] A diagram showing a computer system according to at least one embodiment. [Figure 15D] A diagram showing a computer system according to at least one embodiment. [Figure 15E] A diagram showing a shared programming model according to at least one embodiment. [Figure 15F] A diagram showing a shared programming model according to at least one embodiment. [Figure 16] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 17A] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 17B] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 18A] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 18B] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 19] A diagram showing a computer system according to at least one embodiment. [Figure 20A] A diagram showing a parallel processor according to at least one embodiment. [Figure 20B] [[ID=四十一]]A diagram showing a partition unit according to at least one embodiment. [Figure 20C] A diagram showing a processing cluster according to at least one embodiment. [Figure 20D] A diagram showing a graphics multiprocessor according to at least one embodiment. [Figure 21] A diagram showing a multi-graphics processing unit (GPU) system according to at least one embodiment. [Figure 22] A diagram showing a graphics processor according to at least one embodiment. [Figure 23] A diagram showing the micro-architecture of a processor according to at least one embodiment. [Figure 24] A diagram showing a deep learning application processor according to at least one embodiment. [Figure 25] A diagram showing an exemplary neuromorphic processor according to at least one embodiment. [Figure 26] A diagram showing at least a part of a graphics processor according to at least one embodiment. [Figure 27] A diagram showing at least a part of a graphics processor according to at least one embodiment. [Figure 28] A diagram showing at least a part of a graphics processor core according to at least one embodiment. [Figure 29A] A diagram showing at least a part of a graphics processor core according to at least one embodiment. [Figure 29B] A diagram showing at least a part of a graphics processor core according to at least one embodiment. [Figure 30] A diagram showing a parallel processing unit (PPU) according to at least one embodiment. [Figure 31] A diagram showing a general-purpose processing cluster ("GPC") according to at least one embodiment. [Figure 32] A diagram showing a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment. [Figure 33] A diagram showing a streaming multiprocessor according to at least one embodiment. <
[0006] In at least one embodiment, a sequence of video frames 100 can be received on a video stream, as shown in Figure 1A. In at least one embodiment, the video frames from this sequence are generated by a game engine 102 that shows video frames representing gameplay within the current game session for at least one player. In at least one embodiment, the video frames can be received from another source, such as a video hosting site, and can be received at any time after the video content has been hosted by that video hosting site. In at least one embodiment, consecutive video frames may include changes from the previous video frame due to a change in the state of gameplay. In at least one embodiment, the sequence 100 generated by the game engine 102 may have a default or specific resolution or display size. In at least one embodiment, the resolution of the video frames in sequence 100 may be preferably as high as possible, or lower than the current resolution setting of the display 104 used to view sequence 100, such as a monitor, touchscreen, or television, used to display the gameplay video shown by the game engine 102.
[0007] In at least one embodiment, an upsampling system 152 (or a service, module, or device) can be used to upscale individual frames of sequence 100, as shown in Figure 150 of Figure 1B. In at least one embodiment, frames from game engine 102 can be fed to the upsampling system 152 to increase the resolution of individual frames in order to generate a higher resolution sequence that can be displayed on display 104 at a higher resolution. In at least one embodiment, the amount of upsampling performed may depend on the initial resolution of sequence 100 and the target resolution of display 104, ranging from 1080p to 4k. In at least one embodiment, additional processing may be performed as part of the upsampling process, such as anti-aliasing and transient smoothing. In at least one embodiment, any suitable upsampling algorithm can be used, such as one that utilizes a Gaussian filter. In at least one embodiment, the upsampling process takes jitter into consideration, which may be applied on a frame-by-frame basis.
[0008] In at least one embodiment, deep learning can be used to infer upsampled video frames of a sequence. In at least one embodiment, a super-sampling algorithm that does not rely on machine learning can be used to upsample the current input frame of a video sequence. In at least one embodiment, a Temporary Anti-aliasing Upsampling (TAAU) algorithm can be used to provide initial anti-aliasing and upsampling in a combined manner. In at least one embodiment, information from video frames of the corresponding sequence can be used to infer a higher quality upsampled image. In at least one embodiment, one or more heuristics based on prior knowledge of the rendering pipeline that do not require learning from the data can be used. In at least one embodiment, this can include jitter-aware upsampling and accumulating samples at the upsampled resolution. In at least one embodiment, as shown in Figure 200 of Figure 2A, this previous process data 208 can be provided as input to an upsampler system 210 including at least one neural network, along with the current input video frame 202 and the previous inferred frame 206, in order to infer a higher quality upsampled output image 204 generated by the upsampling algorithm alone.
[0009] In at least one embodiment, the upsampling system 210 provides deep learning for temporary supersampling to provide anti-aliasing and superresolution on a stream (or other sequence or file) of image or video frames. In at least one embodiment, a basic upsampling approach can be used as shown in Figure 250 of Figure 2B. In at least one embodiment, a low-resolution pixel 252 can be segmented into a number of higher-resolution (or lower-resolution) pixels 254. In at least one embodiment, the upsampling can be 4x upsampling as shown in Figure 2B, where each pixel of the input image is segmented into four higher-resolution pixels. In at least one embodiment, the position of a sample 256 at the low-resolution pixel 252 can be used to compute an upsampling kernel for one or more corresponding high-resolution pixels. In at least one embodiment, this kernel provides at least one of blurring, embossing, sharpening, or edge detection.
[0010] In at least one embodiment, the system 300 can upsample a sequence of image frames, as shown in Figure 3. In at least one embodiment, an input image 302 corresponding to a video frame in a sequence or stream is received. In at least one embodiment, the input image 302 is a lower-resolution, denser image. In at least one embodiment, an upsampling module 304 (or system, component, device, or service) can apply an upsampling algorithm, such as those discussed above and illustrated with reference to Figure 2B, to provide subpixel offset-aware upsampling. In at least one embodiment, this upsampled image can be fed to a trained neural network 320. In at least one embodiment, the trained network 320 can accept additional input to attempt to infer a higher-quality upsampled image or video frame. In at least one embodiment, the trained network 320 also accepts data from a previously inferred frame as an input video frame. In at least one embodiment, a dense, large historical image 328 inferred for a previous frame in the sequence can be used to provide historical input data to the trained network 320. In at least one embodiment, a motion warp module 330 or process can be applied to generate a bicubic warped history image 308. In at least one embodiment, motion warping can be used to apply a small offset to data to satisfy one or more constraints. In at least one embodiment, the offset is at least partially due to a determined or predicted motion of a portion of the image. In at least one embodiment, the history image 308 can be processed using a color space translation module 310 to generate a bicubic warped image 312 in a specific color space, such as a YCoCg color space containing lumen values and two chroma values.In at least one embodiment, a bicubic warping image 312 can be supplied to a Luma determination module 318 to provide Luma determination image data as input to a trained network 320. In at least one embodiment, the Luma determination module 318 can also accept an anti-aliased image 316 generated by a temporary anti-aliasing module 314 to provide an anti-aliased Luma value to smooth the results of upsampling on the processed image. In at least one embodiment, the historical image provided as input to the neural network 320 can be partially integrated with the current frame 306 based on an applied decision jitter offset, which can help in temporary convergence to a superior, sharp, high-resolution image.
[0011] In at least one embodiment, the trained neural network 320 generates a unification factor and a number of kernels that can be used to unify the input image 302 and the historical image 328 together to produce an inferred output image 326. In at least one embodiment, the output image 326 has the same resolution as the upscaled image 306. In at least one embodiment, a colorizer module 324 can be used to perform other color space conversions, such as placing the output image 326 into the RGB color space, even if the trained network 320 is operating on image data in the YCoCg color space. In at least one embodiment, the kernels inferred by the trained model 320 can help improve the cognitive quality of the output image 326, which also acts as the historical image 328 for the next input video frame of the corresponding sequence. In at least one embodiment, kernel factors output from the trained network 320 can be applied to improve the inferred upsampled image 326 of various qualities, including sharpness and reduction of ghosting or processing artifacts. In at least one embodiment, at least some of this kernel data can be provided as additional input 322 to a trained network 320 on subsequent image or video frames in an attempt to improve the quality on one or more subsequently processed frames of a sequence.
[0012] In at least one embodiment, the neural network 320 is trained using a dataset containing annotated images or video frames. In at least one embodiment, pairs of images are used for training, including an upsampled image and a corresponding anti-aliased and upsampled higher-resolution image. In at least one embodiment, the neural network 320 can be trained to learn appropriate mappings between these pairs of images. In at least one embodiment, the neural network 320 can also be trained to determine appropriate integration factors and one or more kernel factors to apply. In at least one embodiment, a multi-factor loss function can be utilized to optimize the neural network 320 during training, such as by optimizing network parameters to minimize the corresponding loss values. In at least one embodiment, a multi-factor loss function is utilized because modeling human perception of image quality can be mathematically complex to capture. In at least one embodiment, the loss function used to train a network such as the neural network 320 can utilize both style and temporal components, as well as other losses such as L2 loss to minimize errors. In at least one embodiment, the spatial component helps minimize other occurrences such as ghosting or artifacts, and the transient component helps smooth the motion between frames of the output sequence. In at least one embodiment, sequences of these frame pairs are used for training to improve transient smoothing.
[0013] In at least one embodiment, the neural network 320 predicts various factors for each pixel. In at least one embodiment, the network 320 predicts or infers 10 factors, including a unification factor and 9 elements of a kernel applied to the corresponding image input. In at least one embodiment, when generating predictions, these 9 factors can be applied to the current upsampled frame data. In at least one embodiment, the determined unification factor can be used to unify such processed and upsampled frames with data from previously inferred frames. In at least one embodiment, only one luma channel can be used for this processing and constantization, and a full-color image can be used, but a similar result can be provided with much less data management and processing.
[0014] In at least one embodiment, the loss can be weighted by a per-pixel weighting factor. In at least one embodiment, the per-pixel weighting can draw more attention to areas where disocclusion may exist, or areas that were previously present but are no longer obscured, such that one or more objects suddenly appear or are presented within a video frame of a sequence. In at least one embodiment, good disocclusion management can help reduce the presence of ghosting artifacts. In at least one embodiment, this weighting factor is calculated by comparing a previous warped reference frame with the current reference frame. In at least one embodiment, if a pixel in this previously warped reference frame lies within the bounding box of the color distribution of the corresponding current reference frame, it can be assumed that there is a high probability that there is no disocclusion at this location. In at least one embodiment, if it is determined that there is a significant difference in color between the previously warped reference frame and the current reference frame, a high weighting can be added to this spatial loss. In at least one embodiment, such a high weighting of the spatial loss can cause the spatial loss to be more influenced by areas where there is a large difference in color between the current and previous reference frames.
[0015] In at least one embodiment, only the last warped frame prediction has the current frame as input, instead of the previous set of predictions. In at least one embodiment, this last prediction is based on information from past frames and includes more recent information to minimize artifacts and provide better clarity to the inferred image. In at least one embodiment, errors during prediction during training are implicitly managed by the use of a loss function, because a bad frame or a frame with artifacts will have a high loss value during evaluation and cause the prediction to be discarded. In at least one embodiment, abrupt changes due to scene changes or camera pans may also cause the last prediction to be discarded and not used for upsampling, because there is a large change in color values or positions that are irrelevant to, or at least substantially different from, the current frame.
[0016] In at least one embodiment, as described with reference to Figure 6, supersampling can be performed by a content provider or a cloud resource provider at various locations, such as on a client device. In at least one embodiment, a client device having at least one graphics processor receives or acquires lower-resolution data and then upsamples this data before displaying or presenting the upsampled data. In at least one embodiment, the lower-resolution data may include video data received on a stream, generated by a game or rendering engine, produced by a camera or sensor, or contained within a file. In at least one embodiment, upsampling can occur in near real-time or offline for subsequent viewing or presentation. In at least one embodiment, applications such as gaming may require quick upsampling to enable players to view upscaled content in near real-time with no perceptible lag, in order to enjoy the gaming experience and avoid the disadvantages of significant lag.
[0017] In at least one embodiment, one or more other inputs 322 may include different information determined between the current frame and a previous predicted frame. In at least one embodiment, these inputs may help identify pixels or regions of pixels with large differences in pixel values. In at least one embodiment, this information may be used to the advantage of training or inference time to determine how much to weight specific pixel values in different regions of the image. In at least one embodiment, hidden historical data may also be generated by the network 320 and used as input for subsequent frames, thereby enabling the network 320 to impose information that is useful for subsequent frames or can serve as a starting point for analyzing or inferring subsequent frames.
[0018] In at least one embodiment, upsampling of video frames can be performed using process 400 shown in Figure 4. In at least one embodiment, a stream of lower-resolution video is received (402). In at least one embodiment, individual frames of this stream can be received and analyzed to provide such a stream for a higher-resolution version of the display. In at least one embodiment, the current video frame of this stream can be upsampled using an upsampling algorithm (404). In at least one embodiment, a previously warped video frame prediction is obtained (406) at the same resolution as obtained by upsampling. In at least one embodiment, these frames are appropriately converted to a target color space and a single channel of that target space used for the representation of these frames being processed (408). In at least one embodiment, these frames have at least some additional information, where applicable, as input to a trained neural network to determine an integration factor and one or more kernel factors (410). In at least one embodiment, these inferred factors and input frames are used to generate an output version of the corresponding current input video frame with high image quality and a target upsampling resolution (412). In at least one embodiment, this output video frame can be provided for display as part of a video stream (414), thereby allowing a video stream received at a first lower resolution to be displayed at a second higher resolution with superior image quality and fewer artifacts from the upsampling.
[0019] In at least one embodiment, upsampling of video frames can be performed using process 400 shown in Figure 4. In at least one embodiment, the current frame of video data is received (502). In at least one embodiment, this current video frame of video data is upsampled to a higher resolution of the target using an upscaling process (504). In at least one embodiment, such an upsampled current frame is provided as input to a trained neural network with a previously inferred frame at this higher resolution of the target (506). In at least one embodiment, an output version of this current video frame is inferred at least in part based on the integration of pixel values from this upsampled current frame and the previously inferred frame (508). In at least one embodiment, this output version can be provided for display and processing of subsequent video frames received at a lower resolution (510).
[0020] Neural network training and development A growing variety of industries and applications are leveraging machine learning. In at least one embodiment, deep neural networks (DNNs) developed on processors have been used in various use cases ranging from autonomous vehicles to faster drug development, automated image analysis for security systems, and smart real-time language translation within video chat applications. In at least one embodiment, deep learning is a technique that models the neural learning process of the human brain, which learns continuously, becomes continuously smarter, and delivers faster and more appropriate results over time. Children are initially taught by adults to accurately identify and classify various shapes, and gradually become able to identify shapes without any coaching. Similarly, in at least one embodiment, a deep learning or neural learning system designed to accomplish a similar task needs to be trained to become smarter and more efficient at identifying basic objects, closed objects, etc., while assigning content to objects.
[0021] In at least one embodiment, neurons in the human brain observe various incoming inputs, assign a significant level to each of these inputs, and the output is passed on to other neurons acting upon them. An artificial neuron, or perception, is the most basic model of a neural network. In at least one embodiment, perception may receive one or more inputs representing various characteristics of an object that perception is trained to recognize and classify, each of which is assigned a specific weight based on its importance in defining the shape of the object.
[0022] Deep neural networks (DNNs) consist of many layers of connected perceptions (e.g., nodes) that can be trained with large amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into various sections and sees basic patterns such as lines and angles. The second layer assembles the lines to see higher-level patterns such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the last few layers generate labels for the input image to identify a model of a specific car brand. Once trained, this DNN can be developed and used to identify and classify objects or patterns in a process known as inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include identifying handwritten digits on a check deposited in an ATM machine, identifying a friend's image in a photograph, delivering movie recommendations, identifying and classifying different types of cars, pedestrians, and road obstacles in an autonomous vehicle, or translating human conversation in near real time.
[0023] During training, data flows through the DNN in a forward propagation phase until predictions are generated that show labels corresponding to the inputs. If the neural network does not accurately label the inputs, the error between the correct label and the predicted label is analyzed, and weights are adjusted for each characteristic during the backward propagation phase until the DNN accurately labels the inputs and other inputs in the training dataset. Training complex neural networks requires a massive amount of parallel computing performance, including supported floating-point multiplication and addition. Inference is less numerically computational than training and is a latency-sensitive process in which the trained neural network classifies images, translates conversations, and infers new information from new inputs that were not seen before.
[0024] Neural networks rely heavily on matrix operations, and complex multi-layer networks require enormous amounts of floating-point performance and bandwidth for both efficiency and speed. With thousands of processing cores optimized for matrix operations and carrying tens to hundreds of TFLOPS of performance, the computing platform can carry the performance required for deep neural network-based artificial intelligence and machine learning applications.
[0025] Figure 6 shows components of a system 600 that can be used to train and utilize machine learning in at least one embodiment. As will be discussed, the various components can be provided by various combinations of computing devices and resources, or by a single computing system, which may be under the control of a single entity or a number of entities. Furthermore, embodiments can be triggered, initialized, or requested by different entities. In at least one embodiment, the training of a neural network can be commanded by a provider associated with a provider environment 606, and in at least one embodiment, the training can be requested by a customer or other user who has access to the provider environment through a client device 602 or other such resource. In at least one embodiment, the training data (or data to be analyzed by the trained neural network) can be provided by a provider, a user, or a third-party content provider 624. In at least one embodiment, the client device 602 may be a vehicle or object being navigated on behalf of a user, for example, which can issue requests and / or receive commands to help the device navigate.
[0026] In at least one embodiment, a request can be initiated across at least one network 604 that is received by the provider environment 606. In at least one embodiment, the client device may be any suitable electronic and / or computing device that enables a user to generate and transmit such requests, including desktop computers, notebook computers, computer servers, smartphones, tablet computers, gaming consoles (portable or otherwise), computer processors, computing logic, and set-top boxes. The (one or more) network 604 may include any suitable network for transmitting requests or other such data, including the Internet, intranets, Ethernet®, mobile networks, local area networks (LANs), and networks of direct wireless connections between peers.
[0027] In at least one embodiment, the interface layer 608 can receive requests, which in this example can transfer data to the training and inference manager 610. This manager may be a system or service including hardware and software for managing requests and services corresponding to data or content. In at least one embodiment, this manager may receive requests to train a neural network and can provide data for the request to the training manager 612. In at least one embodiment, the training manager 612 may select an appropriate model or network to use if not specified by the request and can train the model using the relevant training data. In at least one embodiment, the training data may be a batch of data stored in the training data repository 614, received from the client device 602 or obtained from a third-party provider 624. In at least one embodiment, the training manager 612 may be responsible for training the data, for example, by using a LARC-based approach as discussed herein. The network may be any appropriate network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Once the network is trained and well evaluated, the trained network can be stored in a model repository 616, which can store different models or networks for, for example, users, applications, or services. In at least one embodiment, there may be multiple models for a single application or entity that can be used based on a number of different factors.
[0028] In at least one embodiment, at a later point, a client device 602 (or another such device) may receive content (e.g., path decisions) or data that has been at least partially determined or added by the trained neural network. This request may include, for example, input data that is processed using the neural network to obtain one or more inferences or other output values, classifications, or predictions. In at least one embodiment, the input data may be received by the interface layer 608 and directed to the inference module 618, but different systems or services may also be used. In at least one embodiment, the inference module 618 may obtain a suitable trained network, such as a deep neural network (DNN) trained as discussed herein, from the model repository 616 if it is not already stored locally in the inference module 618. The inference module 618 may provide data as input to the trained network and then generate one or more inferences as outputs. This may include, for example, classifications of instances of the input data. In at least one embodiment, the inferences may then be transmitted to the client device 602 for display to the user or other communication. In at least one embodiment, user context data may also be stored in a user context data repository 622, which may include data about the user that could be useful as input to the network when determining data to generate inferences or data to return to the user after an instance has been obtained. In at least one embodiment, related data that may include at least some of the input or inference data may also be stored in a local database 620 for processing further requests. In at least one embodiment, the user may use an account or other information to access resources or functions in the provider environment. In at least one embodiment, if permitted and available, user data may be collected and used to further train the model in order to provide more accurate inferences for further requests.In at least one embodiment, requests are received through a user interface to a machine learning application 626 running on a client device 602, and results can be displayed through the same interface. The client device may include resources such as a processor 628 and memory 630 for generating requests and processing results or responses, and at least one data storage element 632 for storing data for the machine learning application 626.
[0029] In at least one embodiment, processor 628 (or the processor of training manager 612 or inference module 618) is a central processing unit (CPU). However, as described, in such an environment, resources can utilize GPUs to process data for at least certain types of requests. With thousands of cores, GPUs are designed to handle virtually parallel workloads and have therefore become popular in deep learning for training neural networks and generating predictions. While the use of GPUs for offline builds has enabled faster training of larger and more complex models, generating predictions offline implies that request-time input properties cannot be used, and predictions must be generated for all permutations of properties and stored in a lookup table for real-time requests. If the deep learning framework supports CPU mode and the model is small and simple enough to perform feedforward on the CPU with reasonable latency, a service on a CPU instance can host the model. In this case, training can be performed offline on the GPU, and inference can be performed in real time on the CPU. If the CPU approach is not feasible, the service can run on a GPU instance. While GPUs have different performance and cost characteristics than CPUs, running services that offload runtime algorithms to the GPU may need to be designed differently from CPU-based services.
[0030] In at least one embodiment, video data may be provided from a client device 602 for enhancement within the provider environment 606. In at least one embodiment, video data may be processed on the client device 602 for enhancement. In at least one embodiment, video data may be streamed from a third-party content provider 624 and enhanced by the third-party provider 624, the provider environment 606, or the client device 602.
[0031] Figure 7 shows a system 700 that can be used in at least one embodiment to classify data or to generate inferences. In at least one embodiment, both supervised and unsupervised training can be used in at least one embodiment discussed herein. In at least one embodiment, a set of training data 702 (e.g., classified or labeled data) is provided as input to act as training data. In at least one embodiment, the training data may include instances of at least one type of object on which the neural network is trained, and information that identifies objects of that type. In at least one embodiment, the training data may include a set of images, each containing or associated with a representation of an object type, such as a label, metadata, classification, or other information that identifies the type of object presented in each image. Various other types of data can also be used as training data, such as text data, audio data, video data, etc. In at least one embodiment, the training data 702 is provided to a training manager 704 as training input. In at least one embodiment, the training manager 704 may be a system or service including hardware and software, such as one or more computing devices, that run a training application to train a neural network (or other model or algorithm, etc.). In at least one embodiment, the training manager 704 receives instructions or requests indicating the type of model to be used for training. In at least one embodiment, the model may be any suitable statistical model, network, or algorithm useful for such purposes, such as including artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, etc.In at least one embodiment, the training manager 704 may select an initial model or other untrained model from a suitable repository 706, train the model using training data 702, and generate a trained model 708 (e.g., a trained deep neural network) that can be used to classify similar types of data or to generate other such inferences. In at least one embodiment where training data is not used, a suitable initial model may further be selected by each training manager 704 for training on input data.
[0032] In at least one embodiment, a model can be trained in a number of different ways, depending in part on the type of model chosen. In at least one embodiment, a machine learning algorithm may comprise a set of training data, and the model is a model artifact produced by the training process. In at least one embodiment, each instance of the training data contains a correct answer (e.g., classification), which can be called a target or target attribute. In at least one embodiment, the learning algorithm discovers patterns in the data as it trains, mapping input data attributes to targets, predicts answers, and outputs a machine learning model that captures these patterns. In at least one embodiment, the machine learning model can then be used to obtain predictions on new data where the target is not identified.
[0033] In at least one embodiment, the training manager 704 can select from a set of machine learning models, including binary classification, multiclass classification, and regression models. In at least one embodiment, the type of model used may depend at least in part on the type of target being predicted. In at least one embodiment, a machine learning model for a binary classification problem predicts a binary outcome, such as one of two possible classes. In at least one embodiment, a binary classification model can be trained using a learning algorithm such as logistic regression. In at least one embodiment, a machine learning model for a multiclass classification problem allows for the generation of predictions for multiple classes, such as predicting one of three or more outcomes. Polynomial logistic regression may be useful for training multiclass models. A machine learning model for a regression problem predicts numerical values. Linear regression may be useful for training regression models.
[0034] In at least one embodiment, to train a machine learning model according to one embodiment, the training manager must determine the input training data source, the name of the data attribute containing the predicted target, the required data transformation instructions, and other information such as training parameters to control the learning algorithm. In at least one embodiment, during the training process, the training manager 704 can automatically select an appropriate learning algorithm based on the type of target identified in the training data source. In at least one embodiment, the machine learning algorithm can accept parameters used to control the training process and specific characteristics of the resulting machine learning model. These are referred to herein as training parameters. In at least one embodiment, if the training parameters are not specified, the training manager can utilize default values known to work well for a wide range of machine learning tasks. Examples of training parameters for which values can be specified include the maximum model size, the maximum number of passes on the training data, the shuffle type, the regularization type, the learning rate, and the regularization amount. Default settings can be specified with the option to adjust values for fine-tuning performance.
[0035] In at least one embodiment, the maximum model size is the total size in bytes of the patterns produced during model training. In at least one embodiment, a model can be produced with a size specified by default, such as a 100MB model. If the training manager cannot determine enough patterns to fill the model size, a smaller model can be produced. If the training manager finds more patterns to fit a given size, the maximum cutoff can be achieved by trimming patterns that have at least an impact on the quality of the trained model. By selecting a model size, you control the trade-off between the predictive quality of the model and the cost of use. In at least one embodiment, a smaller model allows the training manager to remove many patterns to fit within the maximum size limit that affects predictive quality. In at least one embodiment, a larger model may be more expensive to query for real-time predictions. In at least one embodiment, a larger input dataset does not necessarily lead to a larger model, because the model stores patterns rather than input data. In at least one embodiment, if the patterns are fewer and simpler, the resulting model will be smaller. Input data with a large number of raw attributes (input columns) or inductive properties (outputs of data transformations) are more likely to have more patterns discovered and stored during the training process.
[0036] In at least one embodiment, the training manager 704 may perform a large number of passes or iterations on the training data in an attempt to discover patterns. In at least one embodiment, there may be a default number of passes, such as 10 passes, and in at least one embodiment, a maximum number of passes may be set, such as up to 100 passes. In at least one embodiment, there may be no maximum set, or there may be a convergence criterion or other set of factors that trigger the termination of the training process. In at least one embodiment, the training manager 704 may monitor the quality of patterns (such as for model convergence) during training and may automatically stop training if there are no more data points or patterns to discover. In at least one embodiment, a dataset with only a few observations may require more passes on the data to obtain a sufficiently high model quality. A larger dataset may contain many similar data points, which can reduce the demand for a large number of passes. The potential impact of selecting more data passes on the data is that model training will be longer and more expensive in terms of resource and system utilization.
[0037] In at least one embodiment, the training data is shuffled before training or between training passes. In at least one embodiment, the shuffle is random or pseudo-random to generate a truly random order, but may have some constraints instead of guaranteeing that there are no groupings of a particular type of data, or the shuffled data may be reshuffled if such groupings exist. In at least one embodiment, the shuffle alters the order or arrangement in which the data is used for training so that the training algorithm does not encounter groupings of similar types of data, or encounter too many consecutive observations of a single type of data. In at least one embodiment, the model may be trained to predict objects. In at least one embodiment, the data may be classified by object type before uploading. In at least one embodiment, the algorithm then processes the data alphabetically by object type so that it initially encounters only data for a particular object type. In at least one embodiment, the model begins to learn patterns of objects for that type. In at least one embodiment, the model then encounters only data for a second object type and attempts to adjust the model to fit that object type, potentially degrading the patterns that fit the first object type. Such abrupt switching between object types can create models that do not learn how to accurately predict object types. In at least one embodiment, the training dataset can be shuffled before it is split into training and evaluation subsets so that a relatively uniform distribution of data types is available at both stages. In at least one embodiment, the training manager 704 can shuffle the data using, for example, a pseudo-random shuffle technique.
[0038] In at least one embodiment, when creating a machine learning model in at least one embodiment, the training manager 704 may allow the user to identify settings or apply custom options. In at least one embodiment, the user may identify one or more evaluation settings that indicate a portion of the input data reserved for evaluating the predictive quality of the machine learning model. In at least one embodiment, the user may identify policies that indicate which attributes and attribute transformations are available for model training. In at least one embodiment, the user may also identify various training parameters that control the training process and specific characteristics of the resulting model.
[0039] In at least one embodiment, once the training manager determines that the model has been trained, for example by using at least one termination criterion discussed herein, the trained model 708 can be made available for use by the classifier 714 when classifying (or generating inferences thereto) the effectiveness data 712. In at least one embodiment, this requires a logical transition between the training mode for the model and the inference mode for the model. However, in at least one embodiment, the trained model 708 is first passed to an evaluator 710 which may include an application, process, or service running on at least one computing resource (e.g., the CPU or GPU of at least one server) to evaluate the quality (or other such embodiment) of the trained model. In at least one embodiment, the model is evaluated to determine whether this model provides at least a minimum acceptable or threshold level of performance when predicting targets on new and further data. If not, the training manager 704 may continue training this model. In at least one embodiment, future data instances often do not know the target value, so it may be desirable to check the precise metrics of machine learning on data where the target answer is known and use this judgment as a proxy for predictive accuracy on future data.
[0040] In at least one embodiment, the model is evaluated using a subset of the training data 702 provided for training. This subset can be determined using the shuffle and split approach, as discussed above. In at least one embodiment, this evaluation data subset can be labeled with targets and thus serve as a source of ground truth for evaluation. It is not useful to evaluate the predictive accuracy of a machine learning model on the same data used for training, as this may generate a positive evaluation for a model that remembers the training data, rather than regularizing it. In at least one embodiment, once training is complete, the evaluation data subset is processed using the trained model 708, and the evaluator 710 can determine the accuracy of this model by comparing the ground truth data to the corresponding output (or prediction / observation) of this model. In at least one embodiment, the evaluator 710 can provide a summary or performance metric indicating how well the prediction and the true value match. In at least one embodiment, if the trained model does not meet at least a minimum performance criterion or other such accuracy threshold, the training manager 704 is instructed to perform further training, or in some examples, may attempt to train a new or different model. In at least one embodiment, if the trained model 708 meets the relevant criteria, the trained model can be made available for use by the classifier 714.
[0041] In at least one embodiment, when generating and training a machine learning model, it may be desirable to identify model settings or training parameters that lead to a model capable of making accurate predictions. In at least one embodiment, parameters include the number of passes performed (forward and / or backward), regularization or refinement, model size, and shuffle type. In at least one embodiment, selecting model parameter settings that produce the best predictive performance on evaluation data may lead to model overfitting. In at least one embodiment, overfitting occurs when a model remembers patterns that occur in the training and evaluation data sources but which it could not generalize from the patterns in the data. Overfitting often occurs when the training data includes all the data used for evaluation. In at least one embodiment, an overfitted model may perform well during evaluation but fail to make accurate predictions on new or validity data. In at least one embodiment, to avoid selecting an overfitted model as the best model, the training manager may reserve additional data to enable the model's performance. For example, the training data set may be divided into 60% for training and 40% for evaluation or enablement, which can be divided into two or more stages. In at least one embodiment, after selecting model parameters that work well for evaluation data leading to convergence on a subset of the effectiveness data, such as half of this effectiveness data, a second validation can be performed on the remainder of this effectiveness data to ensure the performance of this model. If this model meets expectations on the effectiveness data, then this model is not overfitting data. In at least one embodiment, a test set or held-out set can be used to test the parameters. In at least one embodiment, using a second validation or test step helps in selecting appropriate model parameters to prevent overfitting.However, providing more data from the training process for activation reduces the amount of data available for training. This is problematic with smaller datasets, as there may not be enough data available for training. In at least one embodiment, the approach in such a situation is to perform mutual activation, as discussed elsewhere in this specification.
[0042] In at least one embodiment, there are many metrics or insights that can be used to examine and evaluate the predictive accuracy of a given model. In at least one embodiment, the evaluation results include predictive accuracy metrics to report on the overall success of the model and visualizations to help leverage the model's accuracy beyond the predictive accuracy metrics. The results can also provide the ability to examine the impact of setting score thresholds, for example, for binary classification, and can generate alerts on metrics to check the effectiveness of the evaluation. The choice of metrics and visualizations may depend at least in part on the type of model being evaluated.
[0043] In at least one embodiment, once trained and evaluated to a satisfactory degree, a trained machine learning model can be used to build or support a machine learning application. In one embodiment, building a machine learning application is an iterative process requiring a sequence of steps. In at least one embodiment, one or more core machine learning problems can be constructed in terms of what is observed and what answers the model predicts. In at least one embodiment, data can then be collected, filtered, and prepared to make it suitable for consumption by a machine learning model training algorithm. This data can be visualized and analyzed to enable data quality and to perform sanity checks to understand the data. Raw data (e.g., input variables) and response data (e.g., target) may not be presented in a way that can be used to train an advanced predictive model. Therefore, it may be desirable to construct more predictive input representations or characteristics from the raw variables. The resulting characteristics can be fed into a learning algorithm to build a model and evaluate the model's quality on the data provided from model building. The model can then be used to generate predictions of target answers for new data instances.
[0044] In at least one embodiment, in the system 700 of Figure 7, the trained model 710 is provided to or made available to a classifier 714 that can use the trained model to process validity data after evaluation. In at least one embodiment, this may include unclassified data received from users or third parties, such as query images looking for information about what is exemplified in these images. In at least one embodiment, the validity data can be processed by the classifier using the trained model, and the resulting results 716 (such as classification or prediction) can be sent back to their respective sources, or processed or stored. In at least one embodiment, where such use is permitted, these now-classified data instances can be stored in a training data repository that can be used by the training manager for further training of the trained model 708. In at least one embodiment, the models are trained continuously as new data becomes available, but in at least one embodiment, these models are retrained periodically, such as once a day or once a week, depending on factors such as the size of the dataset or the complexity of the model.
[0045] In at least one embodiment, the classifier 714 may include appropriate hardware and software for processing the enable data 712 using the trained model. In at least one embodiment, the classifier includes one or more computer servers, each having one or more graphics processing units (GPUs) capable of processing the data. In at least one embodiment, the configuration and design of the GPU may be preferable to that of a CPU or other such component when processing machine learning data. In at least one embodiment, the trained model may be loaded into GPU memory and into received data instances provided to the GPU for processing. The GPU has a much larger number of cores than the CPU, and the GPU cores may also be much less complex. In at least one embodiment, a given GPU may be capable of processing thousands of data instances simultaneously through different hardware threads. In at least one embodiment, the GPU may also be configured to maximize floating-point throughput, which can provide a considerable additional processing advantage for large datasets.
[0046] In at least one embodiment, even when using GPUs, accelerators, and other such hardware to accelerate tasks such as training a model or classifying data using such a model, such tasks still require considerable time, resource allocation, and cost. In at least one embodiment, if a machine learning model is trained using 700 passes and the dataset contains 1,000,000 data instances used for training, all million instances must be processed for each pass. Different parts of the architecture can also be supported by different types of devices. In at least one embodiment, training can be performed using a set of servers at a logically centralized location so that it can be provided as a service, and classification of raw data can be performed by such a service or on client devices. These devices can also be owned, operated, or controlled by the same entity or a number of entities.
[0047] In at least one embodiment, the exemplary neural network 800 shown in Figure 8 can be trained or, in at least one embodiment, made available. In at least one embodiment, the statistical model is an artificial neural network (ANN) with many layers of nodes, including an input layer 802, an output layer 806, and many layers of intermediate nodes 804, the internal layers and nodes are typically not visible or accessible within the neural network and are often called “hidden” layers. In at least one embodiment, only some intermediate layers are shown for illustrative purposes, but it should be understood that there is no limit to the number of intermediate layers that can be available, and any limitation on layers is a resource and time factor required to process using the model. In at least one embodiment, there may be additional types of models, networks, algorithms, or processes used, including other numbers or selections of nodes and layers. In at least one embodiment, the enable data can be processed by the layers of the network to generate a set of inferences or inference scores, which can then be fed into a loss function 808.
[0048] In at least one embodiment, all nodes in a given layer are interconnected with all nodes in adjacent layers. In at least one embodiment, the nodes of an intermediate layer are then each connected to nodes in two adjacent layers. In at least one embodiment, nodes are also called neurons or connection units in some models, and the connections between nodes are called edges. Each node can perform a function on an incoming input, for example, by using a specific function. In at least one embodiment, nodes and edges can acquire different weights during training, and individual layers of nodes can perform specific types of transformations on incoming inputs, and these transformations can also be learned or modulated during training. In at least one embodiment, learning can be supervised or unsupervised, depending at least in part on the type of information contained in the training dataset. In at least one embodiment, various types of neural networks can be utilized, including a convolutional neural network (CNN) with a set of several convolutional and pooling layers, which has proven useful in applications such as image recognition. CNNs can also be trained more easily than other networks due to the relatively small number of parameters that are determined.
[0049] In at least one embodiment, such a complex machine learning model can be trained using a variety of tuning parameters. Selecting parameters, fitting the model, and evaluating the model are part of a model tuning process often called hyperparameter optimization. Such tuning may, in at least one embodiment, require introspecting the underlying model or data. In the training or generation setting, a stable workflow may be important to avoid hyperparameter overfitting, as discussed elsewhere in this specification. Cross-enabling and adding Gaussian noise to the training dataset are techniques that may be useful to avoid overfitting to either one of the datasets. In hyperparameter optimization, it may be desirable to fix the training and enabling sets. In at least one embodiment, hyperparameters can be tuned in specific categories, including data preprocessing (e.g., translating words into vectors), CNN architecture definition (e.g., filter size, number of filters), stochastic gradient descent (SGD) parameters (e.g., learning rate), and regularization or refinement (e.g., dropout probability).
[0050] In at least one embodiment, instances of a dataset can be embedded in a lower-dimensional space of a specific size during preprocessing. In at least one embodiment, the size of this space is a tuned parameter. In at least one embodiment, the architecture of the CNN includes many tuned parameters. A parameter for the filter size can indicate the interpretation of information corresponding to the size of the instances being analyzed. In mathematical linguistics, this is known as the n-gram size. An exemplary CNN uses three different filter sizes, which potentially exhibit different n-gram sizes. The number of filters per filter size may correspond to the filter depth. Each filter attempts to learn something different from the structure of the instances, such as sentence structure for text data. In the convolutional layers, the activation function may be a pooling type set as modified linear units and max pooling. The results can then be concatenated into a single dimensional vector, and the final layer is fully connected on a 2D output. This corresponds to binary classification, to which an optimization function can be applied. One such function is an implementation of the root mean square (RMS) propagation method for gradient descent, where exemplary hyperparameters may include the learning rate, batch size, maximum gradient normal, and epoch. In neural networks, regularization can be a critical consideration. In at least one embodiment, the input data may be relatively sparse. A key hyperparameter in such situations might be dropout at the second-to-last layer, indicating the percentage of nodes that do not "fire" in each training cycle. An exemplary training process can propose different hyperparameter configurations based on feedback to the performance of previous configurations. The model can then be trained with the proposed configurations, evaluated on a specified enable set and performance report. This process can be repeated, for example, to trade off exploration (learning more about different configurations) and exploitation (leveraging previous knowledge to achieve better results).
[0051] By parallelizing the training CNN and leveraging GPU-enabled computing resources, numerous optimization strategies can be attempted for different scenarios. Complex scenarios allow for synchronized model architectures, as well as pre-processed and stochastic gradient descent parameters. This expands the model configuration space. In basic scenarios, only pre-processed and stochastic gradient descent parameters are synchronized. More complex scenarios may have a larger number of configuration parameters than basic scenarios. Synchronization within the joint space can be achieved using a linear or exponential number of steps, iterated through the model optimization loop. The cost of such a synchronization process can be considerably less than that of synchronization processes such as random search and grid search, without significant performance loss.
[0052] In at least one embodiment, backpropagation can be used to calculate the gradient used to determine the weights for a neural network. Backpropagation is a form of differential calculus and, as discussed above, can be used by gradient descent optimization algorithms to adjust the weights assigned to various nodes or neurons. The weights can be determined using the gradient of the associated loss function. Backpropagation can utilize the derivative of the loss function with respect to the output produced by the statistical model. As described, various nodes may have associated activation functions that define the output of each node. Various activation functions can be used appropriately, including radial basis functions (RBFs) and sigmoid functions, which can be used by various support vector machines (SVMs) for data transformation. The activation functions of the intermediate layers of a node are also referred to herein as the inner product kernels. These functions may include, for example, discriminant functions, step functions, sigmoid functions, ramp functions, etc. Activation functions may also be linear or nonlinear.
[0053] In at least one embodiment, an untrained neural network is trained using a training dataset. In at least one embodiment, the training framework is the PyTorch framework, Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework trains the untrained neural network and enables it to be trained using the processing resources described herein to produce a trained neural network. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0054] In at least one embodiment, the untrained neural network is trained using supervised learning, where the training dataset includes inputs paired with desired outputs for each input, or the training dataset includes inputs with known outputs, and the neural network's outputs are manually scored. In at least one embodiment, the untrained neural network is trained in a supervised manner, processing inputs from the training dataset and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, the error is then backpropagated through the untrained neural network. In at least one embodiment, the training framework adjusts the weights controlling the untrained neural network. In at least one embodiment, the training framework includes tools to monitor how well the untrained neural network is converging toward a model, such as a trained neural network, that is better suited to producing the correct answer in the results, based on known input data, such as new data. In at least one embodiment, the training framework iteratively trains the untrained neural network while adjusting the weights to refine the output of the untrained neural network using a loss function and tuning algorithms, such as stochastic gradient descent. In at least one embodiment, the training framework trains an untrained neural network until it reaches a desired accuracy. In at least one embodiment, the trained neural network can then be introduced to implement any number of machine learning operations.
[0055] In at least one embodiment, an untrained neural network is trained using unsupervised learning, where the untrained neural network attempts to train itself using unlabeled data. In at least one embodiment, the training dataset for unsupervised learning includes input data with no associated output data or "ground truth" data. In at least one embodiment, the untrained neural network can learn grouping within the training dataset and determine how individual inputs relate to the untrained dataset. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network that can perform useful actions to reduce the dimensionality of new data. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for the identification of data points in a new dataset that deviate from the normal patterns of the new dataset.
[0056] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled and unlabeled data are mixed in the training dataset. In at least one embodiment, incremental learning, such as a transmission learning technique, may be performed using the training framework. In at least one embodiment, incremental learning enables the trained neural network to adapt to new data without forgetting the knowledge taught to the network during initial training.
[0057] Logic of reasoning and training Figure 9A shows the inference and / or training logic 915 used to perform the inference and / or training operations associated with one or more embodiments. Further details regarding the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B.
[0058] In at least one embodiment, the inference and / or training logic 915 may include, but not limited to, code and / or data storage 901 for storing forward and / or output weights, and / or input / output data, and / or other parameters for constituting layers of neurons or neural networks used for training and / or inference in one or more embodiments. In at least one embodiment, the training logic 915 may include, or be coupled to, code and / or data storage 901 for storing graph code or other software to control the order in which timing and / or weight and / or other parameter information is loaded to constitute logic involving integer and / or floating-point units (collectively, integer arithmetic logic units (ALUs)). In at least one embodiment, the code, such as graph code, loads weight or other parameter information into the processor ALU based on the architecture of the corresponding neural network. In at least one embodiment, the code and / or data storage 901 stores the weight parameters and / or input / output data of each layer of the neural network that is trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using embodiments of one or more embodiments. In at least one embodiment, any part of the code and / or data storage 901 may be included in on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0059] In at least one embodiment, any part of the code and / or data storage 901 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 901 may be cache memory, dynamic random addressable memory (DRAM), static random addressable memory (SRAM), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 901 is internal or external to the processor, or whether it consists of, for example, DRAM, SRAM, flash, or some other storage type, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference being performed, the batch size of the data used for the neural network inference capabilities and / or training, or some combination of these factors.
[0060] In at least one embodiment, the inference and / or training logic 915 may include, but not limited to, code and / or data storage 905 for storing backward and / or output weights, and / or input / output data, corresponding to neurons or layers of a neural network used for training and / or inference in one or more embodiments. In at least one embodiment, the code and / or data storage 905 stores the weight parameters and / or input / output data of each layer of the neural network used for training or in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, the training logic 915 may include, or be coupled to, code and / or data storage 905 for storing graph code or other software to control the order in which timing and / or weight and / or other parameter information is loaded to constitute logic involving integer and / or floating-point units (collectively, integer arithmetic logic units (ALUs)). In at least one embodiment, the code, such as graph code, loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any part of the code and / or data storage 905 may be included in on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any part of the code and / or data storage 905 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 905 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 905 is internal or external to the processor, or whether it consists of, for example, DRAM, SRAM, flash, or other storage types, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size used for neural network inference and / or training, or some combination of these factors.
[0061] In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be different storage structures. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be the same storage structure. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be partially the same storage structure and partially different storage structures. In at least one embodiment, any part of the code and / or data storage 901 and the code and / or data storage 905 may be included in on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0062] In at least one embodiment, the inference and / or training logic 915 may include one or more arithmetic logic units (ALUs) 910, including integer and / or floating-point units, to perform logical and / or mathematical operations, at least in part on or based on training and / or inference code (e.g., graph code), or to perform logical and / or mathematical operations as shown therein, which may generate activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 920, the results of which are functions of input / output data and / or weight parameter data stored in code and / or data storage 901 and / or code and / or data storage 905. In at least one embodiment, the activation stored in the activation storage 920 is generated by linear algebra and / or matrix-based arithmetic performed by (one or more) ALUs 910 in response to instructions or other code, and the weight values stored in the code and / or data storage 905 and / or code and / or data storage 901 are used as operands along with other values such as bias values, gradient information, moment values, or other parameters or hyperparameters, any or all of which can be stored on-chip or off-chip in the code and / or data storage 905 or code and / or data storage 901, or in another storage.
[0063] In at least one embodiment, (one or more) ALU910 are contained within one or more processors or other hardware logic devices or circuits, and in another embodiment, (one or more) ALU910 may be external to the processor or other hardware logic devices or circuits that use them (e.g., coprocessors). In at least one embodiment, ALU910 may be contained within an execution unit of a processor, or within a bank of ALUs accessible by execution units of a processor distributed among different types of processors (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, the code and / or data storage 901, the code and / or data storage 905, and the activation storage 920 may be in the same processor or other hardware logic devices or circuits, and in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activated storage 920 may include on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, the inference and / or training code may be stored in other code accessible by the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuits.
[0064] In at least one embodiment, the activated storage 920 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activated storage 920 may be entirely or partially located within or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activated storage 920 is internal or external to the processor, or whether it consists of, for example, DRAM, SRAM, flash, or some other storage type, may depend on the available storage-on-chip vs. off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or some combination of these factors. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9A can be used in conjunction with application-specific integrated circuits ("ASICs") such as Google's Tensorflow® processing unit, Graphcore® inference processing unit (IPU), or Intel's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9A can be used in conjunction with other hardware such as a central processing unit ("CPU"), graphics processing unit ("GPU"), or field-programmable gate array ("FPGA").
[0065] Figure 9B shows inference and / or training logic 915 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 915 may include hardware logic that has dedicated computational resources or is used exclusively in conjunction with weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9B can be used in conjunction with application-specific integrated circuits (ASICs) such as Google's Tensorflow® processing unit, Graphcore® inference processing unit (IPU), or Intel's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9B can be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field-programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 915 includes code and / or data storage 901 and code and / or data storage 905, which can be used to store code (e.g., graph code), weight values, and / or other information including bias values, gradient information, moment values, and / or other parameter or hyperparameter information. In at least one embodiment shown in Figure 9B, code and / or data storage 901 and code and / or data storage 905 are associated with dedicated computing resources such as computing hardware 902 and computing hardware 906, respectively. In at least one embodiment, computing hardware 902 and computing hardware 906 each include one or more ALUs that perform arithmetic functions, such as linear algebraic functions, only on the information stored in code and / or data storage 901 and code and / or data storage 905, the results of which are stored in activation storage 920.
[0066] In at least one embodiment, each code and / or data storage 901, 905 and corresponding computing hardware 902, 906 corresponds to a different layer of a neural network, thereby, for the mirror concept organization of the neural network, the activation obtained from one "storage / computation pair 901 / 902" of the code and / or data storage 901 and computing hardware 902 is provided as input to one "storage / computation pair 905 / 906" of the code and / or data storage 905 and computing hardware 906. In at least one embodiment, each of the storage / computation pairs 901 / 902 and 905 / 906 can correspond to two or more neural network layers. In at least one embodiment, additional storage / computation pairs (not shown) after or in parallel with the storage / computation pairs 901 / 902 and 905 / 906 may be included in the inference and / or training logic 915.
[0067] Data center Figure 10 shows an exemplary data center 1000 in which at least one embodiment may be used. In at least one embodiment, the data center 1000 includes a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and an application layer 1040.
[0068] In at least one embodiment, as shown in Figure 10, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources ("node CRs") 1016(1) to 1016(N), where "N" represents any positive integer. In at least one embodiment, the node CRs 1016(1) to 1016(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., semiconductor drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more of the nodes CR1016(1) to 1016(N) may be servers having one or more of the computing resources described above.
[0069] In at least one embodiment, the grouped computing resources 1014 may include separate groups of node CRs housed in one or more racks (not shown), or a number of racks housed in a data center in various graphical locations (also not shown). Separate groups of node CRs within the grouped computing resources 1014 may include grouped compute resources, network resources, memory resources, or storage resources that are configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped in one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0070] In at least one embodiment, the resource orchestrator 1012 may constitute or otherwise control one or more nodes CR1016(1) to 1016(N) and / or grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1012 may include a software design infrastructure ("SDI") management entity for the data center 1000. In at least one embodiment, the resource orchestrator may include hardware, software, or any combination thereof.
[0071] In at least one embodiment shown in Figure 10, the framework layer 1020 includes a job scheduler 1022, a configuration manager 1024, a resource manager 1026, and a distribution file system 1028. In at least one embodiment, the framework layer 1020 may include a framework for supporting software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. In at least one embodiment, the software 1032 or application 1042 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1020 may be, but is not limited to, a type of free, open-source software web application framework, such as Apache Spark® ("Spark"), which can use the distribution file system 1028 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 1022 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1000. In at least one embodiment, the configuration manager 1024 may be capable of configuring different layers, such as the software layer 1030 and the framework layer 1020, which includes Spark and a distribution file system 1028 to support large-scale data processing. In at least one embodiment, the resource manager 1026 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support the distribution file system 1028 and the job scheduler 1022. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1014 located in the data center infrastructure layer 1010.In at least one embodiment, the resource manager 1026 may work in conjunction with the resource orchestrator 1012 to manage these mapped or allocated computing resources.
[0072] In at least one embodiment, the software 1032 included in the software layer 1030 may include software used by at least a portion of the nodes CR1016(1) to 1016(N), the grouped computing resources 1014, and / or the distribution file system 1028 of the framework layer 1020. One or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0073] In at least one embodiment, application 1042 included in application layer 1040 may include one or more types of applications used by at least a portion of nodes CR1016(1) to 1016(N), grouped computing resources 1014, and / or distribution file system 1028 of framework layer 1020. One or more types of applications may include, but are not limited to, any number of genomics applications, recognition compute, and machine learning applications including training or inference software, machine learning framework software (e.g., PyTorch, Tensorflow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0074] In at least one embodiment, any of the configuration manager 1024, resource manager 1026, and resource orchestrator 1012 may implement any number and type of self-correcting measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting measures may enable the data center operator of data center 1000 to avoid determining potentially faulty configurations and to eliminate underutilized and / or underperforming portions of the data center.
[0075] In at least one embodiment, the data center 1000 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 1000. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1000 by using weight parameters computed by one or more techniques described herein.
[0076] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to perform training or inference on information such as image recognition, speech recognition, or other artificial intelligence services.
[0077] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 10 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0078] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0079] Computer system Figure 11A is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or any combination thereof 1100, formed together with a processor which may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, the computer system 1100 may include, without limitation, components such as a processor 1102 for using an execution unit which includes logic for executing algorithms for processing data in accordance with the Disclosure, such as in the embodiments described herein. In at least one embodiment, the computer system 1100 may include a processor such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core®, or Intel® Nervana® microprocessors, available from Intel Corporation in Santa Clara, California, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, the computer system 1100 may run a version of the WINDOWS® operating system available from Microsoft Corporation in Redmond, Washington, but other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0080] The embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-a-chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0081] In at least one embodiment, the computer system 1100 may include, without limitation, a processor 1102 which may include, without limitation, one or more execution units 1108 for training and / or inference of machine learning models by the techniques described herein. In at least one embodiment, the computer system 1100 is a single-processor desktop or server system, but in another embodiment, the computer system 1100 may be a multi-processor system. In at least one embodiment, the processor 1102 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1102 may be coupled to a processor bus 1110 which may transmit digital signals between the processor 1102 and other components in the computer system 1100.
[0082] In at least one embodiment, the processor 1102 may include, without limitation, a level 1 ("L1") internal cache memory ("cache") 1104. In at least one embodiment, the processor 1102 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to the processor 1102. Other embodiments may include a combination of both internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, the register file 1106 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, state registers, and instruction pointer registers.
[0083] In at least one embodiment, the processor 1102 also includes an execution unit 1108 which includes, without limitation, logic for performing integer and floating-point arithmetic. In at least one embodiment, the processor 1102 may also include a microcode ("u-code") read-only memory ("ROM") for storing microcode for certain macro instructions. In at least one embodiment, the execution unit 1108 may include logic for handling a packed instruction set 1109. In at least one embodiment, by including the packed instruction set 1109, along with the associated circuitry for executing the instructions, in the instruction set of the general-purpose processor 1102, arithmetic used by many multimedia applications can be performed using the packed data of the general-purpose processor 1102. In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by performing packed data arithmetic using the full width of the processor's data bus, thereby eliminating the need to transfer smaller units of data across the processor's data bus to perform one or more arithmetic operations on a single data element at a time.
[0084] In at least one embodiment, the execution unit 1108 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, the computer system 1100 may include, without limitation, memory 1120. In at least one embodiment, memory 1120 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 1120 may store instructions 1119 and / or data 1121, which may be represented by data signals executed by the processor 1102.
[0085] In at least one embodiment, a system logic chip may be coupled to a processor bus 1110 and memory 1120. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub ("MCH") 1116, and the processor 1102 may communicate with the MCH 1116 via the processor bus 1110. In at least one embodiment, the MCH 1116 may provide a high-bandwidth memory path 1118 to memory 1120 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1116 may lead data signals between the processor 1102, memory 1120, and other components of the computer system 1100, and may bridge data signals between the processor bus 1110, memory 1120, and system I / O 1122. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH1116 may be coupled to memory 1120 via a high-bandwidth memory path 1118, and the graphics / video card 1112 may be coupled to the MCH1116 via an Accelerated Graphics Port ("AGP") interconnect 1114.
[0086] In at least one embodiment, the computer system 1100 may use a system I / O 1122, which is a proprietary hub interface bus, to connect the MCH 1116 to the I / O controller hub ("ICH") 1130. In at least one embodiment, the ICH 1130 may provide direct connectivity to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1120, a chipset, and a processor 1102. Examples may include, but are not limited to, an audio controller 1129, a firmware hub ("Flash BIOS") 1128, a wireless transceiver 1126, data storage 1124, a legacy I / O controller 1123 including a user input and keyboard interface 1125, a serial expansion port 1127 such as a Universal Serial Bus ("USB"), and a network controller 1134. The data storage 1124 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0087] In at least one embodiment, Figure 11A shows a system including interconnected hardware devices or “chips,” while in other embodiments, Figure 11A may show an exemplary system-on-a-chip (“SoC”). In at least one embodiment, devices may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of the computer system 1100 may be interconnected using a compute express link (CXL) interconnect.
[0088] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 11A for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0089] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0090] Figure 12 is a block diagram showing an electronic device 1200 for utilizing a processor 1210, according to at least one embodiment. In at least one embodiment, the electronic device 1200 may be, for example, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device, without limitation.
[0091] In at least one embodiment, the system 1200 may include, without limitation, a processor 1210 communicatively coupled to any number or type of preferred components, peripherals, modules, or devices. In at least one embodiment, the processor 1210 may be coupled using a bus or interface such as an I°C bus, a System Management Bus ("SMBus"), a Low Pin Count (LPC) bus, a Serial Peripheral Interface ("SPI"), a High Definition Audio ("HDA") bus, a Serial Advance Technology Attachment ("SATA") bus, a Universal Serial Bus ("USB") (versions 1, 2, or 3), or a Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 12 shows a system including interconnected hardware devices or “chips,” while in other embodiments, Figure 12 may show an exemplary system-on-a-chip (“SoC”). In at least one embodiment, the devices shown in Figure 12 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of Figure 12 may be interconnected using a Compute Express Link (CXL) interconnect.
[0092] In at least one embodiment, Figure 12 shows a display 1224, a touch screen 1225, a touch pad 1230, a Near Field Communications unit ("NFC") 1245, a sensor hub 1240, a thermal sensor 1246, an Express Chipset ("EC") 1235, a Trusted Platform Module ("TPM") 1238, a BIOS / firmware / flash memory ("BIOS, FW flash") 1222, a DSP 1260, a drive 1220 such as a Solid State Disk ("SSD") or Hard Disk Drive ("HDD"), a Wireless Local Area Network Unit ("WLAN") 1250, a Bluetooth unit 1252, and a Wireless Wide Area Network Unit ("WWAN"). The components may include a USB 3.0 camera ("USB 3.0 camera") 1254, and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 1215, for example, implemented in the LPDDR3 standard. Each of these components may be implemented in any preferred manner.
[0093] In at least one embodiment, other components may be communicatively coupled to the processor 1210 via the components described above. In at least one embodiment, the accelerometer 1241, ambient light sensor ("ALS") 1242, compass 1243, and gyroscope 1244 may be communicatively coupled to the sensor hub 1240. In at least one embodiment, the thermal sensor 1239, fan 1237, keyboard 1246, and touchpad 1230 may be communicatively coupled to the EC 1235. In at least one embodiment, the speaker 1263, headphones 1264, and microphone ("mic") 1265 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1262, which may be communicatively coupled to the DSP 1260. In at least one embodiment, the audio unit 1264 may include, for example, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, the SIM card ("SIM") 1257 may be communicatively coupled to the WWAN unit 1256. In at least one embodiment, components such as the WLAN unit 1250 and the Bluetooth unit 1252, as well as the WWAN 1256, may be implemented in a Next Generation Form Factor ("NGFF").
[0094] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 12 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0095] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0096] Figure 13 shows a computer system 1300 according to at least one embodiment. In at least one embodiment, the computer system 1300 is configured to implement various processes and methods described throughout this disclosure.
[0097] In at least one embodiment, the computer system 1300 includes, but is not limited to, at least one central processing unit ("CPU") 1302, which is connected to a communications bus 1310 implemented using any preferred protocol, such as PCI:Peripheral Component Interconnect ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP:Accelerated Graphics Port ("Accelerated Graphics Port"), Hypertransport, or any other bus or point-to-point communications protocol. In at least one embodiment, the computer system 1300 includes, but is not limited to, main memory 1304 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in the main memory 1304, which may take the form of random access memory ("RAM"). In at least one embodiment, the network interface subsystem ("Network Interface") 1322 provides an interface with other computing devices and networks for receiving data from other systems and transmitting data from the computer system 1300 to other systems.
[0098] In at least one embodiment, the computer system 1300 includes, in at least one embodiment without limitation, an input device 1308, a parallel processing system 1312, and a display device 1306, the display device which can be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from the input device 1308, such as a keyboard, mouse, touchpad, or microphone. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.
[0099] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the training logic 915 may be used in the system of Figure 13 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0100] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0101] Figure 14 shows a computer system 1400 according to at least one embodiment. In at least one embodiment, the computer system 1400 may include, without limitation, a computer 1410 and a USB stick 1420. In at least one embodiment, the computer system 1410 may include, without limitation, any number and type of processors (not shown), as well as memory. In at least one embodiment, the computer 1410 may include, without limitation, a server, a cloud instance, a laptop, and a desktop computer.
[0102] In at least one embodiment, the USB stick 1420 includes, but is not limited to, a processing unit 1430, a USB interface 1440, and a USB interface logic 1450. In at least one embodiment, the processing unit 1430 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1430 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1430 comprises an application-specific integrated circuit ("ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1430 is a tensor processing unit ("TPC") optimized to perform machine learning inference operations. In at least one embodiment, the processing core 1430 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.
[0103] In at least one embodiment, the USB interface 1440 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1440 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1440 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1450 may include any amount and type of logic that enables the processing unit 1430 to interface with or to a device (e.g., a computer 1410) via the USB connector 1440.
[0104] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 14 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0105] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0106] Figure 15A shows an exemplary architecture in which multiple GPUs 1510-1513 are communicably coupled to multiple multi-core processors 1505-1506 via high-speed links 1540-1543 (e.g., bus, point-to-point interconnect). In one embodiment, high-speed links 1540-1543 support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. Various interconnection protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.
[0107] Furthermore, in one embodiment, two or more of the GPUs 1510-1513 may be interconnected via high-speed links 1529-1530, which may be implemented using the same or different protocols / links as those used for high-speed links 1540-1543. Similarly, two or more of the multi-core processors 1505-1506 may be connected via high-speed link 1528, which can be a symmetric multiprocessor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various system components shown in Figure 15A may be implemented using the same protocol / link (for example, via a common interconnection fabric).
[0108] In one embodiment, each multi-core processor 1505-1506 is communicatively coupled to processor memory 1501-1502 via memory interconnects 1526-1527, and each GPU 1510-1513 is communicatively coupled to GPU memory 1520-1523 via GPU memory interconnects 1550-1553. The memory interconnects 1526-1527 and 1550-1553 may utilize the same or different memory access technologies. For example, but not limited to, the processor memories 1501-1502 and GPU memories 1520-1523 may be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or non-volatile memory such as 3D XPoint or Nano-Ram. In one embodiment, (for example, using a two-level memory (2LM) hierarchy), some portions of processor memory 1501-1502 may be volatile memory and other portions may be non-volatile memory.
[0109] As described below, various processors 1505-1506 and GPUs 1510-1513 may be physically coupled to specific memory locations 1501-1502 and 1520-1523, respectively, but an integrated memory architecture may be implemented in which the same virtual system address space (also called the "effective address" space) is distributed among the various physical memory locations. For example, processor memory locations 1501-1502 may each have a 64GB system memory address space, and GPU memory locations 1520-1523 may each have a 32GB system memory address space (in this example, a total of 256GB of addressable memory is obtained).
[0110] Figure 15B shows further details of the interconnection between a multi-core processor 1507 and a graphics acceleration module 1546 in one exemplary embodiment. The graphics acceleration module 1546 may include one or more GPU chips integrated on a line card coupled to the processor 1507 via a high-speed link 1540. Alternatively, the graphics acceleration module 1546 may be integrated on the same package or chip as the processor 1507.
[0111] In at least one embodiment, the illustrated processor 1507 includes a plurality of cores 1560A to 1560D, each having a translation lookaside buffer 1561A to 1561D and one or more caches 1562A to 1562D. In at least one embodiment, cores 1560A to 1560D may include various other components (not shown) for executing instructions and processing data. Caches 1562A to 1562D may have Level 1 (L1) and Level 2 (L2) caches. Furthermore, one or more shared caches 1556 may be included in caches 1562A to 1562D and shared by the set of cores 1560A to 1560D. For example, one embodiment of processor 1507 includes 24 cores, each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. The processor 1507 and the graphics acceleration module 1546 are connected to system memory 1514, which may include processor memories 1501-1502 in Figure 15A.
[0112] Coherence is maintained for data and instructions stored in various caches 1562A-1562D, 1556, and system memory 1514 through inter-core communication via coherence bus 1564. For example, each cache may have its own associated cache coherence logic / circuitry to communicate via coherence bus 1564 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via coherence bus 1564 to monitor cache access.
[0113] In one embodiment, the proxy circuit 1525 connects the graphics acceleration module 1546 to the coherence bus 1564 in a communicative manner, allowing the graphics acceleration module 1546 to participate in the cache coherence protocol as a peer of cores 1560A-1560D. Specifically, interface 1535 provides a connection to the proxy circuit 1525 via a high-speed link 1540 (e.g., PCIe bus, NVLink, etc.), and interface 1537 connects the graphics acceleration module 1546 to link 1540.
[0114] In one implementation, the accelerator integration circuit 1536 provides cache management, memory access, content management, and interrupt management services on behalf of the multiple graphics processing engines 1531, 1532, N of the graphics acceleration module 1546. Each of the graphics processing engines 1531, 1532, N may comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1531, 1532, N may comprise different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a bullet engine. In at least one embodiment, the graphics acceleration module 1546 may be a GPU having multiple graphics processing engines 1531-1532, N, or the graphics processing engines 1531-1532, N may be individual GPUs integrated on a common package, line card, or chip.
[0115] In one embodiment, the accelerator integration circuit 1536 includes a memory management unit (MMU) 1539 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 1514. The MMU 1539 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, the cache 1538 stores commands and data so that they can be efficiently accessed by graphics processing engines 1531-1532, N. In one embodiment, the data stored in the cache 1538 and graphics memories 1533-1534, M is kept coherent with the core caches 1562A-1562D, 1556, and system memory 1514. As mentioned above, this may be achieved via a proxy circuit 1525 instead of cache 1538 and memory 1533-1534, M (for example, by sending updates regarding cache line modifications / access in processor caches 1562A-1562D, 1556 to cache 1538 and receiving updates from cache 1538).
[0116] A set of registers 1545 stores context data for threads executed by graphics processing engines 1531-1532, N, and a context management circuit 1548 manages the thread contexts. For example, the context management circuit 1548 may perform save and restore operations to save and restore the contexts of various threads during a context switch (for example, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1548 may store the current register values in a designated area of memory (for example, identified by a context pointer). Then, when returning to the context, the context management circuit 1548 may restore the register values. In one embodiment, an interrupt management circuit 1547 receives and processes interrupts received from system devices.
[0117] In one implementation, the virtual / effective address from the graphics processing engine 1531 is translated by the MMU 1539 to the real / physical address of the system memory 1514. One embodiment of the accelerator integration circuit 1536 supports multiple (e.g., 4, 8, or 16) graphics accelerator modules 1546 and / or other accelerator devices. The graphics accelerator modules 1546 may be dedicated to a single application running on the processor 1507, or they may be shared among multiple applications. In one embodiment, there exists a virtualized graphics execution environment in which the resources of graphics processing engines 1531-1532, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices," which are allocated to different VMs and / or applications based on processing requirements and the priority associated with the VMs and / or applications.
[0118] In at least one embodiment, the accelerator integration circuit 1536 functions as a bridge to the system for the graphics acceleration module 1546 and provides address translation and system memory caching services. Furthermore, the accelerator integration circuit 1536 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1531-1532, N.
[0119] The hardware resources of the graphics processing engines 1531-1532, N are explicitly mapped to the real address space seen by the host processor 1507, so that any host processor can directly address these resources using effective address values. In one embodiment, one function of the accelerator integration circuit 1536 is to physically isolate the graphics processing engines 1531-1532, N so that they appear as independent units to the system.
[0120] In at least one embodiment, one or more graphics memories 1533-1534,M are each coupled to one of the graphics processing engines 1531-1532,N. The graphics memories 1533-1534,M store instructions and data processed by each of the graphics processing engines 1531-1532,N. The graphics memories 1533-1534,M may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memory such as 3D XPoint or Nano-Ram.
[0121] In one embodiment, a biasing technique is used to reduce data traffic over link 1540, such that the data stored in graphics memory 1533-1534, M is the data that will be most frequently used by graphics processing engines 1531-1532, N, and preferably the data that will not be used (or at least not frequently used) by cores 1560A-1560D. Similarly, the biasing mechanism attempts to keep the data that the cores need (and therefore preferably not needed by graphics processing engines 1531-1532, N) in the core caches 1562A-1562D, 1556, and system memory 1514.
[0122] Figure 15C shows another exemplary embodiment in which the accelerator integration circuit 1536 is integrated within the processor 1507. In this embodiment at least, the graphics processing engines 1531-1532, N communicate directly with the accelerator integration circuit 1536 via the high-speed link 1540 through interfaces 1537 and 1535 (in this case, any form of bus or interface protocol can be used). The accelerator integration circuit 1536 may perform the same operations as described with respect to Figure 15B, but may potentially operate at higher throughput given its proximity to the coherence bus 1564 and caches 1562A-1562D, 1556. At least one embodiment supports different programming models, including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1536 and a programming model controlled by the graphics acceleration module 1546.
[0123] In at least one embodiment, the graphics processing engines 1531-1532,N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can achieve virtualization within a VM / partition by directing other application requests to the graphics processing engines 1531-1532,N.
[0124] In at least one embodiment, the graphics processing engines 1531-1532,N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 1531-1532,N to allow access by each operating system. In a single-partition system without a hypervisor, the graphics processing engines 1531-1532,N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1531-1532,N to provide access to each process or application.
[0125] In at least one embodiment, the graphics acceleration module 1546 or the individual graphics processing engines 1531-1532, N select a process element using a process handle. In at least one embodiment, the process element is stored in system memory 1514 and is addressable using the effective address-to-actual address translation technique described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the host process context with the graphics processing engines 1531-1532, N (i.e., calling system software to add the process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element link list.
[0126] Figure 15D shows an exemplary accelerator integration slice 1590. As used herein, a “slice” comprises a designated portion of the processing resources of the accelerator integration circuit 1536. The application effective address space 1582 in system memory 1514 stores a process element 1583. In one embodiment, the process element 1583 is stored in response to a GPU call 1581 from an application 1580 running on processor 1507. The process element 1583 contains the process state of the corresponding application 1580. A work descriptor (WD) 1584 contained in the process element 1583 may be a single job requested by the application, or it may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1584 is a pointer to a job request queue in the application's address space 1582.
[0127] The graphics acceleration module 1546 and / or individual graphics processing engines 1531-1532, N can be shared by all or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD1584 to the graphics acceleration module 1546 to start jobs in a virtualized environment.
[0128] In at least one embodiment, the dedicated process programming model is implementation-specific. In this model, a single process owns either the graphics acceleration module 1546 or the individual graphics processing engines 1531. Since the graphics acceleration module 1546 is owned by a single process, when the graphics acceleration module 1546 is allocated, the hypervisor initializes the accelerator integration circuit 1536 for the owning partition, and the operating system initializes the accelerator integration circuit 1536 for the owning process.
[0129] During operation, the WD fetch unit 1591 in the accelerator integrated slice 1590 fetches the next WD 1584 containing a representation of the work to be performed by one or more graphics processing engines of the graphics acceleration module 1546. As illustrated, the data from WD 1584 is stored in register 1545 and may be used by the MMU 1539, interrupt management circuit 1547, and / or context management circuit 1548. For example, one embodiment of the MMU 1539 includes a segment / page walk circuit for accessing the segment / page table 1586 in the OS virtual address space 1585. The interrupt management circuit 1547 may process interrupt events 1592 received from the graphics acceleration module 1546. When performing graphics operations, the effective address 1593 generated by the graphics processing engines 1531-1532, N is translated to a real address by the MMU 1539.
[0130] In one embodiment, the same set of registers 1545 may be duplicated for each graphics processing engine 1531-1532, N, and / or graphics acceleration module 1546 and initialized by the hypervisor or operating system. Each of these duplicated registers may be included in the accelerator integration slice 1590. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]
[0131] Table 2 shows exemplary registers that may be initialized by the operating system. [Table 2]
[0132] In one embodiment, each WD1584 is specific to a particular graphics acceleration module 1546 and / or graphics processing engines 1531-1532, N. The WD1584 can contain all the information necessary for the graphics processing engines 1531-1532, N to perform their work, or it can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0133] Figure 15E provides further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor real address space 1598 in which the process element list 1599 is stored. The hypervisor real address space 1598 is accessible via a hypervisor 1596 that virtualizes the graphics acceleration module engine of the operating system 1595.
[0134] In at least one embodiment, a shared programming model allows all or a subset of processes from all or a subset of partitions in the system to use the graphics acceleration module 1546. There are two programming models in which the graphics acceleration module 1546 is shared by multiple processes and partitions: time-slice sharing and graphics-directed sharing.
[0135] In this model, the system hypervisor 1596 owns the graphics acceleration module 1546 and makes its functionality available to all operating systems 1595. In order for the graphics acceleration module 1546 to support virtualization by the system hypervisor 1596, the graphics acceleration module 1546 may comply with the following: 1) Application job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 1546 must provide a mechanism for saving and restoring context. 2) Application job requests must be guaranteed by the graphics acceleration module 1546 to be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1546 must provide a function to preempt job processing. 3) When the graphics acceleration module 1546 is operating in a specified shared programming model, fairness between processes must be guaranteed.
[0136] In at least one embodiment, application 1580 needs to make a system call to operating system 1595 with the type of graphics acceleration module 1546, a work descriptor (WD), an authorization mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1546 describes the acceleration function desired in the system call. In at least one embodiment, the type of graphics acceleration module 1546 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1546 and can be in the form of a command for graphics acceleration module 1546, an effective address pointer to a user-defined structure, an effective address pointer to a queue of commands, or any other data structure for describing the work performed by graphics acceleration module 1546. In one embodiment, the AMR value is the AMR state for use in the current process. In at least one embodiment, the value passed to the operating system is the same as that of the application setting the AMR. If the implementation of the accelerator integration circuit 1536 and the graphics acceleration module 1546 does not support the User Authority Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR to the hypervisor call. The hypervisor 1596 may optionally apply the current Authority Mask Override Register (AMOR) value before placing the AMR into the process element 1583. In at least one embodiment, CSRP is one of the registers 1545 that contains the effective address of an area in the application's effective address space 1582 for the graphics acceleration module 1546 to save and restore context state. This pointer is optional if no state needs to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0137] Upon receiving the system call, operating system 1595 may verify that application 1580 is registered and authorized to use graphics acceleration module 1546. Operating system 1595 then calls hypervisor 1596 with the information shown in Table 3. [Table 3]
[0138] Upon receiving a hypervisor call, hypervisor 1596 verifies that operating system 1595 is registered and authorized to use graphics acceleration module 1546. Hypervisor 1596 then places process element 1583 into a process element link list of the corresponding graphics acceleration module 1546 type. The process element may contain the information shown in Table 4. [Table 4]
[0139] In at least one embodiment, the hypervisor initializes register 1545 of multiple accelerator integration slices 1590.
[0140] As shown in Figure 15F, in at least one embodiment, integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1501-1502 and GPU memories 1520-1523. In this implementation, operations performed on GPUs 1510-1513 utilize the same virtual / effective memory address space as accessing processor memories 1501-1502, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1501, a second portion to a second processor memory 1502, a third portion to GPU memory 1520, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes called the effective address space) is distributed across processor memory 1501-1502 and GPU memory 1520-1523, respectively, so that either processor or GPU can access either physical memory, with virtual addresses mapped to physical memory.
[0141] In one embodiment, bias / coherence management circuits 1594A-1594E in one or more of the MMUs 1539A-1539E ensure cache coherence between the cache of one or more host processors (e.g., 1505) and the caches of GPUs 1510-1513, implement bias techniques to indicate the physical memory where a particular type of data should be stored. Multiple instances of the bias / coherence management circuits 1594A-1594E are shown in Figure 15F, but the bias / coherence circuits may be implemented within the MMU of one or more host processors 1505 and / or within the accelerator integration circuit 1536.
[0142] One embodiment allows GPU-enabled memory 1520-1523 to be mapped as part of system memory and made accessible using shared virtual memory (SVM) techniques without the performance degradation associated with full system cache coherence. In at least one embodiment, the accessibility of GPU-enabled memory 1520-1523 as system memory without cumbersome cache coherence overhead provides a beneficial operating environment for GPU offloading. This configuration allows host processor 1505 software to set operands and access computation results without the overhead of conventional I / O DMA data copying. Such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) access, all of which are less efficient than simple memory access. In at least one embodiment, the ability to access GPU-enabled memory 1520-1523 without cache coherence overhead may be essential for the execution time of offloaded computations. For example, in the case of significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 1510-1513. In at least one embodiment, the efficiency of operand configuration, the efficiency of accessing results, and the efficiency of GPU computation can be helpful in determining the effectiveness of GPU offloading.
[0143] In at least one embodiment, the choice between GPU bias and host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page-granular structure containing 1 or 2 bits per GPU-enabled memory page (i.e., controlled by memory page granularity). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU-enabled memories 1520-1523, with or without a bias cache (for example, to cache frequently used / recently used entries in the bias table) located in or without GPUs 1510-1513. Alternatively, the entire bias table may be maintained within the GPU.
[0144] In at least one embodiment, the bias table entries associated with each access to the GPU-biased memory 1520-1523 are accessed before the actual access to the GPU memory, resulting in the following actions: Firstly, local requests from GPUs 1510-1513 to find their pages in the GPU bias are forwarded directly to the corresponding GPU memories 1520-1523. Local requests from GPUs to find their pages in the host bias are forwarded to processor 1505 (for example, via the high-speed link described above). In one embodiment, a request from processor 1505 to find the requested page in the host processor bias completes the request in the same way as a normal memory read. Alternatively, a request directed to a GPU-biased page may be forwarded to GPUs 1510-1513. In at least one embodiment, the GPU may then move the page to the host processor bias if the page is not currently in use. In at least one embodiment, the bias state of a page can be altered by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply by a hardware-based mechanism.
[0145] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL) which calls the GPU's device driver, which sends a message to the GPU (or queues a command descriptor) to change the bias state and, for some transitions, directs the GPU to perform a cache-flushing operation on the host. In at least one embodiment, the cache-flushing operation is used for transitions from a host processor bias to a GPU bias, but not for transitions in the opposite direction.
[0146] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 1505. To access these pages, processor 1505 may request access from GPU 1510, and GPU 1510 may immediately grant access or not. Therefore, to reduce communication between processor 1505 and GPU 1510, it is beneficial to have GPU-biased pages requested by the GPU but not by the host processor 1505, or vice versa.
[0147] The inference and / or training logic 915 is used to carry out one or more embodiments. Further details regarding the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B.
[0148] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0149] Figure 16 shows an exemplary integrated circuit and associated graphics processor that can be manufactured using one or more IP cores according to various embodiments described herein. In addition to those shown, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0150] Figure 16 is a block diagram illustrating an exemplary system on a chip-integrated circuit 1600 that can be manufactured using one or more IP cores, according to at least one embodiment. In at least one embodiment, the integrated circuit 1600 may include one or more application processors 1605 (e.g., CPUs), at least one graphics processor 1610, and in addition, an image processor 1615 and / or a video processor 1620, either of which is a module IP core. In at least one embodiment, the integrated circuit 1600 may include a USB controller 1625, a UART controller 1630, an SPI / SDIO controller 1635, and I 2 S / I 2 The integrated circuit includes peripheral or bus logic, including a C controller 1640. In at least one embodiment, the integrated circuit 1600 may include a display device 1645 coupled to one or more High Definition Multimedia Interface (HDMI®) controller 1650 and Mobile Industry Processor Interface (MIPI) display interfaces 1655. In at least one embodiment, storage may be provided by a flash memory subsystem 1660, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1665 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits may also include an embedded security engine 1670.
[0151] The inference and / or training logic 915 is used to perform inference and / or training operations associated with one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 can be used within an integrated circuit 1600 for inferring or predicting operations based at least in part on a neural network training operation, a neural network function, and / or architecture, or weight parameters calculated using a neural network use case described herein.
[0152] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0153] Figures 17A–17B show exemplary integrated circuits and associated graphics processors that can be manufactured using one or more IP cores according to various embodiments described herein. In addition to those shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0154] Figures 17A and 17B are block diagrams illustrating exemplary graphics processors for use in a SoC according to embodiments described herein. Figure 17A shows an exemplary graphics processor 1710, a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. Figure 17B shows a further exemplary graphics processor 1740, a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 1710 in Figure 17A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1740 in Figure 17B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1710 and 1740 can be a variation of the graphics processor 1610 in Figure 16.
[0155] In at least one embodiment, the graphics processor 1710 includes a vertex processor 1705 and one or more fragment processors 1715A-1715N (e.g., 1715A, 1715B, 1715C, 1715D-1715N-1, and 1715N). In at least one embodiment, the graphics processor 1710 can execute different shader programs via separate logic, thereby optimizing the vertex processor 1705 to perform operations for a vertex shader program, while one or more fragment processors 1715A-1715N perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. In at least one embodiment, the vertex processor 1705 executes the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, the fragment processors 1715A to 1715N use primitive and vertex data generated by the vertex processor 1705 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 1715A to 1715N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform similar operations to pixel shader programs provided in the Direct 3D API.
[0156] In at least one embodiment, the graphics processor 1710 further includes one or more memory management units (MMUs) 1720A-1720B, caches 1725A-1725B, and circuit interconnects 1730A-1730B. In at least one embodiment, one or more MMUs 1720A-1720B include vertex processors 1705 and / or fragment processors 1715A-1715N, providing virtual-to-physical address mappings for the graphics processor 1710, which may reference vertex or image / text data stored in memory, in addition to vertex or image / text data stored in one or more caches 1725A-1725B. In at least one embodiment, one or more MMUs 1720A-1720B may be synchronized with other MMUs in the system, including one or more MMUs associated with one or more application processors 1605, image processors 1615, and / or video processors 1620 in Figure 16, so that each processor 1605-1620 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1730A-1730B allow the graphics processor 1710 to interface with other IP cores in the SoC via the SoC's internal bus or via a direct connection.
[0157] In at least one embodiment, the graphics processor 1740 includes one or more MMUs 1720A-1720B, caches 1725A-1725B, and circuit interconnects 1730A-1730B of the graphics processor 1710 in Figure 17A. In at least one embodiment, the graphics processor 1740 includes one or more shader cores 1755A-1755N (e.g., 1755A, 1755B, 1755C, 1755D, 1755E, 1755F-1755N-1, and 1755N), which provide an integrated shader core architecture in which a single core, or type, or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1740 includes an intercore task manager 1745 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1755A-1755N, and a tiling unit 1758 for accelerating tiling operations for tile-based rendering, where the rendering operation of the scene is subdivided in image space, for example, to take advantage of local space coherence in the scene or to optimize the use of an internal cache.
[0158] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in integrated circuits 17A and / or 17B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0159] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0160] Figures 18A and 18B illustrate further exemplary graphics processor logic according to embodiments described herein. Figure 18A shows a graphics core 1800, which in at least one embodiment may be included in the graphics processor 1610 of Figure 16, and in at least one embodiment may be an integrated shader core 1755A to 1755N, as shown in Figure 17B. Figure 18B shows a highly parallel general-purpose graphics processing unit 1830 suitable for deployment in a multi-chip module in at least one embodiment.
[0161] In at least one embodiment, the graphics core 1800 includes a shared instruction cache 1802, a texture unit 1818, and a cache / shared memory 1820, which are common to the execution resources within the graphics core 1800. In at least one embodiment, the graphics core 1800 may include multiple slices 1801A-1801N, or per-core partitions, and the graphics processor may include multiple instances of the graphics core 1800. The slices 1801A-1801N may include support logic including local instruction caches 1804A-1804N, thread schedulers 1806A-1806N, thread dispatchers 1808A-1808N, and sets of registers 1810A-1810N. In at least one embodiment, slices 1801A to 1801N may include a set of additional function units (AFU1812A to 1812N), floating-point units (FPU1814A to 1814N), integer arithmetic and logical operation units (ALU1816 to 1816N), address calculation units (ACU1813A to 1813N), double-precision floating-point units (DPFPU1815A to 1815N), and matrix processing units (MPU1817A to 1817N).
[0162] In at least one embodiment, the FPU1814A-1814N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPU1815A-1815N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU1816A-1816N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured to perform mixed-precision operations. In at least one embodiment, the MPU1817A-1817N can also be configured to perform mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPU1817A-1817N can perform various matrix operations to accelerate machine learning application frameworks, including supporting General-Purpose Matrix Multiplication (GEMM) acceleration. In at least one embodiment, AFU1812A~1812N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric function operations (e.g., sine, cosine, etc.).
[0163] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics core 1800 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0164] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0165] Figure 18B shows a General Purpose Processing Unit (GPGPU) 1830, which can be configured in at least one embodiment to enable highly parallel computational operations by an array of graphics processing units. In at least one embodiment, the GPGPU 1830 can be directly linked to other instances of the GPGPU 1830 to generate multiple GPU clusters to improve the training speed of deep neural networks. In at least one embodiment, the GPGPU 1830 includes a host interface 1832 for enabling connectivity with a host processor. In at least one embodiment, the host interface 1832 is a PCI Express interface. In at least one embodiment, the host interface 1832 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the GPGPU 1830 receives commands from the host processor and uses a global scheduler 1834 to distribute the execution threads associated with these commands to a set of compute clusters 1836A-1836H. In at least one embodiment, compute clusters 1836A to 1836H share a cache memory 1838. In at least one embodiment, the cache memory 1838 can act as a high-level cache for the cache memory within compute clusters 1836A to 1836H.
[0166] In at least one embodiment, the GPGPU 1830 includes memories 1844A to 1844B coupled to compute clusters 1836A to 1836H via a set of memory controllers 1842A to 1842B. In at least one embodiment, the memories 1844A to 1844B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.
[0167] In at least one embodiment, each of the compute clusters 1836A to 1836H includes a set of graphics cores, such as the graphics core 1800 in Figure 18A, which may include multiple types of integer and floating-point logic units capable of performing computational operations at varying precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 1836A to 1836H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units may be configured to perform 64-bit floating-point operations.
[0168] In at least one embodiment, multiple instances of GPGPU1830 can be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 1836A-1836H differs across embodiments. In at least one embodiment, multiple instances of GPGPU1830 communicate via a host interface 1832. In at least one embodiment, GPGPU1830 includes an I / O hub 1839, which couples GPGPU1830 to a GPU link 1840, enabling direct connections to other instances of GPGPU1830. In at least one embodiment, the GPU link 1840 is coupled to a dedicated GPU-to-GPU bridge enabling communication and synchronization between multiple instances of GPGPU1830. In at least one embodiment, the GPU link 1840 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1830 are located in separate data processing systems and communicate via a network device accessible through the host interface 1832. In at least one embodiment, the GPU link 1840 can be configured to enable connection to a host processor in addition to, or instead of, the host interface 1832.
[0169] In at least one embodiment, the GPGPU 1830 can be configured to train a neural network. In at least one embodiment, the GPGPU 1830 can be used within an inference platform. In at least one embodiment where the GPGPU 1830 is used for inference, the GPGPU may include fewer compute clusters 1836A-1836H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with memory 1844A-1844B may differ between the inference configuration and the training configuration, with high-bandwidth memory technology being used in the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1830 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more 8-bit integer dot product instructions, which may be used during the inference operation of a deployed neural network.
[0170] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the GPGPU 1830 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0171] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0172] Figure 19 is a block diagram showing a computing system 1900 according to at least one embodiment. In at least one embodiment, the computing system 1900 includes a processing subsystem 1901 having one or more processors 1902 and system memory 1904 communicating via an interconnection path which may include a memory hub 1905. In at least one embodiment, the memory hub 1905 may be a separate component within a chipset component or may be integrated within one or more processors 1902. In at least one embodiment, the memory hub 1905 is coupled to an I / O subsystem 1911 via a communication link 1906. In at least one embodiment, the I / O subsystem 1911 includes an I / O hub 1907 which can enable the computing system 1900 to receive input from one or more input devices 1908. In at least one embodiment, the I / O hub 1907 may enable a display controller, which may be included in one or more processors 1902 and provide output to one or more display devices 1910A. In at least one embodiment, one or more display devices 1910A coupled to the I / O hub 1907 may include local, internal, or embedded display devices.
[0173] In at least one embodiment, the processing subsystem 1901 includes one or more parallel processors 1912 coupled to a memory hub 1905 via a bus or other communication link 1913. In at least one embodiment, the communication link 1913 may be one of any number of communication link technologies or protocols based on standards such as PCI Express, or it may be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 1912 form a computation-intensive parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, one or more parallel processors 1912 form a graphics processing subsystem, which can output pixels to one of one or more display devices 1910A coupled via an I / O hub 1907. In at least one embodiment, one or more parallel processors 1912 may also include a display controller and a display interface (not shown) that enable direct connection to one or more display devices 1910B.
[0174] In at least one embodiment, the system storage unit 1914 can be connected to the I / O hub 1907 to provide storage functionality for the computing system 1900. In at least one embodiment, an I / O switch 1916 can be used to provide an interface mechanism for enabling communication between the I / O hub 1907 and other components such as a network adapter 1918 and / or a wireless network adapter 1919, which may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 1920. In at least one embodiment, the network adapter 1918 may be an Ethernet® adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1919 may include one or more other network devices, including Wi-Fi, Bluetooth, Near Field Communication (NFC), or one or more wireless radios.
[0175] In at least one embodiment, the computing system 1900 may include other components not explicitly shown, such as USB or other port connections, optical storage drives, and video capture devices, which may also be connected to the I / O hub 1907. In at least one embodiment, the communication paths interconnecting the various components of Figure 19 may be implemented using any preferred protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other bus or point-to-point communication interfaces, such as NV-Link High-Speed Interconnect, or other interconnection protocols.
[0176] In at least one embodiment, one or more parallel processors 1912 incorporate circuits optimized for graphics and video processing, such as video output circuits, to constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1912 incorporate circuits optimized for general-purpose processing. In at least one embodiment, the components of the computing system 1900 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1912, a memory hub 1905, a processor 1902, and an I / O hub 1907 can be integrated into a system-on-a-chip (SoC) integrated circuit. In at least one embodiment, the components of the computing system 1900 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 1900 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0177] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 1900 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0178] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0179] Processor Figure 20A shows a parallel processor 2000 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2000 may be implemented using one or more integrated circuit devices such as a programmable processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 2000 is a variation of one or more parallel processors 1912 shown in Figure 19 according to an exemplary embodiment.
[0180] In at least one embodiment, the parallel processor 2000 includes a parallel processing unit 2002. In at least one embodiment, the parallel processing unit 2002 includes an I / O unit 2004 that enables communication with other devices, including other instances of the parallel processing unit 2002. In at least one embodiment, the I / O unit 2004 may be directly connected to the other devices. In at least one embodiment, the I / O unit 2004 is connected to the other devices via the use of a hub or switch interface, such as a memory hub 1905. In at least one embodiment, the connection between the memory hub 1905 and the I / O unit 2004 forms a communication link 1913. In at least one embodiment, the I / O unit 2004 is connected to a host interface 2006 and a memory crossbar 2016, where the host interface 2006 receives commands targeting the execution of processing operations and the memory crossbar 2016 receives commands targeting the execution of memory operations.
[0181] In at least one embodiment, when the host interface 2006 receives a command buffer via the I / O unit 2004, the host interface 2006 can direct work operations to the front end 2008 to execute these commands. In at least one embodiment, the front end 2008 is coupled to a scheduler 2010, which is configured to distribute commands or other work items to the processing cluster array 2012. In at least one embodiment, the scheduler 2010 ensures that the processing cluster array 2012 is properly configured and enabled before tasks are distributed to the processing cluster array 2012. In at least one embodiment, the scheduler 2010 is implemented via firmware logic running on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2010 can be configured to perform complex scheduling and work distribution operations with both coarse and fine granularity, enabling rapid preemption and context switching of threads running on the processing array 2012. In at least one embodiment, host software can prove scheduling workloads on processing array 2012 through one of several graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across processing array 2012 by scheduler 2010 logic in a microcontroller, including scheduler 2010.
[0182] In at least one embodiment, the processing cluster array 2012 may contain up to "N" processing clusters (e.g., cluster 2014A, cluster 2014B to cluster 2014N). In at least one embodiment, each cluster 2014A to 2014N of the processing cluster array 2012 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 2010 may allocate work to clusters 2014A to 2014N of the processing cluster array 2012 using various scheduling and / or work distribution algorithms, which may differ depending on the workload arising for each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 2010 or partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 2012. In at least one embodiment, different clusters 2014A to 2014N of the processing cluster array 2012 may be allocated to process different types of programs or to execute different types of computations.
[0183] In at least one embodiment, the processing cluster array 2012 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2012 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, the processing cluster array 2012 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations including physical operations, and performing data transformations.
[0184] In at least one embodiment, the processing cluster array 2012 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2012 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as mosaic logic and other vertex processing logic. In at least one embodiment, the processing cluster array 2012 can be configured to execute graphics processing-related shader programs, including but not limited to vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2002 can transfer data from system memory via I / O unit 2004 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2022) during processing and then written back to system memory.
[0185] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2002, the scheduler 2010 can be configured to divide the processing workload into tasks of roughly equal size in order to better distribute the graphics processing operations among multiple clusters 2014A to 2014N of the processing cluster array 2012. In at least one embodiment, parts of the processing cluster array 2012 can be configured to perform different types of processing. For example, in at least one embodiment, to generate and display a rendered image, a first part may be configured to perform vertex shading and topology generation, a second part may be configured to perform mosaic and geometry shading, and a third part may be configured to perform pixel shading or other screen-space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2014A to 2014N may be stored in a buffer so that the intermediate data can be transmitted between the clusters 2014A to 2014N for further processing.
[0186] In at least one embodiment, the processing cluster array 2012 may receive processing tasks to be executed via the scheduler 2010, which receives commands defining the processing tasks from the front-end 2008. In at least one embodiment, the processing task may include an index of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters, and commands defining how the data should be processed (e.g., which program to run). In at least one embodiment, the scheduler 2010 may be configured to fetch the index corresponding to the task, or to receive the index from the front-end 2008. In at least one embodiment, the front-end 2008 may be configured to ensure that the processing cluster array 2012 is configured to be in a valid state before the workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is started.
[0187] In at least one embodiment, each of one or more instances of the parallel processing unit 2002 can be coupled to a parallel processor memory 2022. In at least one embodiment, the parallel processor memory 2022 can be accessed via a memory crossbar 2016, which can receive memory requests from the processing cluster array 2012 and the I / O unit 2004. In at least one embodiment, the memory crossbar 2016 can access the parallel processor memory 2022 via a memory interface 2018. In at least one embodiment, the memory interface 2018 may include a plurality of partition units (e.g., partition unit 2020A, partition unit 2020B to partition unit 2020N), each of which can be coupled to a portion of the parallel processor memory 2022 (e.g., a memory unit). In at least one embodiment, the number of partition units 2020A to 2020N is configured to be equal to the number of memory units, so that the first partition unit 2020A has a corresponding first memory unit 2024A, the second partition unit 2020B has a corresponding memory unit 2024B, and the Nth partition unit 2020N has a corresponding Nth memory unit 2024N. In at least one embodiment, the number of partition units 2020A to 2020N does not have to be equal to the number of memory devices.
[0188] In at least one embodiment, memory units 2024A-2024N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2024A-2024N may also include, but not limited to, high-bandwidth memory (HBM) and 3D stacked memory. In at least one embodiment, to efficiently use the available bandwidth of parallel processor memory 2022, render targets such as frame buffers or texture maps may be stored across memory units 2024A-2024N, allowing partition units 2020A-2020N to write portions of each render target in parallel. In at least one embodiment, local instances of parallel processor memory 2022 may be excluded to favor an integrated memory design using system memory and local cache memory together.
[0189] In at least one embodiment, any one of the clusters 2014A to 2014N of the processing cluster array 2012 can process data that will be written to any of the memory units 2024A to 2024N in the parallel processor memory 2022. In at least one embodiment, the memory crossbar 2016 can be configured to forward the output of each cluster 2014A to 2014N to any partition units 2020A to 2020N, or to another cluster 2014A to 2014N, which can perform further processing operations on the output. In at least one embodiment, each cluster 2014A to 2014N can communicate with the memory interface 2018 through the memory crossbar 2016 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2016 has a connection to a memory interface 2018 for communicating with the I / O unit 2004, and a connection to a local instance of the parallel processor memory 2022, enabling processing units in different processing clusters 2014A to 2014N to communicate with system memory or other memory not local to the parallel processing unit 2002. In at least one embodiment, the memory crossbar 2016 can use virtual channels to separate traffic streams between the clusters 2014A to 2014N and the partition units 2020A to 2020N.
[0190] In at least one embodiment, multiple instances of the parallel processing unit 2002 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2002 can be configured to interact with each other, even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2002 may include higher-precision floating-point units than other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2002 or parallel processor 2000 can be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0191] Figure 20B is a block diagram of partition unit 2020 according to at least one embodiment. In at least one embodiment, partition unit 2020 is an instance of one of the partition units 2020A to 2020N in Figure 20A. In at least one embodiment, partition unit 2020 includes an L2 cache 2021, a frame buffer interface 2025, and a raster operations unit ("ROP") 2026. The L2 cache 2021 is a read / write cache configured to perform load and store operations received from the memory crossbar 2016 and the ROP 2026. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 2021 to the frame buffer interface 2025 for processing. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 2025 for processing. In at least one embodiment, the frame buffer interface 2025 interfaces with one of the memory units of the parallel processor memory, such as memory units 2024A to 2024N (for example, within the parallel processor memory 2022) in Figure 20.
[0192] In at least one embodiment, ROP2026 is a processing unit that performs raster operations such as stenciling, z-testing, and blending. In at least one embodiment, ROP2026 then outputs the processed graphics data stored in graphics memory. In at least one embodiment, ROP2026 includes compression logic for compressing depth or color data written to memory and for decompressing depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The compression logic performed by ROP2026 may be modified based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0193] In at least one embodiment, ROP2026 is located within each processing cluster (for example, clusters 2014A-2014N in Figure 20) rather than within the partition unit 2020. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via the memory crossbar 2016. In at least one embodiment, the processed graphics data may be displayed on a display device, such as one of the one or more display devices 1910 in Figure 19, routed for further processing by a processor 1902, or routed for further processing by one of the processing entities in the parallel processor 2000 in Figure 20A.
[0194] Figure 20C is a block diagram of a processing cluster 2014 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 2014A to 2014N in Figure 20A. In at least one embodiment, one or more of the processing clusters 2014 may be configured to run a large number of threads in parallel, where “thread” means an instance of a particular program running on a particular set of input data. In at least one embodiment, a single-instruction, multiple-data (SIMD) instruction issuing technique is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, a single-instruction, multiple-thread (SIMT) technique is used to support the parallel execution of a large number of threads in a globally synchronized manner, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0195] In at least one embodiment, the operation of the processing cluster 2014 can be controlled via a pipeline manager 2032 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2032 receives instructions from the scheduler 2010 in Figure 20A and manages the execution of these instructions via the graphics multiprocessor 2034 and / or texture unit 2036. In at least one embodiment, the graphics multiprocessor 2034 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2014. In at least one embodiment, one or more instances of the graphics multiprocessor 2034 may be included within the processing cluster 2014. In at least one embodiment, the graphics multiprocessor 2034 can process data, and a data crossbar 2040 may be used to distribute the processed data to one of several possible destinations, including other shader units. In at least one embodiment, the pipeline manager 2032 can facilitate the distribution of processed data by specifying the destination of the processed data to be distributed through the data crossbar 2040.
[0196] In at least one embodiment, each graphics multiprocessor 2034 within the processing cluster 2014 may contain an identical set of function execution logic (e.g., arithmetic logic units, load / store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner, allowing new instructions to be issued before the previous instruction is completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be performed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0197] In at least one embodiment, instructions sent to processing cluster 2014 constitute a thread. In at least one embodiment, a set of threads running across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program for different input data. In at least one embodiment, each thread in a thread group can be assigned to a different processing engine in the graphics multiprocessor 2034. In at least one embodiment, a thread group may contain fewer threads than the number of processing engines in the graphics multiprocessor 2034. In at least one embodiment, if a thread group contains fewer threads than the number of processing engines, one or more processing engines may be idle during the cycle in which the thread group is being processed. In at least one embodiment, a thread group may also contain more threads than the number of processing engines in the graphics multiprocessor 2034. In at least one embodiment, if a thread group contains more threads than the number of processing engines in the graphics multiprocessor 2034, processing can be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups can run simultaneously on the graphics multiprocessor 2034.
[0198] In at least one embodiment, the graphics multiprocessor 2034 includes internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2034 can abandon its internal cache and use cache memory within the processing cluster 2014 (e.g., L1 cache 2048). In at least one embodiment, each graphics multiprocessor 2034 may also access L2 cache within a partition unit (e.g., partition units 2020A-2020N in Figure 20A), and these caches may be shared among all processing clusters 2014 and used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2034 may also access off-chip global memory, which may include one or more of the local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of the parallel processing unit 2002 may be used as global memory. In at least one embodiment, the processing cluster 2014 includes multiple instances of a graphics multiprocessor 2034 that can share common instructions and data, which may be stored in an L1 cache 2048.
[0199] In at least one embodiment, each processing cluster 2014 may include a memory management unit ("MMU") 2045 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2045 may be located within the memory interface 2018 in Figure 20A. In at least one embodiment, the MMU 2045 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles and, optionally, cache line indices. In at least one embodiment, the MMU 2045 may include a translation lookaside buffer (TLB) or cache for addresses, which may be located within a graphics multiprocessor 2034 or an L1 cache, or within the processing cluster 2014. In at least one embodiment, physical addresses are processed to distribute surface data access locally, enabling efficient interleaving of requests between partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0200] In at least one embodiment, each graphics multiprocessor 2034 may be coupled to a texture unit 2036 to configure a processing cluster 2014 so that texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data, are performed. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2034 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2034 outputs processed tasks to a data crossbar 2040 to provide processed tasks to another processing cluster 2014 for further processing, or stores processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 2016. In at least one embodiment, a pre-ROP2042 (pre-raster arithmetic unit) is configured to receive data from a graphics multiprocessor 2034 and direct the data to an ROP unit, which may be located within a partition unit (for example, partition units 2020A-2020N in Figure 20A) as described herein. In at least one embodiment, the pre-ROP2042 unit can perform color blending optimization, organize pixel color data, and perform address translation.
[0201] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics processing cluster 2014 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0202] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0203] Figure 20D shows a graphics multiprocessor 2034 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 2034 is coupled with a pipeline manager 2032 of a processing cluster 2014. In at least one embodiment, the graphics multiprocessor 2034 has an execution pipeline that includes, but is not limited to, an instruction cache 2052, an instruction unit 2054, an address mapping unit 2056, a register file 2058, one or more general-purpose graphics processing unit (GPGPU) cores 2062, and one or more load / store units 2066. The GPGPU cores 2062 and the load / store units 2066 are coupled to cache memory 2072 and shared memory 2070 via a memory and cache interconnect 2068.
[0204] In at least one embodiment, the instruction cache 2052 receives a stream of instructions to be executed from the pipeline manager 2032. In at least one embodiment, the instructions are cached in the instruction cache 2052 and dispatched to be executed by the instruction unit 2054. In at least one embodiment, the instruction unit 2054 may dispatch the instructions as thread groups (e.g., warps), each thread group being assigned to a different execution unit within the GPGPU core 2062. In at least one embodiment, the instructions may access a local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, an address mapping unit 2056 may be used to translate an address in the unified address space to a separate memory address accessible by the load / store unit 2066.
[0205] In at least one embodiment, the register file 2058 provides a set of registers to the functional units of the graphics multiprocessor 2034. In at least one embodiment, the register file 2058 provides temporary storage for operands connected to the data paths of the functional units of the graphics multiprocessor 2034 (e.g., GPGPU core 2062, load / store unit 2066). In at least one embodiment, the register file 2058 is divided among the functional units such that each functional unit is allocated a dedicated portion of the register file 2058. In one embodiment, the register file 2058 is divided among different warps being executed by the graphics multiprocessor 2034.
[0206] In at least one embodiment, each GPGPU core 2062 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 2034. The GPGPU cores 2062 may have similar architectures or different architectures. In at least one embodiment, a first part of the GPGPU core 2062 includes a single-precision FPU and an integer ALU, and a second part of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point operations, or enable variable-precision floating-point operations. In at least one embodiment, the graphics multiprocessor 2034 may further include one or more fixed-function units or special-function units for performing specific functions, such as rectangular copying or pixel blending operations. In at least one embodiment, one or more GPGPU cores may also include fixed or special-function logic.
[0207] In at least one embodiment, the GPGPU core 2062 includes SIMD logic that can execute a single instruction for multiple data sets. In at least one embodiment, the GPGPU core 2062 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core may be generated at compile time by the shader compiler, or they may be automatically generated when running a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logic unit.
[0208] In at least one embodiment, the memory and cache interconnect 2068 is an interconnect network connecting each functional unit of the graphics multiprocessor 2034 to the register file 2058 and shared memory 2070. In at least one embodiment, the memory and cache interconnect 2068 is a crossbar interconnect that allows the load / store unit 2066 to implement load and store operations between shared memory 2070 and register file 2058. In at least one embodiment, the register file 2058 can operate at the same frequency as the GPGPU core 2062, and therefore data transfer between the GPGPU core 2062 and register file 2058 is very low latency. In at least one embodiment, shared memory 2070 can be used to enable communication between threads running in functional units within the graphics multiprocessor 2034. In at least one embodiment, cache memory 2072 can be used, for example, as a data cache to cache texture data communicated between functional units and texture unit 2036. In at least one embodiment, the shared memory 2070 can also be used as a program-managed cache. In at least one embodiment, a thread running on the GPGPU core 2062 can programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 2072.
[0209] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into the same package or chip as the core and communicatively coupled to the core via an internal (i.e., internal to the package or chip) processor bus / interconnection. In at least one embodiment, regardless of how the GPU is connected, the processor core may allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0210] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics multiprocessor 2034 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0211] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0212] Figure 21 shows a multi-GPU computing system 11100 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 11100 may include a processor 11102 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 11106A-D via a host interface switch 11104. In at least one embodiment, the host interface switch 11104 is a PCI Express switch device that couples the processor 11102 to a PCI Express bus, through which the processor 11102 can communicate with the GPGPUs 11106A-D. The GPGPUs 11106A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 11116. In at least one embodiment, the GPU-to-GPU links 11116 are connected to each of the GPGPUs 11106A-D via dedicated GPU links. In at least one embodiment, the P2P GPU link 11116 enables direct communication between each of the GPGPUs 11106A-D without requiring communication via the host interface bus 11104 to which the processor 11102 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 11116, the host interface bus 11104 is kept available to allow access to system memory or to communicate with other instances of the multi-GPU computing system 11100, for example, via one or more network devices. In at least one embodiment, the GPGPUs 11106A-D are connected to the processor 11102 via the host interface switch 11104, and in at least one embodiment, the processor 11102 can be directly connected to the GPGPUs 11106A-D, including direct support for the P2P GPU link 11116.
[0213] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in a multi-GPU computing system 11100 for inference or prediction operations, at least in part, based on the training operations of the neural network described herein, the functionality and / or architecture of the neural network, or weight parameters calculated using the use cases of the neural network.
[0214] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0215] Figure 22 is a block diagram of a graphics processor 2200 according to at least one embodiment. In at least one embodiment, the graphics processor 2200 includes a ring interconnect 2202, a pipeline front end 2204, a media engine 2237, and graphics cores 2280A to 2280N. In at least one embodiment, the ring interconnect 2202 connects the graphics processor 2200 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2200 is one of a number of processors integrated within a multi-core processing system.
[0216] In at least one embodiment, the graphics processor 2200 receives batches of commands via a ring interconnect 2202. In at least one embodiment, incoming commands are interpreted by a command streamer 2203 on a pipeline front end 2204. In at least one embodiment, the graphics processor 2200 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 2280A-2280N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2203 feeds the commands to a geometry pipeline 2236. In at least one embodiment, for at least some media processing commands, the command streamer 2203 feeds the commands to a video front end 2234, which is coupled to a media engine 2237. In at least one embodiment, the media engine 2237 includes a Video Quality Engine (VQE) 2230 for post-processing of video and images, and a multi-format encoding / decoding (MFX) 2233 engine for encoding and decoding hardware-accelerated media data. In at least one embodiment, the geometry pipeline 2236 and the media engine 2237 each generate execution threads for thread execution resources provided by at least one graphics core 2280A.
[0217] In at least one embodiment, the graphics processor 2200 includes a scalable thread execution resource featuring modular cores 2280A-2280N (sometimes called core slices), each modular core having multiple sub-cores 2250A-2250N, 2260A-2260N (sometimes called core sub-slices). In at least one embodiment, the graphics processor 2200 may have any number of graphics cores 2280A-2280N. In at least one embodiment, the graphics processor 2200 includes a graphics core 2280A having at least a first sub-core 2250A and a second sub-core 2260A. In at least one embodiment, the graphics processor 2200 is a low-power processor having a single sub-core (e.g., 2250A). In at least one embodiment, the graphics processor 2200 includes a plurality of graphics cores 2280A to 2280N, each of which includes a first set of subcores 2250A to 2250N and a second set of subcores 2260A to 2260N. In at least one embodiment, each of the first subcores 2250A to 2250N includes at least a first set of execution units 2252A to 2252N and media / texture samplers 2254A to 2254N. In at least one embodiment, each of the second subcores 2260A to 2260N includes at least a second set of execution units 2262A to 2262N and samplers 2264A to 2264N. In at least one embodiment, each sub-core 2250A-2250N, 2260A-2260N shares a set of shared resources 2270A-2270N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.
[0218] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics processor 2200 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0219] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0220] Figure 23 is a block diagram showing the microarchitecture of a processor 2300, which may include logic circuits for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2300 may execute instructions including x86 instructions, AMR instructions, and special instructions for application-specific integrated circuits (ASICs). In at least one embodiment, the processor 2300 may include registers for storing packed data, such as 64-bit wide MMX™ registers in a microprocessor enabled by MMX technology, as provided by Intel Corporation in Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point formats, may operate with packed data elements accompanied by Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit wide XMM registers relating to SSE2, SSE3, SSE4, AVX, or higher technologies (collectively referred to as "SSEx") may hold operands of such packed data. In at least one embodiment, the processor 2300 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0221] In at least one embodiment, the processor 2300 includes an in-order front-end ("front-end") 2301 that fetches instructions to be executed and prepares instructions to be used later in the processor pipeline. In at least one embodiment, the front-end 2301 may include several units. In at least one embodiment, an instruction prefetcher 2326 fetches instructions from memory and supplies them to an instruction decoder 2328, which decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2328 decodes the received instruction into one or more operations called "microinstructions" or "microoperations" (also called "microops" or "uops") that the machine can execute. In at least one embodiment, the instruction decoder 2328 may parse the instruction into opcodes and corresponding data, as well as control fields, which are used by the microarchitecture to perform the operation according to at least one embodiment. In at least one embodiment, the trace cache 2330 may assemble the decoded uops into a program-order sequence or trace in the uop queue 2334 so that they can be executed. In at least one embodiment, when the trace cache 2330 encounters a complex instruction, the microcode ROM 2332 provides the uops necessary to complete the operation.
[0222] In at least one embodiment, some instructions can be converted into a single micro-ops, while others require several micro-ops to complete the entire operation. In at least one embodiment, if five or more micro-ops are required to complete an instruction, the instruction decoder 2328 may access the microcode ROM 2332 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-ops so that it can be processed by the instruction decoder 2328. In at least one embodiment, if many micro-ops are required to complete the operation, the instruction may be stored in the microcode ROM 2332. In at least one embodiment, the trace cache 2330 refers to the entry-point programmable logic array ("PLA") to determine the correct microinstruction pointer for reading the microcode sequence in order to complete one or more instructions from the microcode ROM 2332 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2332 has finished sequencing microops for instructions, the machine's front-end 2301 may resume fetching microops from the trace cache 2330.
[0223] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2303 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has a large number of buffers to smooth the flow of instructions and change their order, optimizing performance as instructions are pipelined and scheduled for execution. In at least one embodiment, the out-of-order execution engine 2303 includes, without limitation, an allocator / register renamer 2340, a memory uop queue 2342, an integer / floating-point uop queue 2344, a memory scheduler 2346, a fast scheduler 2302, a slow / general-purpose floating-point scheduler ("slow / general-purpose FP: floating-point scheduler") 2304, and a simple floating-point scheduler ("simple FP scheduler") 2306. In at least one embodiment, the fast scheduler 2302, the slow / general-purpose floating-point scheduler 2304, and the simple floating-point scheduler 2306 are collectively referred to herein as "uop schedulers 2302, 2304, and 2306". In at least one embodiment, the allocator / register renamer 2340 allocates the machine buffers and resources that each uop needs to run. In at least one embodiment, the allocator / register renamer 2340 renames logical registers upon entry into the register file. In at least one embodiment, the allocator / register renamer 2340 also allocates the entries of each uop to one of two uop queues, namely the memory uop queue 2342 for memory operations and the integer / floating-point uop queue 2344 for non-memory operations, prior to the memory scheduler 2346 and the uop schedulers 2302, 2304, 2306. In at least one embodiment, the uop schedulers 2302, 2304, 2306 determine when uops are ready to execute based on whether the sources for their dependent input register operands are ready and whether the execution resources required by the uop to complete their operations are available.In at least one embodiment, the high-speed scheduler 2302 of at least one embodiment may schedule every half of the main clock cycle, while the slow / general-purpose floating-point scheduler 2304 and the simple floating-point scheduler 2306 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2302, 2304, and 2306 arbitrate dispatch ports to schedule uops so that they can be executed.
[0224] In at least one embodiment, the execution block 2311 includes, without limitation, an integer register file / bypass network 2308, a floating-point register file / bypass network ("FP register file / bypass network") 2310, address generation units ("AGUs") 2312 and 2314, fast arithmetic logic units (ALUs) ("fast ALUs") 2316 and 2318, a slow arithmetic logic unit ("slow ALU") 2320, a floating-point ALU ("FP") 2322, and a floating-point move unit ("FP move") 2324. In at least one embodiment, the integer register file / bypass network 2308 and the floating-point register file / bypass network 2310 are also referred to herein as "register files 2308, 2310". In at least one embodiment, AGU2312 and 2314, high-speed ALU2316 and 2318, low-speed ALU2320, floating-point ALU2322, and floating-point movement unit 2324 are also referred to herein as “execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324”. In at least one embodiment, execution block b11 may include, without limitation, any number and type of register files (including zero), bypass networks, address generation units, and execution units in any combination.
[0225] In at least one embodiment, register files 2308, 2310 may be located between the uop schedulers 2302, 2304, 2306 and the execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324. In at least one embodiment, the integer register file / bypass network 2308 performs integer arithmetic. In at least one embodiment, the floating-point register file / bypass network 2310 performs floating-point arithmetic. In at least one embodiment, each of the register files 2308, 2310 may include, without limitation, a bypass network which may bypass or transfer newly completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, the register files 2308, 2310 may communicate data with each other. In at least one embodiment, the integer register file / bypass network 2308 may include, without limitation, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, since floating-point instructions typically have operands of 64 to 128 bits in width, the floating-point register file / bypass network 2310 may include, without limitation, 128-bit wide entries.
[0226] In at least one embodiment, execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324 may execute instructions. In at least one embodiment, register files 2308 and 2310 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2300 may include, without limitation, any number and combination of execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324. In at least one embodiment, floating-point ALU 2322 and floating-point movement unit 2324 may perform floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, the floating-point ALU2322 may include, without limitation, 64-bit floating-point dividers and perform division, square root, and the remaining micro-operations. In at least one embodiment, instructions involving floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to high-speed ALU2316, 2318. In at least one embodiment, high-speed ALU2316, 2318 may perform high-speed operations with an effective latency of half a clock cycle. In at least one embodiment, the slow ALU2320 may include, without limitation, integer execution hardware for long-latency types of operations such as multipliers, shifts, flag logic, and branching, so that most complex integer operations are passed to the slow ALU2320. In at least one embodiment, memory load / store operations may be performed by AGUS2312, 2314. In at least one embodiment, the high-speed ALU2316, high-speed ALU2318, and low-speed ALU2320 may perform integer arithmetic with 64-bit data operands. In at least one embodiment, the high-speed ALU2316, high-speed ALU2318, and low-speed ALU2320 may be implemented to support a variety of data bit sizes, including 16, 32, 128, 256, and so on. In at least one embodiment, the floating-point ALU2322 and floating-point movement unit 2324 may be implemented to support a wide range of operands with various bit widths.In at least one embodiment, the floating-point ALU 2322 and the floating-point movement unit 2324 may operate in conjunction with SIMD and multimedia instructions as 128-bit wide packed-data operands.
[0227] In at least one embodiment, the uop schedulers 2302, 2304, and 2306 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, uops may be speculatively scheduled and executed in processor 2300, so processor 2300 may also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed the scheduler with temporarily inaccurate data. In at least one embodiment, a replay mechanism tracks and redelivers instructions that use inaccurate data. In at least one embodiment, dependent operations may need to be replayed, while independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0228] In at least one embodiment, the term “register” may refer to an onboard processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be accessible from outside the processor (from the programmer’s perspective). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuitry within the processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, or a combination of dedicated and dynamically allocated physical registers. In at least one embodiment, an integer register stores 32-bit integer data. The register file in at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0229] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the execution block 2311 and other memories or registers, whether illustrated or not. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in the execution block 2311. Furthermore, weight parameters may be stored in on-chip or off-chip memories and / or registers (illustrated or not illustrated) that constitute the ALUs of the execution block 2311 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0230] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0231] Figure 24 shows a deep learning application processor 2400 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2400 uses instructions that cause the deep learning application processor 2400 to perform some or all of the processes and techniques described throughout this disclosure when executed by the deep learning application processor 2400. In at least one embodiment, the deep learning application processor 2400 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2400 performs a matrix multiplication operation which is "hardwired" to hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2400 includes, without limitation, processing clusters 2410(1) to 2410(12), inter-chip links ("ICL") 2420(1) to 2420(12), inter-chip controllers ("ICC") 2430(1) to 2430(2), memory controllers ("Mem Ctrlrs") 2442(1) to 2442(4), high-bandwidth memory physical layers ("HBM PHY") 2444(1) to 2444(4), a management-controller central processing unit ("management-controller CPU") 2450, serial peripheral interfaces, inter-integrated and general-purpose input / output blocks ("SPI, I2C, GPIO"), peripheral component interconnect express controllers and direct memory access blocks ("PCIe controllers and DMA") 2470, and a 16-lane peripheral component interconnect express port ("PCIe"). Includes ExpressX16 (2480).
[0232] In at least one embodiment, the processing cluster 2410 may perform deep learning operations, including inference or prediction operations, based on weight parameters computed using one or more training techniques, including the techniques described herein. In at least one embodiment, each processing cluster 2410 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2400 may include any number and type of processing clusters 2400. In at least one embodiment, the inter-chip link 2420 is bidirectional. In at least one embodiment, the inter-chip link 2420 and the inter-chip controller 2430 enable multiple deep learning application processors 2400 to exchange information, including activation information obtained as a result of executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2400 may include any number and type of ICL2420 and ICC2430 (including zero).
[0233] In at least one embodiment, the HBM2 2440 provides a total of 32 gigabytes (GB) of memory. The HBM2 2440(i) is associated with both the memory controller 2442(i) and the HBM PHY 2444(i). In at least one embodiment, any number of HBM2 2440s may provide any type and total amount of high-bandwidth memory and may be associated with any number and type of memory controllers 2442 and HBM PHY 2444 (including zero). In at least one embodiment, SPI, I2C, GPIO 2460, PCIe controller and DMA 2470, and / or PCIe 2480 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically viable way.
[0234] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, the deep learning application processor 2400 is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2400. In at least one embodiment, the deep learning application processor 2400 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2400. In at least one embodiment, the processor 2400 may be used to perform one or more neural network use cases described herein.
[0235] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0236] Figure 25 is a block diagram of a neuromorphic processor 2500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2500 receives one or more inputs from an external source. In at least one embodiment, these inputs may be transmitted to one or more neurons 2502 within the neuromorphic processor 2500. In at least one embodiment, the neurons 2502 and their components may be implemented using circuits or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2500 may include thousands or millions of instances of neurons 2502, but any preferred number of neurons 2502 may be used. In at least one embodiment, each instance of neuron 2502 may include a neuron input 2504 and a neuron output 2506. In at least one embodiment, neuron 2502 may produce an output which may be transmitted to the input of other instances of neuron 2502. For example, in at least one embodiment, the neuron input 2504 and the neuron output 2506 may be interconnected via a synapse 2508.
[0237] In at least one embodiment, neuron 2502 and synapse 2508 may be interconnected so that the neuromorphic processor 2500 can operate to process or analyze the information it receives. In at least one embodiment, neuron 2502 may transmit an output pulse (or “fire” or “spike”) when the input received via neuron input 2504 exceeds a threshold. In at least one embodiment, neuron 2502 may sum or integrate the signals received at neuron input 2504. For example, in at least one embodiment, neuron 2502 may be implemented as a leaky integrate-and-fire neuron, where neuron 2502 may generate an output (or “fire”) using a transfer function such as a sigmoid function or a threshold function when the sum (called “membrane potential”) exceeds a threshold. In at least one embodiment, the leakage integral firing neuron may sum the signals received at neuron input 2504 to obtain the membrane potential, or it may apply a collapse factor (or leak) to reduce the membrane potential. In at least one embodiment, the leakage integral firing neuron may fire if multiple input signals are received at neuron input 2504 quickly enough to exceed a threshold (i.e., before the membrane potential collapse is too small to fire). In at least one embodiment, neuron 2502 may be implemented using a circuit or logic that receives the input, integrates the input to obtain the membrane potential, and collapses the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 2502 may include, without limitation, a comparator circuit or logic that generates an output spike at neuron 2506 when the result of applying the transfer function to neuron 2504 exceeds a threshold. In at least one embodiment, once neuron 2502 fires, it may ignore previously received input information, for example, by resetting the membrane potential to 0 or another suitable default value.In at least one embodiment, once the membrane potential is reset to 0, neuron 2502 may resume normal operation after a suitable period (or refractory period).
[0238] In at least one embodiment, neurons 2502 may be interconnected through synapses 2508. In at least one embodiment, synapses 2508 may operate to transmit a signal from the output of a first neuron 2502 to the input of a second neuron 2502. In at least one embodiment, neuron 2502 may transmit information through two or more instances of synapses 2508. In at least one embodiment, one or more instances of neuron output 2506 may be connected through an instance of synapse 2508 to an instance of neuron input 2504 of the same neuron 2502. In at least one embodiment, an instance of neuron 2502 that generates an output to be transmitted through an instance of synapse 2508 may be called a “presynaptic neuron” with respect to that instance of synapse 2508. In at least one embodiment, an instance of neuron 2502 that receives an input to be transmitted through an instance of synapse 2508 may be called a “postsynaptic neuron” with respect to that instance of synapse 2508. In at least one embodiment, an instance of neuron 2502 may receive input from one or more instances of synapse 2508 and transmit output through one or more instances of synapse 2508, so that a single instance of neuron 2502 may be both a "presynaptic neuron" and a "postsynaptic neuron" with respect to various instances of synapse 2508.
[0239] In at least one embodiment, the neurons 2502 may be organized into one or more layers. Each instance of neuron 2502 may have one neuron output 2506 that can fan out to one or more neuron inputs 2504 through one or more synapses 2508. In at least one embodiment, the neuron output 2506 of a neuron 2502 in a first layer 2510 may be connected to the neuron input 2504 of a neuron 2502 in a second layer 2512. In at least one embodiment, layer 2510 may be called a “feedforward” layer. In at least one embodiment, each instance of neuron 2502 in an instance of the first layer 2510 may fan out to each instance of neuron 2502 in the second layer 2512. In at least one embodiment, the first layer 2510 may be called a “fully connected feedforward layer”. In at least one embodiment, each instance of neuron 2502 in an instance of the second layer 2512 may be fanned out to fewer instances of neuron 2502 in the third layer 2514 than the total number of instances of neuron 2502 in the third layer 2514. In at least one embodiment, the second layer 2512 may be called a “loosely connected feedforward layer”. In at least one embodiment, neurons 2502 in the second layer 2512 may be fanned out to neurons 2502 in multiple other layers, including neurons 2502 in the (same) second layer 2512. In at least one embodiment, the second layer 2512 may be called a “regression layer”. In at least one embodiment, the neuromorphic processor 2500 may include, without limitation, any preferred combination of regression layers and feedforward layers, including, without limitation, both loosely connected feedforward layers and fully connected feedforward layers.
[0240] In at least one embodiment, the neuromorphic processor 2500 may include, without limitation, a reconfigurable interconnect architecture or dedicated hardwired interconnect for connecting synapses 2508 to neurons 2502. In at least one embodiment, the neuromorphic processor 2500 may include, without limitation, circuits or logic that, based on the neural network topology and the fan-in / fan-out of neurons, allow synapses to be allocated to different neurons 2502 as needed. For example, in at least one embodiment, synapse 2508 may be connected to neuron 2502 using an interconnect fabric such as a network-on-a-chip or using a dedicated connection. In at least one embodiment, synaptic interconnects and their components may be implemented using circuits or logic.
[0241] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0242] Figure 26 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2600 includes one or more processors 2602 and one or more graphics processors 2608, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 2602 or processor cores 2607. In at least one embodiment, system 2600 is a processing platform embedded in a system-on-a-chip (SoC) integrated circuit for use in a mobile device, portable device, or embedded device.
[0243] In at least one embodiment, system 2600 may include, or be incorporated into, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable game console, or an online game console. In at least one embodiment, system 2600 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2600 may also include, be coupled to, or be integrated into wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2600 is a television or set-top box device having one or more processors 2602 and a graphical interface produced by one or more graphics processors 2608.
[0244] In at least one embodiment, each of the one or more processors 2602 includes one or more processor cores 2607 for processing instructions that, when executed, perform actions for the system and user software. In at least one embodiment, each of the one or more processor cores 2607 is configured to process a particular instruction set 2609. In at least one embodiment, the instruction set 2609 may facilitate computing via composite instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction words (VLIW). In at least one embodiment, each of the processor cores 2607 may process a different instruction set 2609, which may include instructions that facilitate the emulation of other instruction sets. In at least one embodiment, the processor cores 2607 may also include other processing devices, such as a digital signal processor (DSP).
[0245] In at least one embodiment, the processor 2602 includes cache memory 2604. In at least one embodiment, the processor 2602 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of the processor 2602. In at least one embodiment, the processor 2602 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which may be shared among processor cores 2607 using known cache coherence techniques. In at least one embodiment, the processor 2602 further includes a register file 2606, which may contain different types of registers for storing different types of data (e.g., integer registers, floating-point registers, state registers, and instruction pointer registers). In at least one embodiment, the register file 2606 may contain general-purpose registers or other registers.
[0246] In at least one embodiment, one or more processors 2602 are coupled to one or more interface buses 2610 to transmit communication signals, such as addresses, data, or control signals, between the processors 2602 and other components in the system 2600. In at least one embodiment, the interface bus 2610 may be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface 2610 is not limited to a DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2602 includes an integrated memory controller 2616 and a platform controller hub 2630. In at least one embodiment, the memory controller 2616 facilitates communication between memory devices and other components of the system 2600, while the platform controller hub (PCH) 2630 provides connectivity to I / O devices via a local I / O bus.
[0247] In at least one embodiment, the memory device 2620 may be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase-change memory device, or any other memory device having performance suitable for acting as process memory. In at least one embodiment, the memory device 2620 may operate as system memory for the system 2600 and store data 2622 and instructions 2621 for use by one or more processors 2602 when executing applications or processes. In at least one embodiment, the memory controller 2616 may also be coupled to an optional external graphics processor 2612, which may communicate with one or more graphics processors 2608 within the processor 2602 to perform graphics and media operations. In at least one embodiment, a display device 2611 may be connected to the processor 2602. In at least one embodiment, the display device 2611 may include one or more internal display devices, such as a mobile electronic device or laptop device, or external display devices that are attached via a display interface (e.g., a DisplayPort). In at least one embodiment, the display device 2611 may include a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.
[0248] In at least one embodiment, the platform controller hub 2630 allows peripheral devices to connect to the memory device 2620 and processor 2602 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2646, a network controller 2634, a firmware interface 2628, a wireless transceiver 2626, a touch sensor 2625, and a data storage device 2624 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2624 may be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a peripheral component interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2625 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2626 may be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2628 may enable communication with system firmware, which may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2634 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) may be coupled to the interface bus 2610. In at least one embodiment, the audio controller 2646 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2600 may include an optional legacy I / O controller 2640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 2630 can also connect to one or more connected input devices of the Universal Serial Bus (USB) controller 2642, such as a keyboard and mouse combination 2643, a camera 2644, or other USB input devices.
[0249] In at least one embodiment, instances of the memory controller 2616 and the platform controller hub 2630 may be integrated with a separate external graphics processor, such as an external graphics processor 2612. In at least one embodiment, the platform controller hub 2630 and / or the memory controller 2616 may be external to one or more processors 2602. For example, in at least one embodiment, the system 2600 may include an external memory controller 2616 and a platform controller hub 2630, which may be configured as a memory controller hub and peripheral controller hub within a system chipset that communicate with the processor 2602.
[0250] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the graphics processor 2600. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor 2612. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or 9B. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (illustrated or not illustrated) that constitute the ALUs of the graphics processor 2600 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0251] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0252] Figure 27 is a block diagram of a processor 2700 having one or more processor cores 2702A-2702N, an integrated memory controller 2714, and an integrated graphics processor 2708, according to at least one embodiment. In at least one embodiment, the processor 2700 may include no more than a number of additional cores, including additional cores 2702N represented by dashed rectangles. In at least one embodiment, each of the processor cores 2702A-2702N includes one or more internal cache units 2704A-2704N. In at least one embodiment, each processor core also has access to one or more shared cache units 2706.
[0253] In at least one embodiment, the internal cache units 2704A-2704N and the shared cache unit 2706 represent a cache memory hierarchy within the processor 2700. In at least one embodiment, the cache memory units 2704A-2704N may include at least one level of instruction and data cache within each processor core, as well as one or more levels of shared intermediate level caches such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, where the highest level cache prior to external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2706 and 2704A-2704N.
[0254] In at least one embodiment, the processor 2700 may also include a set of one or more bus controller units 2716 and a system agent core 2710. In at least one embodiment, one or more bus controller units 2716 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2710 provides management functions for various processor components. In at least one embodiment, the system agent core 2710 includes one or more integrated memory controllers 2714 for managing access to various external memory devices (not shown).
[0255] In at least one embodiment, one or more of the processor cores 2702A - 2702N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2710 includes components for coordinating and operating cores 2702A - 2702N during multithreaded processing. In at least one embodiment, the system agent core 2710 may further include a power control unit (PCU), which includes logic and components for adjusting the power state of one or more of the processor cores 2702A - 2702N and the graphics processor 2708.
[0256] In at least one embodiment, the processor 2700 further includes a graphics processor 2708 for performing graphics processing operations. In at least one embodiment, the graphics processor 2708 is coupled to a shared cache unit 2706 and a system agent core 2710 that includes one or more integrated memory controllers 2714. In at least one embodiment, the system agent core 2710 also includes a display controller 2711 for causing the graphics processor to output to one or more attached displays. In at least one embodiment, the display controller 2711 may also be a separate module coupled to the graphics processor 2708 via at least one interconnect, or may be integrated within the graphics processor 2708.
[0257] In at least one embodiment, a ring - based interconnect unit 2712 is used to couple the internal components of the processor 2700. In at least one embodiment, alternative interconnect units such as point - to - point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 2708 is coupled to the ring interconnect 2712 via an I / O link 2713.
[0258] In at least one embodiment, the I / O link 2713 represents at least one of a variety of I / O interconnects including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2718 such as an eDRAM module. In at least one embodiment, each of the processor cores 2702A - 2702N and the graphics processor 2708 uses the embedded memory module 2718 as a shared last-level cache.
[0259] In at least one embodiment, the processor cores 2702A - 2702N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 2702A - 2702N are heterogeneous from the perspective of the instruction set architecture (ISA), where one or more of the processor cores 2702A - 2702N execute a common instruction set, but one or more of the other cores of the processor cores 2702A - 2702N execute a subset of the common instruction set, or a different instruction set. In at least one embodiment, the processor cores 2702A - 2702N are heterogeneous from the perspective of the microarchitecture, where one or more cores with a relatively high power consumption are coupled with one or more cores with a lower power consumption. In at least one embodiment, the processor 2700 can be implemented on one or more chips or as a SoC integrated circuit.
[0260] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the processor 2700. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor 2612, graphics cores 2702A-2702N, or other components in Figure 27. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or Figure 9B. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (illustrated or not illustrated) that constitute the ALUs of the graphics processor 2700 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0261] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0262] Figure 28 is a block diagram of the hardware logic of a graphics processor core 2800 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 2800 is contained within a graphics core array. In at least one embodiment, the graphics processor core 2800, sometimes referred to as a core slice, can be one or more graphics cores in a modular graphics processor. In at least one embodiment, the graphics processor core 2800 is an example of one graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on the desired power and performance envelope. In at least one embodiment, each graphics core 2800 may include a fixed-function block 2830 coupled with a plurality of subcores 2801A to 2801F, also referred to as sub-slices, which include modular blocks of general-purpose and fixed-function logic.
[0263] In at least one embodiment, the fixed-function block 2830 includes a geometry / fixed-function pipeline 2836 that can be shared by all subcores within the graphics processor 2800, for example, in a low-performance and / or low-power graphics processor implementation. In at least one embodiment, the geometry / fixed-function pipeline 2836 includes a 3D fixed-function pipeline, a video front-end unit, a thread spawner and thread dispatcher, and an integrated return buffer manager that manages an integrated return buffer.
[0264] In at least one embodiment, the fixed function block 2830 also includes a graphics SoC interface 2837, a graphics microcontroller 2838, and a media pipeline 2839. In at least one embodiment, the fixed graphics SoC interface 2837 provides an interface between the graphics core 2800 and other processor cores in the system-on-chip integrated circuit. In at least one embodiment, the graphics microcontroller 2838 is a programmable sub-processor configurable to manage various functions of the graphics processor 2800, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 2839 includes logic to facilitate decoding, encoding, preprocessing, and / or postprocessing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 2839 implements media operations via requests to compute logic or sampling logic within sub-cores 2801-2801F.
[0265] In at least one embodiment, the SoC interface 2837 enables the graphics core 2800 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache memory, system RAM, and / or embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2837 also enables communication with fixed-function devices within the SoC, such as a camera imaging pipeline, and enables and / or implements the use of a global memory atomic that can be shared between the graphics core 2800 and the CPU within the SoC. In at least one embodiment, the SoC interface 2837 can also implement power management control for the graphics core 2800 and interface the clock domain of the graphics core 2800 with other clock domains within the SoC. In at least one embodiment, the SoC interface 2837 is configured to receive command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of the graphics cores in the graphics processor. In at least one embodiment, commands and instructions can be dispatched to the media pipeline 2839 when media operations are performed, or to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 2836, geometry and fixed-function pipeline 2814) when graphics processing operations are performed.
[0266] In at least one embodiment, the graphics microcontroller 2838 can be configured to perform various scheduling and management tasks for the graphics core 2800. In at least one embodiment, the graphics microcontroller 2838 can execute graphics and / or compute workload scheduling on various graphics parallel engines in the execution unit (EU) arrays 2802A-2802F and 2804A-2804F within the sub-cores 2801A-2801F. In at least one embodiment, host software running on the CPU core of the SoC, including the graphics core 2800, can send workloads to one of several graphics processor doorbells, which then invoke scheduling operations to the appropriate graphics engine. In at least one embodiment, scheduling operations include determining which workload should be executed next, sending workloads to command streamers, preempting existing workloads running on engines, managing workload progress, and notifying the host software when a workload is complete. In at least one embodiment, the graphics microcontroller 2838 can also facilitate a low-power or idle state for the graphics core 2800 and provide the graphics core 2800 with the ability to save and restore registers within the graphics core 2800 throughout the transition to a low-power state, independently of the operating system and / or graphics driver software on the system.
[0267] In at least one embodiment, the graphics core 2800 may have up to N modular subcores, more or less than the illustrated subcores 2801A to 2801F. For each set of N subcores, in at least one embodiment, the graphics core 2800 may also include shared function logic 2810, shared and / or cache memory 2812, geometry / fixed function pipeline 2814, and additional fixed function logic 2816 for accelerating various graphics and computing processing operations. In at least one embodiment, the shared function logic 2810 may include logic units (e.g., sampler, mathematical, and / or inter-thread communication logic) that can be shared by each of the N subcores in the graphics core 2800. In at least one embodiment, the fixed shared and / or cached memory 2812 can serve as a last-level cache for N sub-cores 2801A-2801F within the graphics core 2800, and can also serve as shared memory accessible by multiple sub-cores. In at least one embodiment, the geometry / fixed function pipeline 2814 may be included in place of the geometry / fixed function pipeline 2836 within the fixed function block 2830, and may include the same or similar logical units.
[0268] In at least one embodiment, the graphics core 2800 includes an additional fixed-function logic 2816 which can include various fixed-function acceleration logic for use by the graphics core 2800. In at least one embodiment, the additional fixed-function logic 2816 includes an additional geometry pipeline for use with position-only shading. In position-only shading, there are at least two geometry pipelines: a full geometry pipeline and a cull pipeline within geometry / fixed-function pipelines 2816, 2836, the cull pipeline being an additional geometry pipeline which may be included within the additional fixed-function logic 2816. In at least one embodiment, the cull pipeline is a reduced version of the full geometry pipeline. In at least one embodiment, the full pipeline and the cull pipeline can run different instances of the application, each instance having a separate context. In at least one embodiment, position-only shading can hide the long cull run of truncated triangles, allowing shading to complete faster in some instances. For example, in at least one embodiment, the sorting pipeline fetches and shades vertex position attributes without rasterizing and rendering pixels into a frame buffer, so that the sorting pipeline logic within the additional fixed-function logic 2816 can run the position shader in parallel with the main application and produce a critical result overall faster than the full pipeline. In at least one embodiment, the sorting pipeline can use the generated critical result to compute visibility information for all triangles, regardless of whether these triangles are sorted or not. In at least one embodiment, the full pipeline (which may be called the replay pipeline in this instance) can consume the visibility information to shade only the visible triangles, skipping the sorted triangles, which are then passed to the rasterization phase.
[0269] In at least one embodiment, the additional fixed-function logic 2816 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, for implementation forms that include training or inference optimization of machine learning.
[0270] In at least one embodiment, each graphics sub-core 2801A-2801F includes a set of execution resources which may be used to perform graphics operations, media operations, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader program. In at least one embodiment, the graphics sub-cores 2801A-2801F include a plurality of EU arrays 2802A-2802F, 2804A-2804F, thread dispatch and inter-thread communication (TD / IC) logic 2803A-2803F, 3D (e.g., texture) samplers 2805A-2805F, media samplers 2806A-2806F, shader processors 2807A-2807F, and shared local memory (SLM) 2808A-2808F. EU arrays 2802A-2802F and 2804A-2804F each include multiple execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in graphics operations, media operations, or compute operations, including graphics, media, or compute shader programs. In at least one embodiment, TD / IC logic 2803A-2803F perform local thread dispatch and thread control operations for execution units within a subcore, facilitating communication between threads running on the execution units in the subcore. In at least one embodiment, 3D samplers 2805A-2805F can read textures or other 3D graphics-related data into memory. In at least one embodiment, the 3D sampler can read texture data in different ways based on a configured sample state and texture format associated with a given texture. In at least one embodiment, the media samplers 2806A to 2806F can perform similar reading operations based on the type and format associated with the media data.In at least one embodiment, each graphics sub-core 2801A-2801F may alternatively include a 3D and media integrated sampler. In at least one embodiment, threads running on execution units within each sub-core 2801A-2801F may utilize shared local memory 2808A-2808F within each sub-core to enable threads running within a thread group to run using a common pool of on-chip memory.
[0271] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the graphics processor 2810. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the graphics processor 2612, the graphics microcontroller 2838, the geometry and fixed-function pipelines 2814 and 2836, or other ALUs embodied in the logic of Figure 27. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or Figure 9B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (illustrated or not illustrated) that constitute the ALU of the graphics processor 2800 for executing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0272] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0273] Figures 29A and 29B show a thread execution logic 2900 including an array of processing elements of a graphics processor core, according to at least one embodiment. Figure 29A shows at least one embodiment in which the thread execution logic 2900 is used. Figure 29B shows exemplary internal details of an execution unit, according to at least one embodiment.
[0274] As shown in Figure 29A, in at least one embodiment, the thread execution logic 2900 includes a shader processor 2902, a thread dispatcher 2904, an instruction cache 2906, a scalable execution unit array including multiple execution units 2908A-2908N, a sampler 2910, a data cache 2912, and a data port 2914. In at least one embodiment, the scalable execution unit array can be dynamically scaled up or down by enabling or disabling one or more execution units (for example, any of execution units 2908A, 2908B, 2908C, 2908D-2908N-1, and 2908N) based on the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect fabric linked to each execution unit. In at least one embodiment, the thread execution logic 2900 includes one or more connections to memory, such as system memory or cache memory, via an instruction cache 2906, a data port 2914, a sampler 2910, and one or more execution units 2908A to 2908N. In at least one embodiment, each execution unit (e.g., 2908A) is a standalone programmable general-purpose computing unit capable of executing multiple concurrent hardware threads while processing multiple data elements in parallel per thread. In at least one embodiment, the array of execution units 2908A to 2908N is expandable or contractible to include any number of individual execution units.
[0275] In at least one embodiment, execution units 2908A-2908N are primarily used to execute shader programs. In at least one embodiment, shader processor 2902 can process various shader programs and dispatch execution threads associated with shader programs via thread dispatcher 2904. In at least one embodiment, thread dispatcher 2904 includes logic for mediating thread start requests from graphics and media pipelines and for instantiating the requested threads on one or more execution units of execution units 2908A-2908N. For example, in at least one embodiment, the geometry pipeline can dispatch vertex shaders, mosaic shaders, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 2904 can also process runtime thread spawning requests from running shader programs.
[0276] In at least one embodiment, execution units 2908A-2908N support an instruction set that includes native support for many standard 3D graphics shader instructions, thereby enabling shader programs from graphics libraries (e.g., Direct3D and OpenGL) to be executed with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 2908A-2908N, which includes one or more arithmetic logic units (ALUs), can issue multiple single-instruction multiple-data (SIMD) executions, enabling an efficient execution environment despite high memory access latency through multithreaded operation. In at least one embodiment, each hardware thread within each execution unit has its own dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, executions are issued multiple times per clock to a pipeline capable of performing integer operations, single-precision and double-precision floating-point operations, SIMD branching, logical operations, transcendental operations, and various other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, the dependent logic within execution units 2908A-2908N puts the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources may be dedicated to processing other threads. For example, in at least one embodiment, during delays associated with vertex shader operation, the execution unit may execute a pixel shader, a fragment shader, or another type of shader program including a different vertex shader.
[0277] In at least one embodiment, each execution unit 2908A-2908N operates on an array of data elements. In at least one embodiment, the number of data elements is the “execution size,” or the number of channels for an instruction. In at least one embodiment, an execution channel is a logical unit of execution for accessing, masking, and controlling the flow within an instruction for data elements. In at least one embodiment, the number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating-point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 2908A-2908N may support integer and floating-point data types.
[0278] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements may be stored in registers as packed data types, and the execution unit processes various elements based on the data size of the elements. For example, in at least one embodiment, when operating on a 256-bit wide vector, 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quad-word (QW) sized data elements), eight separate 32-bit packed data elements (double-word (DW) sized data elements), sixteen separate 16-bit packed data elements (word (W) sized data elements), or thirty-two separate 8-bit data elements (byte (B) sized data elements). However, in at least one embodiment, different vector widths and register sizes are possible.
[0279] In at least one embodiment, one or more execution units can be combined into fused execution units 2909A - 2909N having thread control logic (2907A - 2907N) common to the fused EUs. In at least one embodiment, multiple EUs can be fused into an EU group. In at least one embodiment, each EU in a fused EU group can be configured to execute separate SIMD hardware threads. The number of EUs in a fused EU group may vary according to different embodiments. In at least one embodiment, various SIMD widths including, but not limited to, SIMD8, SIMD16, and SIMD32 can be executed per EU. In at least one embodiment, each fused graphics execution unit 2909A - 2909N includes at least two execution units. For example, in at least one embodiment, fused execution unit 2909A includes a first EU 2908A, a second EU 2908B, and thread control logic 2907A common to the first EU 2908A and the second EU 2908B. In at least one embodiment, thread control logic 2907A controls the threads executed in fused graphics execution unit 2909A to enable each EU within fused execution units 2909A - 2909N to be executed using a common instruction pointer register.
[0280] In at least one embodiment, one or more internal instruction caches (e.g., 2906) are included in thread execution logic 2900 to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 2912) are included to cache thread data during thread execution. In at least one embodiment, sampler 2910 is included to perform texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 2910 includes special texture or media sampling functions and processes texture or media data during sampling before providing the sampled data to the execution units.
[0281] During execution, in at least one embodiment, the graphics and media pipeline sends a thread start request to the thread execution logic 2900 via the thread spawning and dispatch logic. In at least one embodiment, once a group of geometric objects has been processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) in the shader processor 2902 is invoked to further compute the output information and write the results to the output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates the values of various vertex attributes that will be interpolated between the rasterized objects. In at least one embodiment, the pixel processor logic in the shader processor 2902 then executes a pixel shader program or fragment shader program with an application programming interface (API). In at least one embodiment, to execute a shader program, the shader processor 2902 dispatches threads to execution units (e.g., 2908A) via the thread dispatcher 2904. In at least one embodiment, the shader processor 2902 accesses texture data of a texture map stored in memory using the texture sampling logic of the sampler 2910. In at least one embodiment, arithmetic operations on the texture data and input geometry data are performed to compute or truncate one or more pixels of each geometry fragment so that they are not further processed.
[0282] In at least one embodiment, the data port 2914 provides a memory access mechanism for the thread execution logic 2900 to output processed data to memory for further processing in the graphics processor output pipeline. In at least one embodiment, the data port 2914 includes, or is coupled to, one or more cache memories (e.g., data cache 2912) to cache data for memory access through the data port.
[0283] As shown in Figure 29B, in at least one embodiment, the graphics execution unit 2908 may include an instruction fetch unit 2937, a general register file array (GRF) 2924, an architecture register file array (ARF) 2926, a thread arbitrator (arbiter) 2922, a transmit unit 2930, a branch unit 2932, a set of SIMD floating-point units (FPUs) 2934, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 2935. In at least one embodiment, the GRF 2924 and ARF 2926 include sets of general register files and architecture register files associated with each concurrent hardware thread, which may be active in the graphics execution unit 2908. In at least one embodiment, per-thread architecture state is maintained in the ARF 2926, and data used during thread execution is stored in the GRF 2924. In at least one embodiment, the execution state of each thread, including the instruction pointer for each thread, can be stored in a thread-specific register of the ARF2926.
[0284] In at least one embodiment, the graphics execution unit 2908 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and Interleaved Multi-Threading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the number of simultaneous thread targets and registers per execution unit, where the resources of the execution unit are divided across the logic used to execute multiple simultaneous threads.
[0285] In at least one embodiment, the graphics execution unit 2908 can jointly issue multiple instructions, which may each be a different instruction. In at least one embodiment, the thread arbitrator 2922 of the graphics execution unit thread 2908 can dispatch instructions to one of the transmission unit 2930, branch unit 2942, or SIMD FPU 2934 for execution. In at least one embodiment, each execution thread can access 128 general-purpose registers in the GRF2924, where each register can store 32 bytes accessible as a vector of SIMD8 elements of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4K bytes in the GRF2924, but embodiments are not limited in this way, and other embodiments may provide more or less resources. In at least one embodiment, up to 7 threads can run concurrently, but the number of threads per execution unit can also vary depending on the embodiment. In at least one embodiment, where 7 threads can access 4K bytes, the GRF2924 can store a total of 28K bytes. In at least one embodiment, a flexible addressing mode allows multiple registers to be addressed together to construct a wider range of registers or to represent strided rectangular block data structures.
[0286] In at least one embodiment, memory operations, sampler operations, and other high-latency system communications are dispatched via "send" instructions executed by a message delivery transmission unit 2930. In at least one embodiment, branch instructions are dispatched to a dedicated branch unit 2932 to facilitate SIMD divergence and eventual convergence.
[0287] In at least one embodiment, the graphics execution unit 2908 includes one or more SIMD floating-point units (FPUs) 2934 for performing floating-point operations. In at least one embodiment, the FPU 2934 also supports integer operations. In at least one embodiment, the FPU 2934 can perform up to M 32-bit floating-point (or integer) operations SIMD, or up to 2M 16-bit integer operations or 16-bit floating-point operations SIMD. In at least one embodiment, at least one of the FPUs provides extended mathematical capabilities to support high-throughput transcendental mathematical functions and double-precision 64-bit floating-point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 2935 are also present and may be specifically optimized to perform operations related to machine learning computations.
[0288] In at least one embodiment, an array of multiple instances of the graphics execution unit 2908 may be instantiated in a graphics sub-core group (e.g., a sub-slice). In at least one embodiment, the execution unit 2908 can execute instructions across multiple execution channels. In at least one embodiment, each thread executed by the graphics execution unit 2908 runs on a different channel.
[0289] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the execution logic 2900. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or 9B. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (illustrated or not illustrated) that constitute the ALU of the execution logic 2900 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0290] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0291] Figure 30 shows a parallel processing unit ("PPU") 3000 according to at least one embodiment. In at least one embodiment, the PPU 3000 consists of machine-readable code that, when executed by the PPU 3000, causes the PPU 3000 to execute some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the PPU 3000 is a multithreaded processor, which is implemented on one or more integrated circuit devices and utilizes multithreading as a latency-hiding technique designed to process computer-readable instructions (also called machine-readable instructions or simply instructions) in parallel with multiple threads. In at least one embodiment, a thread refers to an execution thread, which is an instantiation of a set of instructions configured to be executed by the PPU 3000. In at least one embodiment, the PPU 3000 is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data that can be displayed on a display device such as a liquid crystal display ("LCD") device. In at least one embodiment, the PPU 3000 is used to perform calculations such as linear algebra and machine learning operations. Figure 30 shows an exemplary parallel processor for illustrative purposes only and should be interpreted as a non-limiting example of the processor architecture intended within the scope of this disclosure, and it should be interpreted that any suitable processor may be used to add to and / or replace such processor.
[0292] In at least one embodiment, one or more PPU3000s are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPU3000 is configured to accelerate deep learning systems and applications, including but not limited to: autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.
[0293] In at least one embodiment, the PPU 3000 includes, but is not limited to, an input / output ("I / O") unit 3006, a front-end unit 3010, a scheduler unit 3012, a work distribution unit 3014, a hub 3016, a crossbar ("Xbar") 3020, one or more general-purpose processing clusters ("GPC") 3018, and one or more partition units ("memory partition units") 3022. In at least one embodiment, the PPU 3000 is connected to a host processor or other PPU 3000 via one or more high-speed GPU interconnects ("GPU interconnects") 3008. In at least one embodiment, the PPU 3000 is connected to a host processor or other peripheral device via interconnect 3002. In at least one embodiment, the PPU 3000 is connected to local memory comprising one or more memory devices ("memory") 3004. In at least one embodiment, the memory device 3004 includes, without limitation, one or more dynamic random-access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices may be configured and / or configurable as a high-bandwidth memory ("HBM") subsystem in which multiple DRAM dies are stacked within each device.
[0294] In at least one embodiment, the high-speed GPU interconnect 3008 may refer to a wired-based multi-lane communication link used by the system to expand and contract and comprising one or more PPUs 3000 in combination with one or more central processing units ("CPUs"), supporting cache coherence and CPU mastering between the PPUs 3000 and the CPUs. In at least one embodiment, data and / or commands are transmitted by the high-speed GPU interconnect 3008 to / from another unit of the PPU 3000 via the hub 3016, such as one or more copy engines, video encoders, video decoders, power management units, and other components that may not be explicitly shown in Figure 30.
[0295] In at least one embodiment, the I / O unit 3006 is configured to send and receive communications (e.g., commands, data) from a host processor (not shown in Figure 30) via the system bus 3002. In at least one embodiment, the I / O unit 3006 communicates with the host processor directly via the system bus 3002 or via one or more intermediate devices such as memory bridges. In at least one embodiment, the I / O unit 3006 may communicate with one or more other processors, such as one or more of the PPUs 3000, via the system bus 3002. In at least one embodiment, the I / O unit 3006 implements a Peripheral Component Interconnect Express ("PCIe") interface to enable communication via the PCIe bus. In at least one embodiment, the I / O unit 3006 implements an interface for communicating with external devices.
[0296] In at least one embodiment, the I / O unit 3006 decodes packets received via the system bus 3002. In at least one embodiment, at least some packets represent commands configured to cause the PPU 3000 to perform various operations. In at least one embodiment, the I / O unit 3006 transmits the decoded commands to various other units of the PPU 3000 specified by the commands. In at least one embodiment, the commands are transmitted to the front-end unit 3010 and / or to the hub 3016, or to one or more other units of the PPU 3000, such as copy engines, video encoders, video decoders, power management units (not explicitly shown in Figure 30). In at least one embodiment, the I / O unit 3006 is configured to route communication between various logical units of the PPU 3000.
[0297] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides the workload to the PPU 3000 for processing. In at least one embodiment, the workload includes instructions and the data to be processed by these instructions. In at least one embodiment, the buffer is a region in memory accessible (e.g., writable / readable) by both the host processor and the PPU 3000, and the host interface unit may be configured to access a buffer in system memory connected to the system bus 3002 via memory requests sent by the I / O unit 3006 over the system bus 3002. In at least one embodiment, the host processor writes a command stream to the buffer and then sends a pointer to the PPU 3000 pointing to the beginning of the command stream, thereby the front-end unit 3010 receives a pointer to one or more command streams, manages one or more command streams, reads commands from the command streams, and forwards the commands to various units of the PPU 3000.
[0298] In at least one embodiment, the front-end unit 3010 is coupled to a scheduler unit 3012 which comprises various GPCs 3018 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 3012 is configured to track status information related to the various tasks managed by the scheduler unit 3012, where the status information may indicate which GPC 3018 a task is assigned to, whether the task is active or inactive, the priority level associated with the task, etc. In at least one embodiment, the scheduler unit 3012 manages the execution of multiple tasks in one or more of the GPCs 3018.
[0299] In at least one embodiment, the scheduler unit 3012 is coupled to a work distribution unit 3014 configured to dispatch tasks for execution on the GPC 3018. In at least one embodiment, the work distribution unit 3014 tracks the number of scheduled tasks received from the scheduler unit 3012, and the work distribution unit 3014 manages a pending task pool and an active task pool for each of the GPC 3018s. In at least one embodiment, the pending task pool comprises several slots (e.g., 32 slots) containing tasks assigned to be processed by a particular GPC 3018, and the active task pool comprises several slots (e.g., 4 slots) for tasks being actively processed by the GPC 3018, so that when one of the GPC 3018s completes the execution of a task, that task is removed from the GPC 3018's active task pool, and one of the other tasks from the pending task pool is selected and scheduled to be executed on the GPC 3018. In at least one embodiment, if an active task is idle on GPC3018, such as while waiting for data dependencies to be resolved, the active task is removed from GPC3018 and returned to the pending task pool, during which time another task is selected from the pending task pool and scheduled to run on GPC3018.
[0300] In at least one embodiment, the work distribution unit 3014 communicates with one or more GPCs 3018 via the X-bar 3020. In at least one embodiment, the X-bar 3020 is an interconnection network that connects many of the units of the PPU 3000 to other units of the PPU 3000, and can be configured to connect the work distribution unit 3014 to a specific GPC 3018. In at least one embodiment, one or more other units of the PPU 3000 may also be connected to the X-bar 3020 via the hub 3016.
[0301] In at least one embodiment, tasks are managed by a scheduler unit 3012 and dispatched to one of the GPCs 3018 by a work distribution unit 3014. The GPC 3018 is configured to process tasks and produce results. In at least one embodiment, the results may be consumed by other tasks within the GPC 3018, routed to a different GPC 3018 via an X-bar 3020, or stored in memory 3004. In at least one embodiment, the results can be written to memory 3004 via a partition unit 3022, which implements a memory interface for reading and writing data to and from memory 3004. In at least one embodiment, the results can be sent to another PPU 3004 or CPU via a high-speed GPU interconnect 3008. In at least one embodiment, the PPU 3000 includes, without limitation, U partition units 3022 equal to the number of separate individual memory devices 3004 coupled to the PPU 3000. In at least one embodiment, the partition unit 3022 is described in more detail below in conjunction with Figure 32.
[0302] In at least one embodiment, the host processor runs a driver kernel that implements an Application Programming Interface (API) that allows one or more applications running on the host processor to schedule operations to run on the PPU3000. In at least one embodiment, multiple compute applications run concurrently on the PPU3000, and the PPU3000 provides isolation, quality of service ("QoS"), and independent address spaces for the multiple compute applications. In at least one embodiment, the application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks to run on the PPU3000, and the driver kernel outputs the tasks to one or more streams being processed by the PPU3000. In at least one embodiment, each task comprises one or more groups of associated threads, which may be called warps. In at least one embodiment, a warp comprises multiple associated threads (e.g., 32 threads) that can run in parallel. In at least one embodiment, a linked thread may refer to multiple threads that include instructions for executing a task and exchange data via shared memory. In at least one embodiment, threads and linked threads are described in further detail in conjunction with Figure 32.
[0303] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the PPU 3000. In at least one embodiment, the PPU 3000 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the PPU 3000. In at least one embodiment, the PPU 3000 may be used to perform one or more neural network use cases described herein.
[0304] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0305] Figure 31 shows a general-purpose processing cluster ("GPC") 3100 according to at least one embodiment. In at least one embodiment, the GPC 3100 is the GPC 3018 in Figure 30. In at least one embodiment, each GPC 3100 includes, but is not limited to, several hardware units for processing tasks, and each GPC 3100 includes, but is not limited to, a pipeline manager 3102, a pre-raster operations unit ("PROP") 3104, a raster engine 3108, a work distribution crossbar ("WDX") 3116, a memory management unit ("MMU") 3118, one or more data processing clusters ("DPC") 3106, and any preferred combination of parts.
[0306] In at least one embodiment, the operation of the GPC3100 is controlled by a pipeline manager 3102. In at least one embodiment, the pipeline manager 3102 manages the configuration of one or more DPC3106 to handle tasks assigned to the GPC3100. In at least one embodiment, the pipeline manager 3102 configures at least one of the one or more DPC3106 to implement at least a portion of the graphics rendering pipeline. In at least one embodiment, the DPC3106 is configured to run a vertex shader program on a programmable streaming multiprocessor ("SM") 3114. In at least one embodiment, the pipeline manager 3102 is configured to route packets received from the work distribution unit to appropriate logical units within the GPC 3100, some of which may be routed to fixed-function hardware units and / or raster engine 3108 of PROP 3104, and other packets may be routed to DPC 3106 to be processed by primitive engine 3112 or SM 3114. In at least one embodiment, the pipeline manager 3102 configures at least one of the DPC 3106 to implement a neural network model and / or computing pipeline.
[0307] In at least one embodiment, the PROP unit 3104 is configured to route data generated by the raster engine 3108 and DPC 3106 to the raster operation (ROP) unit of the partition unit 3022, which is described in more detail above in conjunction with Figure 30. In at least one embodiment, the PROP unit 3104 is configured to perform color blending optimization, organize pixel data, perform address translation, and perform other operations. In at least one embodiment, the raster engine 3108 includes, but is not limited to, several fixed-function hardware units configured to perform various raster operations in at least one embodiment, and the raster engine 3108 includes, but is not limited to, a setup engine, a coarse raster engine, a sorting engine, a clipping engine, a fine raster engine, a tile merging engine, and any preferred combination thereof. In at least one embodiment, the setup engine receives the transformed vertices, generates a plane equation associated with the geometric primitives defined by the vertices, the plane equation is sent to a coarse raster engine to generate coverage information for the primitives (e.g., x,y coverage mask of tiles), the output of the coarse raster engine is sent to a sorting engine where fragments associated with primitives that have fallen into the z test are sorted, and the output is sent to a clipping engine where fragments outside the viewing frustum are clipped. In at least one embodiment, the fragments that have passed clipping and sorting are passed to a fine raster engine to generate attributes for the pixel fragments based on the plane equation generated by the setup engine. In at least one embodiment, the output of the raster engine 3108 includes fragments that will be processed by any preferred entity, such as a fragment shader implemented within the DPC 3106.
[0308] In at least one embodiment, each DPC3106 included in the GPC3100 includes, but is not limited to, an M-Pipe Controller ("MPC") 3110, a primitive engine 3112, one or more SM3114s, and any preferred combination thereof. In at least one embodiment, the MPC3110 controls the operation of the DPC3106 to route packets received from the pipeline manager 3102 to the appropriate unit within the DPC3106. In at least one embodiment, packets associated with vertices are routed to the primitive engine 3112, which is configured to fetch vertex attributes associated with the vertices from memory, while packets associated with shader programs may be sent to the SM3114.
[0309] In at least one embodiment, the SM3114 includes, but is not limited to, a programmable streaming processor configured to handle tasks represented by several threads. In at least one embodiment, the SM3114 is multithreaded and configured to execute multiple threads (e.g., 32 threads) from a particular group of threads simultaneously, implementing a single-instruction multiple-data (SIMD) architecture where each thread in a group of threads (warp) is configured to process a different data set based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instruction. In at least one embodiment, the SM3114 implements a single-instruction multiple-thread (SIMT) architecture where each thread in a thread group is configured to process a different data set based on the same instruction set, but individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained per warp to enable simultaneous processing between warps and serial execution within warps when threads in a warp diverge. In another embodiment, program counters, call stacks, and execution states are maintained for each individual thread, enabling equal concurrent processing across all threads, within warps, and between warps. In at least one embodiment, execution states are maintained for each individual thread, and threads executing the same instruction may converge and execute in parallel for greater efficiency. At least one embodiment of the SM3114 is described in further detail below.
[0310] In at least one embodiment, the MMU3118 provides an interface between the GPC3100 and a memory partition unit (for example, partition unit 3022 in Figure 30), and the MMU3118 provides virtual address-to-physical address translation, memory protection, and memory request arbitration. In at least one embodiment, the MMU3118 provides one or more translation lookaside buffers ("TLBs") for performing virtual address-to-physical address translation of memory.
[0311] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided below in conjunction with Figures 9A and / or 9B. In at least one embodiment, a deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the GPC3100. In at least one embodiment, the GPC3100 is used to infer or predict information by another processor or system, or based on a trained machine learning model (e.g., a neural network) that has been trained by the GPC3100. In at least one embodiment, the GPC3100 may be used to perform one or more neural network use cases described herein.
[0312] In at least one embodiment, such a component can be used to generate improved video using one or more neural networks, such as generating higher frame-rate video from frames of lower frame-rate video.
[0313] Figure 32 shows a memory partition unit 3200 of a parallel processing unit ("PPU") according to at least one embodiment. In at least one embodiment, the partition unit 3200 includes, but is not limited to, a raster operation ("ROP") unit 3202, a level 2 ("L2") cache 3204, a memory interface 3206, and any preferred combination thereof. In at least one embodiment, the memory interface 3206 is coupled to memory. In at least one embodiment, the memory interface 3206 can provide a 32, 64, 128, 1024-bit data bus, or similar implementations, for high-speed data transfer. In at least one embodiment, the PPU incorporates one memory interface 3206 per pair of partition units 3200, where each pair of partition units 3200 is connected to a corresponding memory device. For example, in at least one embodiment, the PPU may be connected to up to Y memory devices, such as a high-bandwidth memory stack or graphics double data rate, version 5, synchronous dynamic random access memory ("GDDR5 SDRAM").
[0314] In at least one embodiment, the memory interface 3206 implements a high-bandwidth memory second generation ("HBM2") memory interface, where Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located in the same physical package as the PPU, resulting in substantial power and area savings compared to conventional GDDR5 SDRAM systems. In at least one embodiment, each HBM2 stack includes, without limitation, four memory dies, where Y is equal to 4, and each HBM2 stack includes a total of eight channels (two 128-bit channels per die) and a 1024-bit data bus width. In at least one embodiment, the memory supports Single-Error Correcting Double-Error Detecting ("SECDED") error correction code ("ECC") to protect data. In at least one embodiment, ECC provides greater reliability for compute applications susceptible to data corruption.
[0315] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 3200 supports integrated memory to provide a single, unified virtual address space for the central processing unit ("CPU") and PPU memory, enabling data sharing between virtual memory systems. In at least one embodiment, the frequency with which the PPU accesses memory located in other processors is tracked to ensure that memory pages are moved to the PPU's physical memory that is accessing the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 3008 supports an address translation service to allow the PPU to directly access the CPU's page table, enabling full access to CPU memory by the PPU.
[0316] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page error for an address not mapped to a page table, the memory partition unit 3200 then maps the address to the page table in response to the page error, and the copy engine then performs the transfer. In at least one embodiment, memory is pinned (e.g., made page-immobile) for multiple operations of the copy engine across multiple processors, thereby substantially reducing the available memory. In at least one embodiment, if there is a hardware page error, the address can be passed to the copy engine regardless of whether the memory page is resident or not, and the copy process is transparent.
[0317] In at least one embodiment, data from memory 3004 in Figure 30 or other system memory is fetched by memory partition unit 3200 and stored in L2 cache 3204, which is located on-chip and shared among various GPCs. In at least one embodiment, each memory partition unit 3200 includes, without limitation, at least a portion of the L2 cache associated with the corresponding memory device. In at least one embodiment, lower levels of cache are implemented in various units within the GPC. In at least one embodiment, each of the SM3114 may implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM3114, and data from L2 cache 3204 is fetched and stored in each of the L1 caches for processing by the functional units of the SM3114. In at least one embodiment, the L2 cache 3204 is coupled to memory interface 3206 and X-bar 3020.
[0318] In at least one embodiment, the ROP unit 3202 performs graphics raster operations related to pixel color, such as color compression and pixel blending. In at least one embodiment, the ROP unit 3202 implements depth testing in conjunction with the raster engine 3108 to receive the depth of sample locations associated with a pixel fragment from the raster engine 3108's sorting engine. In at least one embodiment, the depth is tested against the corresponding depth in the depth buffer of the sample locations associated with the fragment. In at least one embodiment, if the fragment passes the depth test of the sample locations, the ROP unit 3202 updates the depth buffer and sends the result of the depth test to the raster engine 3108. The number of partition units 3200 may differ from the number of GPCs, and it will be understood that each ROP unit 3202 may be coupled to each of the GPCs in at least one embodiment. In at least one embodiment, the ROP unit 3202 tracks packets received from different GPCs and determines which of the results generated by the ROP unit 3202 should be routed through the X-bar 3020.
[0319] Figure 33 shows a streaming multiprocessor ("SM") 3300 according to at least one embodiment. In at least one embodiment, the SM3300 is the SM3114 in Figure 31. In at least one embodiment, the SM3300 includes, but is not limited to, an instruction cache 3302, one or more scheduler units 3304, a register file 3308, one or more processing cores ("cores") 3310, one or more special function units ("SFUs") 3312, one or more load / store units ("LSUs") 3314, an interconnect network 3316, a shared memory / level 1 ("L1") cache 3318, and any preferred combination thereof. In at least one embodiment, a work distribution unit dispatches tasks to run on a General Purpose Processing Cluster ("GPC") of a Parallel Processing Unit ("PPU"), each task is allocated to a specific Data Processing Cluster ("DPC") within the GPC, and if the task is related to a shader program, the task is allocated to one of the SM3300s. In at least one embodiment, a scheduler unit 3304 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks allocated to the SM3300s. In at least one embodiment, the scheduler unit 3304 schedules thread blocks so that they can run as warps of parallel threads, where each thread block is allocated to at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, the scheduler unit 3304 manages multiple different thread blocks, allocates warps to different thread blocks, and then dispatches instructions from multiple different interlocking groups to various functional units (e.g., processing core 3310, SFU 3312, and LSU 3314) during each clock cycle.
[0320] In at least one embodiment, a linked group refers to a programming model for organizing groups of communicating threads, enabling a richer and more efficient representation of parallel decomposition by allowing developers to represent the granularity at which threads communicate. In at least one embodiment, a linked invocation API supports synchronization between thread blocks to enable the execution of parallel algorithms. In at least one embodiment, applications of the conventional programming model provide a single, simple structure for synchronizing linked threads, namely a barrier across all threads in a thread block (e.g., the syncthreads() function). However, in at least one embodiment, the programmer may define thread groups smaller than the granularity of a thread block and synchronize within the defined groups, enabling higher performance, design flexibility, and software reuse in the form of a functional interface across the collective group as a whole. In at least one embodiment, linked groups allow the programmer to explicitly define groups of threads at sub-block (i.e., the same size as a single thread) and multi-block granularity and perform collective actions such as synchronization of threads within the linked group. In at least one embodiment, the programming model supports clean synthesis across software boundaries, thereby allowing libraries and utility functions to be safely synchronized within their local contexts without the need to make assumptions about convergence. In at least one embodiment, the interdependence group primitives enable novel patterns of interdependence parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire grid of thread blocks without limitation.
[0321] In at least one embodiment, the dispatch unit 3306 is configured to send instructions to one or more functional units, and the scheduler unit 3304 includes, without limitation, two dispatch units 3306, enabling the dispatch of two different instructions from the same warp during each clock cycle. In at least one embodiment, each scheduler unit 3304 includes a single dispatch unit 3306 or an additional dispatch unit 3306.
[0322] In at least one embodiment, each SM3300 includes, in at least one embodiment, a register file 3308 that provides a set of registers to the functional units of the SM3300. In at least one embodiment, the register file 3308 is divided among the functional units such that each functional unit is allocated a dedicated portion of the register file 3308. In at least one embodiment, the register file 3308 is divided among different warps being executed by the SM3300, and the register file 3308 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM3300 includes, in least one, a plurality of L processing cores 3310. In at least one embodiment, each SM3300 includes, in least one, a large number (e.g., 128 or more) of individual processing cores 3310. In at least one embodiment, each processing core 3310 includes, without limitation, fully pipelining single-precision, double-precision, and / or mixed-precision processing units, including, without limitation, floating-point arithmetic logic units and integer arithmetic logic units. In at least one embodiment, the floating-point arithmetic logic units implement the IEEE 754-2008 standard for floating-point operations. In at least one embodiment, the processing core 3310 includes, without limitation, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0323] The tensor core is configured to perform matrix operations according to at least one embodiment. In at least one embodiment, one or more tensor cores are included in the processing core 3310. In at least one embodiment, the tensor core is configured to perform matrix operations for deep learning, such as convolution operations for training and inference of neural networks. In at least one embodiment, each tensor core operates on a 4x4 matrix and performs a matrix multiply and accumulate operation D = A × B + C, where A, B, C, and D are 4x4 matrices.
[0324] In at least one embodiment, the inputs A and B for matrix multiplication are 16-bit floating-point matrices, and the matrices C and D for addition are either 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the Tensor Core operates with 16-bit floating-point input data having a 32-bit floating-point sum. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations, resulting in a full-precision product, which is then added using a 32-bit floating-point addition with other intermediate products of a 4x4x4 matrix multiplication. In at least one embodiment, the Tensor Core is used to perform much larger 2D or even higher-dimensional matrix operations built from these smaller elements. In at least one embodiment, APIs such as the CUDA9 C++ API expose special matrix load, matrix multiplication, sum, and matrix store operations for efficient use of the Tensor Core from CUDA-C++ programs. In at least one embodiment, at the CUDA level, the warp-level interface assumes a 16x16 matrix spanning all 32 threads of the warp.
[0325] In at least one embodiment, each SM3300 includes, without limitation, M SFU3312 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU3312 includes, without limitation, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU3312 includes, without limitation, a texture unit configured to perform texture map filtering operations. In at least one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texels) from memory and sampled texture maps to generate sampled texture values for use in a shader program executed by the SM3300. In at least one embodiment, the texture map is stored in shared memory / level 1 cache 3318. In at least one embodiment, the texture unit implements texture operations, such as filtering operations using mip maps (e.g., texture maps with different levels of detail), according to at least one embodiment. In at least one embodiment, each SM3300 includes, without limitation, two texture units.
[0326] Each SM3300, in at least one embodiment, includes, without limitation, N LSUs 3314 that implement load and store operations between the shared memory / L1 cache 3318 and the register file 3308. Each SM3300, in at least one embodiment, includes, without limitation, an interconnection network 3316 that connects each of the functional units to the register file 3308 and connects the LSUs 3314 to the register file 3308, and the shared memory / L1 cache 3318. In at least one embodiment, the interconnection network 3316 may be a crossbar that connects any of the functional units to any of the registers in the register file 3308 and connects the LSUs 3314 to the memory locations of the register file 3308 and the shared memory / L1 cache 3318.
[0327] In at least one embodiment, the shared memory / L1 cache 3318 is, in at least one embodiment, an array of on-chip memory that enables data storage and communication between the SM3300 and the primitive engine, and between threads of the SM3300. In at least one embodiment, the shared memory / L1 cache 3318 has a storage capacity of 128KB, without limitation, and is located on the path from the SM3300 to the partition unit. In at least one embodiment, the shared memory / L1 cache 3318 is used to cache reads and writes. In at least one embodiment, one or more of the shared memory / L1 cache 3318, the L2 cache, and the memory are auxiliary storage.
[0328] In at least one embodiment, performance is improved for both types of memory access by combining data cache and shared memory functionality into a single memory block. In at least one embodiment, the capacity is used or available as a cache by programs that do not use shared memory, so that if the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. According to at least one embodiment, by integrating within the shared memory / L1 cache 3318, the shared memory / L1 cache 3318 can function as a high-throughput conduit for streaming data while simultaneously providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, a simpler configuration can be used when configured for general-purpose parallel computing compared to graphics processing. In at least one embodiment, a fixed-function graphics processing unit is bypassed, resulting in a much simpler programming model. In a general-purpose parallel computing configuration, the work distribution unit directly allocates and distributes thread blocks to the DPC in at least one embodiment. In at least one embodiment, threads within a block execute the same program using a unique thread ID in the computation to ensure that each thread produces a unique result, use the SM3300 to execute the program and perform computations, use the shared memory / L1 cache 3318 to communicate between threads, and use the LSU3314 to read and write to global memory via the shared memory / L1 cache 3318 and the memory partition unit. In at least one embodiment, when configured for general-purpose parallel computing, the SM3300 writes commands that the scheduler unit 3304 can use to start new work on the DCP.
[0329] In at least one embodiment, the PPU is included in or combined with a desktop computer, laptop computer, tablet computer, server, supercomputer, smartphone (e.g., wireless portable device), personal digital assistant ("PDA"), digital camera, vehicle, head-mounted display, portable electronic device, etc. In at least one embodiment, the PPU is embodied on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-a-chip ("SoC") together with one or more other devices such as additional PPUs, memory, a reduced instruction set computer ("RISC") CPU, a memory management unit ("MMU"), and a digital-to-analog converter ("DAC").
[0330] In at least one embodiment, the PPU may be included in a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on the motherboard of a desktop computer. In at least one embodiment, the PPU may be an integrated graphics processing unit ("iGPU") included in the chipset of...
Claims
1. It is a system-on-a-chip (SoC), The central processing unit (CPU) and Memory and PCI (Peripheral Component Interconnect) communication bus and An upsampler that uses at least one neural network to generate a higher resolution image by using already upsampled higher resolution frames of a lower resolution image integrated with frames inferred before the higher resolution image, Graphics processing unit (GPU) and Equipped with, The aforementioned graphics processing unit (GPU) is equipped with a general-purpose processing cluster (GPC), The aforementioned general-purpose processing cluster (GPC) includes streaming multiprocessors (SMs), The aforementioned streaming multiprocessor (SM) is: Instruction cache and Dispatch unit and, The core and Load / Store Unit (LSU) and Shared memory and L1 cache and A system-on-a-chip (SoC) characterized by having the following features.
2. The system-on-a-chip (SoC) according to claim 1, wherein the SM further comprises a register file.
3. The system-on-a-chip (SoC) according to claim 1, characterized in that the SM further comprises one or more special function units (SFUs).
4. The system-on-a-chip (SoC) according to claim 1, characterized in that each of the SMs further comprises one or more interconnections.
5. The system-on-a-chip (SoC) according to claim 1, wherein the GPC further comprises a raster engine.
6. The system-on-a-chip (SoC) according to claim 1, further comprising a hub that interfaces with one or more GPU interconnects.
7. The system-on-a-chip (SoC) according to claim 1, wherein the GPU further comprises an input / output (I / O) unit that interfaces with the PCI communication bus.
8. The system-on-a-chip (SoC) according to claim 1, wherein the GPU further comprises a crossbar (Xbar).
9. The system-on-a-chip (SoC) according to claim 1, wherein the GPU further comprises a memory partition unit.
10. A system-on-a-chip (SoC) is used to perform an upsampler, which is configured to use at least one neural network to generate a higher-resolution image using already upsampled higher-resolution frames of a lower-resolution image integrated with frames inferred before the higher-resolution image. The aforementioned system-on-a-chip (SoC) is The central processing unit (CPU) and Memory and PCI (Peripheral Component Interconnect) communication bus and Equipped with a graphics processing unit (GPU), The GPU comprises a general-purpose processing cluster (GPC), and the general-purpose processing cluster (GPC) comprises a streaming multiprocessor (SM). The aforementioned streaming multiprocessor (SM) is: Instruction cache and Dispatch unit and, The core and Load / Store Unit (LSU) and Shared memory and L1 cache and A method characterized by being equipped with
11. The method according to claim 10, characterized in that the previously inferred frame is high resolution and is inferred by the at least one neural network.
12. The method according to claim 10, wherein the GPU further comprises a scheduler unit.
13. The method according to claim 10, wherein the GPC further comprises a raster engine.
14. The method according to claim 10, wherein the SoC further comprises a hub that interfaces with one or more GPU interconnects.
15. The method according to claim 10, wherein the SoC further comprises a network interface.
16. The method according to claim 10, wherein the SoC further comprises one or more display devices.