Spatiotemporal noise mask for image processing
By generating a spatiotemporal blue noise mask, the problem of high resource consumption in real-time image rendering is solved, image quality is improved, and rendering effects in motion are optimized, thus achieving efficient image processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies consume significant memory, time, or computing resources in real-time image rendering, and the image quality is poor when viewed in motion, especially when it is difficult to balance spatial and temporal dimensions.
A spatiotemporally optimized blue noise mask is generated by modifying the spatial aggregation algorithm to produce a spatiotemporal blue noise mask for image rendering and enhancement. Combined with temporal anti-aliasing and deep learning supersampling techniques, image quality is optimized.
It improves image quality, reduces memory and computing resource consumption, achieves high-quality image rendering when observing in motion, and provides stable filtering effects over time.
Smart Images

Figure CN115439340B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 196,116, filed June 2, 2021, entitled “Noise Masks for Image Processing,” the contents of which are hereby incorporated herein by reference in their entirety. Technical Field
[0003] At least one embodiment relates to processing resources for performing and facilitating real-time image rendering and enhancement. For example, processors or computing systems for generating blue noise masks used in image enhancement according to various novel techniques described herein. Background Technology
[0004] Image processing (such as imaging rendering and enhancement) can utilize significant memory, time, or computational resources, especially when the processing is performed in real time. The amount of memory, time, or computational resources used for image enhancement can be improved. Attached Figure Description
[0005] Figure 1A An example of a process for generating a blue noise mask that is optimal in both spatial and temporal use, according to at least one embodiment, is shown;
[0006] Figure 1B An example of a process for generating a three-dimensional mask for both spatial and temporal use for a framework, according to at least one embodiment, is shown;
[0007] Figure 2 Exemplary images are shown illustrating the use of a blue noise mask that is optimal in both spatial and temporal terms, according to at least one embodiment;
[0008] Figure 3 An energy assessment according to at least one embodiment is shown, wherein there exists a single dimension (horizontal axis) and time is a compressed space xy dimension on the vertical axis;
[0009] Figure 4 The above is illustrated according to at least one embodiment. Figure 3 The three types of blue noise masks mentioned in the text were compared using Fourier analysis to generate frequency results;
[0010] Figure 5 The convergence rate of the 1D function according to at least one embodiment is shown;
[0011] Figure 6 The DFT of a 2D projection of a 64x64x16x16 4D blue noise mask according to at least one embodiment is shown;
[0012] Figure 7 An autocorrelation image illustrating blue noise texture is shown according to at least one embodiment;
[0013] Figure 8 A 2Dx1D spatiotemporal blue noise mask with various sigma / axis is shown according to at least one embodiment;
[0014] Figure 9 It is shown that the generation time according to at least one embodiment is a function of the number of pixels in the blue noise mask and approximately follows y = x 2 A graph of the curve;
[0015] Figure 10 Random transparency using various types of noise is shown according to at least one embodiment;
[0016] Figure 11 Convergence rates in random alpha of various types of noise according to at least one embodiment are shown;
[0017] Figure 12 The illustration shows color mixing prior to quantization of various types of noise to 1 bit per color channel according to at least one embodiment;
[0018] Figure 13 A graph showing the convergence rate in color mixing of various types of noise according to at least one embodiment is shown.
[0019] Figure 14 Four steps are shown in which noise is used to randomly offset the start portion of the ray travel for each pixel ray travel according to at least one embodiment;
[0020] Figure 15 A graph showing the convergence rate of light-traveling fog with various types of noise according to at least one embodiment is presented;
[0021] Figure 16 The diagram illustrates the use of noise to layer 16 samples of line segments for each pixel through a participating medium, according to at least one embodiment.
[0022] Figure 17 A graph showing the convergence rate of light-traveling fog with various types of noise according to at least one embodiment is presented;
[0023] Figure 18 The diagram illustrates, according to at least one embodiment, the use of two independent noise streams to generate the x and y components of a 2D vector mapped to a cosine-weighted hemisphere for each pixel's single ambient occlusion (AO) sample.
[0024] Figure 19This illustrates how AO convergence relates to various types of noise according to at least one embodiment;
[0025] Figure 20 Images using one or more of a 2D blue noise mask, a 3D blue noise mask, a spatiotemporal blue noise mask, and a 2DGR blue noise mask according to at least one embodiment are shown.
[0026] Figure 21 An image using Sobol sequence offset is shown according to at least one embodiment;
[0027] Figure 22 The Heitz & Belcour technique, according to at least one embodiment, uses staggered gradient noise and a stylized grayscale image for a noise pattern target.
[0028] Figure 23 A graph showing convergence in Monte Carlo integral, leakage integral, and convergent leakage integral according to at least one embodiment is shown;
[0029] Figure 24 This demonstrates how a threshold mask according to at least one embodiment can produce point sets of any density;
[0030] Figure 25 This illustrates how a set of threshold points, according to at least one embodiment, maintains its desired frequency on an axis group;
[0031] Figure 26 Five cumulative frames of pixels sampled from an image using a non-uniform importance map according to at least one embodiment are shown, such that pixels oriented toward the center are more likely to be sampled.
[0032] Figure 27 It is shown that white noise according to at least one embodiment may have redundant sampled pixels for each frame, and that spatial blue noise removes spatially redundant pixels over time, and 2Dx1D spatiotemporal blue noise also removes them over time;
[0033] Figure 28A The inference and / or training logic according to at least one embodiment is illustrated;
[0034] Figure 28B The inference and / or training logic according to at least one embodiment is illustrated;
[0035] Figure 29 The training and deployment of a neural network according to at least one embodiment are illustrated;
[0036] Figure 30 An example data center system according to at least one embodiment is shown;
[0037] Figure 31A A chip-level supercomputer according to at least one embodiment is illustrated;
[0038] Figure 31B A rack-mounted supercomputer according to at least one embodiment is illustrated;
[0039] Figure 31C A rack-mounted supercomputer according to at least one embodiment is shown;
[0040] Figure 31D A supercomputer at the entire system level according to at least one embodiment is shown;
[0041] Figure 32 This is a block diagram illustrating a computer system according to at least one embodiment;
[0042] Figure 33 This is a block diagram illustrating a computer system according to at least one embodiment;
[0043] Figure 34 A computer system according to at least one embodiment is shown;
[0044] Figure 35 A computer system according to at least one embodiment is shown;
[0045] Figure 36A A computer system according to at least one embodiment is shown;
[0046] Figure 36B A computer system according to at least one embodiment is shown;
[0047] Figure 36C A computer system according to at least one embodiment is shown;
[0048] Figure 36D A computer system according to at least one embodiment is shown;
[0049] Figure 36E and Figure 36F A shared programming model according to at least one embodiment is shown;
[0050] Figure 37 An exemplary integrated circuit and a related graphics processor according to at least one embodiment are shown.
[0051] Figure 38A and Figure 38B An exemplary integrated circuit and an associated graphics processor according to at least one embodiment are shown.
[0052] Figure 39A and Figure 39BAdditional exemplary graphics processor logic according to at least one embodiment is shown;
[0053] Figure 40 A computer system according to at least one embodiment is shown;
[0054] Figure 41A A parallel processor according to at least one embodiment is shown;
[0055] Figure 41B A partitioning unit according to at least one embodiment is shown;
[0056] Figure 41C A processing cluster according to at least one embodiment is shown;
[0057] Figure 41D A graphics multiprocessor according to at least one embodiment is shown;
[0058] Figure 42 A multi-graphics processing unit (GPU) system according to at least one embodiment is illustrated;
[0059] Figure 43 A graphics processor according to at least one embodiment is shown;
[0060] Figure 44 It is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment;
[0061] Figure 45 A deep learning application processor according to at least one embodiment is shown;
[0062] Figure 46 A block diagram of an example neuromorphic processor is shown according to at least one embodiment;
[0063] Figure 47 At least a portion of a graphics processor according to one or more embodiments is shown;
[0064] Figure 48 At least a portion of a graphics processor according to one or more embodiments is shown;
[0065] Figure 49 At least a portion of a graphics processor according to one or more embodiments is shown;
[0066] Figure 50 A block diagram of a graphics processing engine of a graphics processor is shown according to at least one embodiment;
[0067] Figure 51 It is a block diagram of at least a portion of a graphics processor core according to at least one embodiment;
[0068] Figure 52A and Figure 52B The diagram illustrates thread execution logic according to at least one embodiment, which includes an array of processing elements of a graphics processor core.
[0069] Figure 53 A parallel processing unit (“PPU”) according to at least one embodiment is shown;
[0070] Figure 54 A general-purpose processing cluster (“GPC”) according to at least one embodiment is illustrated;
[0071] Figure 55 A memory partition unit of a parallel processing unit (“PPU”) according to at least one embodiment is shown;
[0072] Figure 56 A streaming multiprocessor according to at least one embodiment is illustrated;
[0073] Figure 57 This is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;
[0074] Figure 58 This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment;
[0075] Figure 59 Example illustrations include an advanced computing pipeline 5810A for processing imaging data according to at least one embodiment;
[0076] Figure 60A Includes example data flow diagrams of virtual instruments supporting ultrasound equipment according to at least one embodiment;
[0077] Figure 60B Includes example data flow diagrams of virtual instruments supporting CT scanners according to at least one embodiment;
[0078] Figure 61A A data flow diagram of a process for training a machine learning model according to at least one embodiment is shown;
[0079] Figure 61B This is an example illustration of a client-server architecture that utilizes a pre-trained annotation model to enhance an annotation tool according to at least one embodiment;
[0080] Figure 62 A software stack of a programming platform according to at least one embodiment is shown;
[0081] Figure 63 The illustration shows an embodiment according to at least one of the embodiments. Figure 62 The CUDA implementation of the software stack;
[0082] Figure 64 The illustration shows an embodiment according to at least one of the embodiments. Figure 62 The ROCm implementation method of the software stack;
[0083] Figure 65 The illustration shows an embodiment according to at least one of the embodiments. Figure 62 The OpenCL implementation of the software stack;
[0084] Figure 66 Software supported by a programming platform according to at least one embodiment is shown;
[0085] Figure 67 A method for using at least one embodiment is shown. Figures 62-65 Compiled code executed on the programming platform;
[0086] Figure 68 A multimedia system according to at least one embodiment is shown;
[0087] Figure 69 A distributed system according to at least one embodiment is shown;
[0088] Figure 70 An oversampling neural network according to at least one embodiment is shown;
[0089] Figure 71 An architecture of an oversampling neural network according to at least one embodiment is shown;
[0090] Figure 72 An example of streaming using an oversampling neural network according to at least one embodiment is shown;
[0091] Figure 73 Examples of simulations using an oversampled neural network according to at least one embodiment are shown; and
[0092] Figure 74 An example of a device using an oversampling neural network according to at least one embodiment is shown. Detailed Implementation
[0093] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of at least one embodiment. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details.
[0094] The system uses a blue noise mask to provide random numbers for the image rendering algorithm at each pixel level, resulting in a pattern of random noise that is perceptually better than white noise. Blue noise masks are generally limited to high frequencies, making them more thoroughly removed with low-pass filters (making denoising with blurring much more effective).
[0095] However, in real-time image rendering, there is also a time axis that needs to be considered. In some embodiments, sampling of the image based on texture improvements over time improves image quality when viewed in motion and can also be applied to temporal filtering methods such as Temporal Anti-aliasing (TAA) and Deep Learning Supersampling (DLSS). Various methods exist for animate blue noise masks over time, but there is a trade-off between quality on the spatial axis and quality on the temporal axis.
[0096] In at least one embodiment, one or more circuits (which may be part of one or more processors in a computer system) generate a blue noise mask that is optimal both spatially and temporally (sometimes referred to herein as a "spatiotemporal blue noise mask"). One or more circuits (implementing software operations) can generate these spatiotemporal blue noise masks as a set of N blue noise textures, such that each texture is a separate, suitable blue noise, but each pixel is also blue noise over time. Suitable blue noise can contain a sufficient amount of higher frequencies and a low amount of lower frequencies.
[0097] In one embodiment, a spatiotemporal blue noise mask is created by modifying a void and cluster algorithm. The void and cluster algorithm is an algorithm used to generate an N-dimensional blue noise mask. The void and cluster algorithm can be modified to produce a noise pattern that simultaneously addresses desired spatial and temporal constraints for real-time image rendering. In one embodiment, the modified void and cluster algorithm receives one or more images with pixel data, where the pixel data includes N-dimensional (e.g., three or more dimensions, where one dimension is time) data.
[0098] Specifically, the techniques described herein are aimed at generating spatiotemporal blue noise masks for real-time image rendering and enhancement. For example, a spatiotemporal blue noise mask can be a three-dimensional mask, where two dimensions correspond to space (e.g., x and y coordinates) and one dimension corresponds to time. In embodiments, the spatiotemporal blue noise mask is used to benefit various applications or techniques of image rendering, including color mixing, random transparency, region light sampling, and volume rendering. Furthermore, the spatiotemporal blue noise mask can be applied to a variety of sampling techniques, including, for example, soft shadows and path tracing, random alpha, and color mixing.
[0099] Specifically, the techniques described herein target generating spatiotemporal blue noise masks that can handle the temporal domain, for example, adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time). For example, the techniques described herein can be used with temporal anti-aliasing, which is commonly used in games and other interactive applications to amortize the cost of rendering across multiple frames. Furthermore, for static image rendering, image enhancement can be achieved by integrating over multiple samples covering each pixel in time or other dimensions, while still preserving the characteristics of the blue noise error spatially. Human perception (and some computational evidence) performs a certain amount of implicit integration over time, especially at high frame rates, and these cases typically provide values to a well-sampled pattern over time without any explicit filtering. Therefore, a two-dimensional (2D) blue noise pattern is used for each frame, which is well-distributed over time at each pixel and converges quickly for Monte Carlo integration (e.g., numerical integration using random numbers). Accordingly, the techniques described herein pertain to systems and methods for generating spatiotemporal blue noise masks that maintain spatial 2D blue noise characteristics while providing blue noise over time at each pixel. In embodiments, the spatiotemporal blue noise mask is generalized to arbitrary dimensions for use in higher dimensions.
[0100] In at least one embodiment, a time slice is a specific time. More generally, a slice may also be a layer of a specific dimension or a multidimensional layer (e.g., based on the coordinates of pixels falling within a plane having a specific dimension or multidimensionality). The techniques described herein describe a framework that enables each two-dimensional (2D) slice (in the spatial domain) of a three-dimensional image to be associated with blue noise properties such that each pixel includes a one-dimensional sampling property in the temporal dimension. In embodiments, acquiring a slice from a higher spatial dimension results in an image in a lower-dimensional space. Thus, a 2D slice may be an image taken from a 3D image or object. In one embodiment, a 3D image over time may be represented by 2D slices (e.g., an XY plane in the spatial domain) and a Z-axis representing the temporal domain. In embodiments, the framework described herein provides blue noise properties in each spatial 2D slice, which also provides superiority over white noise sequences along the time axis. Furthermore, denoising algorithms can benefit from spatiotemporal sampling techniques (and the use of blue noise). Thus, the techniques described herein cause one or more circuits (which may be portions of one or more processors in a computer system) to execute algorithms for spatiotemporal and higher-dimensional noise mask generation (e.g., blue noise masks). Specifically, spatiotemporal and higher-dimensional blue noise masks are generated by modifying the empty clustering algorithm.
[0101] In an embodiment, the empty clustering algorithm has an energy function used to find the emptiest space in the image to place the next pixel in. In one embodiment, the energy function of the empty clustering algorithm can be modified one or more times. In one modification, for spatiotemporal blue noise, the algorithm can be run in 3D to produce a unique energy function. This energy function makes it so that pixels influence each other in the energy field only if they come from the same slice or are the same pixel at different points in time. In this way, each 2D slice of 3D blue noise will be good 2D blue noise, and each pixel will become 1D blue noise over time. As a result, the error may be well hidden as blue noise, but also smaller only by being able to converge to the correct result better. In an embodiment, another modification that can be made is to extend this to higher dimensions by specifying which axes should be grouped together into N-dimensional blue noise. This allows the techniques described herein to extend beyond spatiotemporal blue noise to four-dimensional (4D) spatiotemporal depth blue noise. This is useful, for example, when rendering fog, but it can also be generalized to any dimension and any grouping of those dimensions that a particular rendering algorithm might need. It is possible to customize floating-point constraint solvers like this one to provide blue noise error in screen space while achieving faster convergence for rendering algorithms.
[0102] Blue noise distributions are well-suited to human perception and minimize unwanted low-frequency noise. Sets of blue noise points are often referred to as blue noise masks or blue noise textures. In image rendering, this typically involves integrating samples over multiple frames to amortize rendering costs, or equivalently, multiple samples collected per frame. Therefore, in one embodiment, the techniques described herein achieve various technical advantages, including (but not limited to) using a 2D blue noise pattern that produces samples well-distributed over time at pixels during animation, converges quickly for Monte Carlo integration, and still retains spatial blue noise characteristics. Some spatial blue noise methods applied at each frame independently produce results that temporally represent the white noise spectrum, and therefore converge slowly for integration across time and are unstable when temporally filtered.
[0103] Therefore, the technique described herein is an extension of the empty-joining algorithm involving the reformulation of the energy function, which produces a spatiotemporal blue noise mask exhibiting the blue noise spectrum in both the spatial and temporal domains. This can result in visually pleasing error patterns, fast convergence speed, and increased stability during temporal filtering. In some embodiments, the technique described herein can also be extended to higher dimensions, as it provides unique sampling characteristics for time integration. By applying the technique described herein, improvements can be achieved in a variety of applications such as color mixing, random transparency, low-sampling-count ambient occlusion, and volume rendering.
[0104] In at least one embodiment, the techniques described herein achieve various technical advantages, including, but not limited to, improved real-time image rendering and enhancement in applications using rendering algorithms that require random numbers per pixel, and any location quantization, because it enables very good color mixing (perceptually good from a filtering point of view, and also more accurate spatial and temporal averages of small regions of pixels compared to the actual average of the unquantized source data), which masks the fact that low-bit counting is used. This is useful for reducing memory usage for geometry buffers (G-buffers), render targets, textures, etc.
[0105] While masks can be generated for two-dimensional and three-dimensional applications, the techniques described herein are applicable to multidimensional applications (e.g., greater than 3, 6, 7, etc.). Furthermore, the techniques are not limited to having only one dimension as the time dimension; rather, the techniques described herein can generate multidimensional masks (e.g., 7 dimensions), where one dimension is time, or no dimension is time. For example, mask generation may involve a 7-dimensional mask where the first three axes (e.g., dimensions) represent 3D blue noise, the next two axes (e.g., dimensions) represent 2D blue noise, and the last two axes (e.g., dimensions) represent 1D blue noise. In at least one embodiment, the computer implementation can select the grouping of dimensions when generating the mask.
[0106] In at least one embodiment, after generating the blue noise mask, other types of filtering operations can be applied to the rendering process, such as red noise filtering, bandpass filtering, or other types of noise filtering (e.g., frequency attenuation filters or other denoising methods).
[0107] Figure 1A An example of a process 100 for generating a blue noise mask optimal for both spatial and temporal use, according to at least one embodiment, is illustrated. In at least one embodiment, some or all of process 100 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs, Computing Unified Device Architecture (CUDA) code) jointly executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transient computer-readable medium. In at least one embodiment, not only transient signals (e.g., propagating transient electrical or electromagnetic transmissions) are used to store at least some of the computer-readable instructions that can be used to execute process 100.
[0108] In at least one embodiment, the non-transient computer-readable medium does not necessarily include non-transient data storage circuitry (e.g., buffers, caches, and queues) within the transceiver of transient signals.
[0109] In at least one embodiment, process 100 is executed at least partially on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, process 100 is executed by one or more circuits to calculate the motion of one or more pixels in the first region of the image based at least partially on the motion of one or more pixels in a second region of the image that overlaps with the first region.
[0110] In at least one embodiment, the system performing at least a portion of process 100 includes executable code for generating a blue noise mask that is optimal both spatially and temporally (e.g., a spatiotemporal blue noise mask). The spatiotemporal blue noise mask may be generated as a set of N blue noise textures, each texture individually having good blue noise (e.g., containing a high amount of higher frequencies and a low amount of lower frequencies), and each pixel individually also having blue noise over time. This can provide the desired quality without compromise on either the spatial or temporal axis. In one embodiment, one or more images are acquired from a computing device, camera, etc. In one embodiment, one or more images may be part of a game application, wherein a real-time image rendering algorithm is used to display the images when the game application is executed by the computing device. The computing device may include one or more graphics cards that use deep learning to upscale lower-resolution images to higher resolutions for display on a computing screen. In an embodiment, one or more processors of the computing device execute instructions to apply the spatiotemporal blue noise mask to one or more acquired images for real-time image rendering 104.
[0111] In embodiments, the systems and methods described herein aim to generate blue noise masks that are optimal both spatially and temporally. This can be performed by using a spatial aggregation algorithm and extending that algorithm to handle the spatiotemporal domain. In embodiments, the spatial aggregation algorithm includes the following steps:
[0112] 1. Initial binary mode
[0113] 2. Phase I - Making the pattern progressive
[0114] 3. Stage II - First Half of the Pixel
[0115] 4. Stage III - The second half of the pixel
[0116] 5. Refine texture
[0117] In one embodiment, one or more computing devices perform the generation of dimensions [d0, d1, ..., d]. nAn algorithm for a blue noise mask M is described. The algorithm requires each pixel to store Boolean logic specifying whether the pixel is activated (emitting energy into an energy field), and an integer index specifying the order in which the pixel is activated. The order in which pixels are activated defines the final output color of the pixels, where the first pixel to be activated is black, and the last pixel to be activated is white. Each activated pixel p = (d x ,d y Energy can be given to each point q in the energy field using the following formula:
[0118]
[0119] Where p and q are integer coordinates, and the distance is calculated at the winding boundary (i.e., the loop winding). σ is a configurable parameter that controls the energy attenuation over the distance, thereby controlling the frequency content. In this embodiment, σ = 1.5.
[0120] In this embodiment, this is similar to a Gaussian blur, which is used to distribute the energy from activated pixels across a distance in an energy field. The energy field E can be discretized onto a grid of the same size as the mask M and is defined as follows:
[0121]
[0122] 1. Initial binary mode
[0123] In this embodiment, the first step in the empty clustering algorithm is to generate an initial binary pattern in which less than or equal to half of the pixels are activated. This can be achieved by using white noise or any other pattern. Before proceeding, these pixels may need to be transformed into a blue noise distribution. In this embodiment, the pixel transformation can be performed by repeatedly deactivating the densest cluster pixels and activating the largest empty pixel. This process can be repeated until the same pixel is found for both operations. At this point, the algorithm has converged, and the initial binary pattern can be a blue noise distribution. Although these pixels have been activated in the energy field, they may not have yet been sorted. The largest empty pixel can be the point with the lowest energy in the energy field and is defined as: inf p∈M F(p). Conversely, the densest cluster can be the point with the highest energy in the energy field and is defined as: sup p∈M F(p).
[0124] 2. Phase I - Ordering the Pattern
[0125] In this embodiment, the initial binary pattern is now a blue noise distribution. In this embodiment, the execution points are an ordered sequence. This can be accomplished by repeatedly removing the densest clusters and sorting the pixels by how many pixels are enabled after a pixel is deactivated. This can be repeated until all pixels are deactivated. Afterward, the initial binary pattern can be reactivated, and together with the order generated as described above, an ordered binary blue noise sequence now exists.
[0126] 3. Stage II - First Half of the Pixel
[0127] In one embodiment, the remaining deactivated pixels are activated one at a time until half of the pixels are activated. In another embodiment, this is done by finding the largest gap and activating that pixel. The order given to the pixel can be the number of pixels that were active before that pixel was activated.
[0128] 4. Stage III - The second half of the pixel
[0129] In this embodiment, the state of all pixels is reversed. In this embodiment, active pixels are deactivated, and vice versa. This phase repeatedly finds and deactivates the maximum value cluster, thus assigning it a sort based on the number of pixels that were inactive before that pixel was deactivated. In this embodiment, this phase completes when no more deactivated pixels exist, and all pixels are sorted. This process is the reverse of Phase II because the process from Phase II works best at sparser points. Reversing the problem while adding the remaining pixels allows for sparse execution of the work.
[0130] 5. Refine texture
[0131] In this embodiment, after sorting all pixels, the sorting is converted into pixel values in the output image. If the output image has a resolution of n×n pixels, then their sorting can be from 0 to n. 2 -1. If the output is a k-bit image, these values may need to be remapped from 0 to n. 2 -1 to 0 to 2 k -1. In some instances, this may produce non-unique values in the output texture, but the histogram will be flat as needed.
[0132] In this embodiment, when rendering frames are integrated over time, each pixel undergoes one-dimensional integration along the time axis. If the same 2D blue noise texture is used for each frame, then each pixel obtains the same result and does not provide any new samples for integration. If an independently generated 2D blue noise texture is used for each frame, then each pixel can become a sequence of white noise over time.
[0133] In some instances, multiple two-dimensional blue noise masks can be used for high quality in the spatial domain; however, each pixel may also individually require a high-quality sampling sequence over time. Therefore, in embodiments, one or more computing devices can execute one or more algorithms to generate a three-dimensional blue noise mask. In embodiments, the spatial aggregation algorithm can be reformulated such that it is driven by a new energy function as shown in Equation 1 above. In embodiments, instead of executing the formula in two dimensions, it is executed in three dimensions, and the energy function is constrained in two ways. The energy can be non-zero if two pixels in the energy function are located in the same two-dimensional layer, or if two pixels have the same (x, y) coordinates. A first condition ensures that each two-dimensional layer can have blue noise characteristics, and a second condition ensures that each pixel can have blue noise characteristics over time (as referenced). Figure 3 (To be described in more detail). Without the first condition, each pixel will be blue noise on the time axis, but will be independent of each other and spatially white noise. Without the second condition, each z-plane slice will be independent, and the result will be white noise along the time axis. Without the constraint that one of these conditions must be satisfied, the result will be three-dimensional blue noise that is not well distributed on the spatial or temporal axes (spatiotemporally), but rather well distributed in the 3D volume.
[0134] In the embodiment, pixels in the three-dimensional spatiotemporal blue noise texture are represented as p = (p xy ,p z )=(p x ,p y ,p z In one embodiment, the modified energy formula is then:
[0135]
[0136] In this embodiment, these two constraints work together to provide all the common benefits of spatiotemporal blue noise in two dimensions, while also ensuring good sampling characteristics for each pixel on the time axis. In this embodiment, the distances fed into the energy function are calculated circumferentially on all axes, meaning that individual texture slices are spatially well-tiled, but the temporal quality is also temporally well-tiled, with no seams when time begins at zero. In this embodiment, when generating spatiotemporal blue noise, an initial binary pattern density of 10% of the pixels is used, and σ = 1.9 for all axes.
[0137] Figure 1BAn example of process 106 for generating a 3D mask for use in both space and time is shown, where two dimensions correspond to space (e.g., x and y coordinates) and one dimension corresponds to time. While 3D masks can be generated, N-dimensional masks can also be generated, as explained in receiving operation 108. In at least one embodiment, process 106 is integrated into process 100. In at least one embodiment, some or all of process 106 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs, CUDA code) jointly executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, a system includes a memory storing instructions that, when executed by one or more processors, cause the system to execute the instructions to perform process 106. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium, the computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transient computer-readable medium. In at least one embodiment, not only transient signals (e.g., propagating transient electrical or electromagnetic transmissions) are used to store at least some computer-readable instructions that can be used to execute process 106. In at least one embodiment, the non-transient computer-readable medium does not necessarily include non-transient data storage circuitry (e.g., buffers, caches, and queues) within the transceiver of the transient signal. In at least one embodiment, process 106 is executed at least partially on a computer system such as those described elsewhere in this disclosure. In at least one embodiment, process 106 begins with a receive operation 108 and continues with a computation operation 110.
[0138] In receiving operation 108, the system or computing device receives pixel data with three dimensions corresponding to one or more images. In one embodiment, the system or computing device receives the pixel data based on sampling one or more images from one or more images obtained in step 102 of process 100. In another embodiment, the system or computing device receives the pixel data from another device or another process (e.g., from an application running on another device). Receiving the pixel data may include receiving N-dimensional data of the pixels. For example, receiving the pixel data may include receiving spatial pixel data (e.g., x and y coordinates) and temporal data, wherein the spatial data corresponds to two dimensions and the temporal data corresponds to the temporal dimension of each pixel.
[0139] At calculation operation 110, the system or computing device performing process 106 calculates energy values for some pixels of one or more images received in receiving operation 108. In at least one embodiment, the energy value of a pixel can be calculated according to Equation 1. In at least one embodiment, the energy value corresponds to an intensity value, where the intensity value indicates the intensity of a pixel, for example, how much a pixel stands out relative to other pixels (e.g., when it is activated). Sigma (σ) is a configurable parameter that controls the energy or intensity decreasing with distance (e.g., diminishing), and this parameter may be referred to as an "energy decay parameter" or an "intensity decay parameter." In some embodiments, the energy decay parameter corresponds to a Gaussian blur function, for example, as shown in Equation 2. As shown in Equation 1, the energy value is based on the coordinates of at least some pixels (e.g., pixels p and q), the distance between at least some pixels (e.g., pixels p and q), and the energy decay parameter. When determining the distance between at least some pixels, a pair of pixel distances can be used. In embodiments, when determining the distance between at least some pixels, multiple pixel pairs (e.g., pixels p and q, where q can be any adjacent pixel for p) can be determined. When determining the energy value of a pixel, the distance between pixels can be calculated in a circular fashion. In one embodiment, the energy value is calculated for each pixel. In another embodiment, the system or computing device calculates the energy values of some pixels based on determining relevant processing portions of the image.
[0140] As shown in Equation 1, there are typically two constraints that determine the energy value. For example, the energy value can be non-zero if two pixels in the energy function are in the same 2D layer, or if two pixels have the same (x, y) coordinates. The first condition ensures that each 2D layer can have blue noise characteristics, and the second condition ensures that each pixel can have blue noise characteristics over time (as shown in the reference). Figure 3 (Described in more detail). If the calculated pixel energy does not satisfy these two constraints, the system or computing device may set the energy value to zero. For example, if a pixel and another pixel are not in the same two-dimensional layer or do not have the same coordinates (e.g., at different time slices), then the system or computing device may set the energy value of at least some pixels to zero.
[0141] To generate a 3D mask, in mask operation 112, the system or computing device generates the mask based on the calculated energy value from computation operation 110. Generating the 3D mask may be part of another digital image processing algorithm (e.g., spatial convergence, color mixing, or error diffusion), where the image processing algorithm uses an energy function (e.g., Equation 1) and corresponding energy values of the pixel data as determined in computation operation 110. In one embodiment, because mask operation 112 considers energy according to Equation 1, the generated mask produces a blue noise mask that provides optimal visual results for human perception, for example, as part of a video game, video, or other digital video process. The 3D mask can be considered a blue noise spatiotemporal mask.
[0142] In an embodiment, the spatiotemporal blue noise mask can be configured to specify different dimensions for each axis (see [link to implementation]). Figure 6 For example, one or more computing devices execute an extension corresponding to the empty aggregation algorithm using an energy function (Equation 1) operating in D dimensions, where D is the number of parameters required to index into the noise function to obtain scalar values. These parameters may be axes of the blue noise mask. The D dimensions can be decomposed into one or more sets G, where each set G contains one or more dimensions. When computing masks with more than 3 dimensions, the system or computing device may use Equation 4.
[0143] In operation 114, the system or computing device provides one or more output images based on applying a 3D mask from masking operation 112 to the one or more images, a sample version of the one or more images, or a processed version of the one or more images. In one embodiment, the one or more output images may be part of a game application, wherein a real-time image rendering algorithm is used to display the images when the game application is executed by the computing device. For example, the rendered output images may be part of an image generation pipeline that includes ray tracing or path tracing. In one embodiment, one or more processors of the computing device execute instructions to apply a 3D mask to one or more images for real-time image rendering.
[0144] Process 106 can be integrated with other image processing techniques. In embodiments, process 106 can be integrated into sampling as part of image processing in motion and temporal filtering methods such as temporal anti-aliasing (TAA) and deep learning supersampling (DLSS). As part of DLSS, process 106 may include applying temporal image upscaling to the one or more images, wherein the upscaling is based on neural network inference from a lower resolution image. In at least one embodiment, a spatial-temporal mask is applied before, after, or both before and after image processing associated with TAA and DLSS.
[0145] Similarly, process 106 can be integrated into color mixing, random transparency, region light sampling, volume rendering, path tracing, and / or random alpha image processing techniques. Furthermore, the operation of process 106 can be repeated (e.g., for multiple images) or performed in a different order as part of another digital image processing algorithm. For example, process 106 can be performed as part of a sampling algorithm.
[0146] In at least one embodiment, process 100 and processor 106 can be applied to video or video game content. The video or video game includes a sequence of images (e.g., frames) that can be displayed at a frequency (e.g., frame rate), where a single video frame is an image. Furthermore, a video frame refers to video information, while an audio frame refers to audio information, and video frames can be processed synchronously with or separately from audio frames.
[0147] Figure 2 An exemplary image is shown illustrating the use of a blue noise mask that is optimal in both space and time, according to at least one embodiment. In the embodiment, the blue noise mask provides a way for the system to hide noise and errors. This is useful in real-time rendering where computational resources are limited to making noise completely disappear (which is the motivation for denoising). Although a blue noise mask does not produce less noise and errors than white noise, it does arrange itself in a more visually pleasing, less perceptible, and easier-to-denoise manner. Blue noise has various uses in rasterization and ray tracing. Figure 2 As shown, at the top, blue noise and white noise are used to render grayscale image points as black and white. The top blue noise is much less noisy and looks more like the source image, although it has the same amount of error as the white noise image below it. At the bottom, two noises are used to mix the color images before they are quantized to one bit per color channel. Both images can contain eight colors: red, green, blue, yellow, cyan, magenta, black, and white, and have the same amount of error as the source image, but the blue noise version at the top has better image quality.
[0148] The pointillist case can be obvious when ray tracing and capturing less than one ray per pixel. A black point can be treated as a pixel when choosing to capture a ray and selecting white or blue noise will produce the same kind of result in 3D rendering. The color mixing case occurs when encoding data in the buffer. Being able to use a single bit per color channel instead of the usual 8 bits per color channel means 3 bits are used for color instead of 24, meaning that only 12% of the previous bit depth can be used to represent the data.
[0149] In the example embodiment, 64 are created. 3 A spatiotemporal blue noise mask with a resolution of (64x64x64) can also be created. Additionally, a 64x64 resolution spatiotemporal blue noise mask can be created. 264 3 A 3D blue noise mask and 64 independent 2D blue noise masks. In this embodiment, the empty clustering algorithm described herein is used to generate the 2D and 3D blue noise masks. In this embodiment, a spatiotemporal mask is also created by using a single 2D blue noise mask and adding a golden ratio to each of the 63 frames to produce 64 different masks.
[0150] Figure 3 An energy evaluation according to at least one embodiment is illustrated, wherein there is compression to a single dimension (horizontal axis) and time is a spatial xy dimension on the vertical axis. In the embodiment, an energy function evaluation of the pixel is provided for the pixel in the middle of the illustration. In the embodiment, the spatial clustering algorithm described herein (such as...) is used. Figure 3 (As shown from left to right), energy function evaluation for the central pixel is determined for 2D blue noise, 3D blue noise, and spatiotemporal blue noise masks. In the embodiment, 2D blue noise ( Figure 3 The leftmost image in the image measures the energy of all pixels within the same 2D layer, while 3D blue noise ( Figure 3 The energy of all pixels within the entire 3D texture is measured using the intermediate image in the image. In this embodiment, a spatiotemporal blue noise mask using the spatial aggregation algorithm described herein is employed. Figure 3 (The rightmost image in the image) measures the distance in the 2D layer and along the time dimension of the current pixel.
[0151] Figure 4 The use of Fourier analysis according to at least one embodiment is shown in Figure 3 The comparison frequency results generated on the three types of blue noise masks mentioned above. That is, in Figure 4 The discrete Fourier transform (DFT) of the 2D projections of various blue noise masks is shown. The frequency comparison results can be performed using procedure 100 (see [link]). Figure 1A ) and / or process 106 (see Figure 1B This is generated. In an embodiment, the spatiotemporal blue noise mask has blue noise in space, thus enabling it to provide better image results than white noise on the z-axis (time). In an embodiment, the DFT is averaged to show the desired spectrum, in addition to the golden ratio animated blue noise, which highlights two ways it destroys spatial frequencies at specific frame numbers. In an embodiment, it is desirable to obtain blue noise characteristics in each spatial 2D slice in order to provide a noise sequence that is better than a white noise sequence along the time axis. Figure 4 As shown, a spatiotemporal blue noise mask, as described in this paper, provides both features simultaneously by having 2D blue noise characteristics on the XY plane and adding the blue noise characteristics to the Z-axis.
[0152] In the embodiments, although the DFT indication is blue both spatially and temporally using a spatiotemporal blue noise mask as described herein, the convergence rate of the blue noise over time is increased compared to other alternative methods for animating blue noise (this is in Figure 5 (This is shown in more detail below). Since time integration is equivalent to integrating multiple samples within the same frame, solving it in one domain is equivalent to solving it in another. Time integration often uses leak integrators instead of Monte Carlo integrators.
[0153] 4.3. Special Characteristics
[0154] like Figure 4 As shown, the right two columns also illustrate that if spatiotemporal blue noise is offset along the time axis, it can have the same convergence properties and is actually asymptotically continuous at any index, while also being cyclically continuous. This cyclical continuity / asymptotic nature of the time axis can be a powerful property used in TAA-style time integration and filtering algorithms. In those algorithms, each pixel is asymptotically integrating the integrand over each frame, but when an individual pixel considers its history to be no longer valid due to occlusion changes or similar factors, the pixel will effectively discard its history and restart the integration.
[0155] Using an animated blue noise mask to drive the 1D integral of those pixels means that the global sequence is driving all pixels. Most progressive sequences will only give a progressive sequence starting from index 0 (the exception to this is Sobol, which is progressive for all powers of 2-sized segments). This is problematic because global sequence-driven sampling for individual pixels discards their history at arbitrary times, and those pixels will start sampling at arbitrary positions in the sampling sequence.
[0156] In the case of continuous / progressive spatiotemporal blue noise on the timeline, after rejecting history at any frame number, each pixel can receive the benefit of starting at the beginning of the progressive sequence without the overhead of having to trace the index of each pixel to make this happen. Furthermore, history rejection is typically not a discrete event but a continuous operation, such as clamping historical data to the minimum and maximum of the colors seen in the local neighborhood of the newly rendered pixel values. In some cases, the sampling index is reset, while in others it is not. Nevertheless, the progressive sequence from any index implies whether a pixel has rejected its history, and taking a sample is a good thing, meaning it also handles this continuous history rejection case.
[0157] 4.4. Generalization to any dimension
[0158] In this embodiment, spatiotemporal blue noise masks are valuable in any animation scenario currently using 2D blue noise textures because they are a solution to the problem of animating blue noise masks. In other methods, blue noise can spatially degrade over time or become white noise. In this embodiment, spatiotemporal blue noise has a high-quality blue noise spectrum spatially, and also a blue noise spectrum temporally.
[0159] In one embodiment, one or more computing devices execute an extension of the empty clustering algorithm running in D dimensions, where D is the number of parameters required to index into the noise function to obtain scalar values. These parameters may be axes of the blue noise mask. The D dimensions can be decomposed into one or more sets G, where each set G contains one or more dimensions. A particular set g of G with a membership count d means that all d-dimensional projections of the D-dimensional blue noise mask should be d-dimensional blue noise when only the axes within that set change and all other axes remain constant. A set of axes h can also be defined as all axes not in units of g.
[0160] Once the dimensions are grouped, each group g can be grouped by using only the dimensions present in the group (denoted as p) within the usual Gaussian energy function, provided that the axes from the corresponding group h are equal between two pixels. g and q g Naturally mapped to the energy function E g The energy function between two pixels is the sum of all E values between those pixels. g The sum of functions, and the energy field F is the sum of the energy from every pixel, every other pixel.
[0161]
[0162] E(p,q)=∑ g∈G E g (p,q)
[0163] F(p)=∑ q∈M E(p,q)
[0164] Note that each dimension can have a different size, and different sigma values can be used to control the frequency content of the result. The original empty-clustering algorithm can be seen as making D any arbitrary value, and having only a single group containing all axes in G. Therefore, the empty-clustering algorithm can create a D-dimensional blue noise mask. When considering spatiotemporal blue noise, D can be 3, and G can have two groups in it: g xy and g z In this sense, spatiotemporal blue noise can also be regarded as a 2Dx1D blue noise mask.
[0165] 4.5.4D Blue Noise Mask Analysis
[0166] In this embodiment, the two 4D blue noise masks are configured as follows: 2Dx1Dx1D and 2Dx2D, each with a size of 64x64x16x16. In this embodiment, it is possible to... Figure 6 The frequency analysis shown illustrates the expected frequency behavior for each pair of axes in a 2D DFT. In this embodiment, two masks represent 2D blue noise in the XY plane, but are different under all other projections. A 2Dx1Dx1D blue noise mask can represent 1D blue noise on the Z and W axes under all projections (including the ZW plane, where they all exist and are shown in a cross pattern). On the other hand, 2Dx2D blue noise represents white noise for all other projections except the ZW plane, where it represents 2D blue noise.
[0167] From observations, the generation time of the blue noise mask is a function of the total pixel count, regardless of how those pixels are partitioned across dimensions, and is close to O(n). 2 ),like Figure 9 As shown in the diagram, where n is the number of pixels. Doubling the number of pixels in a blue noise mask will roughly take four times longer to generate the mask. Blue noise masks can be stored as single-channel 8-bit textures. The chart below shows some texture sizes and their byte sizes as examples. Due to good tiling on each axis, smaller textures, such as 64x64x16 (64KB) and 64x64x16x16 (1MB) for spatiotemporal blue noise, are sufficient for image rendering. The actual size and 4D version used for spatiotemporal blue noise are bolded.
[0168]
[0169]
[0170] In an embodiment, the algorithm used to generate the spatiotemporal blue noise mask can be configured to specify different dimensions for each axis (see [link to implementation]). Figure 6 ) and different energies of sigma (see Figure 8 Furthermore, although all axes are circularly continuous, if this is not desired, it is possible to select features for each axis by non-circularly calculating the distance on that axis.
[0171] 4.6. Extension
[0172] When using blue noise masks, multiple independent masks may be required. For example, when using diffuse and specular buffers for color mixing that are later combined via addition, the same blue noise mask may not be used repeatedly for both buffers because it will increase the difference between pixels when pixels are added together (already using the same mixing mode on each buffer). In some instances, the system may generate and load two independent masks, but these masks can consume more memory than needed, especially if each different color buffer in the rendering pipeline requires an independent blue noise mask. This number can even be dynamic or unbounded, which would be even more problematic.
[0173] An alternative way to approximate independent blue noise sources is by offsetting, where a blue noise mask is read for each desired independent blue noise source. Figure 7 The diagram illustrates the autocorrelation of a blue noise mask, showing that small offsets read from the blue noise texture can lead to correlation or anticorrelation, but larger offsets will result in uncorrelated values. This is because blue noise is correlated at small distances but uncorrelated at large distances, as illustrated in the autocorrelation diagram.
[0174] To generalize this to wanting N distinct, independent data sources, you might need N points on the texture, which are almost always farthest from each other. In other words, these points should have low variance. Since star-shaped variance is not a circular measurement, it can be measured circularly. If the number of independent data sources needed is unknown beforehand, an asymptotically circular sequence of low variance can provide any number of points with this characteristic.
[0175] Since higher-dimensional blue noise mask calculations take longer and require more memory to store them in order to obtain an N-dimensional mask, the system can approximate it by starting with an N-1 dimensional mask, reading the values of the first N-1 axes, and then multiplying the last dimension index by the golden ratio, adding it to the mask value, and using a modulus to keep it between 0 and 1.
[0176] N(a0,a1,…,a n )=(N(a0,a1,…,a n-1 )+φa n-1 )mod 1 (5)
[0177] This is illustrated by comparing spatiotemporal blue noise with 2D blue noise animated using the golden ratio, and also by comparing 2Dx1Dx1D blue noise with spatiotemporal blue noise using the golden ratio to add a fourth dimension. While this may compromise the spatial frequency, it does show good convergence and can help in creating temporal and memory usage. Although other irrational numbers exist to form other rank 1 lattices, which could also be used here, they have lower sampling quality, and this method can only sum groups of 1D axes.
[0178] In this embodiment, a lower-quality, higher-dimensional group is added. For example, interleaved gradient noise or a z-sampler can be used to add a 2D group because they are ways to transform 2D integer coordinates into scalars with desired properties on a 2D plane. This scalar can be added to the value read from the blue noise mask, and the modulus can again be used to make it between 0 and 1.
[0179] Figure 5 The convergence rate of the 1D function according to at least one embodiment is shown. That is, Figure 5 The convergence rates of 1D functions with x∈[0,1] using time axes of various mask types are shown, along with Monte Carlo and leakage integrals. In at least one embodiment, hierarchical sampling demonstrates that potentially better convergence rates exist if only the 1D time axis is considered without regard to the 2D plane of screen space. The offset plots in the two columns on the right show that starting integration from indices other than 0 does not affect the results, and that spatiotemporal blue noise is asymptotic from any index and is also continuous when reaching the end of the sequence and restarting at index 0. The Van der Corput base (VDC) does not exhibit the characteristics shown by the unstable accuracy with a low sample size.
[0180] Figure 6 The DFT of a 2D projection of a 64x64x16x16 4D blue noise mask according to at least one embodiment is shown. For clarity, Figure 6 The projections described herein are averaged to show the expected spectrum. In at least one embodiment, process 100 (see [link to documentation]) can be used. Figure 1A ) or process 106 (see Figure 1B (A portion of) is used to generate the DFT.
[0181] Figure 7 An image according to at least one embodiment is shown to illustrate the autocorrelation of a blue noise texture. In the embodiment, neighboring objects can have very different values, which results in ripples of correlation (red / white) and anticorrelation (blue / black) at small offset centers, but rapidly decays to decorrelation values (white / gray).
[0182] Figure 8 A 2Dx1D spatiotemporal blue noise mask according to at least one embodiment is shown, with each axis having various sigma.
[0183] Figure 9 It is shown that the generation time according to at least one embodiment is a function of the number of pixels in the blue noise mask and approximately follows y = x 2 The graph shows the curve. Doubling the pixel count can roughly quadruple the processing time.
[0184] 5.1. Random transparency
[0185] In this embodiment, random transparency is the process of randomly selecting whether to accept or ignore samples based on the transparency level of the material. Complex algorithms have been developed using alternative methods, but the core idea of randomly accepting or rejecting pixels can remain the same. In this embodiment, the spatiotemporal blue noise mask described herein uses very low sample counts and low computational cost (single texture read and comparison), thus providing the same blue noise distribution error in screen space as 2D blue noise, but converging faster than other methods used by 2D blue noise. Random transparency is useful in situations such as delayed lighting that stores information about how pixels are shaded rather than the shading result itself, and storing multiple or arbitrary numbers of layers to subsequently compute the appropriate transparency is impractical. Random transparency is also useful in contexts requiring path tracing of a single sample per ray vertex, and focusing solely on the average pixel value is correct for things like semi-transparency, rather than incurring computational and memory costs to compute the semi-transparency of a single sample. Random transparency works by generating a random number ξ∈[0,1] and comparing that random number with the opacity α∈[0,1] of the material. If ξ is greater than α, the sample is discarded. When white noise random numbers are used on ξ, if the test is performed an infinite number of times, the percentage of pixels that survive the test will match α, but for a small number of samples, it has a large variance in both space and time.
[0186] Conversely, using a 2D blue noise mask will make the percentage of surviving pixels spatially more accurate for a smaller number of samples, which will also randomize the surviving pixels, but space them roughly evenly. However, as previously mentioned, methods for animating blue noise over time either change the spatial blue noise or turn it into white noise over time, thus causing poor convergence when taking multiple samples per frame or integrating over multiple frames. Using a spatiotemporal blue noise mask as described herein (where each individual frame is good blue noise, but each pixel is also a good sampling sequence over time) means that individual frames will have surviving pixels (which are spatially distributed blue noise), and also means that each frame will have very different surviving pixels, thus allowing for better convergence over time or on multiple samples within a single frame. Figure 10 The rendering comparison is shown in the image, and... Figure 11 The convergence rate is shown in the figure.
[0187] Figure 10 Random transparency using various types of noise is illustrated according to at least one embodiment. In the embodiment, Figure 10 One sample is shown for each pixel. The top image is the original frame, where the bottom image is Gaussian blurred with 2σ. In this embodiment, the spatiotemporal blue noise is spatially as good as 2D blue noise and better than the golden ratio animated blue noise.
[0188] 5.2. Color Mixing
[0189] Color mixing is the process of adding a small amount of noise to data before quantizing it to obtain a noisy result rather than quantization artifacts. This can be used to hide striping artifacts that would otherwise occur due to reduced bit depth, thus allowing for less memory to be used while attempting to maintain image quality. Color mixing causes pixels to be randomly rounded up or down when quantized, where the probability of rounding toward a quantization level is based on how far the value is from that level. If quantizing consecutive values x∈[0,1] yields n distinct values, then the quantized value is obtained. Random numbers ξ∈[0,1) can be used in the following equation:
[0190]
[0191] When white noise is used for color mixing, the result is a white noise pattern. If blue noise is used instead, the result is more visually pleasing to humans or displays, while also having a more accurate average value over small pixel areas in space. When spatiotemporal blue noise is used for color mixing, the result may be spatial blue noise, but it may also be temporal blue noise, where each pixel can have a more accurate average value over smaller samples over time when animated. Figure 12 The rendering comparison is shown in the image, and... Figure 13The convergence rate is shown in the figure.
[0192] Figure 11 Convergence rates in random alpha for various types of noise according to at least one embodiment are shown. In the embodiment, golden ratio animated blue noise converges faster than spatiotemporal blue noise, but its frequency can be spatially varied.
[0193] Figure 12 The image illustrates color mixing before quantization to 1 bit per color channel using various types of noise, according to at least one embodiment. The top image is the original frame, and the bottom image is Gaussian blurred with 2 sigma. Figure 12 As shown, spatiotemporal blue noise is as good as 2D blue noise in terms of space and is improved compared to the blue noise of golden ratio animation.
[0194] Figure 13 A graph showing the convergence rate in color mixing of various types of noise according to at least one embodiment is presented. In the embodiment, the golden ratio animated blue noise converges faster than spatiotemporal blue noise, but spatially varies in frequency.
[0195] 5.3. Raymarched Participating Medium with Spatiotemporal Blue Noise
[0196] In one embodiment, one or more computing devices execute an algorithm to render a single scattering heterogeneous participating medium with a very low sample count. This is a different type of algorithm from random transparency or color mixing because it demonstrates how blue noise masks can be applied to arbitrary rendering problems. While more sophisticated algorithms exist for rendering participating media, the technique described herein is simple, performance-optimized, yields good results with very low sample counts, and works with rasterization or ray tracing. In the embodiment, the algorithm is run after the major hits have been shaded and the surface depth is known. The surface depth d can be the length of a line segment along the camera ray r that must be integrated. In the embodiment, this line segment is sampled at n evenly spaced locations, where the space between each sample is... Units. Then, the position p of sample s∈Z[0,n-1] s Calculated as:
[0197]
[0198] At each sampling point p s At a certain location, the fog density field F is sampled to obtain the density f. s This is assumed to be the density for the entire step of the distance.
[0199] f s =F(p s )
[0200] Still in p sEvaluate the light visibility function V to obtain the visibility value v for all light i∈I. s,i ∈[0,1]. Visibility value v s,i It could be a binary value similar to when shooting a single ray of light towards the light source, or it could be a more continuous value similar to using a filter closer to a percentage to read the shadow image, or it could come from shooting multiple shadow ray samples.
[0201] v s,i =V(p) s ,i)
[0202] To calculate individual fog samples c s The color of the fog in the shadow determines the fog color c. unlit The shadow of the fog illuminated by light i and the color of the fog c lit,i Fog color can be calculated or provided. Visibility value v s,i Can be multiplied by c lit,i To obtain the contribution of that light. Add all the illumination contributions and then calculate c. unlit Add to the result to obtain that sample c s The final color of the fog.
[0203]
[0204] To calculate the opacity of the sample o s The usual beer law absorption formula can be used, with density f and step distance d.
[0205] o s =e -df
[0206] When performing integration, the cumulative result r can be initialized with the shadow surface color p, and then moving backward from the surface toward the camera, the color and opacity of the fog sample are calculated, and the usual alpha blending operation is applied to the cumulative result.
[0207] r0 = p
[0208] r s =r s-1 (1-o s )+c s o s
[0209] Running the algorithm with low values across n samples along a line segment causes noticeable banding. Much like the case of color mixing, random numbers can be used to replace the banding with noise. In this embodiment, each primary hit sample (e.g., each pixel) is offset by a random value ξ ∈ [0; 1) to each sample point p. s The positions of the samples are still evenly spaced; they are only shifted forward or backward in depth.
[0210]
[0211] White noise was used to obtain screen-space white noise results. 2D blue noise was used to improve the error pattern. Spatiotemporal blue noise was used to obtain a screen-space blue noise error pattern, and the error magnitude was smaller. Figure 14 The rendering result can be seen in [the image / image], and [the image / image] can be seen in [the image / image]. Figure 15 The convergence graph can be seen in the image.
[0212] Figure 14 The diagram illustrates the use of noise, according to at least one embodiment, to randomly offset the starting portion of the ray travel by 4 ray travel steps per pixel. The top image is the original frame, and the bottom image is a depth-aware Gaussian blur with a sigma of 2.
[0213] Figure 15 A graph showing the convergence rate of ray-traveling fog with various types of noise according to at least one embodiment is presented. In the embodiment, only 4 ray-traveling steps are performed per pixel.
[0214] 5.4. Participating media with 2Dx1Dx1D blue noise
[0215] In one embodiment, one or more computing devices can execute an algorithm utilizing a 2Dx1Dx1D blue noise mask, where a previous algorithm utilized a 2Dx1D spatiotemporal blue noise mask. This algorithm can be used to indicate how many higher-dimensional blue noise masks can be used in a rendering algorithm. In both this algorithm and the previous algorithm, the goal is to integrate a single scattering participating medium. In the previous algorithm, samples at regular intervals are acquired along a line segment, and noise is used to cancel out the starting points of those samples in exchange for noise stripes. In this algorithm, the line segment can be divided into n evenly spaced portions, but instead of using only a single random offset for the entire sampling sequence, this algorithm reads the random offset for each sample. Then, n random values ξ are obtained. s ∈[0,1), and the sampling position p s The following can be calculated:
[0216]
[0217] The rest of the algorithms remain the same. Figure 16 The rendering result can be seen in [the image / image], and [the image / image] can be seen in [the image / image]. Figure 17 The convergence plot can be seen in the image. This reformulation changes it from a ray-tracing technique to a hierarchical sampling technique, and if this is compared to spatiotemporal blue noise convergence, it improves for the same sample count.
[0218] Figure 16The diagram illustrates the use of noise to layer 16 samples of line segments for each pixel through a participating medium, according to at least one embodiment. The top image is the original frame, and the bottom imager is a depth-aware Gaussian blur using 2 sigma.
[0219] Figure 17 A graph showing the convergence rate of ray-traveling fog with various types of noise according to at least one embodiment is presented. In the embodiment, only 4 ray-traveling steps are performed per pixel.
[0220] 5.5. Ray Tracing Ambient Occlusion (AO)
[0221] In this embodiment, ray tracing AO is another algorithm that can be used. In this embodiment, AO uses a 2D vector for each pixel to acquire each AO sample. In the techniques described herein, the algorithm generates a blue noise mask with scalar values for each entry rather than each vector. In one or more embodiments, multiple independent streams of scalar values can be derived from a single blue noise mask by a readout offset having approximately the maximum distance of each stream. Thus, in the ray tracing AO algorithm, this extension can be used as an independent spatiotemporal blue noise data stream for each axis. While other, more complex ray tracing and rasterization AO algorithms exist, this is intended to give high-quality results with low sample counts, such as 1 ray tracing sample per pixel (spp) – or even lower if running at sub-full resolution.
[0222] In this embodiment, the ray tracing AO algorithm runs after the main ray strikes position p, and the surface normal n is known. N random 3D unit vectors ξ_i can be generated, added to the surface normal n, and normalized to obtain N cosine-weighted hemispherical samples v_i, where the hemispheres are oriented towards the surface normal n.
[0223]
[0224] In this embodiment, each v_i is used to capture a ray from position p to obtain the direction of the hit distance d. Because AO is a local shading phenomenon, the hit length can be limited to the scene-dependent maximum of d_max. The AO shading value a_i of this ray can be calculated as the percentage of how far the ray has traveled relative to the maximum distance, which is attributed to causing more occlusion and thus a closer hit that is shaded. This also allows for more information per sample than a binary hit or miss result, resulting in lower amplitude noise.
[0225]
[0226] The AO shadow values a_i can then be averaged to give the combined AO shadow value a, which can be used as the shadow term in the lighting equation.
[0227]
[0228] If independent random numbers are used to generate each component of ξ_i, the result will be white noise error. If 2D blue noise is used, the noise is spatially eliminated, and if spatiotemporal blue noise is used, the AO data acquires the desired sampling characteristics over time. Figure 18 The rendering result can be seen in [the image / image], and [the image / image] can be seen in [the image / image]. Figure 19 The convergence graph can be seen in the image.
[0229] 5.6. Spatiotemporal Blue Noise in Heitz Belcour Technique
[0230] In embodiments, blue noise masks tend to show benefits when used in algorithms employing scalars, such as point drawing, color mixing, and ray travel as the participating medium. They also show benefits when used in simpler graphics algorithms that want vectors rather than scalars (such as AO in ray tracing) by employing multiple independent noise streams for each axis. In previous methods, blue noise masks also tend to stop working when sample counting or dimensionality increases (e.g., path tracing). However, in some embodiments, the techniques described herein can be used as extensions of algorithms for generating blue noise masks in path tracing (e.g., the Heitz & Belcoour technique). For example, in the Heitz & Belcoour technique, there may be a seed value per pixel generated by any desired means, which is used to render the result for each pixel using any desired algorithm and sampling sequence. After this rendering is complete, the Heitz & Belcoour technique can break the screen down into smaller 4x4-order segments and sort the pixels in each segment from darkest to brightest. Furthermore, the Heitz & Belcoour technique can break down cross-screen stitched blue noise textures into identical small parts and sort them. These two sorted lists serve as a mapping of how to exchange the seeds used for rendering previous frames, so that if rendered again, the result will be closer to the blue noise. An R2 low-difference sequence can be used to offset readings into this blue noise texture each frame, resulting in each frame having 2D blue noise values that are largely unrelated to the previous frame. This gives a spatially blue noise result, but provides white noise over time. By combining the techniques described in this paper, the Heitz & Belcoour technique can provide spatiotemporal blue noise results, thus preserving the spatial quality of the blue noise while obtaining the desired sampling characteristics over time. Figure 21 The rendering results are shown in [the image], and [the image] is also shown in [the image]. Figure 23 The convergence plot is shown in the figure. Therefore, it should be possible to achieve virtually any target error pattern, such as the desired staggered gradient noise for better use in time-anti-aliasing scenarios.
[0231] Figure 18 The diagram illustrates the use of two independent noise streams to generate the x and y components of a 2D vector mapped to a cosine-weighted hemisphere for a single AO sample per pixel, according to at least one embodiment. The top image is the original frame, and the bottom image is a depth-aware Gaussian blurred with 2 sigma.
[0232] Figure 19 The method of ambient occlusion (AO) convergence according to at least one embodiment is shown to be related to various types of noise.
[0233] Figure 20 Images using one or more of a 2D blue noise mask, a 3D blue noise mask, a spatiotemporal blue noise mask, and a 2DGR blue noise mask, according to at least one embodiment, are shown. In the embodiments, images from... Figure 20 The image shows the result rendered using Monte Carlo integral rendering, where each pixel has four samples.
[0234] Figure 21 An image offset using a Sobol sequence is shown according to at least one embodiment. In the embodiment, the Sobol sequence is offset by vec2 from each mg type of each frame. In the embodiment, the top image is the original frame, and the bottom image is a depth-aware Gaussian blur with 2 sigma.
[0235] Figure 22 The Heitz & Belcour technique, according to at least one embodiment, uses interleaved gradient noise and a stylized grayscale image for a noise pattern target. These images are rendered using standard path tracing rendering code, but the seeds used to randomize each pixel are reordered to give a rendering result as the target image.
[0236] Figure 23 A graph showing convergence in Monte Carlo integral, leakage integral, and convergent leakage integral according to at least one embodiment is shown.
[0237] 5.7. Spatiotemporal point set
[0238] In this embodiment, the empty clustering algorithm is modified to generate a spatiotemporal blue noise mask. In this embodiment, the spatiotemporal blue noise has the property of being thresholded to a certain percentage, such that a corresponding percentage of pixels will survive, and the surviving pixels will be distributed within the constraints of the dimensional group in a blue noise sample pattern. More specifically, if all pixels in the spatiotemporal 2Dx1D blue noise mask are thresholded to 10%, each 2D XY slice of the mask will show approximately 10% pixel survival, and they will be blue noise distributed (randomized but roughly uniformly spaced). Furthermore, viewing each pixel in isolation along the 1D Z-axis produces a 1D image where approximately 10% of the pixels will also survive, and they will also be blue noise distributed. These properties can be extended to any dimensional and sub-dimensional grouping used to generate the mask.
[0239] An example use case for this feature is in situations where an importance map for sparse ray tracing can be performed within a scene. Approximately how many rays are needed per frame can be defined, and this per-pixel count, along with a per-pixel random number, is used to determine whether a pixel should emit a ray (per frame). When using random numbers that are spatially and temporally white, clustering and gaps occur, resulting in non-uniform and redundant sampling both spatially and temporally. When using a folded book of independent 2D blue noise textures, the results improve spatially, but redundant sampling still exists over time. When using a spatiotemporal blue noise mask, both time and space are sampled more uniformly because the noise pattern can be the desired blue noise pattern in screen space, but for the same number of frames, more unique pixels will have rays emitted for them, thus maximizing the unique information received per frame, per ray. Figure 27 The graph shows the unique pixel count over time, and Figure 26 The image can also be seen visually in the middle.
[0240] Figure 24 This illustrates how a threshold mask, according to at least one embodiment, can form a point set of any density. That is, in Figure 24 In this study, 1024 blue noise samples were compared to different levels using a 128x128x10 2Dx1D blue noise mask thresholding method with the Best Transmission (BNOT) sample. BNOT is spatially much higher quality but has a fixed density and does not provide temporal processing, thus forcing the individual sample sets to become white noise over time.
[0241] Figure 25 This illustrates how a set of threshold points, according to at least one embodiment, maintains its desired spectrum on the axis group. The threshold pixels, as blue noise sample points, are characteristics of a blue noise mask made using a void-gathering algorithm, but are generally not true for blue noise masks. That is, Figure 25The DFT of the 2D projection of a 64x64x64 2Dx1D blue noise mask with a 1 / 8 threshold is shown to illustrate how the set of threshold points preserves the mask’s inherent blue noise spectrum.
[0242] In the embodiments, this document describes a modification to the empty clustering algorithm, wherein blue noise masks of any dimension with blue noise characteristics limited to a set of subspace axes are generated. In the embodiments, these blue noise masks can be used in various low-sample-count rendering algorithms that aim to obtain the desired blue noise error pattern while also converging faster than other methods using blue noise masks. In the embodiments, these blue noise masks may have a threshold that brings these properties into the blue noise sampling domain.
[0243] In some instances, a limitation is that these blue noise masks have scalar values for each entry, rather than vector values. Many rendering techniques require random vector values to operate on, and while the flow of scalar values shows benefits in simpler algorithms such as ambient occlusion or sampled region lights, they completely break down for more complex algorithms such as path tracing. However, in path tracing, the techniques described in this paper can be used in conjunction with Heitz and Belcour techniques. In some cases, blue noise masks can be extended to include vectors to obtain better results more directly without the problems associated with roughly inverted pixel rendering.
[0244] Figure 26 Five cumulative frames of pixels sampled from an image are shown, according to at least one embodiment, using a non-uniform importance map to make pixels oriented towards the center more likely to be sampled. While both 2D blue noise and spatiotemporal blue noise have desired sampling patterns in space, spatiotemporal blue noise samples more unique pixels in a shorter number of frames.
[0245] Since importance sampling is a subject largely inconsistent with the use of specific sample patterns, blue noise itself more often happens to retain the desired properties when undergoing a distortion function. In embodiments, these blue noise masks can be extended to not only have the desired projection for each axis group, but also allow them to have a specific distribution for each axis group. This makes it possible to generate blue noise in the distorted space, meaning that the blue noise is not corrupted in any way, and important sampled PDFs can be baked into them. While some PDFs may be very specialized for their use and therefore may not bake as expected (such as HDRI skybox images), other PDFs will be more reused, such as GGX for specular reflections.
[0246] Figure 27It is shown that white noise according to at least one embodiment can have redundant sampled pixels for each frame, and that spatial blue noise removes spatially redundant pixels over time, and 2Dx1D spatiotemporal blue noise also removes them over time.
[0247] While the techniques described herein involve blue noise masks, noise of other colors (e.g., red noise) can also be applied to improve real-time image rendering and enhancement.
[0248] Reasoning and training logic
[0249] Figure 28A Inference and / or training logic 2815 is illustrated for performing inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, the inference and / or training logic 2815 may be implemented using process 100 or process 106 (see [link]). Figure 1B For example, using DLSS to render images.
[0250] In at least one embodiment, inference and / or training logic 2815 may include, but is not limited to, code and / or data storage 2801 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network trained for and / or used for inference in one or more embodiments. In at least one embodiment, training logic 2815 may include or be coupled to code and / or data storage 2801 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 2801 stores weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 2801 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0251] In at least one embodiment, any portion of the code and / or data storage 2801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 2801 may be a cache memory, dynamic random-addressable memory (“DRAM”), static random-addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 2801 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip or off-chip storage space, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0252] In at least one embodiment, the inference and / or training logic 2815 may include, but is not limited to, code and / or data storage 2805 to store backpropagation and / or output weights and / or input / output data neural networks corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, the code and / or data storage 2805 stores weight parameters and / or input / output data for each layer of a neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, the training logic 2815 may include or be coupled to code and / or data storage 2805 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, an arithmetic logic unit (ALU)).
[0253] In at least one embodiment, code (such as graph code) causes the architecture of the neural network corresponding to that code to load weights or other parameter information into the processor ALU. In at least one embodiment, any portion of the code and / or data storage 2805 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 2805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 2805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice between the code and / or data storage 2805 being internal or external to the processor, for example, whether it consists of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.
[0254] In at least one embodiment, code and / or data storage 2801 and code and / or data storage 2805 may be separate storage structures. In at least one embodiment, code and / or data storage 2801 and code and / or data storage 2805 may be the same storage structure. In at least one embodiment, code and / or data storage 2801 and code and / or data storage 2805 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 2801 and code and / or data storage 2805 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0255] In at least one embodiment, the inference and / or training logic 2815 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 2810 (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 2820, which are functions of input / output and / or weight parameter data stored in code and / or data storage 2801 and / or code and / or data storage 2805. In at least one embodiment, activation is activated in response to execution instructions or other code, and linear algebraic and / or matrix-based mathematical generation performed by ALU 2810 is stored in activation storage 2820, wherein weight values stored in code and / or data storage 2805 and / or code and / or data storage 2801 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, and any or all of these can be stored in code and / or data storage 2805 or code and / or data storage 2801 or other on-chip or off-chip storage.
[0256] In at least one embodiment, one or more ALUs 2810 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 2810 may be located outside the processor or other hardware logic device or the circuits using them (e.g., coprocessors). In at least one embodiment, one or more ALUs 2810 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., central processing unit, graphics processing unit, fixed-function unit, etc.). In at least one embodiment, code and / or data storage 2801, code and / or data storage 2805, and activation storage 2820 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be located in different processors or other hardware logic devices or circuits, or in some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 2820 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0257] In at least one embodiment, the active memory 2820 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 2820 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 2820 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types.
[0258] In at least one embodiment, Figure 28A The inference and / or training logic 2815 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 28A The inference and / or training logic 2815 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware or other hardware (such as field programmable gate array (“FPGA”)).
[0259] Figure 28B Inference and / or training logic 2815 according to at least one embodiment is illustrated. In at least one embodiment, the inference and / or training logic 2815 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 28B The inference and / or training logic 2815 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 28BThe inference and / or training logic 2815 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 2815 includes, but is not limited to, code and / or data storage 2801 and code and / or data storage 2805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 28B In at least one embodiment shown, each of code and / or data storage 2801 and code and / or data storage 2805 is associated with dedicated computing resources (e.g., computing hardware 2802 and computing hardware 2806), respectively. In at least one embodiment, each of computing hardware 2802 and computing hardware 2806 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on the information stored in code and / or data storage 2801 and code and / or data storage 2805, respectively, and the results of the function execution are stored in activation storage 2820.
[0260] In at least one embodiment, each of the code and / or data storage 2801 and 2805 and the corresponding computing hardware 2802 and 2806 corresponds to a different layer of the neural network, such that activations obtained from one “store / computation pair 2801 / 2802” of the code and / or data storage 2801 and computing hardware 2802 provide input as input to the next “store / computation pair 2805 / 2806” of the code and / or data storage 2805 and computing hardware 2806, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each store / computation pair 2801 / 2802 and 2805 / 2806 may correspond to more than one neural network layer. In at least one embodiment, additional store / computation pairs (not shown) may be included in the inference and / or training logic 2815 following or paralleling the store / computation pairs 2801 / 2802 and 2805 / 2806.
[0261] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28BOne or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0262] Neural network training and deployment
[0263] Figure 29 Training and deployment of a deep neural network according to at least one embodiment are illustrated. In at least one embodiment, an untrained neural network 2906 is trained using a training dataset 2902. In at least one embodiment, the untrained neural network 2906 can be implemented using process 100 or process 106 (see [link to documentation]). Figure 1B For example, DLSS or other neural network operations are used to render images. In at least one embodiment, the training framework 2904 is a PyTorch framework, while in other embodiments, the training framework 2904 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 2904 trains an untrained neural network 2906 and enables it to be trained using the processing resources described herein to generate a trained neural network 2908. In at least one embodiment, the weights may be randomly selected or pre-trained using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0264] In at least one embodiment, supervised learning is used to train an untrained neural network 2906, wherein the training dataset 2902 includes inputs paired with desired outputs for input, or wherein the training dataset 2902 includes inputs with known outputs and the neural network 2906 is manually graded output. In at least one embodiment, the untrained neural network 2906 is trained in a supervised manner, and inputs from the training dataset 2902 are processed, and the resulting outputs are compared with a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained neural network 2906. In at least one embodiment, a training framework 2904 adjusts the weights controlling the untrained neural network 2906. In at least one embodiment, the training framework 2904 includes tools for monitoring the degree to which the untrained neural network 2906 converges to a model (e.g., a trained neural network 2908) adapted to generate the correct answer (e.g., result 2914) based on input data (e.g., a new dataset 2912). In at least one embodiment, the training framework 2904 repeatedly trains the untrained neural network 2906 while adjusting the weights to improve the output of the untrained neural network 2906 using a loss function and tuning algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 2904 trains the untrained neural network 2906 until the untrained neural network 2906 reaches the desired accuracy. In at least one embodiment, the trained neural network 2908 can then be deployed to implement any number of machine learning operations.
[0265] In at least one embodiment, unsupervised learning is used to train an untrained neural network 2906, wherein the untrained neural network 2906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 2902 will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 2906 can learn groupings within the training dataset 2902 and can determine how each input relates to the untrained dataset 2902. In at least one embodiment, unsupervised training can be used to generate a self-organizing graph in a trained neural network 2908, which is capable of performing operations useful for reducing the dimensionality of the new dataset 2912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows the identification of data points in the new dataset 2912 that deviate from the normal patterns of the new dataset 2912.
[0266] In at least one embodiment, semi-supervised learning can be used, a technique in which a mixture of labeled and unlabeled data is included in the training dataset 2902. In at least one embodiment, the training framework 2904 can be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 2908 to adapt to a new dataset 2912 without forgetting the knowledge injected into the trained neural network 2908 during initial training.
[0267] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0268] Data Center
[0269] Figure 30 An example data center 3000 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 3000 includes a data center infrastructure layer 3010, a framework layer 3020, a software layer 3030, and an application layer 3040. The data center 3000 can implement process 100 or process 106 (see...). Figure 1B ).
[0270] In at least one embodiment, such as Figure 30As shown, the data center infrastructure layer 3010 may include a resource coordinator 3012, packet computing resources 3014, and node computing resources (“nodes CR”) 3016(1)-3016(N), where “N” represents a positive integer (which may be an integer “N” different from the integers used in other diagrams). In at least one embodiment, nodes CR 3016(1)-3016(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 3018(1)-3018(N) (e.g., dynamic read-only memory, solid-state drives, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 3016(1)-3016(N) may be servers having one or more of the aforementioned computing resources.
[0271] In at least one embodiment, the grouped computing resource 3014 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographical locations. In at least one embodiment, the individual groups of node CRs within the grouped computing resource 3014 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0272] In at least one embodiment, resource coordinator 3012 may be configured or otherwise control one or more nodes CR3016(1)-3016(N) and / or grouped computing resources 3014. In at least one embodiment, resource coordinator 3012 may include a Software Design Infrastructure (“SDI”) management entity for data center 3000. In at least one embodiment, resource coordinator 2812 may include hardware, software, or some combination thereof.
[0273] In at least one embodiment, such as Figure 30As shown, the framework layer 3020 includes a job scheduler 3022, a configuration manager 3024, a resource manager 3026, and a distributed file system 3028. In at least one embodiment, the framework layer 3020 may include a framework of software 3032 supporting the software layer 3030 and / or one or more applications 3042 supporting the application layer 3040. In at least one embodiment, the software 3032 or application 3042 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 3020 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize the distributed file system 3028 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 3022 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 3000. In at least one embodiment, the configuration manager 3024 may be able to configure different layers, such as the software layer 3030 and the framework layer 3020, which includes Spark and a distributed file system 3028 for supporting large-scale data processing. In at least one embodiment, the resource manager 3026 is able to manage cluster or group computing resources mapped to or allocated to support the distributed file system 3028 and the job scheduler 3022. In at least one embodiment, the cluster or group computing resources may include group computing resources 3014 on the data center infrastructure layer 3010. In at least one embodiment, the resource manager 3026 may coordinate with the resource coordinator 3012 to manage these mapped or allocated computing resources.
[0274] In at least one embodiment, the software 3032 included in the software layer 3030 may include software used by at least a portion of the nodes CR3016(1)-3016(N), the grouped computing resources 3014, and / or the distributed file system 3028 of the framework layer 3020. In at least one embodiment, one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0275] In at least one embodiment, one or more applications 3042 included in application layer 3040 may include one or more types of applications used by at least a portion of nodes CR3016(1)-3016(N), grouped computing resources 3014, and / or the distributed file system 3028 of framework layer 3020. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0276] In at least one embodiment, any of the configuration manager 3024, resource manager 3026, and resource coordinator 3012 can perform any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 3000 and can prevent underutilization and / or poor performance of the data center.
[0277] In at least one embodiment, the data center 3000 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to the data center 3000. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to the data center 3000, by using weight parameters calculated through one or more training techniques described herein.
[0278] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0279] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28BDetails are provided regarding the inference and / or training logic 2815. In at least one embodiment, the inference and / or training logic 2815 can be implemented in the system. Figure 30 Used in this context for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0280] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0281] Supercomputing
[0282] The following figures illustrate, but are not limited to, exemplary supercomputer-based systems that can be used to implement at least one embodiment.
[0283] In at least one embodiment, a supercomputer can refer to a hardware system exhibiting substantial parallelism and comprising at least one chip, wherein the chips in the system are interconnected via a network and housed in a hierarchically organized enclosure. In at least one embodiment, a large hardware system filling a machine room with several racks is a specific example of a supercomputer, each rack containing several board / rack modules, each board / rack module containing several chips all interconnected via a scalable network. In at least one embodiment, a single rack of such a large hardware system is another example of a supercomputer. In at least one embodiment, a single chip exhibiting substantial parallelism and comprising several hardware components can also be considered a supercomputer because as feature size can be reduced, the amount of hardware that can be incorporated into a single chip can also increase.
[0284] Figure 31A A chip-level supercomputer 3100 according to at least one embodiment is illustrated. In at least one embodiment, within an FPGA or ASIC chip, the main computation is executed within a finite state machine (3104) called a thread unit. In one embodiment, the supercomputer 3100 may implement process 100 or process 106 (see...). Figure 1BIn at least one embodiment, a task and synchronization network (3102) connects to a finite state machine and is used to schedule threads and perform operations in the correct order. In at least one embodiment, a memory network (3106, 3110) is used to access a multi-level cache hierarchy (3108, 3112) partitioned on the chip. In at least one embodiment, a memory controller (3116) and an off-chip memory network (3114) are used to access off-chip memory. In at least one embodiment, when the design is not suitable for a single logic chip, an I / O controller (3118) is used for cross-chip communication.
[0285] Figure 31B A supercomputer at the rack module level is illustrated according to at least one embodiment. In at least one embodiment, within the rack module, there are multiple FPGA or ASIC chips (3120) connected to one or more DRAM cells (3122) constituting the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its adjacent FPGA / ASIC chip using a wide on-board bus with differential high-speed signaling (3124). In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.
[0286] Figure 31C A rack-mounted supercomputer according to at least one embodiment is shown. Figure 31D A supercomputer at the entire system level is illustrated according to at least one embodiment. In at least one embodiment, reference is made to... Figure 31C and Figure 31DHigh-speed serial optical or copper cables (3126, 3128) are used to implement a scalable, potentially incomplete, hypercube network between rack modules within the rack and across racks throughout the system. In at least one embodiment, one of the FPGA / ASIC chips in the accelerator is connected to the host system via a PCI-Express connection (3130). In at least one embodiment, the host system includes a host microprocessor (3134) running the software portion of an application and a memory consisting of one or more host memory DRAM cells (3132) consistent with the memory on the accelerator. In at least one embodiment, the host system may be a standalone module on one of the racks or may be integrated with one of the modules of the supercomputer. In at least one embodiment, a cubic-connected loop topology provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, a group of FPGA / ASIC chips on a rack module may act as a single hypercube node, increasing the total number of external links per group compared to a single chip. In at least one embodiment, a group comprises chips A, B, C, and D on a rack module having an internal wide differential bus connecting A, B, C, and D in a toroidal organization. In at least one embodiment, there are 12 serial communication cables connecting the rack module to the outside world. In at least one embodiment, chip A on the rack module is connected to serial communication cables 0, 1, and 2. In at least one embodiment, chip B is connected to cables 3, 4, and 5. In at least one embodiment, chip C is connected to cables 6, 7, and 8. In at least one embodiment, chip D is connected to cables 9, 10, and 11. In at least one embodiment, the entire group {A, B, C, D} constituting the rack module can form a hypercube node within a supercomputer system, with up to 2^12 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, for chip A to send a message on link 4 of group {A, B, C, D}, the message must first be routed to chip B, which has an onboard differential wide bus connection. In at least one embodiment, messages arriving at group {A, B, C, D} (i.e., arriving at B) on link 4, destined for chip A, must also first be routed to the correct destination chip (A) within group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes can also be implemented.
[0287] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0288] Computer System
[0289] Figure 32 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 3200 may include, but is not limited to, components such as processor 3202, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, system 3200 may implement process 100 or process 106 (see...). Figure 1B In at least one embodiment, the computer system 3200 may include a processor, such as one available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 3200 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0290] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0291] In at least one embodiment, the computer system 3200 may include, but is not limited to, a processor 3202, which may include, but is not limited to, one or more execution units 3208, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 3200 is a single-processor desktop or server system, but in another embodiment, the computer system 3200 may be a multiprocessor system. In at least one embodiment, the processor 3202 may include, but is not limited to, a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 3202 may be coupled to a processor bus 3210, which may transmit data signals between the processor 3202 and other components in the computer system 3200.
[0292] In at least one embodiment, processor 3202 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 3204. In at least one embodiment, processor 3202 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 3202. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 3206 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0293] In at least one embodiment, an execution unit 3208, including but not limited to logic for performing integer and floating-point operations, is also located within the processor 3202. In at least one embodiment, the processor 3202 may further include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 3208 may include logic for processing a packaged instruction set 3209. In at least one embodiment, by including the packaged instruction set 3209 in the instruction set of a general-purpose processor, along with the associated circuitry for executing the instructions, the packaged data in the processor 3202 can be used to perform operations used by numerous multimedia applications. In at least one embodiment, many multimedia applications can be executed more quickly and efficiently by using the full width of the processor’s data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor’s data bus to perform one or more operations on one data element at a time.
[0294] In at least one embodiment, execution unit 3208 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, computer system 3200 may include, but is not limited to, memory 3220. In at least one embodiment, memory 3220 may be a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or another storage device. In at least one embodiment, memory 3220 may store instructions 3219 and / or data 3221 represented by data signals that can be executed by processor 3202.
[0295] In at least one embodiment, the system logic chip may be coupled to processor bus 3210 and memory 3220. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 3216, and processor 3202 may communicate with MCH 3216 via processor bus 3210. In at least one embodiment, MCH 3216 may provide a high-bandwidth memory path 3218 to memory 3220 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, MCH 3216 may initiate data signals between processor 3202, memory 3220, and other components in computer system 3200, and bridge data signals between processor bus 3210, memory 3220, and system I / O interface 3222. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 3216 can be coupled to memory 3220 via high-bandwidth memory path 3218, and graphics / video card 3212 can be coupled to MCH 3216 via Accelerated Graphics Port (“AGP”) interconnect 3214.
[0296] In at least one embodiment, the computer system 3200 may use the system I / O interface 3222 as a proprietary hub interface bus to couple the MCH 3216 to the I / O controller hub (“ICH”) 3230. In at least one embodiment, the ICH 3230 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 3220, chipset, and processor 3202. Examples may include, but are not limited to, an audio controller 3229, a firmware hub (“Flash BIOS”) 3228, a wireless transceiver 3226, a data storage 3224, a conventional I / O controller 3223 including a user input and keyboard interface 3225, a serial expansion port 3227 (e.g., a Universal Serial Bus (USB) port), and a network controller 3234. In at least one embodiment, the data storage 3224 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0297] In at least one embodiment, Figure 32 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 32 The SoC can be shown. In at least one embodiment, Figure 32The devices shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 3200 are interconnected using a Compute Fast Link (CXL) interconnect.
[0298] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details are provided regarding the inference and / or training logic 2815. In at least one embodiment, the inference and / or training logic 2815 can... Figure 32 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0299] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0300] Figure 33 This is a block diagram illustrating an electronic device 3300 utilizing a processor 3310 according to at least one embodiment. In at least one embodiment, the electronic device 3300 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, laptop computer, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device. In at least one embodiment, the electronic device 3300 may implement process 100 or process 106 (see...). Figure 1B ).
[0301] In at least one embodiment, the electronic device 3300 may, but is not limited to, a processor 3310 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 3310 is coupled using a bus or interface, such as I... 2C-bus, System Management Bus (“SMBus”), Low Pin Count (LPC) bus, Serial Peripheral Interface (“SPI”), High Definition Audio (“HDA”) bus, Serial Advanced Technology Accessory (“SATA”) bus, Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 33 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 33 An exemplary SoC can be shown. In at least one embodiment, Figure 33 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 33 One or more components are interconnected using Computational Fast Link (CXL) interconnects.
[0302] In at least one embodiment, Figure 33 It may include a display 3324, a touch screen 3325, a touchpad 3330, a near field communication unit (“NFC”) 3345, a sensor hub 3340, a thermal sensor 3346, a fast chipset (“EC”) 3335, a trusted platform module (“TPM”) 3338, a BIOS / firmware / flash (“BIOS, FW Flash”) 3322, a DSP 3360, a drive 3320 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 3350, a Bluetooth unit 3352, a wireless wide area network unit (“WWAN”) 3356, a global positioning system (GPS) unit 3355, a camera (“USB 3.0 camera”) 3354 (e.g., a USB 3.0 camera), and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 3315 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.
[0303] In at least one embodiment, other components may be communicatively coupled to processor 3310 via the components described herein. In at least one embodiment, accelerometer 3341, ambient light sensor (“ALS”) 3342, compass 3343, and gyroscope 3344 may be communicatively coupled to sensor hub 3340. In at least one embodiment, thermal sensor 3339, fan 3337, keyboard 3336, and touchpad 3330 may be communicatively coupled to EC 3335. In at least one embodiment, speaker 3363, earphone 3364, and microphone (“mic”) 3365 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 3362, which in turn may be communicatively coupled to DSP 3360. In at least one embodiment, audio unit 3362 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 3357 may be communicatively coupled to WWAN unit 3356. In at least one embodiment, components such as WLAN unit 3350, Bluetooth unit 3352, and WWAN unit 3356 can be implemented as next-generation form factor (NGFF).
[0304] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details are provided regarding the inference and / or training logic 2815. In at least one embodiment, the inference and / or training logic 2815 can be implemented in the system. Figure 33 It is used in the context of reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0305] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0306] Figure 34A computer system 3400 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 3400 is configured to implement various processes and methods described throughout this disclosure. In at least one embodiment, the computer system 3400 may implement process 100 or process 106 (see...). Figure 1B ).
[0307] In at least one embodiment, the computer system 3400 includes, but is not limited to, at least one central processing unit (“CPU”) 3402 connected to a communication bus 3410 implemented using any suitable protocol, such as PCI (“Peripheral Device Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 3400 includes, but is not limited to, main memory 3404 and control logic (e.g., implemented in hardware, software, or a combination thereof), and data may be stored in main memory 3404 in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“Network Interface”) 3422 provides an interface to other computing devices and networks for receiving data using the computer system 3400 and transferring data to other systems.
[0308] In at least one embodiment, the computer system 3400 includes, but is not limited to, an input device 3408, a parallel processing system 3412, and a display device 3406, which may be implemented using conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light-emitting diode (“LED”) display, plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 3408 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the modules described herein may reside on a single semiconductor platform to form the processing system.
[0309] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details are provided regarding the inference and / or training logic 2815. In at least one embodiment, the inference and / or training logic 2815 can be implemented in the system. Figure 34 It is used to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architecture or neural network use cases described herein.
[0310] In at least one embodiment, Figure 28A and28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0311] Figure 35 A computer system 3500 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 3500 includes, but is not limited to, a computer 3510 and a USB stick 3520. In at least one embodiment, the computer 3510 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 3510 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer. In at least one embodiment, the computer system 3500 may implement process 100 or process 106 (see...). Figure 1B ).
[0312] In at least one embodiment, the USB stick 3520 includes, but is not limited to, a processing unit 3530, a USB interface 3540, and USB interface logic 3550. In at least one embodiment, the processing unit 3530 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 3530 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 3530 includes an application-specific integrated circuit (“ASIC”) optimized to perform any amount and type of operations associated with machine learning. For example, in at least one embodiment, the processing unit 3530 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, the processing unit 3530 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.
[0313] In at least one embodiment, the USB interface 3540 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, the USB interface 3540 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, the USB interface 3540 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 3550 may include any amount and type of logic enabling the processing unit 3530 to interface with a device (e.g., computer 3510) via the USB connector 3540.
[0314] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 2815 are combined herein. Figure 28A And / or 28B is provided. In at least one embodiment, the inference and / or training logic 2815 can be provided in the system. Figure 35 The operation is used to infer or predict based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architecture, or neural network use cases as described herein.
[0315] In at least one embodiment, utilizing Figure 28A and 28B The system described herein implements a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0316] Figure 36AAn exemplary architecture is shown in which multiple GPUs 3610(1)-3610(N) are communicatively coupled to multiple multi-core processors 3605(1)-3605(M) via high-speed links 3640(1)-3640(N) (e.g., bus / point-to-point interconnect, etc.). In at least one embodiment, the high-speed links 3640(1)-3640(N) support communication throughput of 4GB / s, 30GB / s, 80GB / s, or higher. In at least one embodiment, various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. In the various figures, “N” and “M” represent positive integers, the values of which may vary from figure to figure. In at least one embodiment, the multiple GPUs 3610(1)-3610(N) may implement process 100 or process 106 individually or in combination (see Figure 1B ).
[0317] Furthermore, in at least one embodiment, two or more GPUs 3610 are interconnected via high-speed links 3629(1)-3629(2), which can be implemented using a protocol / link similar to or different from that used for high-speed links 3640(1)-3640(N). Similarly, two or more multi-core processors 3605 can be connected via high-speed link 3628, which can be a symmetric multiprocessor (SMP) bus operating at speeds of 20GB / s, 30GB / s, 120GB / s, or higher. Alternatively, similar protocols / links (e.g., via a common interconnect structure) can be used. Figure 36A This shows all communication between the various system components.
[0318] In at least one embodiment, each multi-core processor 3605 is communicatively coupled to processor memories 3601(1)-3601(M) via memory interconnects 3626(1)-3626(M), and each GPU 3610(1)-3610(N) is communicatively coupled to GPU memories 3620(1)-3620(N) via GPU memory interconnects 3650(1)-3650(N). In at least one embodiment, memory interconnects 3626 and 3650 may utilize similar or different memory access technologies. By way of example and not limitation, processor memories 3601(1)-3601(M) and GPU memories 3620 may be volatile memories, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories, such as 3D XPoint or Nano-RAM. In at least one embodiment, some portions of the processor memory 3601 may be volatile memory, while other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0319] As described herein, although the various multi-core processors 3605 and GPUs 3610 can be physically coupled to specific memories 3601 and 3620 respectively, and / or can implement a unified memory architecture, in which the virtual system address space (also known as the “effective address” space) is distributed among the various physical memories. For example, processor memories 3601(1)-3601(M) can each contain 64GB of system memory address space, and GPU memories 3620(1)-3620(N) can each contain 32GB of system memory address space, resulting in a total addressable memory size of 256GB when M=2 and N=4. N and M may also be other values.
[0320] Figure 36B Additional details are shown regarding the interconnection between a multi-core processor 3607 and a graphics acceleration module 3646 according to an exemplary embodiment. In at least one embodiment, the graphics acceleration module 3646 may include one or more GPU chips integrated on a line card coupled to the processor 3607 via a high-speed link 3640 (e.g., PCIe bus, NVLink, etc.). In at least one embodiment, the graphics acceleration module 3646 may optionally be integrated on a package or chip having the processor 3607.
[0321] In at least one embodiment, the processor 3607 includes multiple cores 3660A-3660D, each core having a translation back cover buffer (“TLB”) 3661A-3661D and one or more caches 3662A-3662D. In at least one embodiment, the cores 3660A-3660D may include various other components (not shown) for executing instructions and processing data. In at least one embodiment, the caches 3662A-3662D may include level 1 (L1) and level 2 (L2) caches. Furthermore, one or more shared caches 3656 may be included in the caches 3662A-3662D and shared by the respective groups of cores 3660A-3660D. For example, one embodiment of the processor 3607 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, the processor 3607 and the graphics acceleration module 3646 are connected to a system memory 3614, which may include... Figure 36A The processor memory in the memory is 3601(1)-3601(M).
[0322] In at least one embodiment, consistency of data and instructions stored in the various caches 3662A-3662D, 3656 and system memory 3614 is maintained via inter-core communication through the consistency bus 3664. In at least one embodiment, for example, each cache may have associated cache consistency logic / circuit to communicate via the consistency bus 3664 in response to the detection of a read or write to a particular cache line. In at least one embodiment, a cache snooping protocol is implemented via the consistency bus 3664 to snoop on cache accesses.
[0323] In at least one embodiment, proxy circuitry 3625 communicatively couples graphics acceleration module 3646 to coherence bus 3664, thereby allowing graphics acceleration module 3646 to participate in cache coherence protocols as a peer of cores 3660A-3660D. Specifically, in at least one embodiment, interface 3635 provides connectivity to proxy circuitry 3625 via high-speed link 3640, and interface 3637 connects graphics acceleration module 3646 to high-speed link 3640.
[0324] In at least one embodiment, the accelerator integrated circuit 3636 provides cache management, memory access, context management, and interrupt management services for a plurality of graphics processing engines 3631(1)-3631(N) of the graphics acceleration module 3646. In at least one embodiment, the graphics processing engines 3631(1)-3631(N) may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 3631(1)-3631(N) may optionally include different types of graphics processing engines within the GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module 3646 may be a GPU having a plurality of graphics processing engines 3631(1)-3631(N), or the graphics processing engines 3631(1)-3631(N) may be individual GPUs integrated on a general-purpose package, line card, or chip.
[0325] In at least one embodiment, the accelerator integrated circuit 3636 includes a memory management unit (MMU) 3639 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 3614. In at least one embodiment, the MMU 3639 may also include a translation back buffer (“TLB”) (not shown) for caching virtual / effective-to-physical / real address translations. In at least one embodiment, a cache 3638 may store commands and data for efficient access by graphics processing engines 3631(1)-3631(N). In at least one embodiment, a fetch unit 3644 may be used to keep data stored in cache 3638 and graphics memory 3633(1)-3633(M) consistent with core caches 3662A-3662D, 3656 and system memory 3614. As previously mentioned, this task can be accomplished via proxy circuitry 3625 representing cache 3638 and graphics memory 3633(1)-3633(M) (e.g., sending updates related to the modification / access of cache lines on processor caches 3662A-3662D, 3656 to cache 3638 and receiving updates from cache 3638).
[0326] In at least one embodiment, a set of registers 3645 stores context data of threads executed by graphics processing engines 3631(1)-3631(N), and context management circuitry 3648 manages the thread context. For example, context management circuitry 3648 can perform save and restore operations to save and restore the context of individual threads during context switching (e.g., saving the first thread and storing the second thread so that the second thread can be executed by the graphics processing engine). For example, during context switching, context management circuitry 3648 can store the current register value into a designated area in memory (e.g., identified by a context pointer). The register value can then be restored when returning to the context. In at least one embodiment, interrupt management circuitry 3647 receives and processes interrupts received from system devices.
[0327] In at least one embodiment, the virtual / effective address from the graphics processing engine 3631 is translated into a real / physical address in system memory 3614 via MMU 3639. In at least one embodiment, the accelerator integrated circuit 3636 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 3646 and / or other accelerator devices. In at least one embodiment, the graphics accelerator module 3646 may be dedicated to a single application executing on processor 3607, or may be shared among multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented, wherein the resources of the graphics processing engines 3631(1)-3631(N) are shared with multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into “slices” based on processing requirements and priorities associated with VMs and / or applications, which are allocated to different VMs and / or applications.
[0328] In at least one embodiment, the accelerator integrated circuit 3636 acts as a bridge to the system of the graphics acceleration module 3646 and provides address translation and system memory caching services. Additionally, in at least one embodiment, the accelerator integrated circuit 3636 can provide virtualization facilities for the host processor to manage the virtualization, interrupt, and memory management of the graphics processing engines 3631(1)-3631(N).
[0329] In at least one embodiment, since the hardware resources of the graphics processing engines 3631(1)-3631(N) are explicitly mapped to the real address space seen by the host processor 3607, any host processor can directly address these resources using valid address values. In at least one embodiment, a function of the accelerator integrated circuit 3636 is to physically separate the graphics processing engines 3631(1)-3631(N) so that they appear as independent units to the system.
[0330] In at least one embodiment, one or more graphics memories 3633(1)-3633(M) are coupled to each graphics processing engine 3631(1)-3631(N), and N = M. In at least one embodiment, the graphics memories 3633(1)-3633(M) store instructions and data processed by each graphics processing engine 3631(1)-3631(N). In at least one embodiment, the graphics memories 3633(1)-3633(M) may be volatile memories, such as DRAM (including stacked DRAM), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories, such as 3DXPoint or Nano-RAM.
[0331] In at least one embodiment, to reduce data traffic on the high-speed link 3640, a biasing technique can be used to ensure that the data stored in the graphics memory 3633(1)-3633(M) is the data most frequently used by the graphics processing engine 3631(1)-3631(N) and preferably not used (at least infrequently) by the cores 3660A-3660D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep the data needed by the cores (and preferably not the graphics processing engine 3631(-1)-3631(N)) in the caches 3662A-3662D, 3656 and system memory 3614.
[0332] Figure 36C Another exemplary embodiment is shown, wherein the accelerator integrated circuit 3636 is integrated within the processor 3607. In this embodiment, the graphics processing engines 3631(1)-3631(N) communicate directly with the accelerator integrated circuit 3636 via a high-speed link 3640 through interfaces 3637 and 3635 (which may also be any form of bus or interface protocol). In at least one embodiment, the accelerator integrated circuit 3636 can perform operations related to... Figure 36B The described operation is similar. However, due to its close proximity to the coherence bus 3664 and caches 3662A-3662D, 3656, it may have higher throughput. In at least one embodiment, the accelerator integrated circuit supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integrated circuit 3636 and a programming model controlled by the graphics acceleration module 3646.
[0333] In at least one embodiment, graphics processing engines 3631(1)-3631(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel requests from other applications to graphics processing engines 3631(1)-3631(N), thereby providing virtualization within a VM / partition.
[0334] In at least one embodiment, graphics processing engines 3631(1)-3631(N) can be shared by multiple VM / application partitions. In at least one embodiment, the shared model can use a hypervisor to virtualize graphics processing engines 3631(1)-3631(N) to allow each operating system to access them. In at least one embodiment, for a single-partition system without a hypervisor, the operating system owns graphics processing engines 3631(1)-3631(N). In at least one embodiment, the operating system can virtualize graphics processing engines 3631(1)-3631(N) to provide access to each process or application.
[0335] In at least one embodiment, the graphics acceleration module 3646 or the individual graphics processing engine 3631(1)-3631(N) uses a process handle to select a process element. In at least one embodiment, the process element is stored in system memory 3614 and can be addressed using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering its context with the graphics processing engine 3631(1)-3631(N) (i.e., invoking system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element linked list.
[0336] Figure 36D An exemplary accelerator integration slice 3690 is illustrated. In at least one embodiment, a "slice" includes a designated portion of the processing resources of the accelerator integrated circuit 3636. In at least one embodiment, slicing the accelerator integration slice 3690 implements process 100 or process 106 (see...). Figure 1BIn at least one embodiment, the application is an effective address space 3682 in system memory 3614, which stores process element 3683. In at least one embodiment, process element 3683 is stored in response to a GPU call 3681 from an application 3680 executing on processor 3607. In at least one embodiment, process element 3683 contains the process state of the corresponding application 3680. In at least one embodiment, a job descriptor (WD) 3684 contained in process element 3683 may be a single job requested by the application, or it may contain a pointer to a job queue. In at least one embodiment, WD 3684 is a pointer to a job request queue in the effective address space 3682 of the application.
[0337] In at least one embodiment, the graphics acceleration module 3646 and / or the various graphics processing engines 3631(1)-3631(N) may be shared by all processes or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD 3684 to the graphics acceleration module 3646 to begin operations in a virtualized environment.
[0338] In at least one embodiment, the dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns either the graphics acceleration module 3646 or an individual graphics processing engine 3631. In at least one embodiment, when the graphics acceleration module 3646 is owned by a single process, the hypervisor initializes the accelerator integrated circuit 3636 for the owned partition; when the graphics acceleration module 3646 is assigned, the operating system initializes the accelerator integrated circuit 3636 for the owned process.
[0339] In at least one embodiment, during operation, the WD acquisition unit 3691 in the accelerator integration slice 3690 acquires the next WD 3684, which includes instructions for work to be performed by one or more graphics processing engines of the graphics acceleration module 3646. In at least one embodiment, data from the WD 3684 may be stored in register 3645 and used by the MMU 3639, interrupt management circuitry 3647, and / or context management circuitry 3648, as shown. For example, one embodiment of the MMU 3639 includes segment / page roaming circuitry for accessing segment / page tables 3686 within the OS virtual address space 3685. In at least one embodiment, the interrupt management circuitry 3647 may process an interrupt event 3692 received from the graphics acceleration module 3646. In at least one embodiment, when performing graphics operations, a valid address 3693 generated by the graphics processing engines 3631(1)-3631(N) is translated into a real address by the MMU 3639.
[0340] In one embodiment, register 3645 is copied for each graphics processing engine 3631(1)-3631(N) and / or graphics acceleration module 3646, and register 3645 may be initialized by a hypervisor or operating system. In at least one embodiment, each of these copied registers may be included in an accelerator integration slice 3690. Exemplary registers that may be initialized by a hypervisor are shown in Table 1.
[0341]
[0342]
[0343] Table 2 shows exemplary registers that can be initialized by the operating system.
[0344]
[0345] In at least one embodiment, each WD 3684 is specific to a particular graphics acceleration module 3646 and / or graphics processing engine 3631(1)-3631(N). In at least one embodiment, it contains all the information required for the graphics processing engine 3631(1)-3631(N) to complete its work, or it may be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0346] Figure 36E Additional details of an exemplary embodiment of the shared model are shown. This embodiment includes a hypervisor real address space 3698, in which a list of process elements 3699 is stored. In at least one embodiment, the hypervisor real address space 3698 can be accessed via a hypervisor 3696, which virtualizes the graphics acceleration module engine for an operating system 3695.
[0347] In at least one embodiment, the shared programming model allows all processes or subsets of processes from all partitions or subsets of partitions in the system to use the graphics acceleration module 3646. In at least one embodiment, there are two programming models in which the graphics acceleration module 3646 is shared by multiple processes and partitions, namely, time-slice sharing and graphics-oriented sharing.
[0348] In at least one embodiment, in this model, the hypervisor 3696 owns the graphics acceleration module 3646 and makes its functionality available to all operating systems 3695. In at least one embodiment, for the graphics acceleration module 3646 to support virtualization through the hypervisor 3696, the graphics acceleration module 3646 may comply with certain requirements, such as (1) the job requests of the application must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 3646 must provide a context saving and recovery mechanism, (2) the graphics acceleration module 3646 guarantees that the job requests of the application are completed within a specified amount of time, including any conversion errors, or the graphics acceleration module 3646 provides the ability to preempt job processing, and (3) when operating in a directed shared programming model, fairness between the processes of the graphics acceleration module 3646 must be ensured.
[0349] In at least one embodiment, application 3680 needs to make system calls to operating system 3695 using the graphics acceleration module type, working descriptor (WD), authority mask register (AMR) value, and context save / restore region pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 3646 and can take the form of graphics acceleration module 3646 commands, valid address pointers to user-defined structures, valid address pointers to command queues, or any other data structure describing the work to be performed by graphics acceleration module 3646.
[0350] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to that of the application that sets the AMR. In at least one embodiment, if the implementation of the accelerator integrated circuit 3636 (not shown) and the graphics acceleration module 3646 does not support the User Rights Mask Overwrite Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 3696 may selectively apply the current Rights Mask Overwrite Register (AMOR) value before placing the AMR into the process element 3683. In at least one embodiment, CSRP is one of the registers 3645 that contains the effective address of a region in the effective address space 3682 of the application for the graphics acceleration module 3646 to save and restore the context state. In at least one embodiment, this pointer is optional if it is not necessary to save the state between jobs or when a job is preempted. In at least one embodiment, the context save / restore region may be fixed system memory.
[0351] Upon receiving a system call, the operating system 3695 can verify that the application 3680 has been registered and granted permission to use the graphics acceleration module 3646. Then, in at least one embodiment, the operating system 3695 uses the information shown in Table 3 to invoke the hypervisor 3696.
[0352]
[0353]
[0354] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 3696 verifies that the operating system 3695 has been registered and granted permission to use the graphics acceleration module 3646. Then, in at least one embodiment, the hypervisor 3696 adds the process element 3683 to a linked list of process elements of the corresponding graphics acceleration module 3646 type. In at least one embodiment, the process element may include the information shown in Table 4.
[0355]
[0356] In at least one embodiment, the hypervisor initializes multiple accelerator integration slice 3690 registers 3645.
[0357] like Figure 36F As shown, in at least one embodiment, a unified memory is used, which is addressable via a common virtual memory address space for accessing physical processor memories 3601(1)-3601(N) and GPU memories 3620(1)-3620(N). In this implementation, operations performed on GPUs 3610(1)-3610(N) utilize the same virtual / effective memory address space to access processor memories 3601(1)-3601(M) and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 3601(1), a second portion to second processor memory 3601(N), a third portion to GPU memory 3620(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memory 3601 and GPU memory 3620, thereby allowing any processor or GPU to access that memory using a virtual address mapped to any physical memory.
[0358] In at least one embodiment, the bias / coherence management circuitry 3694A-3694E within one or more MMUs 3639A-3639E ensures cache coherence between the caches of one or more host processors (e.g., 3605) and the GPU 3610, and implements biasing techniques to indicate the physical memory in which certain types of data should be stored. In at least one embodiment, although in Figure 36F Several instances of the bias / coherence management circuitry 3694A-3694E are shown, but the bias / coherence circuitry can be implemented within the MMU of one or more host processors 3605 and / or within the accelerator integrated circuit 3636.
[0359] One embodiment allows GPU memory 3620 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology without suffering the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability to access GPU memory 3620 as system memory without the heavy overhead of cache coherence provides a favorable operating environment for GPU offloading. In at least one embodiment, this arrangement allows the host processor 3605 to software-set operands and access computation results without the overhead of conventional I / ODMA data copying. In at least one embodiment, such conventional copying includes driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 3620 without cache coherence overhead can be critical to the execution time of offloaded computations. In at least one embodiment, for example, in cases with high streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPU 3610. In at least one embodiment, the efficiency of operand setting, the efficiency of result access, and the efficiency of GPU computation can play a role in determining the effectiveness of GPU offloading.
[0360] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which may be a page-granular structure (e.g., controlled at the memory page level) comprising 1 or 2 bits of memory pages attached to each GPU. In at least one embodiment, with or without a bias cache (e.g., for caching frequently / recently used entries in the bias table) in GPU 3610, the bias table can be implemented across one or more stolen memory ranges of GPU memory 3620. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.
[0361] In at least one embodiment, prior to actual access to GPU memory, an access to the bias table entry associated with each access to GPU-attached memory 3620 is performed, resulting in the following operations: In at least one embodiment, a local request from GPU 3610 to find its page in the GPU bias is directly forwarded to the corresponding GPU memory 3620. In at least one embodiment, a local request from GPU to find its page in the host bias is forwarded to processor 3605 (e.g., via the high-speed link described herein). In at least one embodiment, a request from processor 3605 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request to a GPU bias page can be forwarded to GPU 3610. In at least one embodiment, if the GPU is not currently using the page, the GPU may subsequently migrate the page to the host processor bias. In at least one embodiment, the page bias state can be changed through a software-based mechanism, a hardware-assisted software mechanism, or, in limited cases, a purely hardware-based mechanism.
[0362] In at least one embodiment, a mechanism for changing the bias state employs an API call (e.g., OpenCL), which subsequently invokes the GPU's device driver. The device driver then sends a message (or enqueues a command descriptor) to the GPU, instructing the GPU to change the bias state and, in some migration, performs a cache refresh operation on the host. In at least one embodiment, the cache refresh operation is used for migration from the host processor 3605 bias to the GPU bias, but not for the reverse migration.
[0363] In at least one embodiment, cache coherence is maintained by temporarily rendering GPU bias pages that the host processor 3605 cannot cache. In at least one embodiment, to access these pages, the processor 3605 may request access from the GPU 3610, which may or may not immediately grant access. Therefore, in at least one embodiment, to reduce communication between the processor 3605 and the GPU 3610, it is beneficial to ensure that the GPU bias pages are pages required by the GPU rather than those required by the host processor 3605, and vice versa.
[0364] Figure 37 Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are illustrated, which may be manufactured using one or more IP cores. In addition to the illustrations, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0365] Figure 37 This is a block diagram illustrating an exemplary system on a chip integrated circuit 3700 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 3700 includes one or more application processors 3705 (e.g., CPUs), at least one graphics processor 3710, and may additionally include an image processor 3715 and / or a video processor 3720, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 3700 includes peripheral or bus logic, which includes a USB controller 3725, a UART controller 3730, an SPI / SDIO controller 3735, and an I... 2 2S / I 2 2C controller 3740. In at least one embodiment, integrated circuit 3700 may include display device 3745 coupled to one or more of High Definition Multimedia Interface (HDMI) controller 3750 and Mobile Industrial Processor Interface (MIPI) display interface 3755. In at least one embodiment, storage may be provided by flash memory subsystem 3760, including flash memory and flash memory controller. In at least one embodiment, a memory interface may be provided via memory controller 3765 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include embedded security engine 3770.
[0366] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, the inference and / or training logic 2815 may be used in integrated circuit 3700 to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0367] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28BOne or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0368] Figure 38A and 38B Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are illustrated, which may be manufactured using one or more IP cores. In addition to the illustrations, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0369] Figure 38A and 38B This is a block diagram illustrating an exemplary graphics processor used within a SoC according to embodiments described herein. Figure 38A An exemplary graphics processor 3810 of a system-on-a-chip according to at least one embodiment is shown, which can be manufactured using one or more IP cores. Figure 38B Further exemplary graphics processor 3840 of a system-on-a-chip according to at least one embodiment is shown, which can be manufactured using one or more IP cores. In at least one embodiment, Figure 38A The graphics processor 3810 is a low-power graphics processor core. In at least one embodiment, Figure 38B The graphics processor 3840 is a higher-performance graphics processor core. In at least one embodiment, each graphics processor 3810, 3840 may be... Figure 37 A variant of the graphics processor 3710. In at least one embodiment, the graphics processor 3810 implements process 100 and / or process 106 (see...). Figure 1B ).
[0370] In at least one embodiment, the graphics processor 3810 includes a vertex processor 3805 and one or more fragment processors 3815A-3815N (e.g., 3815A, 3815B, 3815C, 3815D to 3815N-1 and 3815N). In at least one embodiment, the graphics processor 3810 can execute different shader programs via separate logic, such that the vertex processor 3805 is optimized to perform operations for the vertex shader program, while one or more fragment processors 3815A-3815N perform fragment (e.g., pixel) shading operations for fragments or pixels or shader programs. In at least one embodiment, the vertex processor 3805 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, one or more fragment processors 3815A-3815N use the primitive and vertex data generated by the vertex processor 3805 to generate a framebuffer for display on a display device. In at least one embodiment, one or more fragment processors 3815A-3815N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs provided in the Direct 3D API.
[0371] In at least one embodiment, the graphics processor 3810 additionally includes one or more memory management units (MMUs) 3820A-3820B, one or more caches 3825A-3825B, and one or more circuit interconnects 3830A-3830B. In at least one embodiment, one or more MMUs 3820A-3820B provide virtual-to-physical address mappings for the graphics processor 3810, including for vertex processors 3805 and / or fragment processors 3815A-3815N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 3825A-3825B. In at least one embodiment, one or more MMUs 3820A-3820B can be synchronized with other MMUs within the system, including with... Figure 37 One or more application processors 3705, graphics processors 3715, and / or video processors 3720 are associated with one or more MMUs, enabling each processor 3705-3720 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 3830A-3830B enable the graphics processor 3810 to connect to other IP cores within the SoC via the SoC's internal bus or via a direct connection.
[0372] In at least one embodiment, the graphics processor 3840 includes one or more shader cores 3855A-3855N (e.g., 3855A, 3855B, 3855C, 3855D, 3855E, 3855F to 3855N-1 and 3855N), such as Figure 38B As shown, it provides a unified shader core architecture, where a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 3840 includes an inter-core task manager 3845, which acts as a thread dispatcher to assign execution threads to one or more shader cores 3855A-3855N and a tile unit 3858 to accelerate tile-based rendering operations, where scene rendering operations are subdivided in image space, for example, to utilize local spatial consistency within the scene or optimize the use of internal caches.
[0373] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, the inference and / or training logic 2815 may be integrated into an integrated circuit. Figure 38A and / or Figure 38B The above is used for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases described herein.
[0374] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0375] Figure 39A and 39B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. In at least one embodiment, Figure 39A It shows that it can be included in Figure 37 The graphics core 3900 within the graphics processor 3710, and in at least one embodiment, may be as follows: Figure 38B The unified shader cores shown are 3855A-3855N. Figure 39B A highly parallel general-purpose graphics processing unit (“GPGPU”) 3930 suitable for deployment on a multi-chip module is illustrated in at least one embodiment. In at least one embodiment, the graphics core 3900 implements process 100 ( Figure 1A ) or process 106 (see Figure 1B ).
[0376] In at least one embodiment, the graphics core 3900 includes a shared instruction cache 3902, texture units 3918, and cache / shared memory 3920, which are common to the execution resources within the graphics core 3900. In at least one embodiment, the graphics core 3900 may include multiple slices 3901A-3901N or partitions of each core, and the graphics processor may include multiple instances of the graphics core 3900. In at least one embodiment, slices 3901A-3901N may include supporting logic including local instruction caches 3904A-3904N, thread schedulers 3906A-3906N, thread dispatchers 3908A-3908N, and a set of registers 3910A-3910N. In at least one embodiment, slices 3901A-3901N may include a set of additional functional units (AFU 3912A-3912N), floating-point units (FPU 3914A-3914N), integer arithmetic logic units (ALU 3916A-3916N), address calculation units (ACU 3913A-3913N), double-precision floating-point units (DPFPU 3915A-3915N), and matrix processing units (MPU 3917A-3917N).
[0377] In at least one embodiment, the FPU 3914A-3914N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 3915A-3915N performs double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 3916A-3916N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 3917A-3917N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 3917-3917N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated generalized matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFU 3912A-3912N can perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0378] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This is combined with... Figure 28A and / or Figure 28B Details regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in graphics core 3900 for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0379] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0380] Figure 39BA general-purpose processing unit (GPGPU) 3930 is illustrated in at least one embodiment, which can be configured to enable highly parallel computational operations to be performed by a set of graphics processing units. In at least one embodiment, the GPGPU 3930 can be directly linked to other instances of the GPGPU 3930 to create a multi-GPU cluster to improve the training speed for deep neural networks. In at least one embodiment, the GPGPU 3930 includes a host interface 3932 for connection to a host processor. In at least one embodiment, the host interface 3932 is a PCI Express interface. In at least one embodiment, the host interface 3932 can be a vendor-specific communication interface or communication structure. In at least one embodiment, the GPGPU 3930 receives commands from the host processor and uses a global scheduler 3934 to allocate execution threads associated with those commands to a set of compute clusters 3936A-3936H. In at least one embodiment, compute clusters 3936A-3936H share a cache memory 3938. In at least one embodiment, cache memory 3938 can be used as a higher-level cache within the cache memory of computing clusters 3936A-3936H.
[0381] In at least one embodiment, the GPGPU 3930 includes memories 3944A-3944B, which are coupled to computing clusters 3936A-3936H via a set of memory controllers 3942A-3942B. In at least one embodiment, memories 3944A-3944B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), which includes graphics double data rate (GDDR) memory.
[0382] In at least one embodiment, each of the computing clusters 3936A-3936H includes a set of graphics cores, for example... Figure 39A The graphics core 3900 may include various types of integer and floating-point logic units that can perform computational operations across a wide range of precisions, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each computing cluster 3936A-3936H may be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units may be configured to perform 64-bit floating-point operations.
[0383] In at least one embodiment, multiple instances of the GPGPU 3930 can be configured as a computing cluster. In at least one embodiment, the communication used for synchronization and data exchange by the computing clusters 3936A-3936H varies between embodiments. In at least one embodiment, the multiple instances of the GPGPU 3930 communicate via a host interface 3932. In at least one embodiment, the GPGPU 3930 includes an I / O hub 3939 that couples the GPGPU 3930 to a GPU link 3940, enabling direct connection to other instances of the GPGPU 3930. In at least one embodiment, the GPU link 3940 is coupled to a dedicated GPU-to-GPU bridge, which enables communication and synchronization between the multiple instances of the GPGPU 3930. In at least one embodiment, the GPU link 3940 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, the multiple instances of the GPGPU 3930 reside in a separate data processing system and communicate via network devices accessible through the host interface 3932. In at least one embodiment, the GPU link 3940 may be configured to connect to a host processor other than or as a replacement for the host interface 3932.
[0384] In at least one embodiment, the GPGPU 3930 can be configured to train a neural network. In at least one embodiment, the GPGPU 3930 can be used within an inference platform. In at least one embodiment, when using the GPGPU 3930 for inference, the GPGPU 3930 may include fewer compute clusters 3936A-3936H compared to when using the GPGPU 3930 to train a neural network. In at least one embodiment, the memory technology associated with the memories 3944A-3944B can differ between inference and training configurations, with higher bandwidth memory technology dedicated to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 3930 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during the inference operation of the deployed neural network.
[0385] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28BDetails regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in GPGPU 3930 for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0386] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0387] Figure 40 A block diagram of a computer system 4000 according to at least one embodiment is shown. In at least one embodiment, the computer system 4000 includes a processing subsystem 4001 having one or more processors 4002 and a system memory 4004 communicating via an interconnect path that may include a memory hub 4005. In at least one embodiment, the memory hub 4005 may be a separate component within a chipset component or may be integrated within one or more processors 4002. In at least one embodiment, the memory hub 4005 is coupled to an I / O subsystem 4011 via a communication link 4006. In one embodiment, the I / O subsystem 4011 includes an I / O hub 4007 that enables the computer system 4000 to receive input from one or more input devices 4008. In at least one embodiment, the I / O hub 4007 enables a display controller to provide output to one or more display devices 4010A, the display controller being included in one or more processors 4002. In at least one embodiment, one or more display devices 4010A coupled to the I / O hub 4007 may include local, internal, or embedded display devices. In at least one embodiment, the computing system 4000 implements process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0388] In at least one embodiment, the processing subsystem 4001 includes one or more parallel processors 4012 coupled to a memory hub 4005 via a bus or other communication link 4013. In at least one embodiment, the communication link 4013 may use any of many standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or may be a vendor-specific communication interface or communication architecture. In at least one embodiment, one or more parallel processors 4012 form a compute-intensive parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a multi-core integrated (MIC) processor. In at least one embodiment, one or more parallel processors 4012 form a graphics processing subsystem that can output pixels to one of one or more display devices 4010A coupled via an I / O hub 4007. In at least one embodiment, the parallel processor 4012 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 4010B.
[0389] In at least one embodiment, system storage unit 4014 may be connected to I / O hub 4007 to provide a storage mechanism for computer system 4000. In at least one embodiment, I / O switch 4016 may be used to provide an interface mechanism to enable connectivity between I / O hub 4007 and other components, such as network adapter 4018 and / or wireless network adapter 4019 that may be integrated into the platform, and various other devices that may be added via one or more additional devices 4020. In at least one embodiment, network adapter 4018 may be an Ethernet adapter or another wired network adapter. In at least one embodiment, wireless network adapter 4019 may include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless devices.
[0390] In at least one embodiment, the computer system 4000 may include other components not explicitly shown, such as USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 4007. In at least one embodiment, the interconnection can be implemented using any suitable protocol (e.g., PCI-based protocols such as PCI-Express or other bus or point-to-point communication interfaces and / or protocols). Figure 40 The communication paths of the various components, such as NV-Link high-speed interconnect or interconnect protocols.
[0391] In at least one embodiment, one or more parallel processors 4012 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constituting a graphics processing unit (GPU). In at least one embodiment, the parallel processor 4012 includes circuitry optimized for general-purpose processing. In at least one embodiment, components of the computer system 4000 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, the parallel processor 4012, memory hub 4005, processor 4002, and I / O hub 4007 may be integrated into a system-on-a-chip (SoC) integrated circuit. In at least one embodiment, components of the computer system 4000 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computer system 4000 may be integrated into a multi-chip module (MCM) that can interconnect with other MCMs to a modular computer system.
[0392] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details are provided regarding the inference and / or training logic 2815. In at least one embodiment, the inference and / or training logic 2815 can... Figure 40 The system 4000 is used for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0393] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0394] processor
[0395] Figure 41AA parallel processor 4100 according to at least one embodiment is illustrated. In at least one embodiment, various components of the parallel processor 4100 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 4100 is according to an exemplary embodiment. Figure 40 Variations of one or more parallel processors 4012 are shown. In at least one embodiment, the parallel processor 4100 implements process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0396] In at least one embodiment, the parallel processor 4100 includes a parallel processing unit 4102. In at least one embodiment, the parallel processing unit 4102 includes an I / O unit 4104 that enables communication with other devices, including other instances of the parallel processing unit 4102. In at least one embodiment, the I / O unit 4104 can be directly connected to other devices. In at least one embodiment, the I / O unit 4104 is connected to other devices using a hub or switch interface (e.g., a memory hub 4105). In at least one embodiment, the connection between the memory hub 4105 and the I / O unit 4104 forms a communication link 4113. In at least one embodiment, the I / O unit 4104 is connected to a host interface 4106 and a memory crossbar switch 4116, wherein the host interface 4106 receives commands for performing processing operations, and the memory crossbar switch 4116 receives commands for performing memory operations.
[0397] In at least one embodiment, when host interface 4106 receives a command buffer via I / O unit 4104, host interface 4106 can direct work operations to execute those commands to front end 4108. In at least one embodiment, front end 4108 is coupled to scheduler 4110, which is configured to assign commands or other work items to processing cluster array 4112. In at least one embodiment, scheduler 4110 ensures that processing cluster array 4112 is correctly configured and in an active state before assigning tasks to processing cluster array 4112. In at least one embodiment, scheduler 4110 is implemented via firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 4110 can be configured to perform complex scheduling and work assignment operations at both coarse and fine granular levels, enabling fast preemption and context switching of threads executing on processing array 4112. In at least one embodiment, host software can demonstrate workloads for scheduling on processing array 4112 via one of multiple graphics processing paths. In at least one embodiment, the workload can then be automatically distributed on the processing array 4112 by the scheduler 4110 logic within the microcontroller, which includes the scheduler 4110.
[0398] In at least one embodiment, the processing cluster array 4112 may include up to "N" processing clusters (e.g., clusters 4114A, 4114B to 4114N), where "N" represents a positive integer (which may be an integer different from the integer "N" used in other diagrams). In at least one embodiment, each cluster 4114A-4114N of the processing cluster array 4112 can execute a large number of concurrent threads. In at least one embodiment, the scheduler 4110 may use various scheduling and / or work allocation algorithms to allocate work to the clusters 4114A-4114N of the processing cluster array 4112, which may vary depending on the workload generated by each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 4110, or may be partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 4112. In at least one embodiment, the different clusters 4114A-4114N of the processing cluster array 4112 may be assigned to process different types of programs or to perform different types of computations.
[0399] In at least one embodiment, the processing cluster array 4112 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 4112 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 4112 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.
[0400] In at least one embodiment, the processing cluster array 4112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 4112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing cluster array 4112 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 4102 may transfer data from system memory via I / O unit 4104 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 4122) and then written back to system memory.
[0401] In at least one embodiment, when the parallel processing unit 4102 is used to perform graphics processing, the scheduler 4110 may be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations among the multiple clusters 4114A-4114N of the processing cluster array 4112. In at least one embodiment, portions of the processing cluster array 4112 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 4114A-4114N may be stored in a buffer to allow intermediate data to be transferred between the clusters 4114A-4114N for further processing.
[0402] In at least one embodiment, the processing cluster array 4112 may receive processing tasks to be executed via a scheduler 4110, which receives commands defining the processing tasks from a front end 4108. In at least one embodiment, the processing task may include an index of data to be processed, such as surface (patch) data, raw data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is processed (e.g., what program to execute). In at least one embodiment, the scheduler 4110 may be configured to acquire an index corresponding to a task, or may receive an index from the front end 4108. In at least one embodiment, the front end 4108 may be configured to ensure that the processing cluster array 4112 is configured to be active before initiating the workload specified by an incoming command buffer (e.g., a batch buffer, push buffer, etc.).
[0403] In at least one embodiment, each of one or more instances of the parallel processing unit 4102 may be coupled to the parallel processor memory 4122. In at least one embodiment, the parallel processor memory 4122 may be accessed via a memory crossbar switch 4116, which may receive memory requests from the processing cluster array 4112 and the I / O unit 4104. In at least one embodiment, the memory crossbar switch 4116 may be accessed via a memory interface 4118. In at least one embodiment, the memory interface 4118 may include a plurality of partition units (e.g., partition units 4120A, 4120B to 4120N), each of which may be coupled to a portion (e.g., a memory cell) of the parallel processor memory 4122. In at least one embodiment, the plurality of partition units 4120A-4120N are configured to be equal to the number of memory units, such that the first partition unit 4120A has a corresponding first memory unit 4124A, the second partition unit 4120B has a corresponding memory unit 4124B, and the Nth partition unit 4120N has a corresponding Nth memory unit 4124N. In at least one embodiment, the number of partition units 4120A-4120N may not be equal to the number of memory units.
[0404] In at least one embodiment, memory cells 4124A-4124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory cells 4124A-4124N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory cells 4124A-4124N, allowing partitioning cells 4120A-4120N to write portions of each rendering target in parallel, to efficiently utilize the available bandwidth of the parallel processor memory 4122. In at least one embodiment, local instances of the parallel processor memory 4122 may be excluded to facilitate a unified memory design that combines system memory with local cache memory.
[0405] In at least one embodiment, any of the clusters 4114A-4114N of the processing cluster array 4112 can process data to be written to any memory cell 4124A-4124N within the parallel processor memory 4122. In at least one embodiment, the memory crossbar switch 4116 can be configured to transfer the output of each cluster 4114A-4114N to any partition cell 4120A-4120N or another cluster 4114A-4114N, and the clusters 4114A-4114N can perform further processing operations on the output. In at least one embodiment, each cluster 4114A-4114N can communicate with the memory interface 4118 via the memory crossbar switch 4116 to read from or write to various external storage devices. In at least one embodiment, the memory crossbar switch 4116 has a connection to a memory interface 4118 for communication with I / O unit 4104, and a connection to a local instance of parallel processor memory 4122, thereby enabling processing units within different processing clusters 4114A-4114N to communicate with system memory or other memory not local to parallel processing unit 4102. In at least one embodiment, the memory crossbar switch 4116 may use virtual channels to separate traffic flows between clusters 4114A-4114N and partition units 4120A-4120N.
[0406] In at least one embodiment, multiple instances of the parallel processing unit 4102 may be provided on a single insert card, or multiple insert cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 4102 may be configured to interoperate, even if the different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 4102 may include higher-precision floating-point units relative to other instances. In at least one embodiment, a system combining one or more instances of the parallel processing unit 4102 or the parallel processor 4100 can be implemented in various configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0407] Figure 41B This is a block diagram of a partitioning unit 4120 according to at least one embodiment. In at least one embodiment, the partitioning unit 4120 is... Figure 41A This is an example of one of the partitioning units 4120A-4120N. In at least one embodiment, the partitioning unit 4120 includes an L2 cache 4121, a frame buffer interface 4125, and a ROP 4126 (raster operation unit). In at least one embodiment, the L2 cache 4121 is a read / write cache configured to perform load and store operations received from the memory crossbar switch 4116 and the ROP 4126. In at least one embodiment, the L2 cache 4121 outputs read misses and urgent write-back requests to the frame buffer interface 4125 for processing. In at least one embodiment, updates can also be sent to the frame buffer for processing via the frame buffer interface 4125. In at least one embodiment, the frame buffer interface 4125 communicates with memory cells in the parallel processor memory (such as...). Figure 41A The memory cells 4124A-4124N (e.g., within the parallel processor memory 4122) interact with one of them.
[0408] In at least one embodiment, ROP 4126 is a processing unit that performs raster operations such as stenciling, z-testing, blending, etc. In at least one embodiment, ROP 4126 then outputs processed graphics data stored in graphics memory. In at least one embodiment, ROP 4126 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic utilizing one or more of a variety of compression algorithms. In at least one embodiment, the type of compression performed by ROP 4126 may vary based on the statistical characteristics of the data to be compressed. For example, in at least one embodiment, incremental color compression is performed based on depth and color data on a per-tile basis.
[0409] In at least one embodiment, ROP 4126 is included within each processing cluster (e.g., Figure 41A Clusters 4114A-4114N are used instead of partition units 4120. In at least one embodiment, read and write requests for pixel data are made via memory crossbar switch 4116 instead of pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device (such as...). Figure 40 Displayed by one or more display devices 4010, routed by processor 4002 for further processing, or by... Figure 41A One of the processing entities within the parallel processor 4100 is routed for further processing.
[0410] Figure 41C This is a block diagram of a processing cluster 4114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is... Figure 41A An instance of one of the processing clusters 4114A-4114N. In at least one embodiment, the processing cluster 4114 can be configured to execute a number of threads in parallel, where a "thread" refers to an instance of a specific program executing on a particular set of input data. In at least one embodiment, a Single Instruction Multiple Data (SIMD) instruction issuing technique is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, a Single Instruction Multiple Threading (SIMT) technique is used to support the parallel execution of a large number of generally synchronous threads, which uses a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0411] In at least one embodiment, the operation of the processing cluster 4114 can be controlled by a pipeline manager 4132 that assigns processing tasks to the SIMT parallel processors. In at least one embodiment, the pipeline manager 4132... Figure 41AThe scheduler 4110 receives instructions and manages the execution of these instructions via the graphics multiprocessor 4134 and / or texture unit 4136. In at least one embodiment, the graphics multiprocessor 4134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, the processing cluster 4114 may include various types of SIMT parallel processors with different architectures. In at least one embodiment, the processing cluster 4114 may include one or more instances of the graphics multiprocessor 4134. In at least one embodiment, the graphics multiprocessor 4134 can process data, and the data crossover switch 4140 can be used to distribute the processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, the pipeline manager 4132 can facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossover switch 4140.
[0412] In at least one embodiment, each graphics multiprocessor 4134 within the processing cluster 4114 may include the same set of functional execution logic (e.g., arithmetic logic units, load-memory units, etc.). In at least one embodiment, the functional execution logic may be configured in a pipelined manner, wherein new instructions may be issued before previous instructions complete. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shift operations, and computation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be used to perform different operations, and any combination of functional units may exist.
[0413] In at least one embodiment, instructions sent to the processing cluster 4114 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a general program on different input data. In at least one embodiment, each thread within the thread group may be assigned to a different processing engine within the graphics multiprocessor 4134. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 4134. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during a loop that is processing the thread group. In at least one embodiment, the thread group may also include more threads than the number of processing engines within the graphics multiprocessor 4134. In at least one embodiment, when the thread group includes more threads than the number of processing engines within the graphics multiprocessor 4134, processing can be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on the graphics multiprocessor 4134.
[0414] In at least one embodiment, the graphics multiprocessor 4134 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 4134 may forgo the internal cache and use a cache memory within the processing cluster 4114 (e.g., L1 cache 4148). In at least one embodiment, each graphics multiprocessor 4134 may also access partition units (e.g., Figure 41A The L2 cache is located within partition units 4120A-4120N, which are shared across all processing clusters 4114 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 4134 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of the parallel processing unit 4102 can be used as global memory. In at least one embodiment, the processing cluster 4114 includes multiple instances of the graphics multiprocessor 4134, which can share common instructions and data that can be stored in the L1 cache 4148.
[0415] In at least one embodiment, each processing cluster 4114 may include a memory management unit (“MMU”) 4145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 4145 may reside in Figure 41A The memory interface 4118 is located within the MMU. In at least one embodiment, the MMU 4145 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles and optionally to cache line indices. In at least one embodiment, the MMU 4145 may include an address translation back buffer (TLB) or a cache that may reside within the graphics multiprocessor 4134, the L1 cache 4148, or the processing cluster 4114. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, cache line indices may be used to determine whether a request for a cache line is a hit or a miss.
[0416] In at least one embodiment, the processing cluster 4114 may be configured such that each graphics multiprocessor 4134 is coupled to a texture unit 4136 to perform texture mapping operations that determine texture sample locations, read texture data, and filter texture data. In at least one embodiment, texture data is read as needed from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 4134, and texture data is also retrieved from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 4134 outputs a processed task to a data crossbar switch 4140 to provide the processed task to another processing cluster 4114 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via a memory crossbar switch 4116. In at least one embodiment, a preROP 4142 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 4134 and direct the data to a ROP unit, which may be associated with a partitioning unit (e.g., [missing information]). Figure 41A The PreROP 4142 unit is located together with the partition units 4120A-4120N. In at least one embodiment, the PreROP 4142 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.
[0417] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in a graphics processing cluster 4114 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0418] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0419] Figure 41D A graphics multiprocessor 4134 according to at least one embodiment is illustrated. In at least one embodiment, the graphics multiprocessor 4134 is coupled to a pipeline manager 4132 of a processing cluster 4114. In at least one embodiment, the graphics multiprocessor 4134 has an execution pipeline including, but not limited to, an instruction cache 4152, an instruction unit 4154, an address mapping unit 4156, a register file 4158, one or more general-purpose graphics processing unit (GPGPU) cores 4162, and one or more load / store units 4166. In at least one embodiment, the GPGPU cores 4162 and the load / store units 4166 are coupled to a cache memory 4172 and a shared memory 4170 via a memory and cache interconnect 4168.
[0420] In at least one embodiment, instruction cache 4152 receives a stream of instructions to be executed from pipeline manager 4132. In at least one embodiment, instructions are cached in instruction cache 4152 and dispatched to instruction unit 4154 for execution. In one embodiment, instruction unit 4154 may dispatch instructions as thread groups (e.g., thread bundles), assigning each thread of the thread group to a different execution unit within GPGPU core 4162. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, address mapping unit 4156 may be used to translate addresses in the unified address space into different memory addresses that can be accessed by load / store unit 4166.
[0421] In at least one embodiment, register file 4158 provides a set of registers for functional units of graphics multiprocessor 4134. In at least one embodiment, register file 4158 provides temporary storage for operands of data paths connected to functional units of graphics multiprocessor 4134 (e.g., GPGPU core 4162, load / store unit 4166). In at least one embodiment, register file 4158 is partitioned among each functional unit, such that a dedicated portion of register file 4158 is allocated to each functional unit. In at least one embodiment, register file 4158 is partitioned among different thread bundles being executed by graphics multiprocessor 4134.
[0422] In at least one embodiment, each of the GPGPU cores 4162 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 4134. In at least one embodiment, the GPGPU cores 4162 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 4162 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point algorithms or enable variable-precision floating-point algorithms. In at least one embodiment, the graphics multiprocessor 4134 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores 4162 may also include fixed-function or special-function logic.
[0423] In at least one embodiment, the GPGPU core 4162 includes SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU core 4162 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time or automatically generated when executing a program written and compiled for a Single Program Multiple Data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed using a single SIMD instruction. For example, in at least one embodiment, eight SIMD threads performing the same or similar operations can be executed in parallel using a single SIMD8 logic unit.
[0424] In at least one embodiment, the memory and cache interconnect 4168 is an interconnect network connecting each functional unit of the graphics multiprocessor 4134 to the register file 4158 and the shared memory 4170. In at least one embodiment, the memory and cache interconnect 4168 is a cross-switch interconnect that allows the load / store unit 4166 to perform load and store operations between the shared memory 4170 and the register file 4158. In at least one embodiment, the register file 4158 can operate at the same frequency as the GPGPU core 4162, resulting in very low latency for data transfer between the GPGPU core 4162 and the register file 4158. In at least one embodiment, the shared memory 4170 can be used to enable communication between threads executing on functional units within the graphics multiprocessor 4134. In at least one embodiment, the cache memory 4172 can be used, for example, as a data cache to cache texture data communicated between functional units and texture units 4136. In at least one embodiment, the shared memory 4170 can also be used as a program-managed cache. In at least one embodiment, in addition to the automatically cached data stored in cache memory 4172, the thread executing on GPGPU core 4162 can also programmatically store data in shared memory.
[0425] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., high-speed interconnects such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated with the core on a package or chip and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). In at least one embodiment, regardless of how the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a job descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0426] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 28A and / or Figure 28BDetails regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in a graphics multiprocessor 4134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0427] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0428] Figure 42A multi-GPU computing system 4200 according to at least one embodiment is illustrated. In at least one embodiment, the multi-GPU computing system 4200 may include a processor 4202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 4206A-D via a host interface switch 4204. In at least one embodiment, the host interface switch 4204 is a PCI Express switch device that couples the processor 4202 to a PCI Express bus, through which the processor 4202 can communicate with the GPGPUs 4206A-D. In at least one embodiment, the GPGPUs 4206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 4216. In at least one embodiment, the GPU-to-GPU links 4216 are connected to each of the GPGPUs 4206A-D via dedicated GPU links. In at least one embodiment, the P2P GPU links 4216 enable direct communication between each GPGPU 4206A-D without communication via the host interface switch 4204 to which the processor 4202 is connected. In at least one embodiment, when GPU-to-GPU traffic is directed to the P2P GPU link 4216, the host interface switch 4204 remains available for system memory access or, for example, communication with other instances of the multi-GPU computing system 4200 via one or more network devices. While in at least one embodiment, the GPGPUs 4206A-D are connected to the processor 4202 via the host interface switch 4204, in at least one embodiment, the processor 4202 includes direct support for the P2P GPU link 4216 and can be directly connected to the GPGPUs 4206A-D. In at least one embodiment, the multi-GPU computing system 4200 executes process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0429] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in a multi-GPU computing system 4200 for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0430] In at least one embodiment, Figure 28A and 28BOne or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0431] Figure 43 This is a block diagram of a graphics processor 4300 according to at least one embodiment. In at least one embodiment, the graphics processor 4300 includes a ring interconnect 4302, a pipeline front end 4304, a media engine 4337, and graphics cores 4380A-4380N. In at least one embodiment, the ring interconnect 4302 couples the graphics processor 4300 to other processing units, said processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 4300 is one of many processors integrated within a multi-core processing system.
[0432] In at least one embodiment, the graphics processor 4300 receives multiple batches of commands via a ring interconnect 4302. In at least one embodiment, the input commands are interpreted by a command streamer 4303 in a pipeline front-end 4304. In at least one embodiment, the graphics processor 4300 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 4380A-4380N. In at least one embodiment, for 3D geometry processing commands, the command streamer 4303 provides the commands to the geometry pipeline 4336. In at least one embodiment, for at least some media processing commands, the command streamer 4303 provides the commands to a video front-end 4334, which is coupled to a media engine 4337. In at least one embodiment, the media engine 4337 includes a video quality engine (VQE) 4330 for video and image post-processing, and a multi-format encoding / decoding (MFX) engine 4333 for providing hardware-accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 4336 and the media engine 4337 each generate an execution thread for thread execution resources provided by at least one graphics core 4380.
[0433] In at least one embodiment, the graphics processor 4300 includes scalable thread execution resources featuring graphics cores 4380A-4380N (which may be modular and sometimes referred to as core slices), each graphics core having multiple sub-cores 4350A-4350N, 4360A-4360N (sometimes referred to as core sub-slices). In at least one embodiment, the graphics processor 4300 may have any number of graphics cores 4380A. In at least one embodiment, the graphics processor 4300 includes graphics cores 4380A having at least a first sub-core 4350A and a second sub-core 4360A. In at least one embodiment, the graphics processor 4300 is a low-power processor with a single sub-core (e.g., 4350A). In at least one embodiment, the graphics processor 4300 includes multiple graphics cores 4380A-4380N, each graphics core including a set of first sub-cores 4350A-4350N and a set of second sub-cores 4360A-4360N. In at least one embodiment, each of the first sub-cores 4350A-4350N includes at least a first set of execution units 4352A-4352N and media / texture samplers 4354A-4354N. In at least one embodiment, each of the second sub-cores 4360A-4360N includes at least a second set of execution units 4362A-4362N and samplers 4364A-4364N. In at least one embodiment, each of the sub-cores 4350A-4350N and 4360A-4360N shares a set of shared resources 4370A-4370N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0434] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding inference and / or training logic 2815 are provided. In at least one embodiment, inference and / or training logic 2815 may be used in graphics processor 4300 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0435] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0436] Figure 44 This is a block diagram illustrating a microarchitecture for a processor 4400 according to at least one embodiment, the processor 4400 including logic circuitry for executing instructions. In at least one embodiment, the processor 4400 can execute instructions, including x86 instructions, ARM instructions, and special-purpose instructions for application-specific integrated circuits (ASICs). In at least one embodiment, the processor 4400 may include registers for storing packaged data, such as the 64-bit wide MMX registers used in Intel Corporation's Santa Clara, California-enabled MMX technology microprocessors. TM Registers. In at least one embodiment, MMX registers available in integer and floating-point forms can operate with packaged data elements accompanying Single Instruction Multiple Data (“SIMD”) and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, a 128-bit wide XMM register associated with SSE2, SSE3, SSE4, AVX, or later (generally referred to as “SSEx”) technologies can hold such packaged data operands. In at least one embodiment, processor 4400 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference. In at least one embodiment, processor 4400 executes process 100 (see...) Figure 1A ) or process 106 (see Figure 1B ).
[0437] In at least one embodiment, processor 4400 includes an ordered front end (“front end”) 4401 to fetch instructions to be executed and prepare instructions for later use in the processor pipeline. In at least one embodiment, front end 4401 may include several units. In at least one embodiment, instruction prefetcher 4426 fetches instructions from memory and provides the instructions to instruction decoder 4428, which in turn decodes or interprets the instructions. For example, in at least one embodiment, instruction decoder 4428 decodes the received instructions into one or more machine-executable so-called “micro-instructions” or “micro-operations” (also referred to as “micro-operations” or “micro-instructions”). In at least one embodiment, instruction decoder 4428 parses the instructions into opcodes and corresponding data and control fields, which can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, trace cache 4430 may assemble the decoded micro-instructions into a program-ordered sequence or trace in micro-instruction queue 4434 for execution. In at least one embodiment, when the trace cache 4430 encounters complex instructions, the microcode ROM 4432 provides the microinstructions required to complete the operation.
[0438] In at least one embodiment, some instructions may be converted into a single micro-operation, while others require several micro-operations to complete the entire operation. In at least one embodiment, if more than four micro-instructions are required to complete an instruction, the instruction decoder 4428 may access the microcode ROM 4432 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-instructions for processing at the instruction decoder 4428. In at least one embodiment, if multiple micro-instructions are required to complete the operation, the instructions may be stored in the microcode ROM 4432. In at least one embodiment, the trace cache 4430 references an entry point programmable logic array (“PLA”) to determine the correct micro-instruction pointer for reading a microcode sequence from the microcode ROM 4432 to complete one or more instructions, according to at least one embodiment. In at least one embodiment, after the microcode ROM 4432 has completed the micro-operation ordering of the instructions, the machine front end 4401 may resume fetching micro-operations from the trace cache 4430.
[0439] In at least one embodiment, the out-of-order execution engine (“out-of-order engine”) 4403 can prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions descend the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 4403 includes, but is not limited to, an allocator / register renamer 4440, a memory microinstruction queue 4442, an integer / floating-point microinstruction queue 4444, a memory scheduler 4446, a fast scheduler 4402, a slow / general-purpose floating-point scheduler (“slow / general-purpose FP scheduler”) 4404, and a simple floating-point scheduler (“simple FP scheduler”) 4406. In at least one embodiment, the fast scheduler 4402, the slow / general-purpose floating-point scheduler 4404, and the simple floating-point scheduler 4406 are also collectively referred to as “microinstruction schedulers 4402, 4404, 4406”. In at least one embodiment, the allocator / register renamer 4440 allocates the machine buffers and resources required for the sequential execution of each microinstruction. In at least one embodiment, the allocator / register renamer 4440 renames logical registers to entries in a register file. In at least one embodiment, the allocator / register renamer 4440 also allocates entries for each microinstruction in one of two microinstruction queues, a memory microinstruction queue 4442 for memory operations and an integer / floating-point microinstruction queue 4444 for non-memory operations, preceding the memory scheduler 4446 and microinstruction schedulers 4402, 4404, and 4406. In at least one embodiment, the microinstruction schedulers 4402, 4404, and 4406 determine when they are ready to execute a microinstruction based on the readiness of their dependent input register operand sources and the availability of the execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 4402 can schedule on each half of the master clock cycle, while the slow / general-purpose floating-point scheduler 4404 and the simple floating-point scheduler 4406 can schedule once per master processor clock cycle. In at least one embodiment, microinstruction schedulers 4402, 4404, and 4406 arbitrate the scheduling port to schedule microinstructions for execution.
[0440] In at least one embodiment, execution block 4411 includes, but is not limited to, integer register file / branch network 4408, floating-point register file / branch network (“FP register file / branch network”) 4410, address generation units (“AGU”) 4412 and 4414, fast arithmetic logic units (“fast ALU”) 4416 and 4418, slow arithmetic logic unit (“slow ALU”) 4420, floating-point ALU (“FP”) 4422, and floating-point move unit (“FP move”) 4424. In at least one embodiment, integer register file / branch network 4408 and floating-point register file / bypass network 4410 are also referred to herein as “register files 4408, 4410”. In at least one embodiment, AGUs 4412 and 4414, fast ALUs 4416 and 4418, slow ALU 4420, floating-point ALU 4422, and floating-point movement unit 4424 are also referred to herein as "execution units 4412, 4414, 4416, 4418, 4420, 4422, and 4424". In at least one embodiment, execution block 4411 may include, but is not limited to, any number (including zero) and type of register files, branch networks, address generation units, and execution units (in any combination).
[0441] In at least one embodiment, register networks 4408, 4410 may be arranged between microinstruction schedulers 4402, 4404, 4406 and execution units 4412, 4414, 4416, 4418, 4420, 4422, and 4424. In at least one embodiment, integer register file / branch network 4408 performs integer operations. In at least one embodiment, floating-point register file / branch network 4410 performs floating-point operations. In at least one embodiment, each of register networks 4408, 4410 may include, but is not limited to, a branch network that can bypass or forward recently completed results not yet written to a register file to a new dependent object. In at least one embodiment, register networks 4408, 4410 can communicate data with each other. In at least one embodiment, integer register file / branch network 4408 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, the floating-point register file / branch network 4410 may include, but is not limited to, entries with a width of 128 bits, since floating-point instructions typically have operands with a width of 64 to 128 bits.
[0442] In at least one embodiment, execution units 4412, 4414, 4416, 4418, 4420, 4422, and 4424 can execute instructions. In at least one embodiment, register networks 4408 and 4410 store integer and floating-point data operation values that the microinstructions need to execute. In at least one embodiment, processor 4400 can be, but is not limited to, any number of execution units 4412, 4414, 4416, 4418, 4420, 4422, and 4424, and combinations thereof. In at least one embodiment, floating-point ALU 4422 and floating-point move unit 4424 can perform floating-point, MMX, SIMD, AVX, and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 4422 can be, but is not limited to, a 64-bit multiplication-64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, floating-point hardware can be used to process instructions involving floating-point values. In at least one embodiment, ALU operations can be passed to fast ALUs 4416 and 4418. In at least one embodiment, fast ALUs 4416 and 4418 can perform fast operations with an effective delay of half a clock cycle. In at least one embodiment, most complex integer operations are routed to slow ALU 4420, because slow ALU 4420 can include, but is not limited to, integer execution hardware for long-latency type operations, such as multipliers, shifters, flag logic, and branching. In at least one embodiment, memory load / store operations can be performed by ALUs 4412 and 4414. In at least one embodiment, fast ALU 4416, fast ALU 4418, and slow ALU 4420 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 4416, fast ALU 4418, and slow ALU 4420 can be implemented to support various data bit sizes, including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating-point ALU 4422 and the floating-point moving unit 4424 can be implemented to support a range of operands with various bit widths, for example, they can be combined with SIMD and multimedia instructions to operate on 128-bit wide packaged data operands.
[0443] In at least one embodiment, microinstruction schedulers 4402, 4404, and 4406 schedule dependent operations before the parent load completes execution. In at least one embodiment, since microinstructions can be speculatively scheduled and executed within processor 4400, processor 4400 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be a dependent operation running in the pipeline that temporarily deprives the scheduler of the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and may allow independent operations to be completed. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0444] In at least one embodiment, "register" can refer to an onboard processor storage location that can be used as part of an instruction that identifies an operand. In at least one embodiment, a register can be one that can be used externally to the processor (from a programmer's perspective). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques via circuitry within the processor, such as dedicated physical registers, dynamically allocated physical registers renamed using register renaming, a combination of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for encapsulating data.
[0445] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, some or all of the inference and / or training logic 2815 may be incorporated into execution block 4411 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs shown in execution block 4411. Furthermore, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of execution block 4411 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0446] In at least one embodiment, Figure 28A and28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0447] Figure 45 A deep learning application processor 4500 according to at least one embodiment is illustrated. In at least one embodiment, the deep learning application processor 4500 uses instructions, which, if executed by the deep learning application processor 4500, cause the deep learning application processor 4500 to perform some or all of the processes and techniques described herein. In at least one embodiment, the deep learning application processor 4500 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 4500 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 4500 includes, but is not limited to, a processing cluster 4510(1)-4510(12), an inter-chip link (“ICL”) 4520(1)-4520(12), an inter-chip controller (“ICC”) 4530(1)-4530(2), a second-generation high-bandwidth memory (“HBM2”) 4540(1)-4540(4), a memory controller (“Mem Ctrlr”) 4542(1)-4542(4), a high-bandwidth memory physical layer (“HBM PHY”) 4544(1)-4544(4), a management controller central processing unit (“management controller CPU”) 4550, a serial peripheral interface, internal integrated circuits and general purpose input / output blocks (“SPI, I2C, GPIO”) 4560, a peripheral component interconnect fast controller and direct memory access block (“PCIe controller and DMA”) 4570, and a sixteen-channel peripheral component interconnect fast port (“PCI Express”). x 16”)4580.
[0448] In at least one embodiment, processing cluster 4510 can perform deep learning operations, including inference or prediction operations based on weight parameters computed using one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 4510 can include, but is not limited to, any number and type of processors. In at least one embodiment, deep learning application processor 4500 can include any number and type of processing cluster 4500. In at least one embodiment, inter-chip link 4520 is bidirectional. In at least one embodiment, inter-chip link 4520 and inter-chip controller 4530 enable multiple deep learning application processors 4500 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 4500 can include any number (including zero) and type of ICL 4520 and ICC 4530. In at least one embodiment, processor 4500 executes process 100 (see... Figure 1A ) or process 106 (see Figure 1B ).
[0449] In at least one embodiment, the HBM2 4540 provides a total of 32GB of memory. In at least one embodiment, the HBM2 4540(i) is associated with both the memory controller 4542(i) and the HBM PHY 4544(i), where “i” is any integer. In at least one embodiment, any number of HBM2 4540s can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controller 4542 and HBM PHY 4544. In at least one embodiment, any number and type of SPI, I2C, GPIO 4560, PCIe controller, and DMA 4570 and / or PCIe 4580 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.
[0450] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28BDetails regarding the inference and / or training logic 2815 are provided. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 4500. In at least one embodiment, the deep learning application processor 4500 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 4500. In at least one embodiment, the processor 4500 may be used to perform one or more neural network use cases described herein.
[0451] In at least one embodiment, utilizing Figure 28A and 28B The system described herein implements a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0452] Figure 46 This is a block diagram of a neuromorphic processor 4600 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 4600 may receive one or more inputs from a source external to the neuromorphic processor 4600. In at least one embodiment, these inputs may be transmitted to one or more neurons 4602 within the neuromorphic processor 4600. In at least one embodiment, the neurons 4602 and their components may be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 4600 may include, but is not limited to, thousands upon thousands of instances of neurons 4602, but any suitable number of neurons 4602 may be used. In at least one embodiment, each instance of neuron 4602 may include a neuron input 4604 and a neuron output 4606. In at least one embodiment, the neuron 4602 may generate an output that can be transmitted to the inputs of other instances of the neuron 4602. In at least one embodiment, the neuron input 4604 and the neuron output 4606 may be interconnected via synapses 4608.
[0453] In at least one embodiment, neuron 4602 and synapse 4608 may be interconnected, causing neuromorphic processor 4600 to operate to process or analyze information received by neuromorphic processor 4600. In at least one embodiment, neuron 4602 may send an output pulse (or “trigger” or “peak”) when the input received through neuron input 4604 exceeds a threshold. In at least one embodiment, neuron 4602 may sum or integrate the signal received at neuron input 4604. For example, in at least one embodiment, neuron 4602 may be implemented as a leaky integral-triggered neuron, wherein if the summation (referred to as a “membrane potential”) exceeds a threshold, neuron 4602 may use a transfer function such as a sigmoid or threshold function to generate an output (or “trigger”). In at least one embodiment, the leaky integral-triggered neuron may sum the signal received at neuron input 4604 to a membrane potential and may apply an attenuation factor (or leak) to reduce the membrane potential. In at least one embodiment, a leaking integral-triggered neuron may trigger if multiple input signals are received at neuron input 4604 quickly enough to exceed a threshold (i.e., before the membrane potential decays too low to trigger). In at least one embodiment, neuron 4602 may be implemented using circuitry or logic that receives input, integrates the input to the membrane potential, and decays the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 4602 may include, but is not limited to, comparator circuitry or logic that generates an output spike at neuron output 4606 when the result of applying the transfer function to neuron input 4604 exceeds a threshold. In at least one embodiment, once neuron 4602 is triggered, it can ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 4602 may resume normal operation after a suitable period of time (or recovery period).
[0454] In at least one embodiment, neurons 4602 can be interconnected via synapses 4608. In at least one embodiment, synapses 4608 can be operated to transmit signals from the output of a first neuron 4602 to the input of a second neuron 4602. In at least one embodiment, neurons 4602 can transmit information on more than one instance of synapse 4608. In at least one embodiment, one or more instances of neuron output 4606 can be connected via instances of synapses 4608 to instances of neuron input 4604 in the same neuron 4602. In at least one embodiment, an instance of neuron 4602 that produces an output to be transmitted on the instance of synapse 4608 may be referred to as a "presynaptic neuron". In at least one embodiment, an instance of neuron 4602 that receives input transmitted via an instance of synapse 4608 may be referred to as a "postsynaptic neuron". In at least one embodiment, regarding various instances of synapse 4608, since instances of neuron 4602 can receive input from one or more instances of synapse 4608 and can also transmit output through one or more instances of synapse 4608, a single instance of neuron 4602 can be both a "presynaptic neuron" and a "postsynaptic neuron".
[0455] In at least one embodiment, neurons 4602 may be organized into one or more layers. In at least one embodiment, each instance of neuron 4602 may have a neuron output 4606, which may fan out to one or more neuron inputs 4604 via one or more synapses 4608. In at least one embodiment, the neuron output 4606 of neuron 4602 in the first layer 4610 may be connected to the neuron input 4604 of neuron 4602 in the second layer 4612. In at least one embodiment, layer 4610 may be referred to as a “feedforward layer.” In at least one embodiment, each instance of neuron 4602 in an instance of the first layer 4610 may fan out to each instance of neuron 4602 in the second layer 4612. In at least one embodiment, the first layer 4610 may be referred to as a “fully connected feedforward layer.” In at least one embodiment, each instance of neuron 4602 in an instance of the second layer 4612 fan out to fewer than all instances of neuron 4602 in the third layer 4614. In at least one embodiment, the second layer 4612 may be referred to as a “sparsely connected feedforward layer.” In at least one embodiment, neurons 4602 in the second layer 4612 may fan out to neurons 4602 in multiple other layers, including neurons 4602 fan out to the second layer 4612. In at least one embodiment, the second layer 4612 may be referred to as a "recurrent layer". In at least one embodiment, the neuromorphic processor 4600 may be any suitable combination of recurrent layers and feedforward layers, including but not limited to sparsely connected feedforward layers and fully connected feedforward layers.
[0456] In at least one embodiment, the neuromorphic processor 4600 may include, but is not limited to, a reconfigurable interconnect architecture or dedicated hardwired interconnects to connect synapses 4608 to neurons 4602. In at least one embodiment, the neuromorphic processor 4600 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 4602 as needed, depending on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 4608 may be connected to neurons 4602 using interconnect structures (such as on-chip networks) or via dedicated connections. In at least one embodiment, synaptic interconnects and their components may be implemented using circuitry or logic.
[0457] In at least one embodiment, utilizing Figure 28A and 28B The system described herein implements a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0458] Figure 47 A processing system according to at least one embodiment is illustrated. In at least one embodiment, system 4700 includes one or more processors 4702 and one or more graphics processors 4708, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 4702 or processor cores 4707. In at least one embodiment, system 4700 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices. In at least one embodiment, system 4700 executes process 100 (see...). Figure 1A ) and / or process 106 (see Figure 1B ).
[0459] In at least one embodiment, system 4700 may include or be integrated into a server-based gaming platform, including a game console, mobile game console, handheld game console, or online game console, which are game and media consoles. In at least one embodiment, system 4700 is a mobile phone, smartphone, tablet computing device, or mobile internet device. In at least one embodiment, processing system 4700 may also include components coupled to or integrated into a wearable device, such as a smartwatch, smart glasses, augmented reality, or virtual reality device. In at least one embodiment, processing system 4700 is a television or set-top box device having one or more processors 4702 and a graphical interface generated by one or more graphics processors 4708.
[0460] In at least one embodiment, one or more processors 4702 each include one or more processor cores 4707 to process instructions that, when executed, perform operations against the system and user software. In at least one embodiment, each of the one or more processor cores 4707 is configured to process a specific instruction sequence 4709. In at least one embodiment, the instruction sequence 4709 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). In at least one embodiment, each processor core 4707 may process a different instruction sequence 4709, which may include instructions that facilitate the emulation of other instruction sequences. In at least one embodiment, the processor core 4707 may also include other processing devices, such as a digital signal processor (DSP).
[0461] In at least one embodiment, processor 4702 includes cache memory 4704. In at least one embodiment, processor 4702 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of processor 4702. In at least one embodiment, processor 4702 also uses an external cache (e.g., a Level 3 (L3) cache or a last-level cache (LLC)) (not shown), which can be shared among processor cores 4707 using known cache coherence techniques. In at least one embodiment, processor 4702 further includes a register file 4706, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 4706 may include general-purpose registers or other registers.
[0462] In at least one embodiment, one or more processors 4702 are coupled to one or more interface buses 4710 to transmit communication signals, such as address, data, or control signals, between the processors 4702 and other components in the system 4700. In at least one embodiment, the interface bus 4710 may be a processor bus, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 4710 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 4702 includes an integrated memory controller 4716 and a platform controller hub 4730. In at least one embodiment, the memory controller 4716 facilitates communication between memory devices and other components of the processing system 4700, while the platform controller hub (PCH) 4730 provides connectivity to input / output (I / O) devices via a local I / O bus.
[0463] In at least one embodiment, memory device 4720 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 4720 may be used as system memory of processing system 4700 to store data 4722 and instructions 4721 for use when one or more processors 4702 execute an application or process. In at least one embodiment, memory controller 4716 is also coupled to an optional external graphics processor 4712, which may communicate with one or more graphics processors 4708 of processor 4702 to perform graphics and media operations. In at least one embodiment, display device 4711 may be connected to processor 4702. In at least one embodiment, display device 4711 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 4711 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.
[0464] In at least one embodiment, the platform controller hub 4730 enables peripheral devices to connect to the storage device 4720 and the processor 4702 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 4746, a network controller 4734, a firmware interface 4728, a wireless transceiver 4726, a touch sensor 4725, and a data storage device 4724 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 4724 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 4725 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 4726 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 4728 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 4734 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 4710. In at least one embodiment, audio controller 4746 is a multi-channel high-definition audio controller. In at least one embodiment, processing system 4700 includes an optional legacy I / O controller 4740 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to system 4700. In at least one embodiment, platform controller hub 4730 may also be connected to one or more Universal Serial Bus (USB) controllers 4742 that connect input devices, such as a keyboard and mouse combination 4743, a camera 4744, or other USB input devices.
[0465] In at least one embodiment, instances of the memory controller 4716 and platform controller hub 4730 may be integrated into a discrete external graphics processor, such as external graphics processor 4712. In at least one embodiment, the platform controller hub 4730 and / or the memory controller 4716 may be external to one or more processors 4702. For example, in at least one embodiment, system 4700 may include an external memory controller 4716 and a platform controller hub 4730, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset communicating with processor 4702.
[0466] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28BDetails regarding the inference and / or training logic 2815 are provided. In at least one embodiment, some or all of the inference and / or training logic 2815 may be incorporated into the graphics processor 4708. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in a 3D pipeline. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, in addition to Figure 28A or Figure 28B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 4708 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0467] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0468] Figure 48 This is a block diagram of a processor 4800 having one or more processor cores 4802A-4802N, an integrated memory controller 4814, and an integrated graphics processor 4808 according to at least one embodiment. In at least one embodiment, the processor 4800 may include additional cores, up to and including additional cores 4802N indicated by dashed boxes. In at least one embodiment, each processor core 4802A-4802N includes one or more internal cache units 4804A-4804N. In at least one embodiment, each processor core may also access one or more shared cache units 4806. In at least one embodiment, the processor 4800 executes process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0469] In at least one embodiment, internal cache units 4804A-4804N and shared cache unit 4806 represent a cache memory hierarchy within processor 4800. In at least one embodiment, cache memory units 4804A-4804N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest level of cache preceding external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 4806 and 4804A-4804N.
[0470] In at least one embodiment, the processor 4800 may further include a set of one or more bus controller units 4816 and a system agent core 4810. In at least one embodiment, the one or more bus controller units 4816 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 4810 provides management functions for various processor components. In at least one embodiment, the system agent core 4810 includes one or more integrated memory controllers 4814 to manage access to various external memory devices (not shown).
[0471] In at least one embodiment, one or more processor cores 4802A-4802N include support for multi-threaded concurrent processing. In at least one embodiment, system agent core 4810 includes components for coordinating and operating cores 4802A-4802N during multi-threaded processing. In at least one embodiment, system agent core 4810 may additionally include a power control unit (PCU) including logic and components for regulating one or more power states of processor cores 4802A-4802N and graphics processor 4808.
[0472] In at least one embodiment, processor 4800 further includes a graphics processor 4808 for performing graph processing operations. In at least one embodiment, graphics processor 4808 is coupled to a shared cache unit 4806 and a system proxy core 4810 including one or more integrated memory controllers 4814. In at least one embodiment, system proxy core 4810 further includes a display controller 4811 for driving graphics processor outputs to one or more coupled displays. In at least one embodiment, display controller 4811 may also be a separate module coupled to graphics processor 4808 via at least one interconnect, or it may be integrated within graphics processor 4808.
[0473] In at least one embodiment, ring-based interconnect unit 4812 is used to couple internal components of processor 4800. In at least one embodiment, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, may be used. In at least one embodiment, graphics processor 4808 is coupled to ring interconnect 4812 via I / O link 4813.
[0474] In at least one embodiment, I / O link 4813 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory module 4818 (e.g., eDRAM module). In at least one embodiment, each of processor cores 4802A-4802N and graphics processor 4808 uses embedded memory module 4818 as a shared last-level cache.
[0475] In at least one embodiment, processor cores 4802A-4802N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 4802A-4802N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 4802A-4802N executing a common instruction set, while one or more other processor cores 4802A-4802N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 4802A-4802N are heterogeneous in terms of microarchitecture, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, processor 4800 may be implemented on one or more chips or implemented as a SoC integrated circuit.
[0476] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, some or all of the inference and / or training logic 2815 may be incorporated into the processor 4800. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in... Figure 48 The 3D pipeline, graphics core 4802, shared functional logic, or other logic are included. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use, except... Figure 28A or Figure 28BThe logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of processor 4800 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0477] In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0478] Figure 49 This is a block diagram of a graphics processor 4900, which may be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores. In at least one embodiment, the graphics processor 4900 communicates with registers on the graphics processor 4900 and commands placed in memory via a memory-mapped I / O interface. In at least one embodiment, the graphics processor 4900 includes a memory interface 4914 for accessing memory. In at least one embodiment, the memory interface 4914 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.
[0479] In at least one embodiment, the graphics processor 4900 further includes a display controller 4902 for driving display output data to the display device 4920. In at least one embodiment, the display controller 4902 includes a combination of hardware for one or more overlay planes of the display device 4920 and multi-layer video or user interface elements. In at least one embodiment, the display device 4920 may be an internal or external display device. In at least one embodiment, the display device 4920 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In at least one embodiment, the graphics processor 4900 includes a video codec engine 4906 for encoding, decoding, or transcoding media into, from, or between one or more media encoding formats, including but not limited to Moving Picture Experts Group (MPEG) formats (e.g., MPEG-2), Advanced Video Coding (AVC) formats (e.g., H.264 / MPEG-4 AVC, and SMPTE 421M / VC-1), Joint Picture Experts Group (JPEG) formats (e.g., JPEG), and MotionJPEG (MJPEG) formats. In at least one embodiment, the graphics processor 4900 executes process 100 (see... Figure 1A ) or process 106 (see Figure 1B ).
[0480] In at least one embodiment, the graphics processor 4900 includes a block image transfer (BLIT) engine 4904 to perform two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfer. However, in at least one embodiment, one or more components of a graphics processing engine (GPE) 4910 are used to perform 2D graphics operations. In at least one embodiment, the GPE 4910 is a computational engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0481] In at least one embodiment, GPE 4910 includes a 3D pipeline 4912 for performing 3D operations, such as rendering 3D images and scenes using processing functions that manipulate 3D primitive shapes (e.g., rectangles, triangles, etc.). In at least one embodiment, 3D pipeline 4912 includes programmable and fixed function elements that perform various tasks and / or generate execution threads to 3D / media subsystem 4915. While 3D pipeline 4912 can be used to perform media operations, in at least one embodiment, GPE 4910 also includes a media pipeline 4916 for performing media operations such as video post-processing and image enhancement.
[0482] In at least one embodiment, the media pipeline 4916 includes fixed-function or programmable logic units for performing one or more specialized media operations, such as video decoding acceleration, video deinterlacing, and video encoding acceleration, replacing or representing the video codec engine 4906. In at least one embodiment, the media pipeline 4916 also includes a thread generation unit for generating threads to execute on the 3D / media subsystem 4915. In at least one embodiment, the generated threads perform computations of media operations on one or more graphics execution units included in the 3D / media subsystem 4915.
[0483] In at least one embodiment, the 3D / media subsystem 4915 includes logic for executing threads generated by the 3D pipeline 4912 and the media pipeline 4916. In at least one embodiment, the 3D pipeline 4912 and the media pipeline 4916 send thread execution requests to the 3D / media subsystem 4915, which includes thread dispatch logic for arbitrating various requests and dispatching them to available thread execution resources. In at least one embodiment, the execution resources include an array of graphics execution units for processing 3D and media threads. In at least one embodiment, the 3D / media subsystem 4915 includes one or more internal caches for thread instructions and data. In at least one embodiment, the subsystem 4915 also includes shared memory, including registers and addressable memory, for sharing data between threads and storing output data.
[0484] The inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A and / or Figure 28B Details regarding the inference and / or training logic 2815 are provided. In at least one embodiment, some or all of the inference and / or training logic 2815 may be incorporated into the processor 4900. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs included in the 3D pipeline 4912. Furthermore, in at least one embodiment, the inference and / or training operations described herein may use, except for... Figure 28A or Figure 28B The logic other than that shown is used to perform this task. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 4900 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0485] In at least one embodiment, Figure 28A and 28BOne or more systems described herein are used to implement a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and 28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0486] Figure 50 This is a block diagram of a graphics processing engine 5010 of a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processing engine (GPE) 5010 is... Figure 49 The version of GPE 4910 shown. In at least one embodiment, the media pipeline 5016 is optional and may not be explicitly included in GPE 5010. In at least one embodiment, a separate media and / or image processor is coupled to GPE 5010. In at least one embodiment, the graphics processing engine 5010 executes process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0487] In at least one embodiment, GPE 5010 is coupled to or includes command stream converter 5003, which provides command streams to 3D pipeline 5012 and / or media pipeline 5016. In at least one embodiment, command stream converter 5003 is coupled to memory, which may be system memory, or one or more of internal cache memory and shared cache memory. In at least one embodiment, command stream converter 5003 receives commands from memory and sends the commands to 3D pipeline 5012 and / or media pipeline 5016. In at least one embodiment, the commands are instructions, primitives, or micro-operations retrieved from a circular buffer that stores commands for 3D pipeline 5012 and media pipeline 5016. In at least one embodiment, the circular buffer may further include a batch command buffer storing multiple commands in batches. In at least one embodiment, commands for 3D pipeline 5012 may further include references to data stored in memory, such as, but not limited to, vertex and geometry data for 3D pipeline 5012 and / or image data and memory objects for media pipeline 5016. In at least one embodiment, the 3D pipeline 5012 and the media pipeline 5016 process commands and data by performing operations or by dispatching one or more execution threads to the graphics core array 5014. In at least one embodiment, the graphics core array 5014 includes one or more graphics core blocks (e.g., one or more graphics cores 5015A, one or more graphics cores 5015B), each block including one or more graphics cores. In at least one embodiment, each graphics core includes a set of graphics execution resources, which include general-purpose and graphics-specific execution logic for performing graphics and computation operations, and fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic, including... Figure 28A and Figure 28B The reasoning and / or training logic in 2815.
[0488] In at least one embodiment, the 3D pipeline 5012 includes fixed functions and programmable logic for processing one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 5014. In at least one embodiment, the graphics core array 5014 provides a unified execution resource block for processing shader programs. In at least one embodiment, the multipurpose execution logic (e.g., execution units) within the graphics cores 5015A-5015B of the graphics core array 5014 includes support for various 3D API shader languages and can execute multiple concurrently running threads associated with multiple shaders.
[0489] In at least one embodiment, the graphics core array 5014 further includes execution logic for performing media functions, such as video and / or image processing. In at least one embodiment, in addition to graphics processing operations, the execution unit also includes general-purpose logic programmable to perform parallel general-purpose computing operations.
[0490] In at least one embodiment, output data can be output to memory in a unified return buffer (URB) 5018, the output data being generated by a thread executing on the graphics core array 5014. In at least one embodiment, the URB 5018 can store data from multiple threads. In at least one embodiment, the URB 5018 can be used to send data between different threads executing on the graphics core array 5014. In at least one embodiment, the URB 5018 can also be used for synchronization between threads on the graphics core array 5014 and fixed-function logic within shared-function logic 5020.
[0491] In at least one embodiment, the graphics core array 5014 is scalable, such that it includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance level of the GPE 5010. In at least one embodiment, the execution resources are dynamically scalable, such that they can be enabled or disabled as needed.
[0492] In at least one embodiment, the graphics core array 5014 is coupled to shared function logic 5020, which includes multiple resources shared among the graphics cores in the graphics core array 5014. In at least one embodiment, the shared functions performed by the shared function logic 5020 are embodied in hardware logic units that provide dedicated supplementary functions to the graphics core array 5014. In at least one embodiment, the shared function logic 5020 includes, but is not limited to, a sampler unit 5021, a math unit 5022, and inter-thread communication (ITC) logic 5023. In at least one embodiment, one or more caches 5025 are included in or coupled to the shared function logic 5020.
[0493] In at least one embodiment, shared functionality is used if the demand for dedicated functionality is insufficient to be contained within the graphics core array 5014. In at least one embodiment, a single instance of the dedicated functionality is used in shared functionality logic 5020 and shared among other execution resources within the graphics core array 5014. In at least one embodiment, a specific shared functionality may be included within shared functionality logic 5026 within the graphics core array 5014, said specific shared functionality being widely used within shared functionality logic 5020 of the graphics core array 5014. In at least one embodiment, shared functionality logic 5026 within the graphics core array 5014 may include some or all of the logic within shared functionality logic 5020. In at least one embodiment, all logic elements within shared functionality logic 5020 may be replicated within shared functionality logic 5026 of the graphics core array 5014. In at least one embodiment, shared functionality logic 5020 is excluded to support shared functionality logic 5026 within the graphics core array 5014.
[0494] Inference and / or training logic 2815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 28A 28B provides details regarding the inference and / or training logic 2815. In at least one embodiment, some or all of the inference and / or training logic 2815 may be incorporated into the graphics processor 5010. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the 3D pipeline 5012, graphics core 5015, shared function logic 5026, shared function logic 5020, or... Figure 50 In other logic within the [process]. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use [other methods besides...]. Figure 28A or Figure 28B The logic is performed using logic other than that shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 5010 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0495] In at least one embodiment, utilizing Figure 28A and 28B The system described herein implements a framework for image processing (e.g., real-time image rendering and enhancement). In at least one embodiment, Figure 28A and 28B The system described herein is used to perform various processes, such as combination Figures 1A-27 The processes described. In at least one embodiment, Figure 28A and28B One or more systems described herein are used to generate blue noise masks and apply them to images capable of processing the temporal domain, i.e., adding time to the spatial (image) domain to improve image quality when rendering images over multiple frames (e.g., time).
[0496] Figure 51 This is a block diagram of the hardware logic of a graphics processor core 5100 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 5100 is included within a graphics core array. In at least one embodiment, the graphics processor core 5100 (sometimes referred to as a core slice) may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 5100 is an example of a graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 5100 may include a fixed-function block 5130, also referred to as a sub-slice, coupled to a plurality of sub-cores 5101A-5101F, which includes modules of general-purpose and fixed-function logic. In at least one embodiment, the graphics processor core 5100 executes process 100 (see...). Figure 1A ) or process 106 (see Figure 1B ).
[0497] In at least one embodiment, the fixed-function block 5130 includes a geometry and fixed-function pipeline 5136, which, for example, may be shared by all sub-cores of the graphics processor 5100 in a lower-performance and / or lower-power graphics processor implementation. In at least one embodiment, the geometry and fixed-function pipeline 5136 includes a 3D fixed-function pipeline, a video front-end unit, a thread generator and a thread dispatcher, and a unified return buffer manager that manages a unified return buffer.
[0498] In at least one fixed embodiment, the fixed functional block 5130 also includes a graphics SoC interface 5137, a graphics microcontroller 5138, and a media pipeline 5139. In at least one embodiment, the graphics SoC interface 5137 provides an interface between the graphics core 5100 and other processor cores in the on-chip integrated circuit system. In at least one embodiment, the graphics microcontroller 5138 is a programmable subprocessor configurable to manage various functions of the graphics processor 5100, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 5139 includes logic that facilitates decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 5139 implements media operations via requests for computation or sampling logic within subcores 5101-5101F.
[0499] In at least one embodiment, the SoC interface 5137 enables the graphics core 5100 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, the SoC interface 5137 also enables communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline) and enables the use and / or implementation of global memory atoms that can be shared between the graphics core 5100 and the CPU within the SoC. In at least one embodiment, the graphics SoC interface 5137 also implements power management control for the graphics processor core 5100 and enables interfacing between the clock domain of the graphics processor core 5100 and other clock domains within the SoC. In at least one embodiment, the SoC interface 5137 enables the receipt of command buffers from a command stream converter and a global thread dispatcher, configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, when a media operation is to be performed, commands and instructions can be dispatched to the media pipeline 5139, or when a graphics processing operation is to be performed, they can be assigned to the geometry and fixed-function pipeline (e.g., geometry and fixed-function pipeline 5136, and / or geometry and fixed-function pipeline 5114).
[0500] In at least one embodiment, the graphics microcontroller 5138 can be configured to perform various scheduling and management tasks on the graphics core 5100. In at least one embodiment, the graphics microcontroller 5138 can perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 5102A-5102F, 5104A-5104F in subcores 5101A-5101F. In at least one embodiment, host software executing on the CPU core of the SoC including the graphics core 5100 can submit a workload to one of multiple graphics processor paths, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload should be run next, submitting the workload to a command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is completed. In at least one embodiment, the graphics microcontroller 5138 may also facilitate a low-power or idle state of the graphics core 5100, thereby providing the graphics core 5100 with the ability to save and restore registers across low-power state transitions within the graphics core 5100, indepen...
Claims
1. A computer-implemented method for processing image data, the method comprising: The energy value of the at least some pixels is calculated based on the image coordinates corresponding to at least some pixels in one or more images at different times, wherein the energy value of the pixel is a non-zero value because the pixel and another pixel are located in the same two-dimensional layer or the pixel and the other pixel have the same coordinates in different time slices; A multidimensional mask is generated based on the calculated energy values of the at least some pixels and applied to the one or more images. The multidimensional mask includes multiple dimensions of noise values, including at least two dimensions corresponding to the image space and one dimension corresponding to time. An output image is provided by applying the mask to one or more images.
2. The computer-implemented method of claim 1, wherein calculating the energy values of the at least some pixels is further based on the distance between pairs of the at least some pixels, an energy attenuation parameter, and wherein the distance between pairs of the at least some pixels is calculated in a circular manner.
3. The computer-implemented method according to claim 1 further includes: If a pixel among the at least some pixels is not in the same two-dimensional layer as another pixel or does not have the same coordinates in different time slices, then the energy value of the pixel is set to zero.
4. The computer-implemented method of claim 1, wherein providing the output image is a portion of an image generation pipeline that includes ray tracing.
5. The computer-implemented method according to claim 1, further comprising: The time-lapse image is applied to the one or more images, wherein the lapse is based on one or more values of a higher resolution image magnified from a lower resolution image by neural network inference.
6. The computer-implemented method according to claim 1, wherein the method further comprises: Obtain pixel data corresponding to the one or more images; as well as The empty clustering algorithm is applied to the pixel data, and the energy function used for the empty clustering algorithm is based on the calculated energy value.
7. The computer-implemented method according to claim 2, wherein the energy attenuation parameter is a Gaussian blur parameter.
8. The computer-implemented method according to claim 1, further comprising: A low-pass filter is applied to the output image.
9. The computer-implemented method of claim 1, wherein the mask is applied as part of sampling the one or more images.
10. A processor, comprising: One or more processing units are configured to perform multiple operations, said multiple operations including: The energy value of the at least some pixels is calculated based on the coordinates of at least some pixels in one or more images at different times, the distance between pairs of said at least some pixels, and an energy decay parameter, wherein the energy value of the pixel is a non-zero value because a pixel in the at least some pixels and another pixel are located in the same two-dimensional layer or the pixel and the other pixel have the same coordinates at different time slices; Generate a mask applied to the one or more images based on the calculated energy values of the at least some pixels; and The output image is rendered across multiple frames by applying the mask to one or more images.
11. The processor of claim 10, wherein the distance between pairs of the at least some pixels is calculated in a circular manner.
12. The processor of claim 10, wherein the operation further comprises: Receive pixel data having three dimensions corresponding to the one or more images, wherein one of the three dimensions corresponds to a time dimension, wherein the pixel data includes the coordinates of each pixel, and wherein the mask is a three-dimensional mask.
13. The processor of claim 10, wherein the operation further comprises: If a pixel among the at least some pixels is not in the same two-dimensional layer or does not have the same coordinates as another pixel, then the energy value of the pixel is set to zero.
14. The processor of claim 10, wherein providing the output image is a portion of an image generation pipeline that includes ray tracing.
15. The processor of claim 10, wherein the operation further comprises: Time-based image magnification is applied to one or more images, wherein the magnification is based on neural network inference from a lower-resolution image.
16. The processor of claim 10, wherein the operation further comprises: Obtain pixel data corresponding to the one or more images, including more than three dimensions.
17. The processor of claim 16, wherein the operation further comprises: The spatial aggregation algorithm is applied to the pixel data, and the energy function used for the spatial aggregation algorithm is based on the calculated intensity value.
18. The processor of claim 10, wherein the energy decay parameter is a Gaussian blur parameter.
19. The processor of claim 12, wherein receiving the pixel data having three dimensions corresponding to the one or more images comprises: Sampling is performed on one or more of the images.
20. A computer-readable storage medium having one or more instructions stored thereon, which, if executed by one or more processors, cause the one or more processors to perform the following operations: The energy value of the at least some pixels is calculated based on the coordinates of at least some pixels in one or more images, the distance between pairs of said at least some pixels, and an energy decay parameter, wherein the energy value of the pixel is non-zero if a pixel in the at least some pixels is located in the same two-dimensional layer as another pixel or the pixel and the other pixel have the same coordinates at different time slices; A three-dimensional mask is generated based on the calculated energy values of at least some of the pixels and applied to one or more received images; as well as An output image is provided by applying the three-dimensional mask to one or more received images.
21. The computer-readable storage medium of claim 20, wherein the three-dimensional mask adds blue noise to the one or more images.
22. The computer-readable storage medium of claim 20, wherein the distance between pairs of the at least some pixels is calculated in a circular manner.
23. The computer-readable storage medium of claim 20, further comprising: Receive pixel data having three dimensions corresponding to the one or more images, wherein one of the three dimensions corresponds to the time dimension, and wherein the pixel data includes the coordinates of each pixel.
24. The computer-readable storage medium of claim 20, wherein the operation further comprises: If a pixel among the at least some pixels is not in the same two-dimensional layer or does not have the same coordinates as another pixel, then the energy value of the pixel is set to zero.
25. The computer-readable storage medium of claim 20, wherein providing the output image is a portion of an image generation pipeline that includes ray tracing or path tracing.
26. The computer-readable storage medium of claim 20, wherein the operation further comprises: A time-lapse image is applied to the image, wherein the lapse is based on a lower resolution image.
27. The computer-readable storage medium of claim 20, wherein the operation further comprises: Obtain pixel data corresponding to the one or more images; as well as The empty clustering algorithm is applied to the pixel data, and the energy function used for the empty clustering algorithm is based on the calculated energy value.
28. The computer-readable storage medium of claim 20, wherein the operation further comprises: Obtain pixel data including one or more coordinates of one or more pixels of the one or more images, wherein the energy attenuation parameter is a Gaussian blur parameter based on the coordinates of the pixel data.
29. The computer-readable storage medium of claim 20, wherein the operation further comprises: A low-pass filter is applied before the output image is provided.
30. The computer-readable storage medium of claim 23, wherein receiving the pixel data having three dimensions corresponding to the one or more images comprises: Sampling is performed on one or more of the images.
31. The computer-readable storage medium of claim 21, wherein providing the output image further comprises: The 3D mask uses color mixing, random transparency, region light sampling, volume rendering, path tracing, temporal anti-aliasing, and / or random alpha image processing techniques.
Citation Information
Patent Citations
Method for video compression based on line clipping
CN103763562A
Three-dimensional image relocation method
CN108449588A