Avoiding textured pattern artifacts in content generation systems and applications

By randomly sampling the texture patterns, the Moir pattern artifact problem caused by dense patterned textures is solved, and the generation of high-quality display content is achieved, avoiding image data loss and quality degradation.

CN119941545APending Publication Date: 2025-05-06NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411552613.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-11-01
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When generating high-quality display content, densely patterned textures or object rendered images are prone to artifacts such as moiré patterns, affecting the perceived quality of the image.

Method used

Reduce the possibility of artifacts appearing when rendered images are displayed by random but constrained sampling of texture patterns. Specific methods include introducing randomization during the sampling process, ensuring that the sampling position is within the pixel boundary, and eliminating rendering jitter using first-order approximation.

Benefits of technology

Effectively reduce or prevent the occurrence of artifacts, improve the perceived quality of the image, and avoid image data loss and quality degradation caused by adjustments after sampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941545A_ABST
    Figure CN119941545A_ABST
Patent Text Reader

Abstract

The invention discloses avoiding texture pattern artifacts in content generation systems and applications. The methods presented herein are used to remove or reduce anti-aliasing artifacts, such as Moire patterns or stair steps, in an image to be rendered. In many cases, these artifacts correspond to regular texture patterns with fine details, and addition of randomization in sampling locations may help eliminate the effects of pattern regularity. In at least one embodiment, a first order approximation may be used that introduces a random amount of shift determined using a texture coordinate derivative. The random amount may take into account any jitter offset and shift the texture coordinates by the determined random amount such that the sample selected for that pixel will be selected from among the sample positions corresponding to the random shift, but limited within the boundaries of the pixel.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In various applications, such as, for example, gaming, animation, or virtual reality content generation, it may be desirable to provide a high-quality display of generated content, which may include, for example, fine details at high resolution. However, when rendering images using densely patterned textures or objects, artifacts such as Moiré patterns may appear, due in part to the arrangement of the pattern relative to the pixel grid rendered for the image. When displayed, such artifacts may be so noticeable to a person viewing such an image that the artifact may bother the viewer or may otherwise significantly reduce the perceived quality of the overall image. For images corresponding to frames of a video sequence, Moiré patterns are typically unstable and produce some degree of flicker. Previous approaches have attempted to reduce the presence of these and other such artifacts through post-processing, but making adjustments after sampling may result in a loss of image data and a corresponding decrease in image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:

[0003] Figure 1 shows example texture patterns that, when superimposed relative to a pixel grid, can generate image artifacts in accordance with at least one embodiment;

[0004] Figure 2A-2G illustrates sampling locations that may be selected for respective pixels during an upsampling process according to at least one embodiment;

[0005] Figure 3 An example image rendering and upsampling pipeline is shown in accordance with at least one embodiment;

[0006] Figure 4A and Figure 4B illustrates components of an example content generation system in accordance with at least one embodiment;

[0007] Figure 5 An example process for performing sampling on a texture pattern to reduce the presence or likelihood of image artifacts in accordance with at least one embodiment is shown;

[0008] Figure 6 Components of a distributed system that may be used to generate, modify, and / or provide image content according to at least one embodiment are shown;

[0009] Fig. 7A Inference and / or training logic according to at least one embodiment is shown;

[0010] Figure 7B Inference and / or training logic according to at least one embodiment is shown;

[0011] Figure 8 An example data center system is shown in accordance with at least one embodiment;

[0012] Fig. 9 A computer system according to at least one embodiment is shown;

[0013] Fig.10 A computer system according to at least one embodiment is shown;

[0014] Fig.11 illustrates at least a portion of a graphics processor according to one or more embodiments;

[0015] Fig.12 illustrates at least a portion of a graphics processor according to one or more embodiments;

[0016] Fig.13 is an example data flow diagram of a high-level computing pipeline according to at least one embodiment;

[0017] Fig.14 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline according to at least one embodiment; and

[0018] Fig.15A and Fig. 15B A data flow diagram of a process for training a machine learning model, and a client-server architecture for augmenting an annotation tool with a pre-trained annotation model, according to at least one embodiment are shown. DETAILED DESCRIPTION

[0019] In the following description, various embodiments will be described. For the purpose of explanation, specific configurations and details are set forth to provide a thorough understanding of the embodiments. However, it will also be appreciated by those skilled in the art that the embodiments may be practiced without the specific details. In addition, well-known features may be omitted or simplified to avoid obscuring the described embodiments.

[0020] The systems and methods described herein may be used, without limitation, in electronic and computer gaming systems, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, boats, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones), and / or other vehicle types. In addition, the systems and methods described herein may be used for a variety of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, safety and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, generative artificial intelligence with large language models (LLMs), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0021] The disclosed embodiments may be included in a variety of different systems, such as electronic and computer gaming systems, automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing generative AI operations using LLMs, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems.

[0022] Methods according to various illustrative embodiments are used to reduce or prevent anti-aliasing (and other image) artifacts associated with, for example, sampling textures with fine detail. Specifically, methods are used to perform random but constrained sampling of texture patterns to reduce the likelihood of artifacts such as moiré patterns or stair-stepping appearing in the rendered image when the rendered image is displayed. Rather than attempting to remove artifacts in post-processing (which may result in loss of image data and reduced final image quality), such methods attempt to randomize the texture coordinates of pixels or sample locations in each frame to be rendered, where the randomization occurs at or before the sampling time. In order to avoid problems such as blurring caused by randomization of sampling positions, the sampling positions can be constrained so that the randomization does not cause the sample positions to exceed the boundaries of the corresponding pixels. In at least one embodiment, a first-order approximation (e.g., a first-order Taylor approximation) is used to eliminate rendering jitter (global or otherwise) for a given image or frame to be rendered, and can be in the range of [-0.5, +0.5] 2 A random offset is applied within a range (where the width and height of a pixel are normalized to a value of 1.0). The amount of randomization can depend at least in part on the texture coordinate derivative, taking into account the jitter value. In at least one embodiment, the shifting of the texture coordinates occurs before sampling. This approach can be used with two-dimensional (2D) and three-dimensional (3D) texture coordinates, and can be used with various textures or other such visual components to determine the pixel values ​​of the image to be generated.

[0023] It will be apparent to those skilled in the art based on the teachings and suggestions contained herein that variations of this and other such functionality may also be used within the scope of the various embodiments.

[0024] Figure 1An example grid 100 of pixels of an image to be rendered according to at least one embodiment is shown. The grid 100 shown is a two-dimensional (2D) grid of square pixels, although other types of images or pixel arrangements can benefit from aspects of various embodiments described herein. In addition, it should be understood that the number of pixels shown is very small for ease of explanation, but in various embodiments, the image can have tens of millions or more pixels in the grid. For example, during the rendering process, the pixel value of each individual pixel unit 102 in the grid 100 can be determined. The pixel value (which can correspond to the color value to be used when rendering the image) can be determined using a process such as ray tracing (or other light transport simulation technology) or hit detection and a shading process that analyzes one or more objects identified by ray tracing or hit detection. In some cases, the ray traced for the pixel will intersect with an object associated with a texture. For example, the intersecting object can be represented by a 2D or 3D mesh representing its shape and a texture that can be mapped or projected onto the shape, which provides the appearance of an object with the shape. The color value of a given pixel can then be determined by the position of the pattern sampled for the given pixel.

[0025] The texture of an object may have any suitable appearance and may be regular or irregular in nature. In some cases, the texture may have Figure 1 The texture pattern 110 shown is relatively fine (at least relative to the pixel size of the image to be rendered) and regular in nature. Figure 1 In , the texture pattern is shown as (at least in part) comprising a set of parallel lines, where the spacing between the lines is less than the width of a pixel. As shown, image 120 may substantially correspond to the superposition of texture pattern 110 relative to pixel grid 100. However, as shown, the lines of the texture pattern are slightly angled relative to the boundaries of the pixels in grid 100. As shown in the superimposed image, this may cause the lines of the texture pattern to appear thicker in some areas and thinner in some areas. Although the pixel lines will not actually appear in the final image, the blending or anti-aliasing methods used by various rendering pipelines will produce similar effects. As an example, a reduced image of the overlay is presented, in which a visible pattern that does not correspond to a set of parallel lines is displayed. In this example, the pattern is generally referred to as a moiré pattern 130. A moiré pattern is an artifact that typically occurs due to an interference pattern generated when two or more partially transmitted patterns are superimposed on each other. Such artifacts typically appear when the angles or offsets of the patterns are slightly different, or when the patterns have different spacings or other such variations (usually with at least a certain degree of regularity). Figure 1 Another example of a moiré pattern 140 is shown, which may be produced by different patterns and / or arrangements. Other similar types of artifacts may occur for different types of patterns or textures.

[0026] When generating image data (e.g., by using a rendering pipeline), patterns that cause artifacts may include pixel grids and texture patterns, such as those described in relation to Figure 1 As discussed. Imperfect alignment may occur between the regular pattern and the projected pixel grid, resulting in a type of side signal that was not present in the original image, pattern, or texture. Other textures may also produce other artifacts when projected onto a regular (or at least semi-regular) pixel grid or other such array. The side signals become visible when the image is displayed, but may not be obvious from the image data itself. As discussed above, this may also result in at least some degree of flicker for a series of images or video frames displayed, which may be distracting to the human eye.

[0027] Methods according to various embodiments may attempt to avoid or reduce the presence of such artifacts by changing the way in which texture sampling is performed. In at least one embodiment, the texture coordinates to be sampled may be randomized (or randomly moved) to reduce the effects of regular texture patterns. However, in at least some embodiments, the randomization may be limited to ensure that sampling occurs within locations associated with the same pixel to avoid blurring or other artifacts that may be introduced by sampling from locations associated with other pixel locations.

[0028] In some cases, sampling occurs at the center point of a lower resolution pixel that will be upsampled to a larger number (e.g., 4, 9, or 16) of pixels, and the random sampling can be limited to the boundaries of a given pixel based at least in part on the size of the pixel. However, in order to be able to capture fine details, various upscaling or upsampling systems utilize sub-pixel jitter. Sub-pixel global jitter offset can be advantageously used in image generation and processing techniques such as temporal supersampling, where an image can be rendered at a first resolution and then (intelligently) upsampled to a higher second resolution for presentation by a higher resolution display or rendering mechanism. Supersampling techniques, such as those used by the Deep Learning Supersampling (DLSS) product provided by Nvidia Corporation, can rely on camera jitter, where each frame is rendered from a slightly different camera position and produces a slightly different image. The supersampling process can then attempt to combine these differences to produce a higher resolution image.

[0029] As an example, an upscaling process (e.g., a deep learning-based supersampling or super-resolution process) can be used to increase the resolution of one or more images (e.g., images or video frames in a sequence or video stream). In at least one embodiment, Figure 2AAs shown, this may include zooming in from a set of lower resolution pixels 202 to a set of higher resolution pixels 204, such as by 4x upsampling as shown. In at least one embodiment, this may include a representation of one or more objects in a scene, such as a live game scene. In at least one embodiment, the rendering engine may output an image of one or more objects at a first resolution, which will be zoomed in to one or more higher output resolutions. In at least one embodiment, real-time temporal image reconstruction may be performed at a higher resolution than the initial resolution at which the rendering engine generates the image. In at least one embodiment, the temporal aspect of the process (which may be performed inside the upscaler in at least one embodiment) may involve blending the color values ​​of corresponding points between the current frame and at least one previous or historical frame in the sequence. In at least one embodiment, in order to ensure that such blending occurs for corresponding points on objects in these frames, this previous historical color data may be distorted according to the motion detected between this historical frame and the current frame, such as may be indicated by a set of motion vectors output from the rendering engine or otherwise determined. In at least one embodiment, such distortion may ensure that points (e.g., feature points) of various images are tracked over time and blended using corresponding color values, which may help reduce artifacts such as noise or flickering that occur during playback.

[0030] In at least one embodiment, and as discussed in more detail elsewhere herein, a supersampling algorithm may utilize a neural network to predict a blending factor to determine the amount to weight the color values ​​of a current pixel of a current frame and a corresponding historical pixel from a previously warped historical frame. In at least one embodiment, such an algorithm may also utilize a filter kernel to produce a new, higher resolution output image from a set of inputs. In at least one embodiment, the quality of the output image of such a network may depend at least in part on information available in the input, such as information that may include the current luma frame, historical luma, learned history and color variance masks or motion vector difference buffers. In at least one embodiment, an application may render an aliased 1 sample per pixel (spp) image at 1080p (FullHD) resolution, and the algorithm may reconstruct an anti-aliased 2160p (4k) image from the input image and any such side information sequence provided by the application. In at least one embodiment, such a process may be extended to other resolutions with other upsampling ratios, including the case of pure anti-aliasing with equal input and output resolutions.

[0031] The rendering and / or upsampling process may utilize sample locations 206 within each lower resolution pixel to determine the color of that pixel, as shown in image frame 200. Such sample locations may have at least a certain amount of random offset or jitter applied between frames to enable capture of fine or sub-pixel detail, as discussed in more detail later herein. The color of the image to be rendered may include a set of horizontal color bands, such as Figure 2A The sample position 206 shown will result in white or dark gray being sampled, but will miss sampling any such mid-gray that occupies a large portion of the image space of the current image or frame to be rendered. The color values ​​provided by the sampling may correspond to a subset of higher resolution pixels, such as Figure 2C As shown in the image 220 of FIG. A simple approach is to apply these colors to all (here four) higher resolution pixels corresponding to a single lower resolution pixel. However, this approach will lose at least some of the fine sub-pixel detail obtained using this dithered offset. Then, only or primarily these sampled color values ​​can be considered for the higher resolution pixels 222, 224 that contain one of these sampled points. This approach will cause most pixels (here represented by a cross-fill pattern) to have no color values ​​in the current image, so these pixels will not participate in the dithered perceived blending of the current input color and the distorted previous output color. These color values ​​can be blended with the colors from the previous frame 210, as shown in the image 220 of FIG. Figure 2B , where these colors may have been distorted, filtered, or otherwise processed, as discussed in more detail elsewhere herein. In at least one embodiment, Figure 2C The color value of the current frame is compared with the color value from Figure 2B The blending of distorted color values ​​of previous or historical frames can produce Figure 2D The image 230 shown. In at least one embodiment, this blending can preserve Figure 2C Some fine details shown in .

[0032] For example, in a component such as an amplifier, a spatiotemporal upsampling process may utilize a jittered input image and associated jitter values ​​and a set of low-resolution backward motion vectors for each input image pixel, and possibly other quantities such as exposure values ​​and a depth buffer, as part of an image reconstruction algorithm. In at least one embodiment, using these low-resolution input (backward) motion vectors, a high-resolution output image of a previous frame is warped to align with the geometry of the current time step. In at least one embodiment, based at least in part on the current input image and the warped previous frame output image, a neural network may be employed that may infer a set of anisotropic reconstruction kernel parameters for upsampling and filtering the current input image. In at least one embodiment, the neural network (or a separate neural network) may also infer one or more weighting factors for blending the upsampled input image with at least one warped previous output image. In at least one embodiment, the current input image is upsampled according to these predicted kernel parameters, and the high-resolution output image of the current frame is blended or composed of the current input color, the current input color upsampled by the anisotropic kernel, and the warped previous frame output color.

[0033] However, it is possible that a substantially linear texture pattern has a slight angle or offset relative to the pixel grid, such as Figure 2E 240 or projection of an example. In this example, a global pixel offset is applied so that the sample positions 242 are in the same relative position within each lower resolution pixel. If the pattern is perfectly aligned with the pixel grid, then each sample position in a given row (except at edges or boundaries) should generally return similar values. However, as shown, the slight angle between the texture pattern and the pixel grid row causes the sample positions along a given row of the pixel grid to be sampled from two different lines of the pattern, with some of the sample positions 242 being close to the transition between the line colors. As discussed, this can cause one or more artifacts, such as moiré patterns, to appear, particularly when combined with blending or anti-aliasing processes.

[0034] Methods according to at least one embodiment may attempt to introduce some randomness into the sampling locations of the various low-resolution pixels to avoid some issues with superposition of texture patterns on the pixel grid. However, as previously mentioned, it may be desirable to ensure that the sampling location for a given pixel remains within the boundaries of that pixel to avoid artifacts such as blurring that may result if sampling from adjacent pixels. One approach is to reduce the maximum radius of the jittered sample location from which a value may be sampled. However, in Figure 2E In the example of , the jittered excursions for this particular image or video frame are close to pixel edges, so the radius must be very small, which limits the effectiveness of the method because a lot of randomization cannot be applied from these sample points based in part on proximity to the edge.

[0035] Figure 2F and Figure 2G The effective effect of steps that can be used to introduce per-pixel randomness confined within the boundaries of these pixels is shown. It should be understood that in practice, such images or sample positions need not be determined in this manner, but can be calculated mathematically, in a single step, or in separate steps in a similar or alternative order, as described elsewhere herein. Furthermore, in at least some embodiments, the offsets can be applied to the texture coordinates rather than the sample positions, but the result is effectively a set of random sample positions within the respective pixel positions, such as Figure 2G shown.

[0036] As part of the randomness injection, the global jitter offset for a given image or video frame may be removed or otherwise accounted for in the sampler. This may have the effect of placing the sample position 252 substantially back to the center of the lower resolution pixels (at least the pixels associated with the identified texture), such as Figure 2F 250. If non-uniform dithering is applied to a grid of pixels, this may involve determining an offset for each pixel location and then removing (or otherwise accounting for) the corresponding offset. In the case where a global dithering offset is applied to all pixels of a grid of images, it may be sufficient to track a single dithering offset value for the grid and remove the same offset for each pixel location.

[0037] The advantage of effectively placing the sample positions of these pixels back to the pixel center is that the number of dimensions in which the sample points can then be randomized and still be within the pixel boundaries (assuming symmetric randomization boundaries in at least one embodiment) can be maximized. For example, if the sample points start at the pixel center, the randomness can occur over a function of + / - 0.5 pixel width and + / - 0.5 pixel height. Then, as Figure 2G As shown in the overlay 260 of , the updated sample location 262 can be randomly selected but constrained to be within a function of + / - 0.5 pixel width and depth from the center point, which effectively allows any location within a given pixel to be selected, allowing for maximum randomization while ensuring that it stays within the pixel boundaries. Figure 2G As shown, introducing randomness even on such a small selection of sample points can avoid introducing any minor but unintended patterns caused by superposition, because the randomness eliminates at least some of the effects of the regularity of fine patterns that might otherwise cause image artifacts.

[0038] In at least one embodiment, texture coordinates can be thought of as functions in screen space as follows:

[0039]

[0040] A method according to at least one embodiment may attempt to significantly randomize texture coordinates while constraining them to those locations (or at least primarily to those locations) within a given pixel boundary. This may be extended to a first-order polynomial function at the jittered pixel, such as a first-order Taylor polynomial function, as shown below:

[0041]

[0042] Here, the texture coordinate (tc) is determined in two pixel grid (or "screen") coordinates x and y. The global dither offset is given by x j and j Given, we can therefore pass (xx j ) and (yy j ) term to eliminate the offset. D represents the derivatives that can be taken from these terms to determine the appropriate offset value for a given pixel. Obtaining texture coordinate derivatives is relatively simple in a raster renderer. However, for other techniques such as ray tracing, it can be slightly more challenging, although still possible under various real-time rendering constraints.

[0043] As mentioned above, this approach can first effectively "de-jitter" the sample positions, or otherwise remove or account for global rendering jitter. A random offset can then be applied, such as applying a random offset to the texture coordinates that is within the pixel boundaries, such as by applying [-0.5, 0.5] 2 Random offsets within the range (normalized pixel width and height is 0.5 in each direction). In at least one embodiment, uniform random variables may be used, although various other distributions may be used within the scope of various embodiments.

[0044] In at least one embodiment, texture coordinates may be determined using a function such as:

[0045]

[0046] A texture coordinate is a location (x, y) within a pixel. The derivative is the derivative of the texture coordinate in screen / image space. In this example formula, α is a weight that can be applied when determining the appropriate texture coordinate. The sign used may vary in different cases depending in part on factors such as motion vector conventions and jitter. In at least one embodiment, the weight α can be determined by using a clamped logarithmic fit of the following form:

[0047]

[0048] in:

[0049] pixels = target height (targetHeight) * target width (targetWidth)

[0050] pixelsThreshold = 802300u (or another appropriate threshold)

[0051]

[0052] A=0.25

[0053] pixels=min(max(pixels,pixelsThreshold),pixelsThreshold<<4)

[0054] and

[0055] α=log2(pixels*B)*A

[0056] The image can be rendered at different resolutions; for example, the image can be rendered at a resolution selected based on the identified weights that retains the most detail while significantly removing image artifacts, and these weight values ​​can be fitted into a graph. The graph can correspond to a logarithmic function or other relatively simple graph function. The values ​​of the graph can then be used to select values ​​such as pixel thresholds. Although the example formula is described as being applicable to two-dimensional (2D) texture coordinates, this approach can be extended to three-dimensional (3D) texture coordinates, as well as directional coordinates that can be used for operations such as cubemap sampling, and other such options. Other methods that take into account jitter and limit the sampling points to pixel boundaries can also be used to select random points within the scope of various embodiments, although methods such as quad sampling may be less efficient for at least some operations. This functionality can be implemented when evaluating texture coordinates from barycentric coordinates, for example, a ray tracer will return the barycentric coordinates of a hit triangle mesh, which then need to be converted to texture coordinates using, for example, a linear relationship. This correction or randomization can be implemented in the process of obtaining the corresponding texture coordinates. This functionality can be implemented elsewhere for other types of rendering systems, such as when reading texture coordinates from a buffer in a raster-based engine.

[0057] This approach can attempt to completely eliminate or prevent artifacts (such as moiré patterns) at their source by randomizing the texture samples. The amount of randomization can be associated with the derivative of the texture coordinates to help ensure that values ​​that should correspond to neighboring pixels are not sampled. As mentioned above, the amount of randomization can also be (global) dither-aware. In the case of small amounts of noise introduced, many rendering pipelines or other such systems will be able to significantly reduce or eliminate the noise before providing the final image. Even in the presence of small amounts of noise, the noise is still much less disturbing than a moiré pattern in the same image.

[0058] In at least one embodiment, this randomization of sampling positions can be performed primarily in hardware. For example, the sampling hardware can be provided with appropriate gradients or derivatives and jitter values. The hardware can then determine the randomized texture coordinate shifts to use, and if the derivatives become too large, a weighting can be derived and applied, which may correspond to a lower resolution image generation state.

[0059] As described above, the advantages of artifact removal or reduction can be obtained for a variety of systems and use cases, such as temporal upsampling systems for deep learning based supersampling or super-resolution performance. In at least one embodiment, Figure 3 Components of one such system 300 are shown in , which can be used to perform image reconstruction or other such tasks. Components of such a system 300 can be implemented on one or more processing components, which are of similar or different types, including any of the types discussed herein. In at least one embodiment, a renderer 302, a rendering engine, or other such content generator can be used to generate content such as video game content or animation. The renderer 302 can receive an input of one or more frames of a sequence, and can use a storage content 304 (e.g., a map and graphics assets) modified at least in part based on the input to generate an image or video frame. The renderer 302 can be part of a rendering pipeline, such as rendering software (e.g., Unreal Engine 4 of Epic Games, Inc.), which can provide functions such as deferred shading, global illumination, lit translucency, post-processing, and graphics processing unit (GPU) particle simulation using vector fields.

[0060] The amount of processing required to render a complete high-resolution image may make it difficult to render these images or video frames to meet current frame rates, such as at least 60 frames per second (fps). In at least one embodiment, the renderer 302 may instead be used to generate a rendered image 310 having a resolution lower than one or more final output resolutions, for example to meet timing requirements and reduce processing resource requirements. The low-resolution rendered image 310 may be processed using an amplifier 312 to generate an upscaled image 316 that represents the content of the low-resolution rendered image 310 at a resolution equal to (or at least close to) the target output resolution. As described above, such a system may use (as part of the renderer 302 or as a separate component) a denoiser 306 to reduce the amount of noise that may be present in the lower resolution image 310 generated by the renderer, particularly for images generated using ray tracing or other such processes.

[0061] In this example, an amplifier 312 (which may take the form of a service, system, module, or device) may be used to upscale individual frames of a video or animation sequence. In at least one embodiment, the amount of upscaling to be performed may depend on the initial resolution of the rendered image and the target resolution of the display, such as from 1080p to 4k resolution. Additional processing may also be performed as part of the upsampling process, which may include anti-aliasing and temporal smoothing, for example. In at least one embodiment, an appropriate reconstruction filter may be used, which may involve a filter such as an anisotropic Gaussian filter or a dynamic filter network (DFN). An upsampling process may be used that will take into account sub-pixel jitter that may be applied on a per-frame basis.

[0062] In at least one embodiment, deep learning can be used to reason about a sequence of upsampled video frames. For example, temporal reconstruction can be used to provide anti-aliasing and super-resolution in a combined manner. Information from the corresponding sequence of video frames can be used to reason about a higher quality upsampled image. One or more heuristics based on prior knowledge of the rendering pipeline can be used without the need to learn from data. In at least one embodiment, this can include jitter-aware upsampling and accumulating samples at the upsampled resolution. The jitter offset data can be provided as input to an amplifier 312 including at least one neural network along with the current input video frame and the previously inferred frame to reason about a higher quality upscaled image 316 than would be produced by an upsampling algorithm alone. This upsampling essentially shifts the jitter offset 314 and per-frame samples so that they align with a history buffer that may be at a higher resolution.

[0063] In at least one embodiment, the enlarged image 316 can be provided as an input to a neural network 318 to determine one or more blending factors or blending weights. The neural network 318 can also receive as an input a previous high-resolution image in the sequence, which is distorted and provided to the neural network 318 together with the enlarged image 316. The neural network 318 can also receive other input features, which may be related to the spatial variation and temporal variation discussed herein. Deep learning can be used to reconstruct images for real-time rendering, with a resolution that is many times higher than the actual rendering resolution (e.g., two to nine times). The quality of the image reconstructed from such a process can be comparable to or even exceed the original resolution rendering in terms of at least detail, temporal stability, and lack of general artifacts (e.g., ghosting or lag). The neural network 318 can also determine at least some filtering to be applied when reconstructing or blending the current image with the previous image. In at least one embodiment, this information can then be provided to a blending component 320 together with the enlarged image 316 to be blended with at least one previous image of the sequence. A jitter offset 314 can also be provided as an input to the blending component 320. In at least one embodiment, this blending of the current image with the previous (or historical) image 322 of the sequence can help temporally converge to a nice, sharp, high-resolution output image 324, which can then be provided for presentation via a display 328 or other such presentation mechanism. In at least one embodiment, a copy of the high-resolution output image 324 can also be stored to a history buffer 326 or other such storage location for blending with subsequently generated images in the sequence. Such a process can utilize deep learning to reconstruct images for real-time rendering at a resolution that is several times (e.g., 2x, 4x, or 8x) higher than the actual rendering resolution, and the quality of the reconstructed image can be comparable to the original resolution rendering in terms of at least detail, temporal stability, and the absence of general artifacts such as ghosting or lag. The use of tensor cores can speed up the reconstruction, and the use of the methods described herein can make this rendering process more sample-efficient, thereby greatly improving the frames per second of various applications.

[0064] In such Figure 3 Using buffered information in the described system may involve, for example, Figure 4A. In at least one embodiment, three main input sources are utilized, including a color buffer 402, a motion vector buffer 404, and a depth buffer 406. In at least one embodiment, a preprocessor 408 (e.g., which may involve one or more processes running on one or more processors on one or more computing devices) may receive as input the color information of the current frame generated by a rendering engine or application, and the output of a warper 410 (e.g., a warping function or application executing one or more processors of one or more devices). In at least one embodiment, the warper 410 receives as input the motion vector information of the current frame stored in the motion vector buffer 404 and the depth information of the current frame stored in the depth buffer 406. In at least one embodiment, the warper 410 may receive the data directly from the application or renderer, and may not use a dedicated buffer. The temporal process may also use the high-resolution color data of the previous image in the sequence stored in the history buffer 414 as input to the warper 410. Information for each final output image may also be stored in the history buffer 414 for use in generating subsequent images or frames in the sequence. The warper 410 may utilize the motion vectors and depth data to warp the pixel data or color data of a particular feature of the previous image to a corresponding pixel location in the current image frame, effectively using the motion vectors to map the corresponding pixel locations of the features in the two images so that the color values ​​of similar features can be compared and blended. The preprocessor 408 may perform any relevant processing on the current color data from the color buffer 402 or the warped previous color data from the warper 410. After any preprocessing, the data may be provided as input to a neural network, such as a deep learning (DL) based generator 412, which may analyze the data to determine a pixel-specific weight for each pixel location in the image to be generated. The generated data may be processed by a post-processor 416, which may include one or more processes executed on one or more processors of one or more computing devices, and may output a final high-resolution color image 418. In at least one embodiment, the post-processor may also output information to be stored in a high-resolution color and history buffer 414 for use in generating subsequent images in the current sequence.

[0065] In at least one embodiment, generating a frame using this approach may involve an application providing a low-resolution jittered input image and associated jitter values, a low-resolution backward motion vector for each individual input image pixel, and other quantities (e.g., exposure values ​​and a depth buffer) to a reconstruction algorithm. These low-resolution input (backward) motion vectors may be used to warp the previous frame output image to align with the geometry in the current time step. An upsampling algorithm may be used to upsample the low-resolution current frame image (after any denoising and detail enhancement discussed herein) to the resolution of the high-resolution color image 418. A deep learning (DL)-based generator 412 may be used to infer a weighted value w (at the output resolution) for each output pixel. In at least one embodiment, a high-resolution output image of the current frame may be created as follows:

[0066] Output = w*(upsampled current frame input image) +

[0067] (1-w)*(the distorted previous output image)

[0068] In this type of temporal image reconstruction algorithm, an important factor in the quality (IQ) of the resulting image may be the weighting factor w described above. In at least one embodiment, w should accommodate various criteria, including at least that when an area in the output image is unoccluded due to object motion in the rendered scene, the weighting factor should favor the current input image, or give greater weight to the color values ​​of the current image, such as when w=1.0. When an area in the output image is visible (and similarly colored) in previous frames, the optimal weighting factor may result in an appropriate blend between these previous output images and the current input image. In at least one embodiment, this blending may be more favorable to historical data, such as when the value of w is close to zero because more frames have made the area visible.

[0069] The network can make the prediction weighting based at least in part on the current frame input image and the distorted previous frame output image. In at least one embodiment, whenever the upsampled current image has significantly different values ​​from the distorted previous frame output image, and therefore will appear very different when displayed, the neural network can predict a high value weighting factor w, thereby giving greater importance to the upsampled current frame input image. When the current image has similar values ​​to the distorted previous frame output image, and therefore will appear very similar when displayed, the neural network can predict a low value weighting factor w, thereby giving greater importance to the distorted previous frame output image.

[0070] In at least one embodiment, motion vector difference information can be used as an additional modality or input, as described herein. Additional buffers, such as motion buffer 420, can be used as another input source in such a system 400. In at least one embodiment, other buffers can be used, as described herein, for example, including at least one motion data buffer or depth data buffer. Motion buffer 420, also referred to herein as a history motion buffer, can store new or additional motion vector data, which can persist across frames. In at least one embodiment, the current motion vector from motion vector buffer 404 can be stored in one or more forms, such as corresponding to a transform process, for use in subsequent frames. This motion vector information can be provided as an additional input to warper 410. In at least one embodiment, the warping function of warper 410 can now warp not only the high-resolution color history data from history buffer 414, but also the previous motion vector data from motion buffer 420. Time calculator 422 can perform calculations in which the warped motion vector data from warper 410 is processed together with the current motion vector data from motion vector buffer 404 to determine the difference or difference area. This calculation may involve determining the difference, then determining the norm, and applying a correlation function as previously described. This temporal calculation may then be provided as an additional input to a preprocessor 408, which may then be passed to a deep learning (DL) based generator 412 for determining more accurate pixel specific weightings as described herein, which enables the DL based network to produce higher quality results.

[0071] Figure 4B An example image generation pipeline 450 is shown that can be used in system 400 (e.g. Figure 4A) to render one or more images, such as video frames in a sequence. In this example, an input frame (pixel data) 452 (which may include G-buffer data for a primary surface) for a current frame to be rendered may be received as input to a reflection and refraction component 454 of the rendering system. The reflection and refraction component 454 may use the data to attempt to determine data for any determined reflections and / or refractions in the pixel data, and may provide the data to a backprojection and G-buffer patching component 456, which may perform the backpropagation described herein to locate corresponding points for these reflections and refractions, and use the data to patch a G-buffer 468, which may provide updated input for subsequent frames to be rendered. The data may then be provided to a light sample generation component 458 to perform light sampling, the data may be provided to a ray tracing lighting component 460 to perform ray tracing lighting, and the data may be provided to one or more shaders 462, which may set the pixel color of individual pixels of the frame based at least in part on the determined lighting information (as well as other information, such as color, texture, etc.). The results may be accumulated by an accumulation module 464 or component to generate an output frame 466 of a desired size, resolution, or format.

[0072] In at least one embodiment, the shader 462 may perform a backprojection step. As described herein, randomization of texture coordinates of finely detailed textures may be performed in the shader 462. Once the backward projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer may perform a lighting pass. Using information from the lighting pass and the lighting results from the previous frame, gradients may be calculated and then filtered and used for history rejection. This approach may be used to calculate robust temporal gradients between the current frame and the previous frame in a temporal denoiser of a ray tracing renderer. This backprojection-based approach may also work through reflections and refractions and may work with a rasterized G-buffer. Previous backprojection approaches omitted any G-buffer patching and instead relied on the original current G-buffer samples, which also resulted in false positive gradients. Patching the surface parameters may eliminate false positives in the vast majority of cases, making the denoised image very stable but still able to react quickly to lighting changes. Once the backward projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer may perform a lighting pass. Using information from the lighting pass and the lighting result from the previous frame, gradients are computed and then filtered and used for history rejection.

[0073] Figure 5An example process 500 is shown for reducing the likelihood of anti-aliasing artifacts in a rendered image, which can be performed according to at least one embodiment. It should be understood that for this process and other processes described herein, more, fewer or alternative steps or similar or alternative orders, or at least partially in parallel, can be performed within the scope of various embodiments, unless otherwise explicitly stated. In addition, although this example will be discussed with respect to a patterned texture, there can be various other objects or components that can be sampled to determine pixel values ​​that can also be used within the scope of various embodiments. In this example, a texture to be sampled for a pixel of an image to be rendered (or otherwise generated) is identified 502. In a system that applies jittering to preserve fine details, a jitter offset applied to a current pixel can be determined 504. In many systems, a global jitter value will be applied to all pixels of the image to be rendered, while in other systems, multiple jitter values ​​may be used, such as different jitter values ​​for each pixel, and other such options. In this example, any such jitter offset can be removed 506, or otherwise considered when determining a randomized sample position that remains within a pixel boundary, because failure to consider the jitter offset may cause the sample position to return a value that should be associated with a neighboring pixel.

[0074] In addition to accounting for jitter, one or more texture coordinates of the texture to be sampled may be shifted 508 by a random amount. This shifting of the texture coordinates effectively shifts a given pixel relative to the sample position of the texture. The shift may be constrained to remain within the boundaries of the pixel, which may be determined at least in part based on the derivative of the texture coordinates. The texture may then be sampled 510 at the determined sample position relative to the shifted texture coordinates to determine a sampled pixel value for the pixel. It may be determined 512 whether more pixel values ​​are to be determined for the image, and if so, the process may continue with the next pixel to be sampled. A different randomized sample position may be selected for the next pixel, which will still be constrained to be within the boundaries of the pixel. Once samples of all pixel positions of the image are obtained, the image may be rendered 514 using these sample values. In at least some embodiments, other processes may also be performed during post-processing, such as anti-aliasing, noise reduction, etc. If the image is an image in a sequence (e.g., a sequence of video frames), the process may be repeated for the next image or frame.

[0075] Aspects of the various methods described herein may be lightweight enough to be performed in a variety of locations, such as in real time on a client device, such as a personal computer or a gaming console. Such processing may be performed on content generated on or received by the client device or received from an external source, such as streaming data or other content received from a cloud server 620 or a third-party service 660 over at least one network, among other such options. In some cases, at least a portion of the processing, generation, synthesis, and / or determination of the content may be performed by one of these other devices, systems, or entities and then provided to the client device (or another such recipient) for presentation or other such use.

[0076] As an example, Figure 6An example network configuration 600 is shown that can be used to provide, generate, modify, encode, process and / or transmit image data or other such content. In at least one embodiment, a client device 602 can generate or receive data for a session using components of a content application 604 on the client device 602 and data stored locally on the client device. In at least one embodiment, a content application 624 executed on a cloud server 620 (e.g., a cloud server or an edge server) can initiate a session associated with at least one client device 602, can utilize a session manager and user data stored in a user database 636, and can cause a content manager 626 to determine content (e.g., one or more digital assets (e.g., implicit and / or explicit object representations)) from a content repository 634. The content manager 626 can collaborate with a rendering engine 628 to generate or select objects, digital assets, or other such content to be placed in a virtual environment and allow it to be moved or manipulated in the environment. Views of these objects can be rendered by the rendering engine 628 and provided by the client device 602 for presentation. In at least one embodiment, the rendering engine 628 may work with (or include) an image processing module 630 that may perform processing on the rendered image, which may also call a texture sampling process to randomly shift the texture coordinates of the texture pattern 632 to avoid introducing artifacts into the image to be rendered by the rendering engine 628. At least a portion of the rendered and / or processed content may be transmitted to the client device 602 using an appropriate transmission manager 622 for transmission via download, streaming, or another such transmission channel. An encoder may be used to encode and / or compress at least a portion of such data before transmitting it to the client device 602. In at least one embodiment, the client device 602 receiving such content may provide the content to a corresponding content application 604, which may also or alternatively include a renderer 610, a processor 612, and a mixer 614 for providing, synthesizing, rendering, composing, modifying, or using the content for presentation (or other purposes) on or by the client device 602. The decoder may also be used to decode data received via the network(s) 640 for presentation via the client device 602, such as images or video content presented via a display 606 and audio (e.g., sounds and music) presented via at least one audio playback device 608 (e.g., speakers or headphones). In at least one embodiment, at least a portion of the content may already be stored, rendered, or accessible to the client device 602, so at least the portion of the content does not need to be transmitted via the network 640, such as the content may have been previously downloaded or stored locally on a hard drive or optical disk.In at least one embodiment, the content may be transmitted from the cloud server 620 or user database 636 to the client device 602 using a transmission mechanism (e.g., data streaming). In at least one embodiment, at least a portion of the content may be obtained, enhanced, and / or streamed from another source (e.g., a third-party service 660 or other client device 650), which may also include a content application 662 for generating, enhancing, or providing content. In at least one embodiment, multiple computing devices or multiple processors within one or more computing devices (e.g., which may include a combination of CPUs and GPUs) may be used to perform portions of this functionality.

[0077] In this example, these client devices may include any appropriate computing devices, such as desktop computers, laptops, set-top boxes, streaming devices, game consoles, smartphones, tablet computers, VR headsets, AR goggles, wearable computers, or smart TVs. Each client device may submit a request across at least one wired or wireless network, which may include the Internet, Ethernet, a local area network (LAN), or a cellular network, as well as other such options. In this example, these requests may be submitted to an address associated with a cloud provider, which may operate or control one or more electronic resources in a cloud provider environment, such as a data center or a server farm. In at least one embodiment, the request may be received or processed by at least one edge server located at the edge of the network and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling client devices to interact with servers that are closer, while also improving the security of resources in the cloud provider environment.

[0078] In at least one embodiment, such a system may be used to perform graphics rendering operations. In other embodiments, such a system may be used for other purposes, such as for providing image or video content to test or verify autonomous machine applications, or for performing deep learning operations. In at least one embodiment, such a system may be implemented using an edge device, or may be combined with one or more virtual machines (VMs). In at least one embodiment, such a system may be implemented at least in part in a data center or at least in part using cloud computing resources.

[0079] Reasoning and training logic

[0080] Fig. 7A Inference and / or training logic 715 is shown for performing reasoning and / or training operations associated with one or more embodiments. Fig. 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided.

[0081] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequence, where weights and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores input / output data during training and / or inference using aspects of one or more embodiments and / or weight parameters during forward propagation of each layer of a neural network trained or used in conjunction with one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 may be included within other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0082] In at least one embodiment, any portion of code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 may be cache memory, dynamic random access memory ("DRAM"), static random access memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 701 is internal or external to a processor, for example, or consists of DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training of the neural network, or some combination of these factors.

[0083] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 705 for storing reverse and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, during training and / or inference using aspects of one or more embodiments, the code and / or data storage 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during back propagation of input / output data and / or weight parameters. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or sequence, wherein weights and / or other parameter information is loaded to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 705 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 705 is internal or external to the processor, for example, whether it is composed of DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.

[0084] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0085] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations based at least in part on or as directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from a layer or neuron within a neural network) stored in activation storage 720, which is a function of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 710 to generate activations stored in activation storage 720, where weight values ​​stored in code and / or data storage 701 and / or code and / or data storage 705 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 701 or code and / or data storage 705 or other on-chip or off-chip storage.

[0086] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 710 may be outside a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by an execution unit of a processor, which may be within the same processor or distributed between different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 720 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0087] In at least one embodiment, activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 720 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, activation storage 720 may be selected to be internal or external to a processor, for example, or include DRAM, SRAM, flash memory, or other storage types, depending on the storage available on-chip or off-chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferencing and / or training neural networks, or some combination of these factors. In at least one embodiment, Fig. 7A The inference and / or training logic 715 shown in FIG. 7 may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "LakeCrest") processor. In at least one embodiment, Fig. 7A The illustrated inference and / or training logic 715 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0088] Figure 7B Inference and / or training logic 715 is shown in accordance with at least one or more embodiments. In at least one embodiment, the reasoning and / or training logic 715 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown in FIG. 7 may be used in conjunction with an application specific integrated circuit (ASIC), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "LakeCrest") processor. In at least one embodiment, Figure 7BThe inference and / or training logic 715 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 7B In at least one embodiment shown in , each of code and / or data storage 701 and code and / or data storage 705 is associated with a dedicated computing resource (e.g., computing hardware 702 and computing hardware 706), respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) solely on the information stored in code and / or data storage 701 and code and / or data storage 705, respectively, and the results of performing the functions are stored in activation storage 720.

[0089] In at least one embodiment, each of the code and / or data stores 701 and 705 and the corresponding computing hardware 702 and 706 corresponds to a different layer of the neural network, such that activations from one "storage / compute pair 701 / 702" of the code and / or data store 701 and computing hardware 702 are provided as inputs to the next "storage / compute pair 705 / 706" of the code and / or data store 705 and computing hardware 706, so as to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 715 after or in parallel with the storage / compute pairs 701 / 702 and 705 / 706.

[0090] Data Center

[0091] Figure 8 An example data center 800 is shown in which at least one embodiment may be used. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.

[0092] In at least one embodiment, Figure 8As shown, the data center infrastructure layer 810 may include a resource coordinator 812, grouped computing resources 814, and node computing resources ("node CRs") 816(1)-816(N), where "N" represents any positive integer. In at least one embodiment, the node CRs 816(1)-816(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NWI / O") devices, network switches, virtual machines ("VMs"), power modules and cooling modules, etc. In at least one embodiment, one or more of the node CRs 816(1)-816(N) may be a server having one or more of the above computing resources.

[0093] In at least one embodiment, the grouped computing resources 814 may include a separate grouping (not shown) of node CRs housed in one or more racks, or many racks (also not shown) housed in data centers at various geographic locations. The separate grouping of node CRs within the grouped computing resources 814 may include computing, networks, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0094] In at least one embodiment, resource coordinator 812 may configure or otherwise control one or more node CRs 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a software design infrastructure ("SDI") management entity for data center 800. In at least one embodiment, resource coordinator 812 may include hardware, software, or some combination thereof.

[0095] In at least one embodiment, Figure 8As shown, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 may include a framework that supports software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. In at least one embodiment, the software 832 or the application 842 may include a web-based service software or application, such as a service or application provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be, but is not limited to, a free and open source software network application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can use the distributed file system 828 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 832 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820 including Spark and a distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 can manage cluster or group computing resources mapped to or allocated to support the distributed file system 828 and the job scheduler 822. In at least one embodiment, the cluster or group computing resources can include group computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 can coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.

[0096] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the node CRs 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0097] In at least one embodiment, one or more applications 842 included in the application layer 840 may include one or more types of applications used by at least a portion of the node CRs 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0098] In at least one embodiment, any of the configuration manager 824, resource manager 826, and resource coordinator 812 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of the data center 800 from making potentially bad configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0099] In at least one embodiment, the data center 800 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 800. In at least one embodiment, by using weight parameters calculated by one or more training techniques described herein, information may be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to the data center 800.

[0100] In at least one embodiment, the data center can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.

[0101] Reasoning and / or training logic 715 is used to perform reasoning and / or training operations associated with one or more embodiments. Fig. 7A and / or Figure 7BProvides details about the reasoning and / or training logic 715. In at least one embodiment, the reasoning and / or training logic 715 may be implemented in the system Figure 8 for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0102] Such a component may be used to provide randomized sampling of fine textures that takes into account jitter offsets and is constrained to positions within the boundaries of each corresponding pixel.

[0103] Computer Systems

[0104] Fig. 9 900, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, in accordance with the present disclosure, such as the embodiments described herein, the computer system 900 may include, but is not limited to, components, such as a processor 902, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, the computer system 900 may include a processor, such as a processor available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 900 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0105] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (Internet Protocol) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0106] In at least one embodiment, the computer system 900 may include, but is not limited to, a processor 902, which may include, but is not limited to, one or more execution units 908 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 900 is a single-processor desktop or server system, but in another embodiment, the computer system 900 may be a multi-processor system. In at least one embodiment, the processor 902 may include, but is not limited to, a complex instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word computing ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 902 may be coupled to a processor bus 910, which may transmit data signals between the processor 902 and other components in the computer system 900.

[0107] In at least one embodiment, processor 902 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache 904 may reside external to processor 902. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, register file 906 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.

[0108] In at least one embodiment, logic execution unit 908 that performs integer and floating point operations is also located in processor 902, including, but not limited to. In at least one embodiment, processor 902 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, execution unit 908 may include logic for processing a packed instruction set 909. In at least one embodiment, by including a packed instruction set 909 in the instruction set of a general purpose processor, and associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in processor 902. In one or more embodiments, many multimedia applications may be executed faster and more efficiently by using the full width of the processor's data bus to perform operations on packed data, which may not require the transfer of smaller units of data on the processor's data bus to perform one or more operations one data element at a time.

[0109] In at least one embodiment, execution unit 908 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 900 may include, but is not limited to, memory 920. In at least one embodiment, memory 920 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 920 may store instructions 919 and / or data 921 represented by data signals that may be executed by processor 902.

[0110] In at least one embodiment, the system logic chip can be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 916, and the processor 902 can communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 can provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 can initiate data signals between the processor 902, the memory 920, and other components in the computer system 900, and bridge data signals between the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 can be coupled to the memory 920 via a high-bandwidth memory path 918, and the graphics / video card 912 can be coupled to the MCH 916 via an Accelerated Graphics Port ("AGP") interconnect 914.

[0111] In at least one embodiment, the computer system 900 may use a system I / O 922, which is a proprietary hub interface bus to couple the MCH 916 to an I / O controller hub ("ICH") 930. In at least one embodiment, the ICH 930 may provide direct connection to certain I / O devices through a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 920, the chipset, and the processor 902. Examples may include, but are not limited to, an audio controller 929, a firmware hub ("Flash BIOS") 928, a wireless transceiver 926, a data store 924, a traditional I / O controller 923 including a user input and keyboard interface 925, a serial expansion port 927 (e.g., a universal serial bus (USB) port), and a network controller 934. The data store 924 may include a hard drive, a floppy drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0112] In at least one embodiment, Fig. 9 A system comprising interconnected hardware devices or "chips" is shown, while in other embodiments, Fig. 9An exemplary system on chip (SoC) may be shown. In at least one embodiment, the devices may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 are interconnected using a compute express link (CXL) interconnect.

[0113] The reasoning and / or training logic 715 is used to perform reasoning and / or training operations related to one or more embodiments. Fig. 7A and / or Figure 7B Provide details about the reasoning and / or training logic 715. In at least one embodiment, the reasoning and / or training logic 715 may be Fig. 9 for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0114] Such a component may be used to provide randomized sampling of fine textures that takes into account jitter offsets and is constrained to positions within the boundaries of each corresponding pixel.

[0115] Fig.10 1 is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0116] In at least one embodiment, system 1000 may include, but is not limited to, a processor 1010 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Fig.10 A system is shown that includes interconnected hardware devices or "chips", while in other embodiments, Fig.10 An exemplary system on chip (SoC) may be shown. In at least one embodiment, Fig.10 The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, Fig.10One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0117] In at least one embodiment, Fig.10 It may include a display 1024, a touch screen 1025, a touch pad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, a fast chipset (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash memory (“BIOS, FW Flash”) 1022, a DSP 1060, a drive 1020 (e.g., a solid state disk (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) 1055, a camera (“USB 3.0 camera”) 1054 (e.g., a USB 3.0 camera), and / or a low power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented with, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.

[0118] In at least one embodiment, other components may be communicatively coupled to the processor 1010 through the components described above. In at least one embodiment, the accelerometer 1041, ambient light sensor (“ALS”) 1042, compass 1043, and gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, the thermal sensor 1039, fan 1037, keyboard 1036, and touchpad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, the speaker 1063, earphone 1064, and microphone (“mic”) 1065 may be communicatively coupled to the audio unit (“audio codec and class D amplifier”) 1062, which in turn may be communicatively coupled to the DSP 1060. In at least one embodiment, the audio unit 1062 may include, for example, but not limited to, an audio encoder / decoder (“codec”) and a class D amplifier. In at least one embodiment, the SIM card (“SIM”) 1057 may be communicatively coupled to the WWAN unit 1056. In at least one embodiment, components such as the WLAN unit 1050 and the Bluetooth unit 1052 and the WWAN unit 1056 may be implemented as a next generation form factor (NGFF).

[0119] The reasoning and / or training logic 715 is used to perform reasoning and / or training operations associated with one or more embodiments. Fig. 7A and / or Figure 7BProvide details about the reasoning and / or training logic 715. In at least one embodiment, the reasoning and / or training logic 715 may be Fig.10 Systems for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0120] Such a component may be used to provide randomized sampling of fine textures that takes into account jitter offsets and is constrained to positions within the boundaries of each corresponding pixel.

[0121] Fig.11 1100 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, processing system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1102 or processor cores 1107. In at least one embodiment, processing system 1100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.

[0122] In at least one embodiment, the processing system 1100 may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console for gaming and media consoles. In at least one embodiment, the processing system 1100 is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1100 may also include a wearable device coupled to or integrated in a wearable device, such as a smart watch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.

[0123] In at least one embodiment, one or more processors 1102 each include one or more processor cores 1107 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a specific instruction group 1109. In at least one embodiment, the instruction group 1109 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or calculate by very long instruction words (VLIW). In at least one embodiment, one or more processor cores 1107 can each process different instruction groups 1109, which can include instructions that help emulate other instruction groups. In at least one embodiment, one or more processor cores 1107 can also include other processing devices, such as digital signal processors (DSPs).

[0124] In at least one embodiment, one or more processors 1102 include cache memory 1104. In at least one embodiment, one or more processors 1102 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared between various components of one or more processors 1102. In at least one embodiment, one or more processors 1102 also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared between one or more processor cores 1107 using known cache coherence techniques. In at least one embodiment, one or more processors 1102 additionally include a register file 1106, and the processor may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 1106 may include general registers or other registers.

[0125] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transmit communication signals, such as address, data, or control signals, between one or more processors 1102 and other components in the processing system 1100. In at least one embodiment, the one or more interface buses 1110 may be a processor bus, such as a version of a direct media interface (DMI) bus, in one embodiment. In at least one embodiment, the one or more interface buses 1110 are not limited to a DMI bus, and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between memory devices and other components of the processing system 1100, while the platform controller hub (PCH) 1130 provides connections to I / O devices through a local I / O bus.

[0126] In at least one embodiment, the memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have appropriate performance to be used as a processor memory. In at least one embodiment, the memory device 1120 may be used as a system memory of the processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute applications or processes. In at least one embodiment, the memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 in the one or more processors 1102 to perform graphics and media operations. In at least one embodiment, a display device 1111 may be connected to the processor 1102. In at least one embodiment, the display device 1111 may include one or more of the internal display devices, such as in a mobile electronic device or laptop device or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1111 may include a head mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0127] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to be connected to the memory device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, a data storage device 1124 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or long-term evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1128 enables communication with the system firmware and can be, for example, a unified extensible firmware interface (UEFI). In at least one embodiment, the network controller 1134 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to one or more interface buses 1110. In at least one embodiment, the audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1100. In at least one embodiment, the platform controller hub 1130 can also be connected to one or more universal serial bus (USB) controllers 1142, which connect input devices such as a keyboard and mouse 1143 combination, a camera 1144, or other USB input devices.

[0128] In at least one embodiment, instances of the memory controller 1116 and the platform controller hub 1130 may be integrated into a discrete external graphics processor, such as the external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 may be external to one or more processors 1102. For example, in at least one embodiment, the processing system 1100 may include the external memory controller 1116 and the platform controller hub 1130, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 1102.

[0129] The reasoning and / or training logic 715 is used to perform reasoning and / or training operations associated with one or more embodiments. Fig. 7Aand / or Figure 7B Detail is provided regarding the inference and / or training logic 715. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the processing system 1100. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in a graphics processor. Additionally, in at least one embodiment, the inference and / or training operations described herein may use a processor other than the ALU. Fig. 7A and / or Figure 7B In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0130] Such a component may be used to provide randomized sampling of fine textures that takes into account jitter offsets and is constrained to positions within the boundaries of each corresponding pixel.

[0131] Fig.12 is a block diagram of a processor 1200 having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208 in accordance with at least one embodiment. In at least one embodiment, the processor 1200 may include additional cores, up to and including the additional core 1202N represented by the dashed box. In at least one embodiment, each of the one or more processor cores 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1206.

[0132] In at least one embodiment, one or more internal cache units 1204A-1204N and one or more shared cache units 1206 represent a cache memory hierarchy within the processor 1200. In at least one embodiment, the one or more cache memory units 1204A-1204N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as LLC. In at least one embodiment, cache coherency logic maintains coherency between the various cache units 1206 and 1204A-1204N.

[0133] In at least one embodiment, the processor 1200 may also include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, the one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1210 provides management functions for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown).

[0134] In at least one embodiment, one or more processor cores 1202A-1202N include support for multiple threads simultaneously. In at least one embodiment, system agent core 1210 includes components for coordinating and operating one or more processor cores 1202A-1202N during multithreaded processing. In at least one embodiment, system agent core 1210 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of one or more processor cores 1202A-1202N and graphics processor 1208.

[0135] In at least one embodiment, the processor 1200 also includes a graphics processor 1208 for performing graphics processing operations. In at least one embodiment, the graphics processor 1208 is coupled to one or more shared cache units 1206 and a system agent core 1210 including one or more integrated memory controllers 1214. In at least one embodiment, the system agent core 1210 also includes a display controller 1211 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1211 may also be a separate module coupled to the graphics processor 1208 via at least one interconnect, or may be integrated within the graphics processor 1208.

[0136] In at least one embodiment, a ring-based interconnect unit 1212 is used to couple the internal components of the processor 1200. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1208 is coupled to the ring interconnect 1212 via an I / O link 1213.

[0137] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 1218 (e.g., eDRAM modules). In at least one embodiment, each of one or more processor cores 1202A-1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last level cache.

[0138] In at least one embodiment, one or more processor cores 1202A-1202N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, one or more processor cores 1202A-1202N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1202A-1202N execute a common instruction set, while one or more other cores in one or more processor cores 1202A-1202N execute a subset or a different instruction set of the common instruction set. In at least one embodiment, one or more processor cores 1202A-1202N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, processor 1200 can be implemented on one or more chips or implemented as a SoC integrated circuit.

[0139] The reasoning and / or training logic 715 is used to perform reasoning and / or training operations associated with one or more embodiments. Fig. 7A and / or Figure 7B Detail is provided regarding the inference and / or training logic 715. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in Fig.12 In addition, in at least one embodiment, the inference and / or training operations described herein may use the inference and / or training operations described herein. Fig. 7A and / or Figure 7B In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1200 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0140] Such a component may be used to provide randomized sampling of fine textures that takes into account jitter offsets and is constrained to positions within the boundaries of each corresponding pixel.

[0141] Virtualized computing platform

[0142] Fig.13 is an example data flow diagram of a process 1300 for generating and deploying an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 1300 can be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. The process 1300 can be executed within a training system 1304 and / or a deployment system 1306. In at least one embodiment, the training system 1304 can be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use with the deployment system 1306. In at least one embodiment, the deployment system 1306 can be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements of the facility 1302. In at least one embodiment, one or more applications in the pipeline can use or call services (e.g., reasoning, visualization, computation, AI, etc.) of the deployment system 1306 during application execution.

[0143] In at least one embodiment, some applications used in the high-level processing and reasoning pipeline may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained at the facility 1302 using data 1308 (e.g., imaging data) generated at the facility 1302 (and stored on one or more picture archiving and communication system (PACS) servers at the facility 1302), the machine learning model may be trained using imaging or sequencing data 1308 from another one or more facilities, or a combination thereof. In at least one embodiment, the training system 1304 may be used to provide applications, services, and / or other resources to generate a working, deployable machine learning model for the deployment system 1306.

[0144] In at least one embodiment, the model registry 1324 can be backed by an object store, which can support versioning and object metadata. In at least one embodiment, the object store can be accessed from within the cloud platform through, for example, a cloud storage compatible application programming interface (API). In at least one embodiment, the machine learning models within the model registry 1324 can be uploaded, listed, modified, or deleted by developers or partners of the system interacting with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate a model with an application so that the model can be executed as part of the execution of a containerized instantiation of the application.

[0145] In at least one embodiment, training pipeline 1304 ( Fig.13 ) may include situations where the facility 1302 is training their own machine learning model, or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device, a sequencing device, and / or other type of device may be received. In at least one embodiment, once the imaging data 1308 is received, the AI-assisted annotation 1310 may be used to help generate annotations corresponding to the imaging data 1308 to be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 1310 may include one or more machine learning models (e.g., a convolutional neural network (CNN)) that may be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, the AI-assisted annotation 1310 may then be used directly, or may be adjusted or fine-tuned using an annotation tool to generate ground truth data. In at least one embodiment, the AI-assisted annotation 1310, the labeled data 1312, or a combination thereof may be used as ground truth data for training a machine learning model. In at least one embodiment, the trained machine learning model may be referred to as output model(s) 1316 and may be used by the deployment system 1306 as described herein.

[0146] In at least one embodiment, the training pipeline may include a scenario where the facility 1302 requires a machine learning model for performing one or more processing tasks for one or more applications in the deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model may be selected from the model registry 1324. In at least one embodiment, the model registry 1324 may include machine learning models that are trained to perform a variety of different reasoning tasks on imaging data. In at least one embodiment, the machine learning models in the model registry 1324 may be trained on imaging data from a different facility (e.g., a facility located far away) rather than the facility 1302. In at least one embodiment, the machine learning model may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training is performed on imaging data from a specific location, the training may be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data from off-site. In at least one embodiment, once a model is trained or partially trained at one location, the machine learning model can be added to the model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in the model registry 1324. In at least one embodiment, the machine learning model can then be selected from the model registry 1324 (and referred to as the output model 1316) and can be deployed in the deployment system 1306 to perform one or more processing tasks for one or more applications of the deployment system.

[0147] In at least one embodiment, the scenario may include a facility 1302 that requires a machine learning model for performing one or more processing tasks for one or more applications in the deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model). In at least one embodiment, the machine learning model selected from the model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at the facility 1302 due to population differences, robustness of the training data used to train the machine learning model, diversity of training data anomalies, and / or other issues with the training data. In at least one embodiment, AI-assisted annotations 1310 can be used to help generate annotations corresponding to the imaging data 1308 for use as ground truth data for training or updating the machine learning model. In at least one embodiment, labeled clinical data 1312 can be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI-assisted annotation 1310, labeled data 1312, or a combination thereof) can be used as ground truth data to retrain or update a machine learning model. In at least one embodiment, the trained machine learning model can be referred to as an output model 1316 and can be used by the deployment system 1306, as described herein.

[0148] In at least one embodiment, the deployment system 1306 may include software 1318, services 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, the deployment system 1306 may include a software "stack" such that the software 1318 may be built on top of the services 1320 and may use the services 1320 to perform some or all of the processing tasks, and the services 1320 and software 1318 may be built on top of the hardware 1322 and use the hardware 1322 to perform the processing, storage, and / or other computing tasks of the deployment system. In at least one embodiment, the software 1318 may include any number of different containers, each of which may perform an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks (e.g., reasoning, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and reasoning pipeline. In at least one embodiment, in addition to receiving and configuring imaging data for use by each container and / or containers used by facility 1302 after processing through the pipeline, a high-level processing and reasoning pipeline can also be defined based on the selection of different containers desired or required to process imaging data 1308 (e.g., to convert the output back into a usable data type. In at least one embodiment, the combination of containers within software 1318 (e.g., which constitute a pipeline) can be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument can utilize services 1320 and hardware 1322 to perform some or all of the processing tasks of an application instantiated in a container.

[0149] In at least one embodiment, the data processing pipeline can receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of the deployment system 1306). In at least one embodiment, the input data can represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data can be pre-processed as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing can be performed on the output of one or more inference tasks or other processing tasks of the pipeline to prepare output data for the next application and / or prepare the output data for transmission and / or use by the user (e.g., as a response to the inference request). In at least one embodiment, the inference task can be performed by one or more machine learning models, such as a trained or deployed neural network, which can include the output model 1316 of the training system 1304.

[0150] In at least one embodiment, tasks of a data processing pipeline may be encapsulated in containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, containers or applications may be published to a private (e.g., limited access) area of ​​a container registry (described in more detail herein), and trained or deployed models may be stored in the model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) may be available in a container registry, and once a user selects an image from a container registry for deployment in a pipeline, the image may be used to generate an instantiated container for the application for use by the user's system.

[0151] In at least one embodiment, a developer (e.g., a software developer, a clinician, a physician, etc.) can develop, publish, and store applications (e.g., as containers) for performing image processing and / or reasoning on provided data. In at least one embodiment, the development, publishing, and / or storage can be performed using a software development kit (SDK) associated with the system (e.g., to ensure that the developed applications and / or containers conform to or are compatible with the system). In at least one embodiment, the developed applications can be tested locally (e.g., at a first facility, on data from the first facility) using the SDK, which is installed as part of the system (e.g., Fig.12 Processor 1200 in the process 1300) may support at least some services 1320. In at least one embodiment, because a DICOM object may contain from one to hundreds of images or other data types, and because the data varies, the developer may be responsible for managing (e.g., setting up constructs for building pre-processing into the application, etc.) the extraction and preparation of the incoming data. In at least one embodiment, once validated by process 1300 (e.g., for accuracy), the application is made available in the container registry for selection and / or implementation by the user to perform one or more processing tasks on the data at the user's facility (e.g., a second facility).

[0152] In at least one embodiment, the developer can then share the application or container over a network for use by a system (e.g., Fig.131300). In at least one embodiment, the completed and validated application or container can be stored in the container registry, and the associated machine learning model can be stored in the model registry 1324. In at least one embodiment, the requesting entity (which provides the reasoning or image processing request) can browse the container registry and / or the model registry 1324 to obtain applications, containers, data sets, machine learning models, etc., select the desired combination of elements to be included in the data processing pipeline, and submit the image processing request. In at least one embodiment, the request may include the input data necessary to execute the request (and in some examples, patient-related data), and / or may include a selection of applications and / or machine learning models to be executed when processing the request. In at least one embodiment, the request can then be passed to one or more components (e.g., a cloud) of the deployment system 1306 to perform processing of the data processing pipeline. In at least one embodiment, the processing performed by the deployment system 1306 may include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or the model registry 1324. In at least one embodiment, once the results are generated by the pipeline, the results may be returned to the user for reference (eg, for viewing in a viewing application suite executing locally, on a local workstation or terminal).

[0153] In at least one embodiment, to assist in processing or executing applications or containers in the pipeline, services 1320 may be utilized. In at least one embodiment, services 1320 may include computing services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, services 1320 may provide functionality common to one or more applications in software 1318, and thus may abstract functionality into services that may be called or utilized by applications. In at least one embodiment, the functionality provided by services 1320 may run dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using parallel computing platform 1230 ( Fig.12)) to scale well. In at least one embodiment, rather than requiring that each application that shares the same functionality provided by the service 1320 must have a corresponding instance of the service 1320, the service 1320 can be shared between and among various applications. In at least one embodiment, as a non-limiting example, the service may include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included that can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data enhancement service may be further included that can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other enhancements. In at least one embodiment, a visualization service may be used that can add image rendering effects (e.g., ray tracing, rasterization, denoising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service may be included that provides beamforming, segmentation, reasoning, imaging, and / or support for other applications within the pipeline of the virtual instrument.

[0154] In at least one embodiment, where the service 1320 includes an AI service (e.g., an inference service), as part of the execution of an application, one or more machine learning models can be executed by calling (e.g., as an API call) an inference service (e.g., an inference server) to execute one or more machine learning models or processing thereof. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can call the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, the software 1318 implementing the high-level processing and inference pipeline, which includes a segmentation application and anomaly detection application, can be pipelined because each application can call the same inference service to perform one or more inference tasks.

[0155] In at least one embodiment, the hardware 1322 may include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient, purpose-built support for the software 1318 and services 1320 in the deployment system 1306. In at least one embodiment, the use of GPU processing may be implemented to perform local processing (e.g., at the facility 1302) within the AI / deep learning system, in the cloud system, and / or other processing components of the deployment system 1306 to improve the efficiency, accuracy, and effectiveness of image processing and generation. In at least one embodiment, as a non-limiting example, with respect to deep learning, machine learning, and / or high-performance computing, the software 1318 and / or services 1320 may be optimized for GPU processing. In at least one embodiment, at least some of the computing environments of the deployment system 1306 and / or training system 1304 may be executed in a data center, one or more supercomputers, or high-performance computer systems with GPU-optimized software (e.g., a combination of hardware and software of an NVIDIA DGX system). In at least one embodiment, as described herein, hardware 1322 may include any number of GPUs that may be called to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU optimized execution for deep learning tasks, GPU processing for machine learning tasks or other computing tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU optimized software (e.g., as provided on NVIDIA's DGX system) may be used as a hardware abstraction and scaling platform to execute a cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform may integrate an application container cluster system or coordination system (e.g., KUBERNETES) on multiple GPUs to achieve seamless scaling and load balancing.

[0156] Fig.14 is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, the system 1400 can be used to implement Fig.13 The process 1300 and / or other processes of the system 1400 may include a high-level processing and reasoning pipeline. In at least one embodiment, the system 1400 may include a training system 1304 and a deployment system 1306. In at least one embodiment, the training system 1304 and the deployment system 1306 may be implemented using software 1318, services 1320, and / or hardware 1322, as described herein.

[0157] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) can be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 can be implemented locally (with respect to a healthcare service facility) or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1426 can be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol can include a network token, which can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, the API of the virtual instrument (described herein) or other instances of system 1400 can be restricted to a set of public IPs that have been audited or authorized for interaction.

[0158] In at least one embodiment, the various components of system 1400 can communicate with each other using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communications between facilities and components of system 1400 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be transmitted via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0159] In at least one embodiment, similar to the present disclosure regarding Fig.13 As described, the training system 1304 can execute one or more training pipelines 1404. In at least one embodiment, where the deployment system 1306 will use one or more machine learning models in one or more deployment pipelines 1410, the one or more training pipelines 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, as a result of the one or more training pipelines 1404, one or more output models 1316 can be generated. In at least one embodiment, the one or more training pipelines 1404 can include any number of processing steps, such as, but not limited to, conversion or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by the deployment system 1306. In at least one embodiment, similar to the description regarding Fig.13 The one or more training pipelines 1404 of the first example described may be used for a first machine learning model, similar to the one described with respect to Fig.13The one or more training pipelines 1404 of the second example described may be used for a second machine learning model, similar to the one or more training pipelines 1404 of the second example described. Fig.13 The one or more training pipelines 1404 of the third example described may be used for a third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 may be used according to the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models may have been trained and are ready for deployment, so the training system 1304 may not perform any processing on the machine learning model, and one or more machine learning models may be implemented by the deployment system 1306.

[0160] In at least one embodiment, the output model 1316 and / or the pre-trained model 1406 may include any type of machine learning model, depending on the implementation or embodiment. In at least one embodiment and without limitation, the machine learning model used by the system 1400 may include using linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., autoencoder, convolution, recursion, perceptron, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machine, etc.), and / or other types of machine learning models.

[0161] In at least one embodiment, one or more training pipelines 1404 may include AI-assisted annotation, as described herein with respect to at least Fig.14In more detail. In at least one embodiment, the labeled clinical data 1312 (e.g., traditional annotations) can be generated by any number of techniques. In at least one embodiment, in some examples, labels or other annotations can be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of application suitable for generating ground-truth annotations or labels, and / or can be hand-drawn. In at least one embodiment, the ground-truth data can be synthetically generated (e.g., generated from a computer model or rendering), truly generated (e.g., designed and generated from real-world data), machine-generated (e.g., using feature analysis and learning to extract features from the data and then generate labels), manually annotated (e.g., a labeler or annotation expert, defining the location of the label), and / or a combination thereof. In at least one embodiment, for each instance of imaging data 1308 (or other data types used by machine learning models), there can be corresponding ground-truth data generated by the training system 1304. In at least one embodiment, AI-assisted annotation 1310 can be performed as part of a deployment pipeline 1410; supplementing or replacing the AI-assisted annotation 1310 included in the training pipeline 1404. In at least one embodiment, the system 1400 may include a multi-layer platform that may include a software layer (e.g., software 1318) of a diagnostic application (or other application type) that may perform one or more medical imaging and diagnostic functions. In at least one embodiment, the system 1400 may be communicatively coupled (e.g., via an encrypted link) to a PACS server network of one or more facilities. In at least one embodiment, the system 1400 may be configured to access and reference data from a PACS server to perform operations such as training a machine learning model, deploying a machine learning model, image processing, reasoning, and / or other operations.

[0162] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or authenticated API through which an application or container may be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, the application may then call or execute one or more services 1320 to perform computational, AI, or visualization tasks associated with the respective application, and the software 1318 and / or services 1320 may utilize hardware 1322 to perform processing tasks in an effective and efficient manner. In at least one embodiment, a pair of DICOM adapters 1402A, 1402B may be used to send communications to or receive communications by the training system 1304 and the deployment system 1306.

[0163] In at least one embodiment, the deployment system 1306 may execute one or more deployment pipelines 1410. In at least one embodiment, the one or more deployment pipelines 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) - including AI-assisted annotations, generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, as described herein, the deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scanning instrument, a virtual sequencing instrument, etc.). In at least one embodiment, there may be more than one deployment pipeline 1410 for a single device, depending on the information desired from the data generated by the device. In at least one embodiment, in the case where an abnormality is desired to be detected from an MRI machine, there may be one or more first deployment pipelines 1410, and in the case where image enhancement is desired from the output of the MRI machine, there may be one or more second deployment pipelines 1410.

[0164] In at least one embodiment, the image generation application may include a processing task that includes the use of a machine learning model. In at least one embodiment, a user may wish to use their own machine learning model, or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in an application that performs a processing task. In at least one embodiment, the application may be selectable and customizable, and by defining the construction of the application, the deployment and implementation of the application for a particular user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of the system 1400 (e.g., services 1320 and hardware 1322), one or more deployment pipelines 1410 may be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.

[0165] In at least one embodiment, deployment system 1306 may include a user interface ("UI") 1414 (e.g., a graphical user interface, a web interface, etc.) that may be used to select applications to be included in deployment pipeline 1410, to arrange applications, to modify or change applications or their parameters or configuration, to use and interact with deployment pipeline 1410 during setup and / or deployment, and / or to otherwise interact with deployment system 1306. In at least one embodiment, although not shown with respect to training system 1304, UI 1414 (or a different user interface) may be used to select models to use in deployment system 1306, to select models for training or retraining in training system 1304, and / or to otherwise interact with training system 1304.

[0166] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 may be used to manage the interaction between the application or container of the deployment pipeline 1410 and the service 1320 and / or the hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate the interaction from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although shown as included in the software 1318, this is not intended to be limiting, and in some examples, the pipeline manager 1412 may be included in the service 1320. In at least one embodiment, the application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) may include a container coordination system that can group applications into containers as a logical unit for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from the deployment pipeline 1410 (e.g., rebuilding applications, splitting applications, etc.) with individual containers, each application can be executed in a self-contained environment (e.g., at the kernel level) to improve speed and efficiency.

[0167] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed separately (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer), which can allow for focus and attention on the tasks of a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can assist in communication and collaboration between different containers or applications. In at least one embodiment, as long as the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container), the application coordination system 1428 and / or the pipeline manager 1412 can facilitate communication and sharing of resources between and among each application or container. In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, load balance, and determine the sharing of services or resources between and among the various applications or containers. In at least one embodiment, the scheduler can be used to track resource requirements of applications or containers, current or planned use of those resources, and resource availability. Thus, in at least one embodiment, the scheduler can allocate resources to different applications and allocate resources between and among applications, taking into account the needs and availability of the system. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0168] In at least one embodiment, the services 1320 utilized and shared by the applications or containers in the deployment system 1306 may include computing services 1416, AI services 1418, visualization services 1420, and / or other service types. In at least one embodiment, an application may call (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize computing services 1416 to perform supercomputing or other high performance computing (HPC) tasks. In at least one embodiment, one or more computing services 1416 may be utilized to perform parallel processing (e.g., using a parallel computing platform 1430) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, a parallel computing platform 1430 (e.g., NVIDIA's CUDA) may implement general computing on a GPU (GPGPU) (e.g., GPU / graphics 1422). In at least one embodiment, the software layer of the parallel computing platform 1430 may provide access to a virtual instruction set and parallel computing elements of a GPU to execute computing kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, memory may be shared between and among multiple containers, and / or between and among different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within a container to enable the same data (e.g., where multiple different stages of an application or multiple applications are processing the same information) to be used for a shared memory segment from the parallel computing platform 1430. In at least one embodiment, instead of copying data and moving the data to different locations in memory (e.g., read / write operations), the same data in the same location of the memory may be used for any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, since the data as a result of processing is used to generate new data, this information of the new location of the data may be stored and shared between various applications. In at least one embodiment, the location of the data and the location of the updated or modified data may be part of the definition of how to understand the payload in the container.

[0169] In at least one embodiment, one or more AI services 1418 may be utilized to perform reasoning services for executing machine learning models associated with an application (e.g., the task is to execute one or more processing tasks of the application). In at least one embodiment, one or more AI services 1418 may utilize an AI system 1424 to execute machine learning models (e.g., neural networks such as CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other reasoning tasks. In at least one embodiment, an application of the deployment pipeline 1410 may use one or more output models 1316 and / or other models of the application for the self-training system 1304 to perform reasoning on imaging data. In at least one embodiment, two or more examples of reasoning using an application coordination system 1428 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high priority / low latency path that can achieve a higher service level agreement, such as for performing reasoning on an urgent request in an emergency, or for a radiologist during a diagnostic process. In at least one embodiment, a second category may include a standard priority path that may be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, application coordination system 1428 can allocate resources (e.g., services 1320 and / or hardware 1322) for different reasoning tasks of AI service 1418 based on priority paths.

[0170] In at least one embodiment, the shared memory may be mounted to one or more AI services 1418 in the system 1400. In at least one embodiment, the shared memory may operate as a cache (or other storage device type) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of the deployment system 1306 may receive the request and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, in order to process the request, the request may be entered into a database, and if it is not already in the cache, the machine learning model may be located from the model registry 1324, a verification step may ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not yet running or there are not enough instances of the application, a scheduler (e.g., a scheduler of the pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if the inference server has not yet been started to execute the model, the inference server may be started. Any number of inference servers may be started for each model. In at least one embodiment, in a pull model that clusters inference servers, the model can be cached whenever load balancing is beneficial. In at least one embodiment, the inference servers can be statically loaded into the corresponding distributed servers.

[0171] In at least one embodiment, inference can be performed using an inference server running in a container. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally with multiple versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request is received to perform inference on a model, a new instance can be loaded. In at least one embodiment, when the inference server is started, the model can be passed to the inference server so that the same container can be used to serve different models as long as the inference server is running as a different instance.

[0172] In at least one embodiment, during application execution, a request for reasoning for a given application may be received, and a container (e.g., an instance of a hosting reasoning server) may be loaded (if not already loaded), and a launcher may be called. In at least one embodiment, preprocessing logic in the container may load, decode, and / or perform any additional preprocessing on the incoming data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is ready for reasoning, the container may perform reasoning on the data as needed. In at least one embodiment, this may include a single reasoning call for one image (e.g., a hand X-ray), or may require reasoning for hundreds of images (e.g., a chest CT). In at least one embodiment, the application may summarize the results before completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities may be assigned to different models or applications. For example, some models may have a real-time (TAT less than 1 minute) priority, while other models may have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, model execution time may be measured from a requesting mechanism or entity and may include collaborative network traversal time as well as execution time of an inference service.

[0173] In at least one embodiment, the transmission of requests between the service 1320 and the reasoning application can be hidden behind a software development kit (SDK), and robust transmission can be provided through queues. In at least one embodiment, requests will be placed in queues through an API for individual application / tenant ID combinations, and the SDK will pull requests from the queue and provide the request to the application. In at least one embodiment, the name of the queue can be provided in the environment from which the SDK will pick up the queue. In at least one embodiment, asynchronous communication through queues may be useful because it can allow any instance of the application to pick up work when it is available. Results can be transmitted back through the queue to ensure that no data is lost. In at least one embodiment, the queue can also provide the ability to split the work, because the highest priority work can enter the queue connected to most instances of the application, and the lowest priority work can enter the queue connected to a single instance, which processes the tasks in the order received. In at least one embodiment, the application can run on a GPU-accelerated instance generated in the cloud 1426, and the reasoning service can perform reasoning on the GPU.

[0174] In at least one embodiment, visualization services 1420 may be utilized to generate visualizations for viewing application and / or deployment pipeline 1410 outputs. In at least one embodiment, visualization services 1420 may utilize GPU / graphics 1422 to generate visualizations. In at least one embodiment, visualization services 1420 may implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations may include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, and the like. In at least one embodiment, a virtualized environment may be used to generate a virtual interactive display or environment (e.g., a virtual environment) for system users (e.g., doctors, nurses, radiologists, etc.) to interact with. In at least one embodiment, visualization services 1420 may include internal visualizers, movies, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).

[0175] In at least one embodiment, hardware 1322 may include GPU / graphics 1422, AI system 1424, cloud 1426, and / or any other hardware for executing training system 1304 and / or deployment system 1306. In at least one embodiment, GPU / graphics 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs that may be used to perform processing tasks for any feature or functionality of compute services 1416, AI services 1418, visualization services 1420, other services, and / or software 1318. For example, for AI services 1418, GPU / graphics 1422 may be used to perform pre-processing on imaging data (or other data types used by machine learning models), post-processing on the output of machine learning models, and / or perform inference (e.g., to execute machine learning models). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 may use GPU / graphics 1422. In at least one embodiment, cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 may use a GPU, and cloud 1426 (or at least part of a task being deep learning or reasoning) may be executed using one or more AI systems 1424. Likewise, while hardware 1322 is shown as discrete components, this is not intended to be limiting, and any component of hardware 1322 may be combined with or utilized by any other component of hardware 1322.

[0176] In at least one embodiment, AI system 1424 may include a purpose-built computing system (e.g., a supercomputer or HPC) configured for reasoning, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPU, RAM, storage, and / or other components, features, or functions, AI system 1424 (e.g., NVIDIA's DGX) may also include software (e.g., a software stack) that can use multiple GPUs / graphics 1422 to perform GPU-optimized tasks. In at least one embodiment, one or more AI systems 1424 may be implemented in a cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of system 1400.

[0177] In at least one embodiment, cloud 1426 may include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC) that may provide a GPU-optimized platform for executing processing tasks of system 1400. In at least one embodiment, cloud 1426 may include an AI system 1424 for executing one or more AI-based tasks of system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1426 may be integrated with an application coordination system 1428 that utilizes multiple GPUs to enable seamless scaling and load balancing between and among applications and services 1320. In at least one embodiment, cloud 1426 may be responsible for executing at least some services 1320 of system 1400, including one or more compute services 1416, one or more AI services 1418, and / or one or more visualization services 1420, as described herein. In at least one embodiment, cloud 1426 can perform large and small batch inference (e.g., executing NVIDIA's TENSORRT), provide a parallel computing platform 1430 (e.g., NVIDIA's CUDA), execute an application coordination system 1428 (e.g., Kubernetes), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher quality movie effects), and / or can provide other functionality for system 1400.

[0178] Fig.15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, a non-limiting example may be used. Fig.14The process 1500 may be performed by the system 1400 of FIG. 1400. In at least one embodiment, the process 1500 may utilize services and / or hardware as described herein. In at least one embodiment, the refined model 1512 generated by the process 1500 may be executed by the deployment system for one or more containerized applications in the deployment pipeline.

[0179] In at least one embodiment, model training 1514 may include retraining or updating an initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data (such as customer data set 1506), and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset, deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously fine-tuned parameters (e.g., weights and / or biases) retained from previous training, so training or retraining may not take as long or require as much processing as training the model from scratch. In at least one embodiment, during model training 1514, by resetting or replacing the output or loss layer of the initial model 1504, when generating predictions on the new customer data set 1506, the parameters for the new data set may be updated and re-adjusted based on the loss calculation associated with the accuracy of the output or loss layer.

[0180] In at least one embodiment, the pre-trained model 1506 may be stored in a data store or registry. In at least one embodiment, the pre-trained model 1506 may have been trained at least in part at one or more facilities other than the facility performing process 1500. In at least one embodiment, in order to protect the privacy and rights of patients, subjects, or customers of different facilities, the pre-trained model 1506 may have been trained locally using locally generated customer or patient data. In at least one embodiment, the cloud and / or other hardware may be used to train the pre-trained model 1306, but confidential, privacy-protected patient data may not be transmitted to, used by, or accessed by any component of the cloud (or other non-local hardware). In at least one embodiment, if the pre-trained model 1506 is trained using patient data from more than one facility, the pre-trained model 1506 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, for example where customer or patient data has been issued privacy issues (e.g., by being relinquished, used for experimental purposes, etc.), or where the customer or patient data is included in a public dataset, customer or patient data from any number of facilities can be used to train pre-trained model 1506 locally and / or externally, such as in a data center or other cloud computing infrastructure.

[0181] In at least one embodiment, when selecting an application to use in a deployment pipeline, a user may also select a machine learning model to use for a particular application. In at least one embodiment, a user may not have a model to use, so the user may select a pre-trained model 1506 to use with the application. In at least one embodiment, the pre-trained model may not be optimized to generate accurate results on the customer dataset 1506 of the user facility (e.g., based on patient diversity, demographics, type of medical imaging equipment used, etc.). In at least one embodiment, the pre-trained model 1506 may be updated, retrained, and / or fine-tuned for use at various facilities before deploying the pre-trained model into a deployment pipeline for use with one or more applications.

[0182] In at least one embodiment, a user may select a pre-trained model 1506 to be updated, retrained, and / or fine-tuned, and the pre-trained model may be referred to as the initial model 1504 for training the system in process 1500. In at least one embodiment, a customer data set 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by equipment at a facility) may be used to perform model training (which may include, but is not limited to, transfer learning) on ​​the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to the customer data set 1506 may be generated by the model training system 1304. In at least one embodiment, the ground truth data may be generated, at least in part, at the facility by a clinician, scientist, physician, practitioner.

[0183] In at least one embodiment, AI-assisted annotation 1310 may be used in some examples to generate ground truth data. In at least one embodiment, AI-assisted annotation 1310 (e.g., implemented using an AI-assisted annotation SDK) may utilize a machine learning model (e.g., a neural network) to generate ground truth data for recommendations or predictions for a customer data set. In at least one embodiment, a user may use the annotation tool within a user interface (graphical user interface (GUI)) on a computing device.

[0184] In at least one embodiment, user 1510 can interact with the GUI via computing device 1508 to edit or fine-tune annotations or automatic annotations. In at least one embodiment, a polygon editing feature can be used to move vertices of a polygon to a more precise or fine-tuned position.

[0185] In at least one embodiment, once the customer dataset 1506 has associated ground truth data, the ground truth data (e.g., from AI-assisted annotations 1310, manual labeling, etc.) can be used during model training to generate a refined model 1512. In at least one embodiment, the customer dataset 1506 can be applied to the initial model 1504 any number of times, and the ground truth data can be used to update the parameters of the initial model 1504 until an acceptable level of accuracy is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, the refined model 1512 can be deployed within one or more deployment pipelines at a facility for use in performing one or more processing tasks with respect to medical imaging data.

[0186] In at least one embodiment, the refined model 1512 can be uploaded to the pre-trained models 1542 in the model registry for selection by another facility. In at least one embodiment, this process can be completed at any number of facilities, so that the refined model 1512 can be further refined any number of times on new data sets to generate a more general model.

[0187] Fig. 15B is an example illustration of a client-server architecture 1532 for enhancing an annotation tool with a pre-trained model 1542 according to at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 can be instantiated based on the client-server architecture 1532. In at least one embodiment, the AI-assisted annotation tool 1536 in an imaging application can assist a radiologist, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application can include a software tool that, as a non-limiting example, helps a user 1510 identify several extreme points on a specific organ of interest in an original image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of the specific organ. In at least one embodiment, the results can be stored in a data store as training data 1538 and used as, for example, but not limited to, ground truth data for training. In at least one embodiment, when the computing device 1508 sends extreme points for AI-assisted annotation 1310, for example, a deep learning model can receive the data as input and return an inference result for segmenting an organ or anomaly. In at least one embodiment, the pre-instantiated annotation tool (e.g., Fig. 15B The AI-assisted annotation tools 1536 in the annotation assistant server 1540 can be enhanced by making an API call (e.g., API call 1544) to a server (such as an annotation assistant server 1540), which may include a set of pre-trained models 1542 stored in, for example, an annotation model registry. In at least one embodiment, the annotation model registry can store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that are pre-trained to perform AI-assisted annotation 1310 on specific organs or abnormalities. In at least one embodiment, these models can be further updated by using a training pipeline. In at least one embodiment, the pre-installed annotation tools can be improved over time as new annotated data is added.

[0188] Various embodiments may be described by the following terms:

[0189] 1. A computer-implemented method comprising:

[0190] identifying a texture to sample corresponding to a pixel of an image to be rendered;

[0191] shifting one or more texture coordinates of the texture by a random amount selected to constrain the texture coordinates at a sample location to remain within the boundaries of the pixel;

[0192] sampling the texture at the sample location using a shader to determine a sample value for the pixel; and

[0193] The image is rendered using the sample values ​​of the pixels of the image.

[0194] 2. A computer-implemented method according to clause 1, wherein the random amount of the shift is determined using a first order Taylor polynomial.

[0195] 3. A computer-implemented method according to clause 2, wherein the first-order Taylor polynomial is given by the following formula:

[0196]

[0197] 4. The computer-implemented method of clause 1, further comprising:

[0198] determining a global dithering amount to apply to the image to be rendered; and

[0199] The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

[0200] 5. The computer-implemented method of clause 1, further comprising:

[0201] The random amount for the shifting is selected based at least in part on one or more texture coordinate derivatives.

[0202] 6. A computer-implemented method according to clause 1, wherein the sampling is performed as part of a light transport simulation process or a rasterization process.

[0203] 7. A computer-implemented method according to clause 1, wherein the sampled hit point is determined by tracing a ray to the pixel, and wherein the shifting of the one or more texture coordinates occurs before determining the texture coordinate corresponding to the hit point to be used.

[0204] 8. The computer-implemented method of clause 1, wherein a weight is applied to one or more texture coordinates to determine the random amount for the shifting.

[0205] 9. The computer-implemented method of clause 1, wherein the shift is calculated using a clamped logarithmic function.

[0206] 10. A processor, comprising:

[0207] One or more processing units for:

[0208] Determines the amount of global dithering applied to the image being rendered;

[0209] shifting one or more texture coordinates of a texture by a random amount that removes the global jitter and ensures that texture sample points of a pixel of the image remain within the boundaries of the pixel; and

[0210] The texture of the pixel is sampled at the texture sample point to determine a pixel value to be used for the image to be rendered.

[0211] 11. The processor of clause 10, wherein the random amount is determined based at least in part on one or more texture coordinate derivatives.

[0212] 12. The processor of claim 10, wherein the random amount of the shift is determined using a first order Taylor polynomial.

[0213] 13. A processor according to clause 10, wherein the one or more processing units are further configured to:

[0214] determining an amount of the global dithering to apply to the image to be rendered; and

[0215] The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

[0216] 14. A processor according to clause 10, wherein a hit point for the sample is determined by tracing a ray for the pixel, and wherein the shifting of the one or more texture coordinates occurs before determining the texture coordinate corresponding to the hit point to be used.

[0217] 15. A processor according to clause 10, wherein the processor is comprised of at least one of:

[0218] A system for performing simulation operations;

[0219] Systems for performing simulated operations to test or validate autonomous machine applications;

[0220] Systems for performing digital twin operations;

[0221] A system for performing light transport simulations;

[0222] Systems for rendering graphical output;

[0223] Systems for performing deep learning operations;

[0224] Systems implemented using edge devices;

[0225] Systems for generating or presenting virtual reality (VR) content;

[0226] Systems for generating or presenting augmented reality (AR) content;

[0227] Systems for generating or presenting mixed reality (MR) content;

[0228] A system comprising one or more virtual machines VM;

[0229] A system implemented at least in part in a data center;

[0230] Systems that perform hardware testing using simulation;

[0231] Systems for synthetic data generation;

[0232] A system that uses large language models (LLMs) to perform generative AI operations.

[0233] A collaborative content creation platform for 3D assets; or

[0234] A system implemented at least in part using cloud computing resources.

[0235] 16. A system comprising:

[0236] One or more processors for partially rendering an image by shifting one or more texture coordinates of a texture to be sampled of the image by a random amount that accounts for global dithering and ensures that texture sampling points for pixels of the image remain within the boundaries of the pixel.

[0237] 17. The system of clause 16, wherein the random amount is determined based in part on one or more texture coordinate derivatives.

[0238] 18. The system of clause 16, wherein said random amount of said shift is determined using a first order Taylor polynomial.

[0239] 19. A system according to clause 16, wherein the one or more processing units are further configured to:

[0240] determining the global dithering amount to apply to the image to be rendered; and

[0241] The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

[0242] 20. A system according to clause 16, wherein the system comprises at least one of the following:

[0243] A system for performing simulation operations;

[0244] Systems for performing simulated operations to test or validate autonomous machine applications;

[0245] Systems for performing digital twin operations;

[0246] A system for performing light transport simulations;

[0247] Systems for rendering graphical output;

[0248] Systems for performing deep learning operations;

[0249] A system for performing generative AI operations using large language models (LLMs),

[0250] Systems implemented using edge devices;

[0251] Systems for generating or presenting virtual reality (VR) content;

[0252] Systems for generating or presenting augmented reality (AR) content;

[0253] Systems for generating or presenting mixed reality (MR) content;

[0254] A system comprising one or more virtual machines VM;

[0255] A system implemented at least in part in a data center;

[0256] Systems that perform hardware testing using simulation;

[0257] Systems for synthetic data generation;

[0258] A collaborative content creation platform for 3D assets; or

[0259] A system implemented at least in part using cloud computing resources.

[0260] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of the present disclosure as defined by the appended claims.

[0261] Unless otherwise noted or clearly contradictory to the context, in the context of describing the disclosed embodiments (particularly in the context of the appended claims), the use of the terms "one" and "an" and "the" and similar references should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include", "have", "include" and "contain" should be interpreted as open terms (meaning "including but not limited to") unless otherwise noted. The term "connected" (which refers to a physical connection when unmodified) should be interpreted as partially or completely included, attached to or connected together, even if there are some interventions. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, respectively, and each individual value is incorporated into the specification as if it were individually described herein. Unless otherwise noted or contradictory to the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set including one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.

[0262] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B and C" are understood in context to be generally used to refer to an item, clause, or the like that may be A or B or C, or any non-empty subset of the set A and B and C. For example, in the illustrative example of a set having three members, the conjunction phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunction language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. In addition, unless expressly indicated otherwise or contradicted by context, the term "plurality" refers to a plural state (e.g., "plurality of items" means a plurality of items). The number of items in a plurality of items is at least two, but may be more if expressly indicated or indicated by context. Further, the phrase "based on" means "based at least in part on" rather than "based solely on" unless otherwise specified or clear from context.

[0263] Unless otherwise indicated herein or clearly contradictory to the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are jointly executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of, for example, a computer program that includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all the code, but the plurality of non-transitory computer-readable storage media stores all the code together. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.

[0264] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. In addition, a computer system that implements at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices that operate in different ways, so that the distributed computer system performs the operations described herein, and so that a single device does not perform all operations.

[0265] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential to practicing the disclosure.

[0266] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0267] In the specification and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. On the contrary, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0268] Unless explicitly stated otherwise, it is to be understood that throughout the specification, terms such as “processing”, “computing”, “calculating”, “determining” and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers or other such information storage, transmission or display devices of the computing system.

[0269] In a similar manner, the term "processor" may refer to any device or part of a memory that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.

[0270] In this document, reference may be made to obtaining, acquiring, receiving or inputting analog or digital data into a subsystem, a computer system or a computer-implemented machine. It is possible to obtain, acquire, receive or input analog and digital data in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. It is also possible to reference to providing, outputting, transmitting, sending or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending or presenting analog or digital data can be accomplished by transmitting data as an input or output parameter of a function call, an application programming interface or an interprocess communication mechanism.

[0271] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. In addition, although specific responsibilities are defined above for discussion purposes, various functions and responsibilities may be allocated and divided in different ways, depending on the circumstances.

[0272] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.

Claims

1. A computer-implemented method comprising: identifying a texture to sample corresponding to a pixel of an image to be rendered; shifting one or more texture coordinates of the texture by a random amount selected to constrain the texture coordinates at a sample location to remain within the boundaries of the pixel; sampling the texture at the sample location using a shader to determine a sample value for the pixel; as well as The image is rendered using the sample values ​​of the pixels of the image.

2. The computer-implemented method of claim 1, wherein the random amount of the shift is determined using a first order Taylor polynomial.

3. The computer-implemented method of claim 2, wherein the first-order Taylor polynomial is given by the following formula:

4. The computer-implemented method of claim 1 , further comprising: determining a global dithering amount to apply to the image to be rendered; as well as The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

5. The computer-implemented method of claim 1 , further comprising: The random amount for the shifting is selected based at least in part on one or more texture coordinate derivatives.

6. The computer-implemented method of claim 1, wherein the sampling is performed as part of a light transport simulation process or a rasterization process.

7. The computer-implemented method of claim 1 , wherein the sampled hit point is determined by tracing a ray for the pixel, and wherein the shifting of the one or more texture coordinates occurs prior to determining the texture coordinate corresponding to the hit point to use.

8. The computer-implemented method of claim 1, wherein a weight is applied to one or more texture coordinates to determine the random amount for the shifting.

9. The computer implemented method of claim 1, wherein the shift is calculated using a clamped logarithmic function.

10. A processor, comprising: One or more processing units for: Determines the amount of global dithering applied to the image being rendered; shifting one or more texture coordinates of a texture by a random amount that removes the global jitter and ensures that a texture sample point of a pixel of the image remains within the boundaries of the pixel; as well as The texture of the pixel is sampled at the texture sample point to determine a pixel value to be used for the image to be rendered.

11. The processor of claim 10, wherein the random amount is determined based at least in part on one or more texture coordinate derivatives.

12. The processor of claim 10, wherein the random amount of the shift is determined using a first order Taylor polynomial.

13. The processor of claim 10, wherein the one or more processing units are further configured to: determining an amount of the global dithering to apply to the image to be rendered; and The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

14. The processor of claim 10, wherein a hit point for the sample is determined by tracing a ray for the pixel, and wherein the shifting of the one or more texture coordinates occurs prior to determining the texture coordinate corresponding to the hit point to use.

15. The processor of claim 10, wherein the processor is included in at least one of the following: A system for performing simulation operations; Systems for performing simulated operations to test or validate autonomous machine applications; Systems for performing digital twin operations; A system for performing light transport simulations; Systems for rendering graphical output; Systems for performing deep learning operations; Systems implemented using edge devices; Systems for generating or presenting virtual reality (VR) content; Systems for generating or presenting augmented reality (AR) content; Systems for generating or presenting mixed reality (MR) content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; Systems that perform hardware testing using simulation; Systems for synthetic data generation; A system that uses large language models (LLMs) to perform generative AI operations. A collaborative content creation platform for 3D assets; or A system implemented at least in part using cloud computing resources.

16. A system comprising: One or more processors for partially rendering an image by shifting one or more texture coordinates of a texture to be sampled of the image by a random amount that accounts for global dithering and ensures that texture sampling points for pixels of the image remain within the boundaries of the pixel.

17. The system of claim 16, wherein the random amount is determined based in part on one or more texture coordinate derivatives.

18. The system of claim 16, wherein the random amount of the shift is determined using a first order Taylor polynomial.

19. The system of claim 16, wherein the one or more processing units are further configured to: determining the global dithering amount to apply to the image to be rendered; and The global jitter amount is removed prior to shifting the one or more texture coordinates of the texture by the random amount.

20. The system of claim 16, wherein the system comprises at least one of the following: A system for performing simulation operations; Systems for performing simulated operations to test or validate autonomous machine applications; Systems for performing digital twin operations; A system for performing light transport simulations; Systems for rendering graphical output; Systems for performing deep learning operations; A system for performing generative AI operations using large language models (LLMs), Systems implemented using edge devices; Systems for generating or presenting virtual reality (VR) content; Systems for generating or presenting augmented reality (AR) content; Systems for generating or presenting mixed reality (MR) content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; Systems that perform hardware testing using simulation; Systems for synthetic data generation; A collaborative content creation platform for 3D assets; or a system implemented at least in part using cloud computing resources.