Adaptive Temporal Image Filtering for Rendering Photorealistic Lighting
By positioning the matching surface and performing parameter repair through reverse projection technology, the problem of difficulty in dealing with reflected and refracted images in the prior art is solved, and high-quality rendering effects are achieved.
Patent Information
- Application Number
- CN202111490344.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-19
- Filing Date
- 2021-12-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The prior art is difficult to effectively reduce artifacts in real-time rendering, especially when processing reflected and refracted images, and conventional methods cannot effectively handle these complex optical effects.
The matching surface is positioned through reverse projection technology, surface parameter loading and patching is performed to calculate a robust time gradient, reduce artifacts and smooth animations.
Effective processing of reflected and refracted images is realized, reducing artifacts, and improving rendering quality and stability.
Smart Images

Figure CN114627234B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 123,540, filed December 10, 2020, entitled "Adaptive Temporal Image Filtering with Gradient Estimation Based on Backward Projection of Surfaces and G-Buffer Patching", the entire content of which is incorporated herein by reference for all purposes. Background Art
[0003] With the effort towards physically based rendering, stochastic sampling of shading (e.g., using path tracing) is becoming increasingly important in real-time rendering. To achieve high performance in real-time and / or under other time-sensitive or processing-sensitive conditions, sampling can be limited to a lower sample count. However, a lower sample count may result in noise, which makes it necessary to use complex reconstruction filters to produce better quality results. Research on such filters has shown dramatic improvements in both quality and performance, as they utilize the coherence of consecutive frames by reusing temporal information to achieve stable, denoised results. However, existing temporal filters often produce objectionable artifacts such as ghosting and lag.
[0004] A temporal gradient corresponding to the difference between the (noisy) shading results of the current and previous frames from a subset of the surfaces visible on the screen can be used to attempt to reduce the number of artifacts in a rendered image sequence or frame sequence. A conventional method for computing the temporal gradient is to use the visibility buffer of the previous frame - a buffer that stores the geometric information of each pixel. Some pixels from the previous frame are forward projected into the current frame, and they replace the regular pixels in the current frame. For these gradient pixels, the primary ray is traced to the reprojected surface, rather than being traced to a direction determined solely by the pixel position for a fixed distance as is usually the case. Unfortunately, this method is ineffective for rasterized geometry buffers or "G-buffers" (i.e., the screen space representation of geometric and material information), because one cannot have the rasterizer use a custom sub-pixel offset for each pixel.
[0005] A method based on forward projection also cannot handle specular and refractive images well, such as those rendered using principal surface replacement (PSR). One reason is that in at least some systems only the true principal surface (i.e., the first vertex of the principal specular reflection path) can be forward projected. Subsequent path vertices cannot be re-projected because they lack the information about the reflection / refraction chain necessary to calculate the new positions and the fact that one needs to trace the entire path, not just the principal visibility ray, to determine whether the secondary surface is visible. Attempting to use forward projection only at the first path vertex in a PSR renderer results in false positive gradients when the camera moves or the secondary surface moves, which invalidates the cumulative history in the denoiser and produces more noise than normal. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various embodiments in accordance with the present disclosure will be described with reference to the drawings, in which:
[0007] Figure 1A and Figure 1B shows an image in an image sequence according to at least one embodiment;
[0008] Figure 2A 、 Figure 2B 、 Figure 2C and Figure 2D shows the illumination difference between two images in a sequence according to at least one embodiment;
[0009] Figure 3A 、 Figure 3B 、 Figure 3C and Figure 3D shows the stages of a projection method according to at least one embodiment;
[0010] Figure 4 shows an example ray tracing pipeline according to at least one embodiment;
[0011] Figure 5A and Figure 5B shows a process for determining pixel data of an image in a sequence according to at least one embodiment;
[0012] Figure 6 shows components of a system for generating image data according to at least one embodiment;
[0013] Figure 7A shows inference and / or training logic according to at least one embodiment;
[0014] Figure 7B shows inference and / or training logic according to at least one embodiment;
[0015] Figure 8Illustrates an example data center system according to at least one embodiment;
[0016] Figure 9 Illustrates a computer system according to at least one embodiment;
[0017] Figure 10 Illustrates a computer system according to at least one embodiment;
[0018] Figure 11 Illustrates at least a portion of a graphics processor according to one or more embodiments;
[0019] Figure 12 Illustrates at least a portion of a graphics processor according to one or more embodiments;
[0020] Figure 13 Is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;
[0021] Figure 14 Is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline; and
[0022] Figure 15A And Figure 15B Illustrates a data flow diagram of a process for training a machine learning model according to at least one embodiment, and a client - server architecture that utilizes a pre - trained annotation model to enhance an annotation tool. DETAILED DESCRIPTION
[0023] Methods according to various embodiments can provide for determining robust temporal gradients for the generation of content such as image sequences, videos, or animations. In particular, various embodiments provide for computing a robust temporal gradient between a current frame and a previous frame in a temporal denoiser for a ray - traced renderer. Methods according to at least some of these embodiments can utilize back - projection to locate matching surfaces for gradient estimation, followed by surface parameter carry - forward and patching.
[0024] In at least one embodiment, an image generation system can generate an image or a sequence of video frames that, when played in sequence, presents a scene of content, such as a scene of an animation or a game. In each such image, there may be various objects represented, such as foreground and background objects that can be either static or dynamic. Whether static or dynamic, the positioning, size, and orientation of any of these objects may vary at least in part based on the movement of a virtual camera that is used to determine factors such as the viewpoint or zoom level of the image. In such an image or video sequence, these changes in view, positioning, size, and orientation can be seen as a set of movements of the individual pixels that represent those objects. This characteristic movement between pixel positions in different frames may appear somewhat jagged, noisy, or unsmooth when considering only the position information for those features in the current frame. To smooth the apparent movement or animation, pixel data from one or more prior frames can be blended with the pixel data of the current frame. However, in order to blend these pixel values, it is necessary to correlate the pixels that represent similar object features in these different images. In at least some embodiments, it is also appropriate to accurately render the lighting or shading of those objects or affect those objects as the objects move between these frames.
[0025] For example, consider Figure 1A an initial image frame 100, which may correspond to a frame of a video sequence. In this example, there are two light sources, including a main light source 102 (such as the sun or moon) and an auxiliary light source 104 such as a street lamp. In other examples, there are fewer or additional light sources of similar or different types. In this example, the light source 102 is considered the main light source because it is the brightest light source in the scene. In this scene, there are also multiple other objects, including foreground and background objects. These objects include dynamic objects such as a person 106, a first vehicle 108, and a second vehicle 110. To make the scene appear at least somewhat realistic, the relative positions of the light sources 102, 104 on the objects in the scene can be considered in order to properly illuminate or shade those objects. In addition to illuminating the objects, this relative position information will be used to calculate the shadows 112, 114, 116 cast by the individual objects with respect to those light sources. Further, there may be reflections 118 of the objects represented in the image, such as the reflections of the objects 106, 108 visible on the shiny surface of another object 110.
[0026] Figure 1BShows the next or subsequent image frame 150 generated for the example sequence. In this image, at least some of the objects may be in different positions, poses, views, or orientations, resulting in the features of these objects being represented at different pixel positions in different frames of the sequence. As described above, in order to provide smooth animation between these frames, it may be desirable to blend at least some of the "historical" pixel data of the initial frame 100 with the pixel data of this subsequent frame 150. To this end, a blending process can attempt to correlate the positions of these features at least to the extent that they are both represented in these two frames in order to perform the blending and to determine the weights for this blending or other such aspects. One way to attempt to correlate this information is to determine the pixel motion between these two images. In Figure 1B it, a number of motion vectors are shown, which represent how certain features of selected objects have moved relative to their positions in the initial frame 100. As shown, one or more of the light sources may have moved, and one or more of the objects illuminated by those light sources have moved. This can not only affect the positions of the features of those objects, but also the positions of the corresponding shadows. In addition, the reflected images of the objects may also change at least in part based on the motion of these other objects. For example, the reflected images 118 of the person 106 and the vehicle 108 on the vehicle 110 will change position based on the motion of these objects in the scene, and the appearance of these objects will change at least in part based on the different illumination of these objects due to the motion.
[0027] As described above, the temporal gradient representing the difference between the shading results of the current frame and the previous frame for a subset of the surfaces visible on the screen can be used to reduce the number of artifacts and smooth the animation in the sequence. Also as described above, conventional methods of determining the temporal gradient (e.g., by forward projection) do not handle reflected and refracted images well. To better understand how to determine reflected images, Figures 2A - 2D shows an exemplary ray-based shading method that can be utilized according to various embodiments. In Figure 2A the example image frame 200, there is a single light source 202 that illuminates two objects in the scene, in this case a box 204 and a cone 210. The light from the light source 202 also produces shadows 206, 208, which are caused by the rays from the lamp intersecting the box 204 and the cone 210 respectively and thus (mostly) being prevented from reaching the shadow regions 206, 208 behind those objects. In this example, the front-facing side of the box 204 includes a highly reflective feature 212, such as a mirror or a smooth metal panel, which can also reflect the light from the light source 202. In this example, the reflected light can include the light reflected from the cone 210, such that the reflected image 214 of the cone appears in the appropriate position (with the appropriate scale and mirror orientation) in the reflective feature 212.
[0028] To determine illumination, shadows, and reflected images, methods such as ray tracing can be utilized. As shown in image 220 of Figure 2B , a set of rays (only two rays are shown here for simplicity, but it can be understood that many projected rays can exist) are projected or traced from light source 202 in different directions. The first ray 222 is incident at point 224 on box 204, so the pixel color of the pixel at this position should be determined at least in part based on this direct illumination from light source 202. In addition, the ray can be traced beyond the intersection point 224 (or the second ray 226 can be traced) to determine the second intersection point 228 where the ray will intersect another surface (in this example, the floor on which the box is placed). The pixel value determined for the pixel at this second intersection point 228 can be determined at least in part based on the light from light source 202 being blocked (or obstructed) by box 204 - which creates a shadow region on the floor. Similarly, a ray 230 projected from the light source can be incident at or intersect the intersection point 232 on cone 210. In this example, the ray can then be directed from the intersection point 232 (or a second ray 234 is projected from it) to determine the corresponding intersection point 236 where it intersects the highly reflective surface 212 of box 204. Then, in addition to the fact that the intersection point 236 on the highly reflective surface 212 in this example is also likely to receive direct light (and potentially other reflected light) from the light source, the pixel value determined for the pixel at intersection point 236 can be determined at least in part based on the reflected light and one or more surface properties of cone 210 at intersection point 232 and the surface properties of the highly reflective surface 212 at intersection point 236.
[0029] Such a ray tracing method can be repeated for each frame in the sequence or at least a subset of the frames. For example, Figure 2C , image 240 of Figure 2DAs shown in image 260, ray tracing can be performed again to determine aspects of the image, such as the current position of the shadow and the new position and / or orientation of the reflected image of the cone. In at least some systems, such as in the G-buffer used for rendering, the reflected image of the cone on a highly reflective surface will replace part of the box data. However, as described above, it may be difficult to determine the determination of ray 262 for the reflected image using conventional methods. For the realism of motion, it may be important to accurately determine the ray and relate it to the appropriate pixels from earlier frames, appropriately weighting the historical data to prevent noise or jerkiness and other potential artifacts such as time lags.
[0030] As described above, at least some of the motion of the objects in the scene may be the result of the animation or motion of the virtual camera (which changes the view direction). Although it may be relatively straightforward to reconstruct the view direction of the primary surface by subtracting the camera position from the surface position, this is not possible in various conventional systems for secondary surfaces such as reflective or refractive surfaces. One option is to utilize the view direction of the current frame and leave it unpatched. Using a slightly mismatched view direction is mainly relevant for very shiny surfaces, so there may be a certain amount of noise on polished metals and the like visible in the reflected image, which is generally more preferable than having a large amount of noise on each reflective or refractive surface during the motion period.
[0031] The method according to various embodiments can provide adaptive temporal image filtering that works for reflected and refracted images as well as for shadows and directly illuminated objects. In at least one embodiment, the method for rendering frames in a sequence can use a conventional rendering process for each video frame such that the geometry buffer (or "G-buffer") is rendered. Then, this can be followed by rendering the reflected and / or refracted images using, for example, primary surface replacement (PSR). Next, a back-projection pass can be performed. To perform the back-projection, the current frame can be divided into a set of strata, where each stratum can include an array of adjacent pixels, such as (but not limited to) a 3x3 square of pixels. For each stratum, a single pixel can be selected that has a matching surface identified in the previous frame. In at least one embodiment, the motion vectors generated during the G-buffer fill and PSR passes can be used to locate the matching surface, which can then be effective for reflective and refractive surfaces. To determine whether a certain surface is the same in the current frame pixel and the previous frame pixel, several methods can be used, among other such options, which can include comparing the depth and normal of the two surfaces, or comparing the visibility buffer data.
[0032] Figures 3A - 3D An example method of using strata for back-projection is shown. In Figure 3AIn [description], a portion of frame pairing 300 is shown, where the pairing includes two adjacent frames in a sequence of frames of a scene. There are two objects represented in these frames, both of which move between the frames. As shown, the moving object and any background objects can have different motion vectors, for example where the background can be static and thus have zero motion. These frames are shown as being composed of (exaggerated) pixels, showing which pixels represent different parts of those objects. As described above, the pixels of each frame can be divided into a grid of layer blocks, for example square layer blocks of 3x3 pixels each. Figure 3A One such layer block 306 is shown in [description]. As described above, in this example, the process will utilize back-projection, such that the layer block is selected from a subsequent (or current) frame 304 rather than the initial frame 302. In this example, each pixel will be included in one layer block. The size of the layer block can be determined based on any of the various factors discussed herein, and the number of pixels included in the layer block can determine the amount of sampling performed for each frame pairing in the sequence. As shown for Figure 3B image pair 320 in [description], back-projection can involve determining the corresponding pixel position in the prior or initial frame 302 for each pixel position in a given layer block. An attempt can be made to determine the matching surface pixel for each pixel of the layer block, if represented in these two frames as discussed above. In at least one embodiment, the motion vector can be calculated by taking the current position of an object in world space, its previous position in world space, and then transforming these positions into screen space using the current and previous view-projection matrices or other such methods. In this case, the motion vector represents the difference between the screen space positions of the same object or surface.
[0033] In many cases, the layer block will contain multiple pixels with matching surfaces. In such cases, a representative pixel can be selected. In at least one embodiment, this can include a pixel 342 that has the brightest illumination on the initial or previous frame 302, but is not used for gradient estimation, as shown for Figure 3C image pair 340 in [description]. Using the brightest illumination heuristic can make the detection of illumination changes more robust and perceptually better. Using pixels that are also not used for gradient estimation helps eliminate biases that may be caused by reusing the same random number sequence for a certain surface over multiple consecutive frames.
[0034] After selecting the gradient pixels in the layer block and locating the matching pixels in the previous frame, G-buffer patching can be performed. G-buffer patching can involve taking certain parameters of a surface from the previous frame G-buffer (or other relevant buffer or cache), and writing them into the gradient pixels of the current frame G-buffer. This can be similar to the forward projection 362 of those parameters, as shown for Figure 3Das shown for frame pair 360. The parameters discussed can include those used to calculate surface illumination for comparison with the previous frame illumination of the same surface for gradient calculation. Such parameters include, but are not limited to, random number generator seeds, normals, metallicity, and roughness. In this example, backprojection can thus be used to find a matching surface for gradient estimation, followed by surface parameter loading and patching. Patching surface parameters in this way can, in the vast majority of cases, eliminate false positive gradients, making the denoised image very stable while allowing for a rapid response to illumination changes.
[0035] In at least one embodiment, parameters that should be very closely matched, but not simply copied from the prior frame, include surface world position and view direction. Typically, deferred renderers do not store these parameters, but rather these parameters are reconstructed from pixel position and camera information during shading. At least some PSR-based renderers do store these parameters, as for reflective or refractive surfaces, the position and view direction may not be reconstructable. In at least some embodiments, these parameters of the gradient surface can be patched into the corresponding G-buffer. However, in at least some embodiments, in contrast to parameters such as roughness mentioned above, these parameters should not be simply copied.
[0036] For example, when rendering a scene, the position information about animated objects as shown and discussed above can vary between frames. This provides a valid reason for the existence of illumination gradients and invalidates at least partially the illumination history information. Instead of copying the position data, visibility buffer information from the previous frame can be used to calculate the new position of the same surface based on the mesh, triangle indices, and barycentric coordinates of the same surface and the updated vertex buffer. The view direction may also vary due to factors such as animation and camera movement. In at least some cases, the view direction for the main surface can be relatively directly reconstructed, for example, by subtracting the camera position from the surface position. For secondary surfaces such as reflective or refractive surfaces, such a method may not be possible, and thus the current frame view direction can be used without patching. Using a slightly mismatched view direction may be important for surfaces such as those that are very bright or highly reflective, and thus slightly more noise can be observed on polished metals and similar objects or features visible in the reflection image.
[0037] Figure 4An example rendering pipeline 400 that can be used to render images or frames in a sequence is shown. In this example, pixel data 402 for the current frame to be rendered (which may include G-buffer data for a primary surface) can be received as an input to a reflection and refraction component 404 of the rendering system. As described above, this reflection and refraction component 404 can use this data to attempt to determine data in the pixel data for any determined reflection images and / or refraction images, and can provide this data to a back-projection and G-buffer patching component 406, which can perform the back-propagation discussed herein to locate corresponding points of those reflection images and refraction images, and use this data to patch the G-buffer 418, which can provide updated inputs for subsequent frames to be rendered. The data can then be provided to a light sample generation component 408 that performs light sampling, a ray tracing illumination component 410 that performs ray tracing illumination, and one or more shaders 412 that can set the pixel colors of the individual pixels for the frame based at least in part on the determined illumination information (along with other information such as color, texture, etc.). The results can be accumulated by an accumulation module 414 or component for generating an output frame 416 of a desired size, resolution, or format.
[0038] In at least one embodiment, the shader 408 can perform a back-projection step. Once the back-projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform an illumination pass. Using the information from the illumination pass and the illumination results from a previous frame, a gradient can be calculated, then filtered and used for history rejection. Such a method can be used to calculate a robust temporal gradient between the current frame and a previous frame in a temporal denoiser for a ray-traced renderer. Such a back-projection-based method can also resolve reflection images and refraction images and can operate on a rasterized G-buffer. Previous methods for back-projection, instead, omitted any G-buffer patching and relied on the original current G-buffer samples, which also led to false-positive gradients. Patching the surface parameters can eliminate false positives in the vast majority of cases, making the denoised image very stable but still reacting quickly to illumination changes. Once the back-projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform an illumination pass. Using the information from the illumination pass and the illumination results from a previous frame, a gradient can be calculated, then filtered and used for history rejection.
[0039] Figure 5AShows an example process 500 that can be used to generate images in a sequence according to at least one embodiment. In this example, for instance, a set of motion vectors for the current frame to be rendered into the G-buffer is generated 502, where the G-buffer includes pixel data for one or more reflection images and / or refraction images. Any reflection or refraction image can be rendered using a process such as primary surface replacement (PSR). A back-projection pass is performed 504 using the previous frame in the sequence. In at least one embodiment, the back-projection can be performed using an adjacent pixel group from the current frame. Based on the back-projection pass, one or more matching surfaces are located 506 between the current frame and the previous frame, e.g., one or more matching surfaces for each pixel group. Then, the G-buffer previously rendered for the current frame can be patched 508 or updated using information from these matching surfaces or at least a selected subset of these matching surfaces (e.g., one surface per pixel group). In at least one embodiment, patching the G-buffer includes writing one or more parameters of the matching surface of the G-buffer from the previous frame into the gradient pixels of the G-buffer of the current frame. The one or more parameters can correspond to parameters used to calculate surface illumination, such as can include at least one of a random generator seed, a normal value, a metallicity value, or a roughness value, and other such options. Next, the patched G-buffer can be used to calculate 510 illumination information. Then, the calculated illumination information can be used to determine 512 one or more light differences between the current frame and the previous frame. Then, the image can be rendered 514 at least partially based on the one or more light differences. In at least one embodiment, the rendering can involve performing one or more illumination passes for the current frame and calculating one or more temporal gradients as representing the one or more light differences based on the output of the illumination pass for the current frame and the calculated illumination information corresponding to the previous frame. Then, these temporal gradients can be filtered and used for history rejection, e.g., to eliminate historical data used for blending or determining the pixels of the current frame. Then, the rendered image can be output 514 for display on a display device or other such presentation, which can be part of a video file or video stream.
[0040] Figure 5B Shows an example process 550 that can perform backpropagation using multiple pixel layer blocks according to at least one embodiment. As regarding Figure 5AAs discussed in the process, such a process can be used to locate one or more matching surfaces. As described above, such a process can enable the forward projection of surfaces from a prior frame to a current frame (as used in various conventional systems) to be replaced with a backward projection from the current frame to the prior frame. Instead of first calculating some subset of pixels in the current frame that will be used for the gradient, the G-buffer for the current frame can first be fully rendered 552, including all reflected and refracted images. The pixels of the current frame can be partitioned 554 into an array of tile blocks of adjacent pixels, such as an array or grid of 3x3 pixel tile blocks. A current tile block to be analyzed can be selected 556. The primary surface can be determined 558 for the pixels of the tile block in the current frame. A backward projection can be performed 560 for each of those pixels in order to attempt 562 to locate one or more matching surfaces in the prior frame. One or more selection criteria that were not used for the gradient estimation in the prior frame - such as the brightest pixels (or pixels corresponding to the shiniest surfaces, etc.) - can be used to select 564 the pixels in the tile block in the prior frame that have an identified matching surface. Then, the matching surfaces for the selected pixels can be connected to the tile block of the current frame and used to patch 566 the G-buffer accordingly. Patching the G-buffer can include, for example, calculating the new position of the first surface depicted in the prior frame based at least on visibility buffer information from the previous frame. The visibility buffer information can include, for example, mesh information, triangle information, barycentric information, or one or more updated vertex buffers corresponding to the first surface. A decision 568 can be made as to whether there are more tile blocks to evaluate. If so, then the process can continue by selecting the next tile block to evaluate. If not, then the lighting data can be calculated 570 using the patched G-buffer as discussed herein, and further the lighting data can be used to determine 572 the light difference between corresponding pixels of the current frame and the prior frame (i.e., calculate the gradient), and as discussed with respect to Figure 5A the process, render the final frame for output.
[0041] As discussed, the various methods presented herein are lightweight enough to be performed in real time on a client device such as a personal computer or a game console. Such processing can be performed on content generated on the client device or received from an external source (e.g., streamed content received via at least one network). The source can be any suitable source, such as a game console, a streaming media provider, a third-party content provider, or another client device and other such options. In some cases, the processing and / or rendering of the content can be performed by one of these other devices, systems, or entities and then provided to the client's device (or another such recipient) for presentation or another such use.
[0042] In at least one embodiment, backprojection can be used to train an image generation network. This can include, for example, a generative neural network (e.g., a generative adversarial network (GAN)). In at least one embodiment, the network can be trained using backprojection to estimate the illumination differences between frames. Then, at inference time, the network can be used to render the current frame of an image or video stream using the pixel data sources of the current frame and a prior frame, in order to accurately represent the motion of reflected and refracted images for a certain scene.
[0043] As an example, Figure 6An example network configuration 600 is shown that can be used to provide, generate, or modify content. In at least one embodiment, the client device 602 can use components of the content application 604 on the client device 602 and data locally stored on the client device to generate the content of a session. In at least one embodiment, the content application 624 (e.g., an image generation or editing application) executed on the content server 620 (e.g., a cloud server or an edge server) can initiate a session associated with at least the client device 602, such as by leveraging a session manager and user data stored in the user database 634, and can cause the content 632 to be determined by the content manager 626. A scene generation module 628, such as may be related to an animation or game application, can generate or obtain the content determined to be provided, where at least part of the content may need to be rendered using the rendering engine 630 if that type of content or platform requires it, and is transmitted to the client device 602 using an appropriate transmission manager 622, for sending via download, streaming, or other such transmission channels. In at least one embodiment, the content 632 can include the materials that can be used by the rendering engine to render a scene based on a determined scene graph or other such rendering guidelines. In at least one embodiment, the client device 602 that receives the content can provide the content to the corresponding content application 604, which may also or alternatively include a scene generation module 612 or a rendering engine 614 (if necessary), which is used to render at least some of the content for presentation via the client device 602, such as presenting image or video content via the display 606, and presenting audio such as sounds and music via at least one audio playback device 608 such as a speaker or headphones. In at least one embodiment, at least some of the content may already be stored on the client device 602, rendered on the client device 602, or accessible to the client device 602, such that at least that portion of the content does not need to be transmitted over the network 640, such as in the case where the content may have been downloaded or locally stored on a hard drive or optical disc previously. In at least one embodiment, a transmission mechanism such as a data stream can be used to transmit the content from the server 620 or the content database 634 to the client device 602. In at least one embodiment, at least a portion of the content can be obtained or streamed from another source such as a third-party content service 660, which may also include a content application 662 for generating or providing content. In at least one embodiment, multiple computing devices or multiple processors within one or more computing devices such as may include a combination of a CPU and a GPU can be used to execute portions of the functionality.
[0044] In at least one embodiment, content application 624 includes a content manager 626 that can determine or analyze the content before transmitting it to client device 602. In at least one embodiment, content manager 626 may also include or cooperate with other components capable of generating, modifying, or enhancing the content to be provided. In at least one embodiment, this may include a rendering engine for rendering image or video content. In at least one embodiment, an image, video, or scene generation component 628 may be used to generate images, videos, or other media content. In at least one embodiment, an enhancement component 630—which may also include a neural network—may perform one or more enhancements to the content as discussed and proposed herein. In at least one embodiment, content manager 626 may cause the content (enhanced or unenhanced) to be transmitted to client device 602. In at least one embodiment, content application 604 on client device 602 may also include components such as a rendering engine, an image or video generator 612, and a content enhancement module 614 such that any or all of this functionality may be additionally or alternatively performed on client device 602. In at least one embodiment, content application 662 on third-party content service system 660 may also include such functionality. In at least one embodiment, the location at which at least some of this functionality is performed may be configurable or may depend on similar factors such as the type of client device 602 or the availability of a network connection with appropriate bandwidth. In at least one embodiment, a system for content generation may include any suitable combination of hardware and software in one or more locations. In at least one embodiment, image or video content at one or more generated resolutions may also be provided to or made available for other client devices 650, such as for downloading or streaming from a media source storing a copy of the image or video content. In at least one embodiment, this may include transmitting images of game content for a multiplayer game, where different client devices may display the content at different resolutions including one or more super-resolutions.
[0045] In this example, these client devices can include any suitable computing device, such as a desktop computer, a laptop computer, a set-top box, a streaming device, a gaming console, a smartphone, a tablet computer, a VR headset, AR goggles, a wearable computer, or a smart TV. Each client device can submit requests across at least one wired or wireless network, which can include the Internet, Ethernet, a local area network (LAN), or a cellular network, among other such options. In this example, these requests can be submitted to an address associated with a cloud provider, which can operate or control one or more electronic resources in a cloud provider environment, such as a data center or a server farm. In at least one embodiment, the request can be received or processed by at least one edge server located at the network edge and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling client devices to interact with a closer server, while also improving the security of resources in the cloud provider environment.
[0046] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as for providing image or video content to test or validate autonomous machine applications, or for performing deep learning operations. In at least one embodiment, such a system can be implemented using edge devices, or can incorporate one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.
[0047] Inference and training logic
[0048] Figure 7A Illustrated is inference and / or training logic 715 for performing inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B provide details regarding inference and / or training logic 715.
[0049] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters that configure neurons or layers of a neural network trained to and / or for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 can include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequencing, where weight and / or other parameter information is loaded to configure the logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores the weight parameters and / or input / output data of each layer of the neural network used or trained in conjunction with one or more embodiments during forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 can be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0050] In at least one embodiment, any portion of the code and / or data storage 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 701 can be cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 701 is internal or external to the processor, e.g., or consists of DRAM, SRAM, flash memory, or some other storage type, can depend on the available storage space on or off the chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0051] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, code and / or data storage 705 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained as and / or used for inference in aspects of one or more embodiments. In at least one embodiment, during training and / or inference using aspects of one or more embodiments, code and / or data storage 705 stores weight parameters and / or input / output data for each layer of the neural network used or trained in conjunction with one or more embodiments during backpropagation of the input / output data and / or weight parameters. In at least one embodiment, the training logic 715 can include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or sequencing, where weights and / or other parameter information are loaded to configure logic that includes integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, any portion of the code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 705 is internal or external to the processor, e.g., whether it consists of DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.
[0052] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0053] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations at least in part based on or as indicated by training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values from layers or neurons within a neural network) stored in the activation store 720, which are a function of input / output and / or weight parameter data stored in the code and / or data store 701 and / or the code and / or data store 705. In at least one embodiment, the activations are generated by linear algebra and / or matrix-based mathematics performed by the ALU 710 in response to executing instructions or other code, where the weight values stored in the code and / or data store 705 and / or the code and / or data store 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data store 705 or the code and / or data store 701 or other on-chip or off-chip storage.
[0054] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 710 can be outside of the processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 can be included within the execution units of a processor or otherwise included in a group of ALUs accessible by the execution units of a processor, which execution units of the processor can be within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, the code and / or data store 701, the code and / or data store 705, and the activation store 720 can be on the same processor or other hardware logic device or circuit, while in another embodiment, they can be on different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation store 720 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, the inference and / or training code can be stored with other code accessible to the processor or other hardware logic or circuit and can be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuits.
[0055] In at least one embodiment, the activation store 720 can be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation store 720 can be wholly or partly inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation store 720 is internal or external to the processor, e.g., or comprises DRAM, SRAM, flash memory, or other storage types, can depend on the on-chip or off-chip available storage, the latency requirements for training and / or inference functions, the batch size of data used in inferring and / or training a neural network, or some combination of these factors. In at least one embodiment, Figure 7A the inference and / or training logic 715 shown in can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the TM processing unit from Google, the inference processing unit (IPU) from Graphcore (e.g., “LakeCrest”) processor from IntelCorp. In at least one embodiment, Figure 7A the inference and / or training logic 715 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (such as a field programmable gate array (“FPGA”)).
[0056] Figure 7B FIG. shows the inference and / or training logic 715 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used along with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B the inference and / or training logic 715 shown in can be used in conjunction with an application specific integrated circuit (ASIC), such as the TM processing unit from Google, the inference processing unit (IPU) from Graphcore (e.g., “LakeCrest”) processor from IntelCorp. In at least one embodiment, Figure 7BThe inference and / or training logic 715 shown in FIG. may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as a field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (such as graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 7B In at least one embodiment shown in FIG., each of the code and / or data storage 701 and the code and / or data storage 705 is respectively associated with dedicated computing resources (such as computing hardware 702 and computing hardware 706). In at least one embodiment, each of the computing hardware 702 and the computing hardware 706 includes one or more ALUs that respectively perform mathematical functions (such as linear algebra functions) only on the information stored in the code and / or data storage 701 and the code and / or data storage 705, and the result of the executed function is stored in the activation storage 720.
[0057] In at least one embodiment, each of the code and / or data storage 701 and 705 and the corresponding computing hardware 702 and 706 respectively corresponds to different layers of a neural network, such that the activation obtained from one "storage / computation pair 701 / 702" of the code and / or data storage 701 and the computing hardware 702 is provided as an input to the next "storage / computation pair 705 / 706" of the code and / or data storage 705 and the computing hardware 706, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 701 / 702 and 705 / 706 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with the storage computation pairs 701 / 702 and 705 / 706.
[0058] Data center
[0059] Figure 8 FIG. shows an example data center 800 in which at least one embodiment may be used. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.
[0060] In at least one embodiment, as Figure 8As shown, the data center infrastructure layer 810 may include a resource coordinator 812, grouped computing resources 814, and node computing resources ("node C.R.") 816(1)-816(N), where "N" represents any positive integer. In at least one embodiment, the node C.R. 816(1)-816(N) may include, but is not limited to, any number of central processing units ("CPU") or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output ("NWI / O") devices, network switches, virtual machines ("VM"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node C.R. 816(1)-816(N) may be a server having one or more of the above computing resources.
[0061] In at least one embodiment, the grouped computing resources 814 may include separate groupings (not shown) of node C.R. housed within one or more racks, or many racks (also not shown) within data centers located in various geographical locations. The separate groupings of node C.R. within the grouped computing resources 814 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. including a CPU or processor may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0062] In at least one embodiment, the resource coordinator 812 may configure or otherwise control one or more of the node C.R. 816(1)-816(N) and / or the grouped computing resources 814. In at least one embodiment, the resource coordinator 812 may include a software design infrastructure ("SDI") management entity for the data center 800. In at least one embodiment, the resource coordinator 108 may include hardware, software, or some combination thereof.
[0063] In at least one embodiment, as Figure 8As shown, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 may include a framework that supports software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. In at least one embodiment, the software 832 or the application 842 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 828 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 832 may include a Spark driver to facilitate scheduling of the workloads supported by the various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820 including Spark and the distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 828 and the job scheduler 822. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 may coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.
[0064] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes C.R. 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.
[0065] In at least one embodiment, one or more applications 842 included in the application layer 840 may include one or more types of applications used by at least a portion of nodes C.R. 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0066] In at least one embodiment, any one of the configuration manager 824, the resource manager 826, and the resource coordinator 812 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 800 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing parts of the data center.
[0067] In at least one embodiment, the data center 800 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 800. In at least one embodiment, by using the weight parameters calculated by one or more training techniques described herein, the resources described above with respect to the data center 800 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks.
[0068] In at least one embodiment, the data center may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the above resources. In addition, one or more of the above software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0069] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. As described herein in connection with Figure 7A and / or Figure 7BProvide details regarding inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be used in a system Figure 8 of the system for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0070] Such components can be used to generate enhanced content, such as images or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0071] A computer system
[0072] Figure 9 is a block diagram showing an exemplary computer system according to at least one embodiment. The exemplary computer system can be a system with interconnected devices and components, a system-on-chip (SOC), or some combination thereof formed with a processor that can include execution units to execute instructions. In at least one embodiment, according to the present disclosure, for example, the embodiments described herein, the computer system 900 can include, but is not limited to, components such as a processor 902, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, the computer system 900 can include a processor, such as a processor family available from Intel Corporation of Santa Clara, California, XeonTM, XScaleTM, and / or StrongARMTM, Core TM or Nervana TM microprocessor, although other systems (including PCs, engineering workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, the computer system 900 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.
[0073] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, an embedded application can include a microcontroller, a digital signal processor (“DSP”), a system-on-chip, a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0074] In at least one embodiment, computer system 900 can include, but is not limited to, a processor 902, which can include, but is not limited to, one or more execution units 908 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 900 is a single-processor desktop or server system, but in another embodiment, computer system 900 can be a multi-processor system. In at least one embodiment, processor 902 can include, but is not limited to, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910, which can transfer data signals between processor 902 and other components in computer system 900.
[0075] In at least one embodiment, processor 902 can include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 904. In at least one embodiment, processor 902 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory can reside external to processor 902. Other embodiments can also include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 906 can store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0076] In at least one embodiment, a logic execution unit 908 including, but not limited to, performing integer and floating point operations is also located in the processor 902. In at least one embodiment, the processor 902 may further include a microcode ("ucode") read-only memory ("ROM") for storing the microcode of certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for processing a packet instruction set 909. In at least one embodiment, by including the packet instruction set 909 in the instruction set of a general-purpose processor and the associated circuitry for the instructions to be executed, operations used by many multimedia applications can be performed using the packet data in the processor 902. In one or more embodiments, operations can be performed on the packet data by using the full width of the processor's data bus to accelerate and more efficiently execute many multimedia applications, which may not require transmitting smaller data units on the processor's data bus to perform one or more operations on one data element at a time.
[0077] In at least one embodiment, the execution unit 908 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 900 may include, but not limited to, a memory 920. In at least one embodiment, the memory 920 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage devices. In at least one embodiment, the memory 920 may store instructions 919 and / or data 921 represented by data signals that can be executed by the processor 902.
[0078] In at least one embodiment, the system logic chip can be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip can include, but is not limited to, a Memory Controller Hub (“MCH”) 916, and the processor 902 can communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 can provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 can initiate data signals among the processor 902, the memory 920, and other components in the computer system 900, and bridge data signals among the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 can be coupled to the memory 920 via the high-bandwidth memory path 918, and the graphics / video card 912 can be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.
[0079] In at least one embodiment, the computer system 900 can use the system I / O 922, which is a proprietary hub interface bus, to couple the MCH 916 to an I / O Controller Hub (“ICH”) 930. In at least one embodiment, the ICH 930 can provide a direct connection to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 920, the chipset, and the processor 902. Examples can include, but are not limited to, an audio controller 929, a Firmware Hub (“FlashBIOS”) 928, a wireless transceiver 926, a data storage 924, a legacy I / O controller 923 that includes a user input and keyboard interface, a serial expansion port 927 (such as a Universal Serial Bus (USB) port), and a network controller 934. The data storage 924 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash device, or other mass storage devices.
[0080] In at least one embodiment, Figure 9 a system including interconnected hardware devices or “chips” is shown, while in other embodiments, Figure 9An exemplary system-on-a-chip (SoC) can be shown. In at least one embodiment, the device can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 are interconnected using a Compute Express Link (CXL) interconnect.
[0081] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B In at least one embodiment, inference and / or training logic 715 can be used in a Figure 9 system for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0082] Such components can be used to generate enhanced content, such as images or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0083] Figure 10 is a block diagram showing an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 can be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, a notebook computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0084] In at least one embodiment, system 1000 can include, but is not limited to, a processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1010 is coupled using a bus or interface, such as an I²C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 10 shows a system that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 10 An exemplary system-on-a-chip (SoC) can be shown. In at least one embodiment, Figure 10 the devices shown in Figure 10One or more components are interconnected using Compute Express Link (CXL) interconnects.
[0085] In at least one embodiment, Figure 10 may include a display 1024, a touch screen 1025, a touchpad 1030, a Near Field Communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, an Embedded Controller (“EC”) 1035, a Trusted Platform Module (“TPM”) 1038, a BIOS / Firmware / Flash (“BIOS, FWFlash”) 1022, a DSP 1060, a drive 1020 (such as a Solid State Drive (“SSD”) or a Hard Disk Drive (“HDD”)), a Wireless Local Area Network unit (“WLAN”) 1050, a Bluetooth unit 1052, a Wireless Wide Area Network unit (“WWAN”) 1056, a Global Positioning System (GPS) 1055, a camera (“USB3.0 camera”) 1054 (such as a USB3.0 camera) and / or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented to, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0086] In at least one embodiment, other components may be communicatively coupled to the processor 1010 via the components described above. In at least one embodiment, an accelerometer 1041, an Ambient Light Sensor (“ALS”) 1042, a compass 1043, and a gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1036, and a touchpad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, a speaker 1063, headphones 1064, and a microphone (“mic”) 1065 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 1062, which may in turn be communicatively coupled to the DSP 1060. In at least one embodiment, the audio unit 1062 may include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a Subscriber Identity Module (“SIM”) 1057 may be communicatively coupled to the WWAN unit 1056. In at least one embodiment, components (such as the WLAN unit 1050, the Bluetooth unit 1052, and the WWAN unit 1056) may be implemented in a Next Generation Form Factor (NGFF).
[0087] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Below in connection with Figure 7A and / or Figure 7BProvide details regarding inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be in Figure 10 a system for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functionality, and / or architectures or neural network use cases described herein.
[0088] Such components can be used to generate enhanced content, such as images or video content having upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0089] Figure 11 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and can be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.
[0090] In at least one embodiment, system 1100 can be included in or incorporated within a server-based gaming platform, including a game console such as a game and media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, system 1100 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1100 can also be coupled to or integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.
[0091] In at least one embodiment, each of one or more processors 1102 includes one or more processor cores 1107 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 1107 is configured to process a particular set of instructions 1109. In at least one embodiment, the set of instructions 1109 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, the processor cores 1107 can each process a different set of instructions 1109, which can include instructions that help to emulate other instruction sets. In at least one embodiment, the processor cores 1107 can also include other processing devices, such as a digital signal processor (DSP).
[0092] In at least one embodiment, the processor 1102 includes a cache memory 1104. In at least one embodiment, the processor 1102 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of the processor 1102. In at least one embodiment, the processor 1102 also uses an external cache (e.g., a level three (L3) cache or a last level cache (LLC)) (not shown), and the external cache can be shared among the processor cores 1107 using known cache coherence techniques. In at least one embodiment, the processor 1102 further includes a register file 1106, and the processor can include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 1106 can include general-purpose registers or other registers.
[0093] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transfer communication signals, such as address, data, or control signals, between the processors 1102 and other components in the system 1100. In at least one embodiment, the interface bus 1110 can be a processor bus, such as a version of the Direct Media Interface (DMI) bus, in one embodiment. In at least one embodiment, the interface bus 1110 is not limited to the DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a Platform Controller Hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between the memory device and other components of the processing system 1100, while the Platform Controller Hub (PCH) 1130 provides connections to I / O devices via a local I / O bus.
[0094] In at least one embodiment, the memory device 1120 can be a Dynamic Random Access Memory (DRAM) device, a Static Random Access Memory (SRAM) device, a flash memory device, a Phase Change Memory device, or have suitable performance to be used as processor memory. In at least one embodiment, the storage device 1120 can be used as the system memory of the processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, the memory controller 1116 is also coupled to an optional external graphics processor 1112, which can communicate with one or more graphics processors 1108 in the processor 1102 to perform graphics and media operations. In at least one embodiment, a display device 1111 can be connected to the processor 1102. In at least one embodiment, the display device 1111 can include one or more of an internal display device, such as in a mobile electronic device or a laptop device, or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1111 can include a Head-Mounted Display (HMD), such as a stereoscopic display device for Virtual Reality (VR) applications or Augmented Reality (AR) applications.
[0095] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to be connected to the storage device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, a data storage device 1124 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1128 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1134 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1110. In at least one embodiment, the audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1100. In at least one embodiment, the platform controller hub 1130 can also be connected to one or more universal serial bus (USB) controllers 1142, which connect input devices, such as a keyboard and mouse 1143 combination, a camera 1144, or other USB input devices.
[0096] In at least one embodiment, instances of the memory controller 1116 and the platform controller hub 1130 can be integrated into a discrete external graphics processor, such as the external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 can be external to one or more processors 1102. For example, in at least one embodiment, the system 1100 can include an external memory controller 1116 and a platform controller hub 1130, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 1102.
[0097] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Below in connection with Figure 7A and / or Figure 7BProvide details regarding inference and / or training logic 715. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the graphics processor 1100. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor. Additionally, in at least one embodiment, the inference and / or training operations described herein may be completed using logic other than the Figure 7A or Figure 7B logic shown. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the graphics processor to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0098] Such components may be used to generate enhanced content, such as images or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0099] Figure 12 is a block diagram of a processor 1200 having one or more processor cores 1202A - 1202N, an integrated memory controller 1214, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, the processor 1200 may include additional cores, up to and including the additional core 1202N shown in dashed boxes. In at least one embodiment, each processor core 1202A - 1202N includes one or more internal cache units 1204A - 1204N. In at least one embodiment, each processor core may also access one or more shared cache units 1206.
[0100] In at least one embodiment, the internal cache units 1204A - 1204N and the shared cache units 1206 represent the cache memory hierarchy within the processor 1200. In at least one embodiment, the cache memory units 1204A - 1204N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest - level cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1206 and 1204A - 1204N.
[0101] In at least one embodiment, the processor 1200 may further include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1210 provides management functions for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown).
[0102] In at least one embodiment, one or more processor cores 1202A-1202N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 1210 includes components for coordinating and operating cores 1202A-1202N during multi-threaded processing. In at least one embodiment, the system agent core 1210 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of the processor cores 1202A-1202N and the graphics processor 1208.
[0103] In at least one embodiment, the processor 1200 further includes a graphics processor 1208 for performing graphics processing operations. In at least one embodiment, the graphics processor 1208 is coupled to the shared cache unit 1206 and the system agent core 1210 that includes one or more integrated memory controllers 1214. In at least one embodiment, the system agent core 1210 further includes a display controller 1211 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1211 may also be a separate module coupled to the graphics processor 1208 via at least one interconnect, or may be integrated within the graphics processor 1208.
[0104] In at least one embodiment, the ring-based interconnect unit 1212 is used to couple the internal components of the processor 1200. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1208 is coupled to the ring interconnect 1212 via an I / O link 1213.
[0105] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including a package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1218 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 1202A - 1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last-level cache.
[0106] In at least one embodiment, processor cores 1202A - 1202N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, processor cores 1202A - 1202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of processor cores 1202A - 1202N execute a common instruction set, while one or more other processor cores 1202A - 1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, in terms of microarchitecture, processor cores 1202A - 1202N are heterogeneous, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, processor 1200 can be implemented on one or more chips or be implemented as a SoC integrated circuit.
[0107] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B In at least one embodiment, some or all of inference and / or training logic 715 can be incorporated into processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs, which are embodied in Figure 12 graphics processor 1512, graphics cores 1202A - 1202N, or other components in. Additionally, in at least one embodiment, the inference and / or training operations described herein can be completed using logic other than Figure 7A or Figure 7B shown. In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALUs of graphics processor 1200 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0108] Such components can be used to generate enhanced content, such as images or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0109] Virtualized computing platform
[0110] Figure 13 It is an example data flow diagram of a process 1300 for generating and deploying an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 1300 may be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. The process 1300 may be executed within a training system 1304 and / or a deployment system 1306. In at least one embodiment, the training system 1304 may be used to perform the training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for the deployment system 1306. In at least one embodiment, the deployment system 1306 may be configured to offload processing and computing resources in a distributed computing environment to reduce the infrastructure requirements of the facilities 1302. In at least one embodiment, one or more applications in the pipeline may use or invoke the services (e.g., inference, visualization, computing, AI, etc.) of the deployment system 1306 during application execution.
[0111] In at least one embodiment, some applications used in an advanced processing and inference pipeline may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, data 1308 (e.g., imaging data) generated at the facility 1302 (and stored on one or more picture archiving and communication system (PACS) servers at the facility 1302) may be used to train a machine learning model at the facility 1302, imaging or sequencing data 1308 from another or more facilities may be used to train a machine learning model, or a combination thereof. In at least one embodiment, the training system 1304 may be used to provide applications, services, and / or other resources to generate a working, deployable machine learning model for the deployment system 1306.
[0112] In at least one embodiment, the model registry 1324 may be supported by an object store that may support version control and object metadata. In at least one embodiment, the object store may be accessed from within a cloud platform via, for example, a cloud storage (e.g., Figure 14 of the cloud 1426) - compatible application programming interface (API). In at least one embodiment, the machine learning models within the model registry 1324 may be uploaded, listed, modified, or deleted by the developers or partners of the systems that interact with the API. In at least one embodiment, the API may provide access to methods that allow users with appropriate credentials to associate a model with an application such that the model may be executed as part of the execution of a containerized instantiation of the application.
[0113] In at least one embodiment, the training pipeline 1404( Figure 14 ) can include the following scenarios: where the facility 1302 is training their own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device, a sequencing device, and / or other types of devices can be received. In at least one embodiment, once the imaging data 1308 is received, AI-assisted annotation 1310 can be used to help generate annotations corresponding to the imaging data 1308 to be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 1310 can include one or more machine learning models (e.g., a convolutional neural network (CNN)), and the machine learning model can be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, the AI-assisted annotation 1310 can then be used directly or can be adjusted or fine-tuned using annotation tools to generate the ground truth data. In at least one embodiment, the AI-assisted annotation 1310, the labeled clinical data 1312, or a combination thereof can be used as the ground truth data for training the machine learning model. In at least one embodiment, the trained machine learning model can be referred to as the output model 1316 and can be used by the deployment system 1306 as described herein.
[0114] In at least one embodiment, the training pipeline 1404( Figure 14) may include the following scenarios: where the facility 1302 requires a machine learning model for performing one or more processing tasks for deploying one or more applications in the deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from the model registry 1324. In at least one embodiment, the model registry 1324 may include machine learning models that are trained to perform various different inference tasks on imaging data. In at least one embodiment, the machine learning models in the model registry 1324 can be trained on imaging data from different facilities (e.g., a facility located remotely) rather than the facility 1302. In at least one embodiment, the machine learning model may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a specific location, it can be trained at that location or at least in a manner that protects the confidentiality of the imaging data or restricts the off-site transfer of the imaging data. In at least one embodiment, once the model or a part of the model has been trained at one location, the machine learning model can be added to the model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in the model registry 1324. In at least one embodiment, a machine learning model (referred to as the output model 1316) can then be selected from the model registry 1324 and can be used in the deployment system 1306 to perform one or more processing tasks for one or more applications of the deployment system.
[0115] In at least one embodiment, in the training pipeline 1404( Figure 14) Among them, the scenario may include a facility 1302 that requires a machine learning model to perform one or more processing tasks for deploying one or more applications in the system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model). In at least one embodiment, due to population differences, robustness of the training data for training the machine learning model, diversity of training data anomalies, and / or other issues with the training data, the machine learning model selected from the model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at the facility 1302. In at least one embodiment, AI-assisted annotation 1310 can be used to help generate annotations corresponding to the imaging data 1308 for use as ground truth data for training or updating the machine learning model. In at least one embodiment, labeled clinical data 1312 can be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof) can be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, the trained machine learning model can be referred to as the output model 1316 and can be used by the deployment system 1306 as described herein.
[0116] In at least one embodiment, the deployment system 1306 may include software 1318, services 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, the deployment system 1306 may include a software “stack” such that the software 1318 may be built on top of the services 1320 and the services 1320 may be used to perform some or all of the processing tasks, and the services 1320 and the software 1318 may be built on top of the hardware 1322 and the hardware 1322 may be used to perform the processing, storage, and / or other computing tasks of the deployment system. In at least one embodiment, the software 1318 may include any number of different containers, where each container may execute an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and inference pipeline. In at least one embodiment, in addition to the containers that receive and configure the imaging data for use by each container and / or by the facility 1302 after being processed through the pipeline, a high-level processing and inference pipeline may be defined based on the selection of different containers desired or required for processing the imaging data 1308 (e.g., to convert the output back to a usable data type). In at least one embodiment, a combination of containers within the software 1318 (e.g., which form a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize the services 1320 and the hardware 1322 to perform some or all of the processing tasks of the applications instantiated in the containers.
[0117] In at least one embodiment, a data processing pipeline may receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of the deployment system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may be preprocessed as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks in the pipeline to prepare the output data for the next application and / or to prepare the output data for transmission and / or use by the user (e.g., as a response to the inference request). In at least one embodiment, the inference tasks may be performed by one or more machine learning models, such as a trained or deployed neural network, which may include the output model 1316 of the training system 1304.
[0118] In at least one embodiment, the tasks of a data processing pipeline can be encapsulated in containers, where each container represents a discrete, fully functional instantiation of an application and a virtualized computing environment that can reference a machine learning model. In at least one embodiment, a container or application can be published to a private (e.g., limited access) area of a container registry (described in more detail herein), and a trained or deployed model can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, the image can be used to generate a container for instantiation of the application for use by the user's system.
[0119] In at least one embodiment, a developer (e.g., a software developer, a clinician, a doctor, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publishing, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, on data from the first facility), where the SDK, as part of the system (e.g., Figure 14 system 1400 in the system) can support at least some services 1320. In at least one embodiment, since a DICOM object can contain from one to hundreds of images or other data types, and due to the variability of the data, the developer is responsible for managing (e.g., setting up constructs for building preprocessing into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by the system 1400 (e.g., for accuracy), the application becomes available in the container registry for a user to select and / or implement to perform one or more processing tasks on data at the user's facility (e.g., a second facility).
[0120] In at least one embodiment, the developer can then share the application or container over a network for use by a system (e.g., Figure 14User access and use of the system 1400). In at least one embodiment, a completed and verified application or container can be stored in a container registry, and a related machine learning model can be stored in the model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) can browse the container registry and / or the model registry 1324 to obtain applications, containers, datasets, machine learning models, etc., select a desired combination of elements to include in a data processing pipeline, and submit an image processing request. In at least one embodiment, the request can include input data necessary to execute the request (and in some examples, patient-related data), and / or can include a selection of an application and / or a machine learning model to be executed when processing the request. In at least one embodiment, the request can then be passed to one or more components (e.g., the cloud) of the deployment system 1306 to perform the processing of the data processing pipeline. In at least one embodiment, the processing performed by the deployment system 1306 can include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or the model registry 1324. In at least one embodiment, once a result is generated through the pipeline, the result can be returned to the user for reference (e.g., for viewing in a viewing application suite executed locally, on a local workstation, or on a terminal).
[0121] In at least one embodiment, to assist in processing or executing an application or container in the pipeline, the service 1320 can be utilized. In at least one embodiment, the service 1320 can include a computing service, an artificial intelligence (AI) service, a visualization service, and / or other service types. In at least one embodiment, the service 1320 can provide functions common to one or more applications in the software 1318, and thus the functions can be abstracted as services that can be invoked or utilized by the applications. In at least one embodiment, the functions provided by the service 1320 can run dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using Figure 14scale well with a parallel computing platform 1430). In at least one embodiment, rather than requiring each application that shares the same functionality provided by the shared service 1320 to have a corresponding instance of the service 1320, the service 1320 can be shared among and within various applications. In at least one embodiment, by way of non-limiting example, the service can include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included, which can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data augmentation service can be further included, which can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other augmentations. In at least one embodiment, a visualization service can be used, which can add image rendering effects (e.g., ray tracing, rasterization, denoising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service can be included, which provides beamforming, segmentation, inference, imaging, and / or support for other applications within the pipeline of a virtual instrument.
[0122] In at least one embodiment, where the service 1320 includes an AI service (e.g., an inference service), as part of the execution of an application, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., an inference server) to perform one or more machine learning models or their processing. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can call the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, the software 1318 that implements an advanced processing and inference pipeline, which includes a segmentation application and an anomaly detection application, can be pipelined because each application can call the same inference service to perform one or more inference tasks.
[0123] In at least one embodiment, the hardware 1322 can include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 can be used to provide efficient, specially built support for the software 1318 and services 1320 in the deployment system 1306. In at least one embodiment, GPU processing can be implemented to perform local processing (e.g., at the facility 1302) within an AI / deep learning system, in a cloud system, and / or in other processing components of the deployment system 1306 to improve the efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, as a non-limiting example, with respect to deep learning, machine learning, and / or high-performance computing, the software 1318 and / or services 1320 can be optimized for GPU processing. In at least one embodiment, at least some of the computing environments of the deployment system 1306 and / or the training system 1304 can be executed in a data center with GPU-optimized software (e.g., the hardware and software combination of an NVIDIA DGX system), one or more supercomputers, or high-performance computer systems. In at least one embodiment, as described herein, the hardware 1322 can include any number of GPUs that can be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform can also include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on an NVIDIA DGX system) can be used as a hardware abstraction and scaling platform to execute a cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform can integrate an application container cluster system or a coordination system (e.g., KUBERNETES) across multiple GPUs to achieve seamless scaling and load balancing.
[0124] Figure 14 is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, the system 1400 can be used to implement Figure 13 process 1300 and / or other processes, including advanced processing and inference pipelines. In at least one embodiment, the system 1400 can include a training system 1304 and a deployment system 1306. In at least one embodiment, the training system 1304 and the deployment system 1306 can be implemented using the software 1318, services 1320, and / or hardware 1322 as described herein.
[0125] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) can be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 can be implemented locally (with respect to a healthcare service facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1426 can be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol can include a network token, which can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, the API of a virtual instrument (described herein) or other instances of system 1400 can be restricted to a set of public IPs that have been audited or authorized for interaction.
[0126] In at least one embodiment, the various components of system 1400 can communicate with each other using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between the facilities and components of system 1400 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be conveyed via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.
[0127] In at least one embodiment, similar to that described herein with respect to Figure 13 the training system 1304 can execute a training pipeline 1404. In at least one embodiment, where the deployment system 1306 will use one or more machine learning models in a deployment pipeline 1410, the training pipeline 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, an output model 1316 can be generated as a result of the training pipeline 1404. In at least one embodiment, the training pipeline 1404 can include any number of processing steps, such as but not limited to the transformation or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by the deployment system 1306. In at least one embodiment, a training pipeline 1404 similar to the first example described with respect to Figure 13 can be used for a first machine learning model, and a training pipeline 1404 similar to the second example described with respect to Figure 13 can be used for a second machine learning model, and a training pipeline 1404 similar to that described with respect to Figure 13The training pipeline 1404 of the described third example can be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used according to the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models may have been trained and ready for deployment, so the training system 1304 may not perform any processing on the machine learning models, and one or more machine learning models can be implemented by the deployment system 1306.
[0128] In at least one embodiment, depending on the implementation or embodiment, the output model 1316 and / or the pre-trained model 1406 can include any type of machine learning model. In at least one embodiment and without limitation, the machine learning models used by the system 1400 can include linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.
[0129] In at least one embodiment, the training pipeline 1404 can include AI-assisted annotation, as described herein with respect to at least Figure 15BMore detailed descriptions are provided. In at least one embodiment, the labeled clinical data 1312 (e.g., traditional annotations) can be generated by any number of techniques. In at least one embodiment, in some examples, the labels or other annotations can be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a markup program, another type of application suitable for generating ground truth annotations or labels, and / or can be hand-drawn. In at least one embodiment, the ground truth data can be synthetically generated (e.g., generated from a computer model or rendering), real-world generated (e.g., designed and generated from real-world data), machine-automatically generated (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a tagger or annotation expert to define the location of the label), and / or combinations thereof. In at least one embodiment, for each instance of the imaging data 1308 (or other data types used by the machine learning model), there can be corresponding ground truth data generated by the training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of the deployment pipeline 1410; supplementing or replacing the AI-assisted annotation included in the training pipeline 1404. In at least one embodiment, the system 1400 can include a multi-layer platform, and the multi-layer platform can include a software layer (e.g., software 1318) of a diagnostic application (or other application type), which can perform one or more medical imaging and diagnostic functions. In at least one embodiment, the system 1400 can be communicatively coupled (e.g., via an encrypted link) to a PACS server network of one or more facilities. In at least one embodiment, the system 1400 can be configured to access and reference data from the PACS server to perform operations such as training a machine learning model, deploying a machine learning model, image processing, inference, and / or other operations.
[0130] In at least one embodiment, the software layer can be implemented as a secure, encrypted, and / or certified API through which an application or container can be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, the application can then call or execute one or more services 1320 to perform computing, AI, or visualization tasks associated with the respective application, and the software 1318 and / or the services 1320 can utilize the hardware 1322 to perform processing tasks in an efficient and effective manner.
[0131] In at least one embodiment, the deployment system 1306 may execute a deployment pipeline 1410. In at least one embodiment, the deployment pipeline 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) - including AI-assisted annotation, where the imaging data is generated by an imaging device, a sequencing device, a genomics device, etc., as described above. In at least one embodiment, as described herein, the deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scan instrument, a virtual sequencing instrument, etc.). In at least one embodiment, for a single device, there may be more than one deployment pipeline 1410, depending on the information desired from the data generated by the device. In at least one embodiment, in the case where an abnormality is expected to be detected from an MRI machine, there may be a first deployment pipeline 1410, and in the case where image enhancement is desired from the output of the MRI machine, there may be a second deployment pipeline 1410.
[0132] In at least one embodiment, an image generation application may include a processing task that includes using a machine learning model. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in an application that performs a processing task. In at least one embodiment, an application may be selectable and customizable, and by defining the construction of the application, the deployment and implementation of the application for a particular user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of the system 1400 (e.g., services 1320 and hardware 1322), the deployment pipeline 1410 may be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.
[0133] In at least one embodiment, the deployment system 1306 may include a user interface 1414 (e.g., a graphical user interface, a web interface, etc.), which may be used to select applications to be included in the deployment pipeline 1410, arrange the applications, modify or change the applications or their parameters or construction, use and interact with the deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with the deployment system 1306. In at least one embodiment, although not shown with respect to the training system 1304, the user interface 1414 (or a different user interface) may be used to select models to be used in the deployment system 1306, to select models for training or retraining in the training system 1304, and / or to otherwise interact with the training system 1304.
[0134] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 may be used to manage the interaction between the applications or containers of the deployment pipeline 1410 and the services 1320 and / or the hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate interactions from application to application, from an application to the service 1320, and / or from an application or service to the hardware 1322. In at least one embodiment, although shown as included in the software 1318, this is not intended to be limiting, and in some examples, the pipeline manager 1412 may be included in the service 1320. In at least one embodiment, the application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) may include a container coordination system that may group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications (e.g., rebuilt applications, split applications, etc.) from the deployment pipeline 1410 with individual containers, each application may execute in a self - contained environment (e.g., at the kernel level) to improve speed and efficiency.
[0135] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed separately (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer), which can allow for tasks focused on and attention paid to a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can assist in the communication and collaboration between different containers or applications. In at least one embodiment, as long as the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container), the application coordination system 1428 and / or the pipeline manager 1412 can facilitate communication and resource sharing between and among each application or container. In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, perform load balancing, and determine the sharing of services or resources between and among the various applications or containers. In at least one embodiment, a scheduler can be used to track the resource requirements of an application or container, the current or planned use of those resources, and the resource availability. Thus, in at least one embodiment, the scheduler can allocate resources to different applications, and distribute resources between and among applications, taking into account the requirements and availability of the system. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or deferred processing), etc.
[0136] In at least one embodiment, the services 1320 utilized and shared by an application or container in the deployment system 1306 may include compute services 1416, AI services 1418, visualization services 1420, and / or other service types. In at least one embodiment, an application may call (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize the compute services 1416 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1416 may be utilized to perform parallel processing (e.g., using the parallel computing platform 1430) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, the parallel computing platform 1430 (e.g., NVIDIA's CUDA) may implement general-purpose computing on a GPU (GPGPU) (e.g., GPU 1422). In at least one embodiment, the software layer of the parallel computing platform 1430 may provide access to the virtual instruction set of the GPU and parallel computing elements to execute compute kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, the memory may be shared between and among multiple containers and / or between and among different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within a container to use the same data for a shared memory segment from the parallel computing platform 1430 (e.g., where multiple different stages of one application or multiple applications are processing the same information). In at least one embodiment, rather than copying data and moving the data to different locations in memory (e.g., read / write operations), the same data in the same location in memory may be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, since the data is used to generate new data as a result of processing, the information of the new location of the data may be stored and shared among various applications. In at least one embodiment, the location of the data and the location of the updated or modified data may be part of the definition of how to understand the payload in a container.
[0137] In at least one embodiment, an AI service 1418 can be utilized to perform an inference service for executing a machine learning model associated with an application (e.g., the task is to perform one or more processing tasks of the application). In at least one embodiment, the AI service 1418 can utilize an AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application of the deployment pipeline 1410 can use one or more output models 1316 from the self-training system 1304 and / or other models of the application to perform inference on imaging data. In at least one embodiment, two or more examples of performing inference using an application coordination system 1428 (e.g., a scheduler) can be available. In at least one embodiment, the first category can include a high-priority / low-latency path, which can implement a higher service level agreement, such as for performing inference on emergency requests in an emergency situation or for a radiologist during a diagnostic process. In at least one embodiment, the second category can include a standard-priority path, which can be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1428 can allocate resources (e.g., service 1320 and / or hardware 1322) based on the priority path for different inference tasks of the AI service 1418.
[0138] In at least one embodiment, a shared memory may be installed in the AI service 1418 in the system 1400. In at least one embodiment, the shared memory may operate as a cache (or other storage device type) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of the deployment system 1306 may receive the request and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be input into a database, and if not already in the cache, the machine learning model may be located from the model registry 1324. A verification step may ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of the pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if an inference server has not been started to execute the model, the inference server may be started. Any number of inference servers may be started for each model. In at least one embodiment, in a pull model where inference servers are clustered, the model may be cached whenever load balancing is beneficial. In at least one embodiment, the inference servers may be statically loaded into the corresponding distributed servers.
[0139] In at least one embodiment, an inference server running in a container may be used to perform inference. In at least one embodiment, an instance of the inference server may be associated with a model (and optionally with multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on a model is received, a new instance may be loaded. In at least one embodiment, when the inference server is started, the model may be passed to the inference server such that the same container may be used to serve different models as long as the inference server runs as different instances.
[0140] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., an instance hosting an inference server) can be loaded (if not already loaded), and a launcher can be invoked. In at least one embodiment, the preprocessing logic in the container can load, decode, and / or perform any additional preprocessing on the incoming data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is ready for inference, the container can perform inference on the data as needed. In at least one embodiment, this can include a single inference call on an image (e.g., a hand X-ray), or may require inference on hundreds of images (e.g., chest CTs). In at least one embodiment, the application can summarize the results before completion, which can include but is not limited to a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities can be assigned to different models or applications. For example, some models can have real-time (TAT less than 1 minute) priority, while other models can have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, the model execution time can be measured from the requesting agency or entity and can include the collaborative network traversal time as well as the execution time of the inference service.
[0141] In at least one embodiment, the transfer of requests between the service 1320 and the inference application can be hidden behind a software development kit (SDK) and can provide a robust transmission via a queue. In at least one embodiment, requests will be placed in the queue via an API for an individual application / tenant ID combination, and the SDK will pull requests from the queue and provide the requests to the application. In at least one embodiment, the name of the queue can be provided in the environment from which the SDK will pick up the queue. In at least one embodiment, asynchronous communication via the queue can be useful as it can allow any instance of the application to pick up work when it is available. Results can be transferred back via the queue to ensure no data is lost. In at least one embodiment, the queue can also provide the ability to split the work, as the highest priority work can go into the queue connected to most instances of the application, while the lowest priority work can go into the queue connected to a single instance that processes tasks in the order received. In at least one embodiment, the application can run on a GPU-accelerated instance that is generated in the cloud 1426, and the inference service can perform inference on the GPU.
[0142] In at least one embodiment, a visualization service 1420 can be utilized to generate visualizations for viewing the outputs of an application and / or a deployment pipeline 1410. In at least one embodiment, the visualization service 1420 can utilize a GPU 1422 to generate visualizations. In at least one embodiment, the visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, a virtualized environment can be used to generate a virtual interactive display or environment (e.g., a virtual environment) for interaction by system users (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, the visualization service 1420 can include an in-house visualizer, movie and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, in-house optics, etc.).
[0143] In at least one embodiment, the hardware 1322 can include a GPU 1422, an AI system 1424, a cloud 1426, and / or any other hardware for executing the training system 1304 and / or the deployment system 1306. In at least one embodiment, the GPU 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) can include any number of GPUs that can be used to perform processing tasks for any features or functions of the computing service 1416, the AI service 1418, the visualization service 1420, other services, and / or software 1318. For example, for the AI service 1418, the GPU 1422 can be used to perform preprocessing on imaging data (or other data types used by machine learning models), perform postprocessing on the outputs of machine learning models, and / or perform inference (e.g., to execute a machine learning model). In at least one embodiment, the cloud 1426, the AI system 1424, and / or other components of the system 1400 can use the GPU 1422. In at least one embodiment, the cloud 1426 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1424 can use GPUs, and one or more AI systems 1424 can be used to execute the cloud 1426 (or at least part of the tasks for deep learning or inference). Similarly, although the hardware 1322 is shown as discrete components, this is not intended to be limiting, and any component of the hardware 1322 can be combined with or utilized by any other component of the hardware 1322.
[0144] In at least one embodiment, the AI system 1424 can include a specially constructed computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to a CPU, RAM, memory, and / or other components, features, or functions, the AI system 1424 (e.g., NVIDIA's DGX) can also include software (e.g., a software stack) that can use multiple GPUs 1422 to perform GPU-optimized operations. In at least one embodiment, one or more AI systems 1424 can be implemented in the cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1400.
[0145] In at least one embodiment, the cloud 1426 can include GPU-accelerated infrastructure (e.g., NVIDIA's NGC), which can provide a GPU-optimized platform for performing the processing tasks of the system 1400. In at least one embodiment, the cloud 1426 can include an AI system 1424 for performing one or more AI-based tasks of the system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the cloud 1426 can be integrated with the application coordination system 1428 that utilizes multiple GPUs to achieve seamless scaling and load balancing between and within the applications and services 1320. In at least one embodiment, as described herein, the cloud 1426 can be responsible for executing at least some of the services 1320 of the system 1400, including the computing service 1416, the AI service 1418, and / or the visualization service 1420. In at least one embodiment, the cloud 1426 can perform inference on large and small batches (e.g., execute NVIDIA's TENSORRT), provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA's CUDA), execute the application coordination system 1428 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher-quality movie effects), and / or can provide other functions for the system 1400.
[0146] Figure 15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, the following can be used as non-limiting examples Figure 14System 1400 to perform process 1500. In at least one embodiment, process 1500 may utilize services 1320 and / or hardware 1322 of system 1400, as described herein. In at least one embodiment, the refined model 1512 generated by process 1500 may be executed by deployment system 1306 for one or more containerized applications in deployment pipeline 1410.
[0147] In at least one embodiment, model training 1314 may include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data such as customer dataset 1506, and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset or deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously fine-tuned parameters (e.g., weights and / or biases) retained from a previous training, so training or retraining 1314 may not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 1314, by resetting or replacing the output or loss layer of the initial model 1504, when generating predictions on a new customer dataset 1506 (e.g., Figure 13 image data 1308), the parameters of the new dataset may be updated and readjusted based on the loss calculation associated with the accuracy of the output or loss layer.
[0148] In at least one embodiment, the pre-trained model 1406 may be stored in a data store or registry (e.g., Figure 13of the model registry 1324). In at least one embodiment, the pre-trained model 1406 may have been trained at least in part at one or more facilities other than the facility where the execution process 1500 is performed. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers of different facilities, the pre-trained model 1406 may have been trained locally using locally generated customer or patient data. In at least one embodiment, the cloud 1426 and / or other hardware 1322 may be used to train the pre-trained model 1406, but confidential, privacy-protected patient data may not be transmitted to, used by, or accessed by any component of the cloud 1426 (or other non-local hardware). In at least one embodiment, if patient data from more than one facility is used to train the pre-trained model 1406, the pre-trained model 1406 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, for example, where privacy issues have been released for customer or patient data (e.g., by waiver, for experimental use, etc.), or where customer or patient data is included in a public dataset, customer or patient data from any number of facilities may be used to train the pre-trained model 1406 locally and / or externally, such as in a data center or other cloud computing infrastructure.
[0149] In at least one embodiment, when selecting an application to use in the deployment pipeline 1410, the user may also select a machine learning model for a particular application. In at least one embodiment, the user may not have a model to use, so the user may select a pre-trained model 1406 to use with the application. In at least one embodiment, the pre-trained model 1406 may not have been optimized to generate accurate results on the customer dataset 1506 of the user facility (e.g., based on patient diversity, demographics, type of medical imaging device used, etc.). In at least one embodiment, before deploying the pre-trained model 1406 into the deployment pipeline 1410 for use with one or more applications, the pre-trained model 1406 may be updated, retrained, and / or fine-tuned for use at various facilities.
[0150] In at least one embodiment, a user may select a pre-trained model 1406 to be updated, retrained, and / or fine-tuned, and the pre-trained model 1406 may be referred to as an initial model 1504 of the training system 1304 in process 1500. In at least one embodiment, a customer dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by devices at a facility) may be used to perform model training 1314 (which may include, but is not limited to, transfer learning) on the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to the customer dataset 1506 may be generated by the training system 1304. In at least one embodiment, the ground truth data may be generated at least in part by clinicians, scientists, doctors, practitioners at the facility (e.g., clinical data 1312 such as marked as Figure 13 in).
[0151] In at least one embodiment, in some examples, AI-assisted annotation 1310 may be used to generate ground truth data. In at least one embodiment, the AI-assisted annotation 1310 (e.g., implemented using an AI-assisted annotation SDK) may utilize a machine learning model (e.g., a neural network) to generate proposed or predicted ground truth data for the customer dataset. In at least one embodiment, a user 1510 may use annotation tools within a user interface (graphical user interface (GUI)) on a computing device 1508.
[0152] In at least one embodiment, the user 1510 may interact with the GUI via the computing device 1508 to edit or fine-tune the annotation or auto-annotation. In at least one embodiment, a polygon editing feature may be used to move the vertices of a polygon to a more precise or fine-tuned position.
[0153] In at least one embodiment, once the customer dataset 1506 has associated ground truth data, the ground truth data (e.g., from AI-assisted annotation, manual marking, etc.) may be used during model training 1314 to generate a refined model 1512. In at least one embodiment, the customer dataset 1506 may be applied to the initial model 1504 any number of times, and the ground truth data may be used to update the parameters of the initial model 1504 until an acceptable accuracy level is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, the refined model 1512 may be deployed within one or more deployment pipelines 1410 at the facility to perform one or more processing tasks regarding medical imaging data.
[0154] In at least one embodiment, the refined model 1512 can be uploaded to the pre-trained model 1406 in the model registry 1324 for selection by another facility. In at least one embodiment, this process can be done at any number of facilities such that the refined model 1512 can be further refined any number of times on new datasets to generate a more general model.
[0155] Figure 15B FIG. is an example illustration of a client-server architecture 1532 for enhancing an annotation tool using a pre-trained annotation model. In at least one embodiment, the AI-assisted annotation tool 1536 can be instantiated based on the client-server architecture 1532. In at least one embodiment, the annotation tool 1536 in the imaging application can assist the radiologist, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application can include software tools, as non-limiting examples, that assist the user 1510 in identifying several extreme points on a specific organ of interest in the original image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of the specific organ. In at least one embodiment, the results can be stored as training data 1538 in the data store and used as (e.g., but not limited to) ground truth data for training. In at least one embodiment, when the computing device 1508 sends the extreme points for AI-assisted annotation 1310, for example, a deep learning model can receive this data as input and return an inference result for segmenting the organ or abnormality. In at least one embodiment, a pre-instantiated annotation tool (e.g., Figure 15B the AI-assisted annotation tool 1536B in ) can be enhanced by making an API call (e.g., API call 1544) to a server, such as the annotation assistant server 1540, which can include a set of pre-trained models 1542 stored, for example, in the annotation model registry. In at least one embodiment, the annotation model registry can store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that are pre-trained to perform AI-assisted annotation for specific organs or abnormalities. In at least one embodiment, these models can be further updated by using the training pipeline 1404. In at least one embodiment, as new labeled clinical data 1312 is added, the pre-installed annotation tools can be improved over time.
[0156] Such components can be used to generate enhanced content, such as images or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.
[0157] Other other variations are within the spirit of the present disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative constructions, certain of its embodiments shown in the drawings have been described in detail above. However, it is to be understood that the intention is not to limit the disclosure to the one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0158] Unless otherwise stated or clearly contradicted by the context, in the context of describing the disclosed embodiments (especially in the context of the appended claims), the use of the terms "a" and "an" and "the" and similar referents should be construed to cover both the singular and the plural, rather than as a definition of the terms. Unless otherwise stated, the terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to"). The term "connected" (when not otherwise modified, referring to a physical connection) should be construed to mean included in whole or in part, attached to, or joined together, even if there are some intervening elements. Unless otherwise indicated herein, references to numerical ranges in this document are only intended as a shorthand method for referring separately to each individual value falling within the range, and each individual value is incorporated into the specification as if it were recited individually herein. Unless otherwise indicated or contradicted by the context, the use of the term "set" (e.g., "set of items") or "subset" should be construed to mean a non-empty set including one or more members. Further, unless otherwise indicated or contradicted by the context, a "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but rather the subset and the corresponding set may be equal.
[0159] Unless otherwise expressly indicated or clearly contradicted by context, a conjunctive phrase such as "at least one of A, B, and C" or "at least one of A, B or C" is understood in context to typically mean that the items, clauses, etc. can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set having three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B or C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by context, the term "plurality" denotes a plural state (e.g., "a plurality of items" means multiple items). The number of items in a plurality is at least two, but can be more if expressly indicated or indicated by context. Further, unless otherwise stated or clear from context, the phrase "based on" means "at least partially based on" rather than "based solely on".
[0160] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored, for example, in the form of a computer program on a computer-readable storage medium that includes multiple instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., propagated transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, and when the executable instructions are executed by one or more processors of a computer system (i.e., as a result of being executed), the computer system performs the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the multiple non-transitory computer-readable storage media lack all of the code, but the multiple non-transitory computer-readable storage media together store all of the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, e.g., the non-transitory computer-readable storage medium stores the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of the instructions.
[0161] Thus, in at least one embodiment, a computer system is configured to implement one or more services that perform, individually or jointly, the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software that enables the implementation of the operations. Additionally, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices operating in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.
[0162] Any and all uses of examples or exemplary language (e.g., "such as") provided herein are for the sole purpose of better illuminating embodiments of the present disclosure and do not limit the scope of the disclosure, unless otherwise required. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0163] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference in their entirety as if each reference were individually and specifically indicated to be incorporated by reference and all of its contents were set forth herein.
[0164] In the specification and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not necessarily intended as synonyms for each other. Instead, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0165] Unless otherwise explicitly stated, it is understood that throughout the specification, terms such as "processing", "computing", "calculating", "determining", etc., refer to actions and / or processes of a computer or computing system or similar electronic computing device that process and / or transform data represented as physical quantities (e.g., electrons) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.
[0166] In a similar manner, the term "processor" may refer to any device or portion of a memory that processes electronic data from registers and / or memories and converts that electronic data into other electronic data that may be stored in registers and / or memories. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes that execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" may be used interchangeably herein, provided that a system can embody one or more methods and a method can be considered a system.
[0167] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. There are various ways to obtain, acquire, receive, or input analog and digital data, such as by receiving data as an argument to a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transferring, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transferring, sending, or presenting analog or digital data may be implemented by transmitting the data as an input or output argument to a function call, an application programming interface, or an interprocess communication mechanism.
[0168] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. Additionally, although specific assignments of responsibilities were defined above for purposes of discussion, the various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.
[0169] Moreover, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A computer-implemented method, comprising: generating one or more motion vectors for a current frame; performing a back-projection pass for a group of two or more adjacent pixels in the current frame using at least one previous frame among one or more previous frames; using the back-projection pass and the motion vectors to locate one or more common matching surfaces between the current frame and the one or more previous frames; patching a G-buffer of the current frame at least in part based on information corresponding to the one or more matching surfaces; determining one or more light differences between the current frame and the at least one previous frame at least in part based on the patched G-buffer; rendering an image at least in part based on the one or more light differences; and outputting the rendered image for display on a display device.
2. The method according to claim 1, wherein the one or more motion vectors are generated during at least one of the following: rendering the G-buffer; rendering one or more reflection images; or rendering one or more refraction images.
3. The method according to claim 2, wherein the rendering of one or more reflection images or the rendering of one or more refraction images is performed using primary surface replacement (PSR).
4. The method according to claim 1, wherein performing the back projection comprises: For at least one tile in the current frame, selecting at least one pixel corresponding to a matching surface in the at least one previous frame.
5. The method according to claim 4, wherein the at least one tile comprises square pixels.
6. The method according to claim 4, wherein selecting the at least one pixel comprises: When multiple pixels correspond to the matching surface, selecting the pixel with the highest light value.
7. The method according to claim 4, wherein selecting the at least one pixel further comprises: Locating a pixel in the at least one previous frame that matches the selected pixel.
8. The method according to claim 1, wherein patching the G-buffer comprises: Writing one or more parameters of the matching surface of the G-buffer from the at least one previous frame into gradient pixels of the G-buffer of the current frame.
9. The method according to claim 8, wherein the one or more parameters correspond to parameters used to calculate surface illumination.
10. The method according to claim 8, wherein the one or more parameters comprise at least one of the following: a random generator seed; a normal value; a metallicity value; or a roughness value.
11. The method according to claim 8, wherein patching the G-buffer comprises: Calculating a new position of a first surface depicted in the at least one previous frame at least based on visibility buffer information from a previous frame.
12. The method according to claim 11, wherein the visibility buffer information comprises at least one of the following: mesh information corresponding to the first surface; triangle information corresponding to the first surface; barycentric coordinate information corresponding to the first surface; or one or more updated vertex buffers corresponding to the first surface.
13. The method according to claim 1, wherein rendering the image comprises: performing one or more lighting passes for the current frame; calculating one or more temporal gradients as representing the one or more light differences at least in part based on the output of the one or more lighting passes for the current frame and calculated lighting information corresponding to the at least one previous frame; filtering the one or more temporal gradients; and Use one or more filtered temporal gradients for historical rejection.
14. A computer-implemented system, comprising: One or more processors; And One or more memory devices that store instructions which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: Generate motion vectors for a current frame; Using at least one previous frame among one or more previous frames, perform a back-projection pass for a group of two or more adjacent pixels in the current frame; Using the back-projection pass, use the motion vectors to locate one or more common matching surfaces between the current frame and the one or more previous frames; Patch the G-buffer of the current frame based on information corresponding to the one or more matching surfaces; Determine one or more light differences between the current frame and the at least one previous frame based at least in part on the patched G-buffer; Render an image based at least in part on the one or more light differences; and Output the rendered image for display on a display device.
15. The system of claim 14, wherein the system comprises at least one of the following: A system for performing simulation operations; A system for performing simulation operations to test or verify autonomous machine applications; A system for rendering graphical output; A system for performing deep learning operations; A system implemented using edge devices; A system incorporating one or more virtual machines (VMs); A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
16. The system of claim 14, wherein the motion vectors are generated during at least one of the following: Rendering the G-buffer; Rendering one or more reflection images; or Rendering one or more refraction images.
17. The system of claim 16, wherein the rendering of one or more reflection images or the rendering of one or more refraction images is performed using primary surface replacement (PSR).
18. The system according to claim 14, wherein performing the backprojection comprises: For at least one tile in the current frame, select at least one pixel corresponding to a matching surface in the previous frame.
19. The system of claim 18, wherein the at least one tile comprises 3x3 square pixels.
20. The system of claim 18, wherein selecting the at least one pixel comprises: When multiple pixels correspond to the matching surface, select the pixel with the highest illumination value.
21. The system according to claim 18, wherein selecting the at least one pixel further comprises: Locate a pixel in the at least one previous frame that matches the selected pixel.
22. The system according to claim 14, wherein patching the G-buffer comprises: Write one or more parameters of the matching surface of the G-buffer from the at least one previous frame into the gradient pixels of the G-buffer of the current frame.
23. The system of claim 22, wherein the one or more parameters correspond to parameters used to calculate surface illumination.
24. The system of claim 22, wherein the one or more parameters comprise at least one of the following: Random generator seed; Normal value; Metallicity value; or Roughness value.
25. The system according to claim 22, wherein patching the G-buffer comprises: Calculate a new position of a first surface depicted in the at least one previous frame based at least on visibility buffer information from the at least one previous frame.
26. The system according to claim 25, wherein the visibility buffer information includes at least one of the following: Mesh information corresponding to the first surface; Triangle information corresponding to the first surface; Centroid information corresponding to the first surface; Or One or more updated vertex buffers corresponding to the first surface.
27. The system according to claim 14, wherein rendering the image includes: Performing one or more lighting passes for the current frame; Calculating one or more temporal gradients as representing the one or more light differences based on the output of the one or more lighting passes for the current frame and the calculated lighting information corresponding to the at least one previous frame; Filtering the one or more temporal gradients; And Using the one or more filtered temporal gradients for history rejection.
Citation Information
Patent Citations
High-speed pixel matching method based on LK optical flow
CN111640084A
Method for photographing panoramic picture
US20090058991A1