Increasing level of detail of neural field using diffusion model
By combining heuristic methods and user prompts in the image super-resolution model and diffusion model, the problem of image blurring and insufficient details during 3D scene generation in the prior art is solved, and image generation at higher resolution and detail levels is achieved.
Patent Information
- Application Number
- CN202411637347.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art can easily lead to blurred image or impractical details when generating complex 3D scenes with high resolution, limiting the level of detail and resolution of objects.
The blur of the image is determined by using one or more heuristics and provided to a pre-trained image super-resolution model to generate higher resolution images. At the same time, combining the diffusion model and the NeRF model, user prompts are used to generate images of higher-level details.
It achieves improving the resolution and level of detail of 3D objects without damaging image quality, providing higher quality images and finer details, and enhancing the effect of content generation.
Smart Images

Figure CN120020885A_ABST
Abstract
Description
BACKGROUND OF THE DISCLOSURE
[0001] In various applications (such as gaming, animation, or virtual reality content generation), it may be beneficial (if not necessary) to render complex three-dimensional (3D) objects in a way that appears substantially realistic or at least consistent to a human viewer. Machine learning has improved the ability to generate new views of complex 3D scenes, such as by using methods based on neural radiance fields (NeRF), which can use a model of a 3D object generated from two-dimensional (2D) images of the object to render new views of the 3D object or environment. Additionally, due to the methods of generating these objects, these objects are theoretically not limited by resolution. However, in practice, the optimization process and the source materials for content generation are limited in terms of resolution, and thus, attempting to reconstruct finer details may result in blurry or unrealistic images. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Various embodiments in accordance with the present disclosure will be described with reference to the accompanying drawings, in which:
[0003] Figure 1 An example network for object generation in accordance with at least one embodiment is shown;
[0004] Figure 2A An example environment for object generation and interaction in accordance with at least one embodiment is shown;
[0005] Figure 2B An example environment for modifying the level of detail of an object in accordance with at least one embodiment is shown;
[0006] Figure 2C An example environment for modifying the level of detail of an object in accordance with at least one embodiment is shown;
[0007] Figure 3A An example environment for generating an object at a higher level of detail than an initial object in accordance with at least one embodiment is shown;
[0008] Figure 3B An example pipeline for generating a higher resolution object in accordance with at least one embodiment is shown;
[0009] Figure 4A An example environment for generating an object at a higher level of detail than an initial object in accordance with at least one embodiment is shown;
[0010] Figure 4B An example pipeline for generating the level of detail of an object in response to a prompt in accordance with at least one embodiment is shown;
[0011] Figure 5A An example process for generating an object with a higher level of detail is shown;
[0012] Figure 5B Illustrates an example process for generating an object with a higher level of detail according to at least one embodiment;
[0013] Figure 5C Illustrates an example process for generating an object with a higher level of detail according to at least one embodiment;
[0014] Figure 6 Illustrates components of a distributed system that can be used to update or perform inference using a machine learning model according to at least one embodiment;
[0015] Figure 7A Illustrates inference and / or training logic according to at least one embodiment;
[0016] Figure 7B Illustrates inference and / or training logic according to at least one embodiment;
[0017] Figure 8 Illustrates an example data center system according to at least one embodiment;
[0018] Figure 9 Illustrates a computer system according to at least one embodiment;
[0019] Figure 10 Illustrates a computer system according to at least one embodiment;
[0020] Figure 11 Illustrates at least a portion of a graphics processor according to one or more embodiments;
[0021] Figure 12 Illustrates at least a portion of a graphics processor according to one or more embodiments;
[0022] Figure 13 Is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;
[0023] Figure 14 Is a system diagram of an example system for training, adapting, instantiating, and deploying a machine learning model in an advanced computing pipeline according to at least one embodiment; and
[0024] Figure 15A and Figure 15B Illustrates a data flow diagram of a process for training a machine learning model according to at least one embodiment, and a client - server architecture that utilizes a pre - trained annotation model to enhance an annotation tool. Detailed Description
[0025] In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, those skilled in the art will also understand that the embodiments may be practiced without specific details. Additionally, well-known features may be omitted or simplified so as not to obscure the described embodiments.
[0026] The systems and methods described herein can be used without limitation by the following: non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in in-vehicle infotainment or digital or driver virtual assistant applications), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones) and / or other vehicle types. Additionally, the systems and methods described herein can be used for a variety of purposes such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or participant simulation and / or digital twins, data center processing, conversational artificial intelligence (AI), generative artificial intelligence with large language models (LLMs), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing and / or any other suitable applications.
[0027] The disclosed embodiments can be incorporated in a variety of different systems such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems that include one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing generative AI operations using LLMs, systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least in part using cloud computing resources and / or other types of systems.
[0028] Methods according to various embodiments can be used to increase the resolution of one or more three-dimensional (3D) objects or scenes represented by a neural radiance field (NeRF). A NeRF can be used to learn a plenoptic function (e.g., a five-dimensional (5D) function taking 3D spatial coordinates and two or more view-related angles as input) to output a four-dimensional (4D) radiance field. The 4D radiance field can be used to display the density and color of the light quality at a given position when viewed from a certain direction. Thus, an object represented by a NeRF can capture view-related effects, such as specular reflections of the object. For faster learning and rendering, a NeRF can be represented or visualized as a voxel grid, where features are stored at the corners of voxels (e.g., using an octree as a non-limiting example). When presenting an object, a user or program associated with the object can provide commands to obtain a further level of detail, such as zooming in on the object and / or obtaining finer details about the components of the object, among other options. As the level of detail increases (e.g., the user "zooms in" on the object or "zooms in" towards the object), the resolution will be limited by the finite resolution of the images in the set used to generate the NeRF. The system and method address this problem by providing improved resolution for different levels of detail. In at least one embodiment, one or more heuristics can be used to determine that an image is "blurry" (e.g., below a threshold sharpness or resolution), and then the image can be provided to a pre-trained image super-resolution model to generate a new high-resolution image at a given level of detail (e.g., higher than the resolution of the input image, higher than the threshold resolution). This new image can be provided to the user for viewing, or added to enhance the set of images used to generate the NeRF, and then the image set can be used to further train or update the NeRF model to generate higher-quality images at different levels of detail. Additionally, the level of detail can also refer to specific components that form the object. For example, in at least one embodiment, a NeRF can be used to hallucinate new images to obtain a finer level of detail, such as by pairing the NeRF with a diffusion model that can accept a prompt and then generate a new image. For example, a user can zoom in on an image and then provide a prompt about the information they want to view, such as a prompt defining the image or an image region to obtain a higher level of detail. The prompt can be provided to various models, such as a large language model or a predefined model or image set with a hierarchy of levels of detail. From there, a diffusion model can be used to generate a new image to provide to the user for viewing, and the image can be added to a set of images to update or fine-tune an existing NeRF model or train or update a new NeRF model, and then the generated NeRF model can be enabled to provide a higher level of detail. In various embodiments, alignment can also be considered to provide multi-view Figure 1 consistency, such as by selecting different features in the image when hallucinating new images. In this way, a NeRF can be used to generate a finer level of detail.
[0029] A variety of systems and methods can be used to form a content generation pipeline, which can include one or more diffusion models and / or image super-resolution models to optimize NeRF for generating one or more objects, which can be represented as 3D volumes. For example, the content generation pipeline can be used to generate 3D objects or scenes in response to requests (such as prompts or commands to provide renders, among other options). Thereafter, the user can interact with the generated objects or scenes. For example, in an interactive environment, the user can perform actions such as panning, rotating, zooming, etc. As described herein, while the theoretical resolution of objects generated using NeRF may be infinite, practical considerations (such as the resolution of the images used to train the model) will effectively limit the resolution, and at some point, the objects and / or associated images may become blurry or otherwise be considered "low (lower) resolution". The systems and methods of the present disclosure address these issues by providing methods to increase the image resolution at certain levels and / or generate new content showing different levels of detail of a particular image or scene. In at least one embodiment, one or more trained models (such as diffusion models, image super-resolution models, etc.) can form part of the pipeline to receive inputs (such as images, prompts, commands, triggers from one or more workflows, etc.) associated with a certain level of detail of an image. Thereafter, one or more models can be used to provide additional levels of detail, such as by increasing the resolution of the image and / or generating new images associated with components of the original image. In this way, different levels of detail can be generated for various objects and scenes, and moreover, it can be used to retrain or improve the training datasets of NeRF and / or other content generation models.
[0030] Figure 1 An example pipeline 100 that can be used to generate an object 102 (such as an object forming at least a part of a scene) is shown. A set of images 104 can be used to train a NeRF network 106 (which can include one or more diffusion models or various types of image and / or object generation models) to generate different output objects, such as outputs based on input prompts or requests. For example, the images can be one or more objects or scenes captured from different angles. The positions and orientations of the viewpoints can be calculated and used to train the network. In this example, the NeRF network 106 includes one or more neural networks, such as a multi-layer perceptron (MLP), but it should be understood that other networks can be incorporated into the NeRF network 106 and / or be accessible to the NeRF network 106. For example, one or more LLMs can be associated with the NeRF network 106 to receive and process input prompts. The NeRF network 106 can be used to generate data related to color 108 and volume density 110, which can be provided to a renderer 112 to generate an object 202, such as a 3D representation of the object 102.
[0031] The object 102 represented by NeRF may be different from typical representations using grid-based methods (e.g., voxel grids, polygon meshes, etc.). However, NeRF can be converted to a mesh using various operations (e.g., marching cubes method). Meshes typically include faces and vertices, and thus may be difficult to manipulate as different networks are needed to estimate or otherwise determine an appropriate number of faces and vertices to provide a discrete representation of the object. On the other hand, NeRF is a neural representation of how different points appear in a given camera view and thus lacks the faces and vertices common to mesh representations. The object 102 can be used in various applications, such as for rendering and manipulation within an interactive environment.
[0032] In at least one embodiment, the neural field generated using the NeRF network 106 can be defined for all spatial and / or temporal coordinates, which can be represented as a mapping from a given coordinate to a quantity (e.g., scalar, vector, or tensor). In operation, the NeRF network 106 can be trained by sampling coordinates of a scene (e.g., from image 104) and providing these coordinates to a neural network to process field quantities, and then sampling these field quantities for a given reconstruction domain. The reconstruction can then be mapped back to the sensor domain (e.g., a domain considering depth and normals), which can be a set of 2D RGB images. An error rate can then be computed to optimize the network. Thus, the pipeline for training can include coordinate sampling, neural network processing for the radiance field and reconstruction domain, volume rendering, differentiable forward mapping, and then optimization for the RGB images and sensor domain. The reconstruction can be a neural field that maps world coordinates to field quantities, and sensor observations (e.g., cameras, microphones, 2D images, etc.) can be transformed into measurements for forward mapping (e.g., volume rendering) to the reconstruction. As previously mentioned, NeRF can receive a single continuous 5D coordinate as input to provide a spatial location and view direction, and then feed the spatial location and view direction through an MLP to output a corresponding color intensity and volume density. The volume density can be an indication of how much radiance or luminance has been accumulated along a ray passing through a given 3D coordinate point, and a measure of the impact of the given 3D coordinate point on the overall scene. That is, the volume density is used to determine the likelihood that the predicted color value should be considered when rendering the scene / object. During training of NeRF, the target density and color may be unknown, and thus, these features are mapped back to the input 2D image, compared with the ground truth image, and then optimized using the computed loss.
[0033] Figure 2AAn example environment 200 that can be used with embodiments of the present disclosure is shown. The illustrated environment 200 includes an object 202 within an interaction environment 204 that is being viewed by one or more users and / or associated with a content generation pipeline, including but not limited to content generation for video games, AR, VR, MR, online shopping, kiosks, etc. The object 202 can be part of a 3D representation of a scene 206, which can include various additional objects or features. In at least one embodiment, a user can interact with the object 202 and / or portions of the scene 206, such as by inputting requests to perform operations such as rotating to a different viewpoint, panning, or zooming in.
[0034] The systems and methods of the present disclosure can be directed to improving the level of detail (LOD) provided for an object and / or a scene. The LOD can correspond to a change in perspective, which can also be referred to as “zooming” or “scaling,” where, for example, the appearance from a digital camera can represent a change in the focal length of the camera relative to an object 202 or a scene 206 at a given location. That is, the LOD can refer to interacting with the fine details of a given object such that at a first LOD, certain features may not be visible, but when the LOD is increased by one or more levels, these same features may become visible. By way of non-limiting example, a pineapple can be visible and recognizable by a user at a first LOD, but the user may not be able to determine fine details, such as the number of small fruits on the pineapple, or read the text on a pineapple label. By increasing the LOD, the user can view the object as “closer” and / or view certain parts of the object as “larger,” and thus, certain features may appear larger, more prominent, clearer, or more precise within a given view area, providing the user with more information and details to distinguish finer characteristics and features of the pineapple. Accordingly, the systems and methods of the present disclosure can be used to generate different LODs for a given object (such as the object 202 rendered using one or more NeRFs), as well as various other options.
[0035] Various embodiments may also refer to the LOD of individual characteristics of one or more objects or scenes, which may be step changes and / or based on different characteristics of the object. Returning to the example of the pineapple, the LOD may refer to different parts of the pineapple, such as the plant, the fruit itself, the small fruits that make up the fruit, the seeds, the cells of the plant, the deoxyribonucleic acid (DNA) that forms the chromosomes, etc. In this way, the LOD can refer to the view at a given predefined level. For example, the first level may correspond to the view at the first level of detail, and the zoom request may move to the second level, which shows finer details associated with the second predefined level, and so on. In other words, the LOD can be presented as a series of nesting dolls or information hierarchies, where different levels provide smaller or finer details in a predefined manner. In at least one embodiment, the system and method may use one or more diffusion or other image generation models to generate these further LODs, for example, by pairing such a system with other models (including but not limited to LLMs) to parse the input from the user, prepare prompts for the image generation model, and then use the image generation model to generate and / or define other LODs. In this way, a predefined set of hierarchical views can be established for different objects, where there may be step changes between different LODs.
[0036] Figure 2B An example sequence 210 is shown, where the LOD changes from the first level 212 to the second level 214 and then to the third level 216. In this example, the LOD may change due to a user input of "zoom" or other commands that modify the perspective. At the first level 212, the entire object 202 is visible, which is still a pineapple in this example. At the second level 214, only a part of the object 202 is visible, but now, finer details of the object 202 are shown. For example, the crown 218 in the first level 212 is no longer shown, but now the individual small fruits 220 are shown with higher resolution and clarity. In other words, certain regions related to the zoom command now appear closer and larger within the view area. Additionally, at the third level 216, more details about the small fruit 220 are shown, such as the seeds 222 shown within the small fruit 220.
[0037] In various embodiments, this increased level of detail may be desirable, but if the initial resolution of the generated object 202 is below a threshold (e.g., the resolution of the 2D images used in the NeRF model), the finer details may be blurred or unclear for other reasons. For example, the third level 216 shows some pixelation / blurring around the seed 222. Embodiments of the present disclosure can overcome this problem by identifying blurriness and / or resolution below the threshold and then modifying the image (e.g., using a trained super-resolution model) to provide a higher-resolution image (e.g., an image with a resolution greater than the threshold). Additionally, the system and method can also use the newly generated image as input training data for NeRF, thereby providing higher-resolution training images to train NeRF to generate higher-resolution output objects.
[0038] Figure 2C An example sequence 230 is shown where the LOD changes from a first level 232 to a second level 234, a third level 236, a fourth level 238, and a fifth level 240. In this example, the different levels can correspond to the nested doll configuration described herein, where the levels are predefined for a given object 202. For example, the first level 232 shows a pineapple plant in a field, the second level 234 shows a single pineapple, the third level 236 shows the small fruits 220, the fourth level 238 shows the cells 242 of the pineapple, and the fifth level 240 shows the DNA 242. Such a sequence can be provided as part of a learning module, by way of example only, to show the different levels of a given object, e.g., in an educational setting. Thus, the process of "zooming in" on the object 202 in the sequence 230 may not change the perspective but rather completely change the image to show different levels within a predefined level of detail. For example, Figure 2C the sequence 230 can be a stored hierarchy provided in response to a request, such as, by way of example, a learning module or within an educational series, and the stored hierarchy can be generated and prepared for interaction with the user. In at least one embodiment, the sequence 230 can also be provided as an option for the user to explore or otherwise interact with the object formed by the various components. By another non-limiting example, if the initial object is a car at the first level, the second level can illustrate the engine, and the third level can illustrate parts of the engine, such as an exploded view of the piston cylinder arrangement, and so on. In this way, a predefined hierarchy can be established for interaction and then provided in response to one or more inputs or prompts.
[0039] In various embodiments, one or more objects generated at levels 232 - 240 can be generated at least in part using one or more generative networks (such as diffusion networks). For example, a first level 232 can be presented to the user, and the user wishes to learn more about the object (e.g., a pineapple) shown at the first level 232, and then an action or prompt can be provided to move to the second level 234. In at least one embodiment, the action or prompt can be an input, such as a scroll wheel to move to the next level or a click on an arrow. In at least one embodiment, the action or prompt can be a text or voice prompt, such as "What is a pineapple made of?", which can generate different levels of detail, such as a third level 236, to illustrate individual small fruits 220. The user can continue to interact with different levels of detail, which can be generated in real - time or near real - time as the user interacts with the object. Additionally, and / or alternatively, different levels can be predetermined and defined for interaction with the user.
[0040] The system and method can also generate in real - time or near real - time (e.g., with no significant delay) to enable the user to interact with a given object and / or ask questions about a given object. For example, the user can converse with an LLM to ask a series of questions about an object, and the answers can be used as prompts to generate an image or object, and one or more image - generation systems can be used to hallucinate the image or object. Thus, the system and method can integrate various additional models and then select one or more rendering techniques based on the type of information the user is seeking.
[0041] Various embodiments of the present disclosure can integrate one or more additional trained models with a NeRF network in order to generate different LODs from the initial generation of the NeRF network. Thus, the system and method can target a content - generation pipeline where an input (e.g., user input) can be used to generate a finer LOD from an initial object. In at least one embodiment, one or more thresholds can be used to determine when an additional network is suitable for generating finer details. However, in various embodiments, a predefined set of levels can be used.
[0042] Figure 3AAn example environment 300 is shown that can be used with embodiments of the present disclosure to increase the resolution of one or more objects generated using a content generator 302. In this example, an input 304 is provided to the content generator 302, such as a text input, an auditory cue, a command request, etc. The input can come from a user interacting with the environment, for example, the user provides a cue to the environment to generate one or more objects or scenes, which can be 3D objects or scenes, and / or can include audio or video. The cue can be a text cue, such as a cue provided to a text-to-image generator, a voice cue, a converted voice cue (e.g., a voice cue converted to text), a command cue (e.g., a command to generate a random image, a command to generate an image specifically for generator training), and / or a combination thereof. In at least one embodiment, the input 304 is provided without direct human-machine interaction, such as within a content generation pipeline. For example, an initial control input can be provided, such as "generate a lake scene", and then a single object can be provided as a sub-input that is not directly generated by a person or user, such as a single input of "lake" or "surrounded by trees" or "add a boat", etc.
[0043] The content generator 302 can include one or more image generation models, which are shown as NeRF models 106 in the example depiction, but can include additional models. Additionally, the NeRF model 106 can be a representation of multiple content generation models as one (e.g., a diffusion model) to generate one or more images that can be used to represent an object as a NeRF. The input 304 can be provided to the NeRF model 106, which can generate one or more representations 306 of the object and / or scene in response to the input. The representation 306 can, for example, correspond to a 3D object viewable from a certain camera view. In at least one embodiment, the representation 306 is directly provided to a renderer 112, which can render the object at least in part based on instructions or information associated with the representation 306. However, in one or more embodiments, the representation 306 can first be evaluated by a resolution evaluator 308 before rendering. Similarly, the output of the renderer 112 can also be evaluated by the resolution evaluator 308 before the object 102 is provided for the user to view or use.
[0044] In at least one embodiment, the resolution evaluator 308 determines whether the resolution or clarity of the representation 306 and / or the object 102 exceeds a threshold level. Resolution can refer to the detail contained in an image and can be measured in a variety of ways that can be used in conjunction with embodiments of the present disclosure. Resolution can include one or more measurements to quantify how close lines can be to each other and still be visibly distinguishable, for example, in lines per millimeter or lines per inch. Additionally, measurements can be evaluated by the overall image size (e.g., lines per picture height) or angular subtense. Additionally, line pairs can be used as a measure of resolution (e.g., line pairs per millimeter). Pixel count is another way to describe resolution (e.g., number of pixel columns times number of pixel rows), where a higher determined value of pixels per inch (e.g., number of pixel columns times number of pixel rows) may indicate higher resolution. Another measure of resolution can be for spatial resolution and its factors, such as the determination of "blurriness" or "sharpness" as described herein, which can also be a factor in pixels per inch determination. Additionally, one or more standards organizations can set different resolutions, such as "standard definition" or "high definition", etc.
[0045] In at least one embodiment, the resolution evaluator 308 can determine one or more measurements of an image or image representation corresponding to the resolution, such as an evaluation of pixels per inch from a given view direction. This information can be determined at least in part on the image used to generate the representation 306. For example, if the initial input image is a low-resolution image, the resulting output NeRF model may also be low-resolution, at least at a finer level of detail, such as when the user zooms in. Thus, the system and method can determine the resolution associated with the representation 306 and then determine a possible scaling level of the representation to maintain a resolution above the threshold. If a command for a further level of detail is received, the resolution evaluator 308 can determine that a new image and / or representation should be generated to maintain image quality, such as by using a trained image super-resolution model 310. In at least one embodiment, the super-resolution model can be used to enhance the resolution from low resolution to high resolution, where "low" and "high" are at least partially based on a comparison between the initial input and output. Various models can use degradation functions as well as one or more neural networks to find the inverse function of the degradation, which can include methods such as pre-upsampling super-resolution, post-upsampling super-resolution, residual networks, multi-level residual networks, recursive networks, progressive reconstruction networks, multi-branch networks, attention-based networks, generative models, etc. Then, the updated representation 312 can be provided to the renderer 112 using the super-resolution model 310 to render it as the object 102.
[0046] As Figure 3AAs shown, various embodiments of the present disclosure may deploy the super-resolution model 310 based on the evaluation of the representation 306 from the NeRF model 106 and / or based on the evaluation of the output rendering from the renderer 112. For example, before providing the output to the user, the renderer 112 may provide the output object 102 to the resolution evaluator 308, which may determine whether the resolution of the output object reaches or exceeds a threshold resolution, and then prompt the super-resolution model 310 to generate an updated representation 312 for rendering and presentation based on that determination. It should be understood that the updated representation 312 may be a single image or a NeRF model or NeRF representation, depending at least on the provided input and the selected method. For example, the updated representation 312 may be provided back to the NeRF model 106 to generate a new representation 306 using a higher-resolution image.
[0047] Various embodiments may also use the images generated by the super-resolution model 310 to improve the NeRF model 106. For example, the output of the model (e.g., the representation 312) may also be provided back to the NeRF model 106 for storage and later use, where it may be used as one of the images for training or generating the representation 306. In this way, high-resolution images can be used for training to generate high-resolution representations 306.
[0048] Figure 3B An example pipeline 320 that may be used with embodiments of the present disclosure is shown. In this example, the input 304A corresponds to a command to render an image and / or an object (in this example, a "pineapple"). The command may come from a user input command, such as a command input to a text-to-image model, or from a part of a workflow to render one or more images or objects to be placed in a scene, such as a graphics pipeline for a game that renders objects in the scene at least partially based on scene information. Thus, the systems and methods may be used in embodiments with direct user input and / or input in response to one or more additional commands, as well as other options and combinations. In this example, the NeRF model 106 may generate an output 102A corresponding to an image of the pineapple. The pineapple may be generated using one or more trained machine learning systems, such as a 2D or 3D object to be placed in a scene.
[0049] Receive another input 304B corresponding to a "zoom" or change in the perspective of object 102A. For example, object 102A can be presented to the user, such as using a renderer that can provide a visible representation of the object in the environment. The user and / or workflow can present a second command 304B to zoom in on object 102A, which can be a command to provide more detailed details for one or more features that form object 102A. The zoom command can be a user command, such as using a combination of a scroll wheel or keyboard commands, or the zoom command can respond to one or more other commands, such as the user selecting an object in a video game and then the workflow performing a zoom in towards that object to provide more detailed details, and other options. As described herein, a second object 102B can be prepared for rendering and presented to the user, but it is also possible to perform one or more resolution evaluations, such as using an evaluator 308, before or after rendering. The evaluator 308 can determine one or more measurable aspects of the second object 102B to determine whether the second object 102B meets and / or exceeds a threshold. For example, the pixels per inch of the second object 102B can be determined and compared to a threshold, which can be determined at least in part based on one or more properties of the computing device used by the user to view the object. For example, if it is determined that the user has a super high-definition monitor, the threshold may be greater than that of a user operating with a standard definition monitor. In another example, the settings selected by the user for the execution of the environment can also be used to determine the threshold, and other options.
[0050] Along a first path (labeled "1"), it can be determined that the resolution of the second object 102B reaches or exceeds the threshold, and thus, the second object can be prepared for rendering and / or presented to the user. Along a second path (labeled "2"), it can be determined that the resolution of the second object 102B does not reach or exceed the threshold, and thus, a super-resolution model 310 can be used to increase or otherwise enhance the resolution and generate a third object 102C (as an image of the NeRF model 106 and / or a single image for presentation), which can be further evaluated (labeled "3"), and then, if the resolution reaches or exceeds the threshold, the third object can be presented to the user (labeled "4") or the third object can be further processed along the second path (labeled "2"). In this way, additional commands can be received, then the resolution can be checked, and additional objects can be generated in response to the results of the resolution evaluation.
[0051] Figure 4AAn example environment 400 is shown that can be used with embodiments of the present disclosure to generate content in response to a request to increase the LOD using one or more models. Various embodiments include the LLM 310 and / or the NeRF model 106, and may also include one or more image generation models, such as diffusion models. In this example, an input 304, such as a question about an object visible to the user, is provided to the LLM 310. For example, the user may view an object and ask a specific question about its composition, such as when paired in an educational program. The LLM 310 may receive the input and generate a prompt 402, which may be passed to the NeRF model 106, which in at least one embodiment includes a diffusion model, to generate the object 102. In this way, the user can provide an input about an image or object, and one or more models can hallucinate additional details to generate additional information that does not exist in the original object in at least some embodiments.
[0052] Figure 4B An example pipeline 420 that can be used in embodiments of the present disclosure is shown. In this example, a starting object 422 is provided to the user, which can be an object generated by one or more trained models and / or can be an object that has been selected as the start of a series of objects at a predetermined level, as described herein. The input 304A in this example is in the form of a question asking what type of field is shown in the starting object 422. The input 304A is provided to the LLM 310, which can determine information related to the image. For example, the LLM 310 is a multimodal model that can evaluate the image and provide a response about one or more objects within the image, and can generate a prompt 402A corresponding to "pineapple" as an answer to the input 304A. Then, the NeRF model 106 can receive the prompt 402A to generate an object 102A corresponding to a pineapple. In at least one embodiment, the object 102A includes a degree of detail to allow the user to interact, such as zooming in or otherwise viewing different features. The user can zoom in on the output 102A to form a second object 102B, which may be from a different perspective direction or angle, and can generate a second input 304B pointing to the features of the second object 102B. For example, the second input 304B can be a question, such as "What is a pineapple made of?", which can be sent back to the LLM 310 to generate a second prompt 304B to answer the question. The second prompt 304B answers that a pineapple is made of small fruits, and then provides it to the NeRF model 106 to generate a third object 102C corresponding to small fruits. Then, other questions may be sent back to the LLM 310, providing different levels of detail for various interactions.
[0053] Figure 5AAn example process 500 is shown, which can be used to generate new frames to obtain a greater level of detail (LOD) and update a set of training images. It should be understood that for this and other processes described herein, within the scope of various embodiments, there may be additional, fewer, or alternative operations that are performed in a similar or alternative order or at least partially in parallel, unless otherwise explicitly stated. Additionally, although this example relates to NeRF and using prompts and using diffusion models to generate content, it should be understood that various other such tasks can benefit from aspects of various embodiments and can also use various different model representations and / or generative models. In this example, a target level of detail 502 for the 3D volume is determined. The 3D volume can be associated with NeRF, but various embodiments can also be used with other 3D volumes that can be converted to NeRF and / or 3D volumes converted from NeRF. The current view can be a frame or view of the 3D volume that represents the 3D volume, and the current view can be provided to an image generation network 504. The image generation network can include one or more trained models, such as a diffusion model, which can take an input image or command and generate one or more images associated with the input image or command. For example, a text-to-image model can take an input text prompt and generate an image associated with that prompt. Similarly, a super-resolution image model can take an input image and then generate an output image with a higher resolution. An updated view representing the 3D volume can be generated using the current view 506 and provided for viewing 508, such as provided to a user. The updated view can then be added to a set of images associated with the 3D volume 510, and then the 3D volume and / or a model associated with the 3D volume can be trained to generate additional images and / or views using the newly generated images 512.
[0054] Figure 5B An example process 520 for generating an object in response to a request is shown. In this example, a request to generate an object using NeRF 522 is received. The request can include an associated prompt or command, such as a specific question associated with the object and / or a command to perform one or more actions on the object. Based on the prompt, a level of detail for the object can be determined 524. For example, the level of detail can be associated with the perspective, features of the object, and / or a combination thereof. An object at the target level of detail can be generated 526, such as by using one or more diffusion models. The generated object can then be provided 528 in response to the request.
[0055] Figure 5CAn example process 540 for generating an object at a higher level of detail in response to a request is shown. In this example, a command 542 to change the level of detail of the object is received. For example, a user may enter a command to zoom in (e.g., change the perspective) on an object, which may be a 3D object represented by a NeRF. Additionally, a command to decrease the level of detail may be provided, such as instructions on how one or more components fit into the system. A representation of the object may be generated at the target level of detail 544, and the resolution of the representation at the target level of detail may be determined 546. For example, the number of pixels per inch may be determined for the representation.
[0056] In at least one embodiment, the resolution may be compared to a threshold 548. If the resolution exceeds the threshold, a visual representation of the object may be presented at the target level of detail 550. If the resolution is less than the threshold, the representation may be provided to a trained super-resolution 552, which may generate a second representation of the object with a higher resolution 554. Optionally, the second representation at the higher resolution may then be evaluated to determine a second resolution of the second representation 556, and the second representation may be compared to the threshold, repeating the process until a stopping condition is reached.
[0057] As discussed, aspects of the various methods presented herein may be lightweight enough to be executed in real time on a device such as a client device like a personal computer or gaming console. Such processing may be performed on content generated on or received by the client device or content received from an external source, such as streamed data or other content received via at least one network. In some cases, the processing and / or determination of the content may be performed by one of these other devices, systems, or entities and then provided to the client device (or another such recipient) for presentation or for other such uses.
[0058] As an example, Figure 6An example network configuration 600 is shown that can be used to provide, generate, modify, encode, process, and / or transmit image data or other such content. In at least one embodiment, a client device 602 can use components of a control application 604 on the client device 602 and data locally stored on the client device to generate or receive session data. In at least one embodiment, a content application 624 executing on a server 620 (e.g., a cloud server or an edge server) can initiate a session associated with at least one client device 602, can utilize a session manager and user data stored in a user database 636, and can cause a content manager 626 to determine content from an asset repository 634, such as one or more digital assets (e.g., object representations). The content manager 626 can work with an image synthesis module 628 to generate or synthesize new objects, digital assets, or other such content to be provided via the client device 602 for rendering. In at least one embodiment, the image synthesis module 628 can use one or more neural networks or machine learning models, which can be trained or updated using a training module 632 or system located on or communicating with the server 620. This can include training and / or using a diffusion model 630 to generate content tiles that can be used by the image synthesis module 628, e.g., applying non-repeating textures to an environmental area of an image or video data to be rendered via the client device 602. At least a portion of the generated content can be transmitted to the client device 602 using an appropriate transmission manager 622, for sending via download, streaming, or another such transmission channel. An encoder can be used to encode and / or compress at least a portion of this data before transmitting it to the client device 602. In at least one embodiment, a client device 602 receiving such content can provide the content to a corresponding control application 604, which can also or alternatively include a graphical user interface 610, a content manager 612, and an image synthesis or diffusion module 614 for providing, synthesizing, modifying, or using the content for rendering on or by the client device 602 (or other purposes). A decoder can also be used to decode data received via the network(s) 640 for rendering via the client device 602, such as image or video content rendered via a display 606 and audio (e.g., sounds and music) rendered via at least one audio playback device 608 (e.g., speakers or headphones). In at least one embodiment, at least a portion of the content may already be stored on, rendered on, or accessible to the client device 602, so at least that portion of the content does not need to be transmitted via the network 640, e.g., the content may have been previously downloaded or locally stored on a hard drive or optical disc.In at least one embodiment, a delivery mechanism (such as data streaming) can be used to deliver the content from server 620 or user database 636 to client device 602. In at least one embodiment, at least a portion of the content can be obtained, enhanced, and / or streamed from another source (such as third-party service 660 or other client device 650), which can also include content application 662 for generating, enhancing, or providing content. In at least one embodiment, multiple computing devices or multiple processors within one or more computing devices (such as a combination of a CPU and a GPU) can be used to perform portions of this functionality.
[0059] In this example, the client devices can include any suitable computing device, such as a desktop computer, laptop computer, set-top box, streaming device, gaming console, smartphone, tablet computer, VR headset, AR goggles, wearable computer, or smart TV. Each client device can submit requests across at least one wired or wireless network, which can include the Internet, Ethernet, local area network (LAN), or cellular network, among other such options. In this example, the requests can be submitted to an address associated with a cloud provider, which can operate or control one or more electronic resources in a cloud provider environment, such as a data center or server farm. In at least one embodiment, the request can be received or processed by at least one edge server located at the edge of the network and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling client devices to interact with a closer server, while also enhancing the security of resources in the cloud provider environment.
[0060] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as providing image or video content for testing or validating autonomous machine applications, or for performing deep learning operations. In at least one embodiment, such a system can be implemented using edge devices, or can incorporate one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.
[0061] Inference and training logic
[0062] Figure 7A Shown is inference and / or training logic 715 for performing inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B to provide details regarding inference and / or training logic 715.
[0063] In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters that configure neurons or layers of a neural network trained to and / or for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or sequencing, where weight and / or other parameter information is loaded to configure the logic, including integer and / or floating point units (collectively arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores the weight parameters and / or input / output data of each layer of the neural network used or trained in conjunction with one or more embodiments during forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0064] In at least one embodiment, any portion of the code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 701 may be cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 701 is internal or external to the processor, e.g., or consists of DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0065] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, code and / or data storage 705 for storing the backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained as and / or used for inference in aspects of one or more embodiments. In at least one embodiment, during training and / or inference using aspects of one or more embodiments, the code and / or data storage 705 stores the weight parameters and / or input / output data of each layer of the neural network used or trained in conjunction with one or more embodiments during the backpropagation of the input / output data and / or weight parameters. In at least one embodiment, the training logic 715 can include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or sequencing, where the weights and / or other parameter information is loaded to configure the logic, which includes integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads the weight or other parameter information into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, any portion of the code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 705 is internal or external to the processor, e.g., whether it consists of DRAM, SRAM, flash memory, or some other type of storage, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.
[0066] In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be separate storage structures. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be the same storage structure. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 can be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of the code and / or data storage 701 and the code and / or data storage 705 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0067] In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations at least in part based on or as indicated by training and / or inference code (e.g., graph code), the result of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in the activation store 720, which are a function of input / output and / or weight parameter data stored in the code and / or data store 701 and / or the code and / or data store 705. In at least one embodiment, the activations are generated by linear algebra and / or matrix-based mathematics performed by the ALU 710 in response to executing instructions or other code, where the weight values stored in the code and / or data store 705 and / or the code and / or data store 701 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data store 705 or the code and / or data store 701 or other on-chip or off-chip storage.
[0068] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 710 can be outside of the processor or other hardware logic devices or circuits that use them (e.g., coprocessors). In at least one embodiment, one or more ALUs 710 can be included within the execution units of a processor or otherwise included in a group of ALUs accessible by the execution units of a processor, which execution units of the processor can be within the same processor or distributed among different types of different processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, the code and / or data store 701, the code and / or data store 705, and the activation store 720 can be on the same processor or other hardware logic device or circuit, while in another embodiment, they can be on different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation store 720 can be included with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory. Additionally, the inference and / or training code can be stored with other code accessible by the processor or other hardware logic or circuit and can be fetched and / or processed using the fetch, decode, schedule, execute, retire, and / or other logic circuits of the processor.
[0069] In at least one embodiment, the activation store 720 can be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation store 720 can be wholly or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, depending on the on-chip or off-chip available storage, the latency requirements for training and / or inference functions, the batch size of the data used in inferring and / or training a neural network, or some combination of these factors, it can be selected whether the activation store 720 is internal or external to the processor, e.g., or includes DRAM, SRAM, flash memory, or other storage types. In at least one embodiment, Figure 7A the inference and / or training logic 715 shown in can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the processing unit from Google, the TM inference processing unit (IPU) from Graphcore or the (e.g., “LakeCrest”) processor from Intel Corp. In at least one embodiment, Figure 7A the inference and / or training logic 715 shown in can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (e.g., field programmable gate array (“FPGA”)).
[0070] Figure 7B Illustrated is the inference and / or training logic 715 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used along with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B the inference and / or training logic 715 shown in can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the processing unit from Google, the TM inference processing unit (IPU) from Graphcore or the (e.g., “LakeCrest”) processor from Intel Corp. In at least one embodiment, Figure 7BThe inference and / or training logic 715 shown in FIG. may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as a field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 715 includes, but is not limited to, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (such as graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 7B In at least one embodiment shown in FIG., each of code and / or data storage 701 and code and / or data storage 705 is respectively associated with dedicated computing resources (such as computing hardware 702 and computing hardware 706). In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that respectively perform mathematical functions (such as linear algebra functions) only on the information stored in code and / or data storage 701 and code and / or data storage 705, and the results of the executed functions are stored in activation storage 720.
[0071] In at least one embodiment, each of code and / or data storage 701 and 705 and the corresponding computing hardware 702 and 706 respectively corresponds to a different layer of a neural network, such that the activation obtained from one “storage / compute pair 701 / 702” of code and / or data storage 701 and computing hardware 702 is provided as an input to the next “storage / compute pair 705 / 706” of code and / or data storage 705 and computing hardware 706, so as to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 701 / 702 and 705 / 706 may correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with the storage compute pairs 701 / 702 and 705 / 706.
[0072] Data center
[0073] Figure 8 FIG. shows an example data center 800 in which at least one embodiment may be used. In at least one embodiment, the data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.
[0074] In at least one embodiment, as Figure 8As shown, the data center infrastructure layer 810 may include a resource coordinator 812, grouped computing resources 814, and node computing resources ("node C.R.") 816(1)-816(N), where "N" represents any positive integer. In at least one embodiment, the node C.R. 816(1)-816(N) may include, but is not limited to, any number of central processing units ("CPU") or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid-state drives or disk drives), network input / output ("NWI / O") devices, network switches, virtual machines ("VM"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node C.R. 816(1)-816(N) may be servers having one or more of the above computing resources.
[0075] In at least one embodiment, the grouped computing resources 814 may include separate groupings (not shown) of node C.R. housed within one or more racks, or many racks (also not shown) within data centers located in various geographical locations. The separate groupings of node C.R. within the grouped computing resources 814 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. including CPUs or processors may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0076] In at least one embodiment, the resource coordinator 812 may configure or otherwise control one or more of the node C.R. 816(1)-816(N) and / or the grouped computing resources 814. In at least one embodiment, the resource coordinator 812 may include a software design infrastructure ("SDI") management entity for the data center 800. In at least one embodiment, the resource coordinator 812 may include hardware, software, or some combination thereof.
[0077] In at least one embodiment, as Figure 8As shown, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 may include a framework that supports software 832 of the software layer 830 and / or one or more application programs 842 of the application layer 840. In at least one embodiment, the software 832 or the application program 842 may respectively include web-based service software or application programs, such as services or application programs provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 828 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 832 may include a Spark driver to facilitate scheduling of workloads supported by the various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820 including Spark and the distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 is capable of managing cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 828 and the job scheduler 822. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 may coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.
[0078] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes C.R. 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.
[0079] In at least one embodiment, one or more applications 842 included in the application layer 840 may include one or more types of applications used by at least a portion of nodes C.R. 816(1)-816(N), the packet computing resources 814, and / or the distributed file system 828 of the framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0080] In at least one embodiment, any one of the configuration manager 824, the resource manager 826, and the resource coordinator 812 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 800 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing portions of the data center.
[0081] In at least one embodiment, the data center 800 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 800. In at least one embodiment, by using the weight parameters calculated by one or more training techniques described herein, the resources described above with respect to the data center 800 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks.
[0082] In at least one embodiment, the data center may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the above resources. In addition, one or more of the above software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0083] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This is described herein in connection with Figure 7A and / or Figure 7BProvide details regarding inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be used in a system Figure 8 of a system for inferring or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0084] Such components can be used for object generation and modification.
[0085] A computer system
[0086] Figure 9 is a block diagram showing an exemplary computer system according to at least one embodiment, which can be a system with interconnected devices and components, a system-on-chip (SOC), or some combination thereof formed with a processor that can include execution units to execute instructions. In at least one embodiment, according to the present disclosure, such as the embodiments described herein, the computer system 900 can include, but is not limited to, components such as a processor 902, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, the computer system 900 can include a processor, such as those available from Intel Corporation of Santa Clara, California processor families, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessors, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) can also be used. In at least one embodiment, the computer system 900 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.
[0087] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular telephones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, an embedded application can include a microcontroller, a digital signal processor (“DSP”), a system-on-chip, a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0088] In at least one embodiment, computer system 900 can include, but is not limited to, a processor 902 that can include, but is not limited to, one or more execution units 908 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 900 is a single-processor desktop or server system, but in another embodiment, computer system 900 can be a multi-processor system. In at least one embodiment, processor 902 can include, but is not limited to, a complex instruction set computing (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) computing microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910 that can transfer data signals between processor 902 and other components in computer system 900.
[0089] In at least one embodiment, processor 902 can include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 904. In at least one embodiment, processor 902 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory can reside external to processor 902. Other embodiments can also include a combination of internal and external caches depending on the specific implementation and requirements. In at least one embodiment, register file 906 can store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0090] In at least one embodiment, a logic execution unit 908 including, but not limited to, performing integer and floating point operations is also located in the processor 902. In at least one embodiment, the processor 902 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for processing a packed instruction set 909. In at least one embodiment, by including the packed instruction set 909 in the instruction set of a general-purpose processor and the associated circuitry for the instructions to be executed, operations used by many multimedia applications can be performed using the packed data in the processor 902. In one or more embodiments, operations can be performed on the packed data by using the full width of the data bus of the processor to accelerate and more efficiently execute many multimedia applications, which may not require transmitting smaller data units on the data bus of the processor to perform one or more operations on one data element at a time.
[0091] In at least one embodiment, the execution unit 908 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 900 may include, but not limited to, a memory 920. In at least one embodiment, the memory 920 can be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage devices. In at least one embodiment, the memory 920 can store instructions 919 and / or data 921 represented by data signals that can be executed by the processor 902.
[0092] In at least one embodiment, a system logic chip may be coupled to a processor bus 910 and a memory 920. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 916, and a processor 902 may communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 may provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 916 may initiate data signals among the processor 902, the memory 920, and other components in the computer system 900, and may bridge data signals among the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 may be coupled to the memory 920 via the high-bandwidth memory path 918, and a graphics / video card 912 may be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.
[0093] In at least one embodiment, the computer system 900 may use a system I / O 922, which is a proprietary hub interface bus, to couple the MCH 916 to an I / O controller hub (“ICH”) 930. In at least one embodiment, the ICH 930 may provide direct connections to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 920, the chipset, and the processor 902. Examples may include, but are not limited to, an audio controller 929, a firmware hub (“Flash BIOS”) 928, a wireless transceiver 926, a data storage 924, a legacy I / O controller 923 including a user input and keyboard interface 925, a serial expansion port 927 (such as a Universal Serial Bus (USB) port), and a network controller 934. The data storage 924 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash device, or other mass storage devices.
[0094] In at least one embodiment, Figure 9 a system including interconnected hardware devices or “chips” is shown, while in other embodiments, Figure 9An exemplary System on Chip (SoC) can be shown. In at least one embodiment, the device can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 are interconnected using Compute Express Link (CXL) interconnects.
[0095] Inference and / or training logic 715 is used to perform inference and / or training operations related to one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B In at least one embodiment, inference and / or training logic 715 can be used in a Figure 9 system for inferring or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0096] Such components can be used for object generation and modification.
[0097] Figure 10 is a block diagram illustrating an electronic device 1000 for utilizing processor 1010 according to at least one embodiment. In at least one embodiment, electronic device 1000 can be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, a notebook computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0098] In at least one embodiment, system 1000 can include, but is not limited to, processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface such as an I²C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 10 a system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 10 an exemplary System on Chip (SoC) can be shown. In at least one embodiment, Figure 10 the devices shown in Figure 10 can be interconnected with proprietary interconnect lines, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment,
[0099] In at least one embodiment, Figure 10 it may include a display 1024, a touch screen 1025, a touchpad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, a fast chipset (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash (“BIOS, FW Flash”) 1022, a DSP 1060, a drive 1020 (such as a solid state disk (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) 1055, a camera (“USB3.0 camera”) 1054 (such as a USB3.0 camera) and / or a low power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented in accordance with, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0100] In at least one embodiment, other components may be communicatively coupled to the processor 1010 via the components described above. In at least one embodiment, an accelerometer 1041, an ambient light sensor (“ALS”) 1042, a compass 1043 and a gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1036 and a touchpad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, a speaker 1063, headphones 1064 and a microphone (“mic”) 1065 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 1062, which may in turn be communicatively coupled to the DSP 1060. In at least one embodiment, the audio unit 1062 may include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a subscriber identity module (“SIM”) 1057 may be communicatively coupled to the WWAN unit 1056. In at least one embodiment, components (such as the WLAN unit 1050, the Bluetooth unit 1052 and the WWAN unit 1056) may be implemented in a next generation form factor (NGFF).
[0101] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 715 are provided below in connection with Figure 7A and / or Figure 7B In at least one embodiment, the inference and / or training logic 715 may be in Figure 10The system infers or predicts operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0102] Such components can be used for object generation and modification.
[0103] Figure 11 FIG. 1100 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0104] In at least one embodiment, system 1100 can be included in or incorporated within a server-based gaming platform, including a game console such as a game and media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, system 1100 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 1100 can also be coupled to or integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.
[0105] In at least one embodiment, each of the one or more processors 1102 includes one or more processor cores 1107 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a specific set of instructions 1109. In at least one embodiment, the set of instructions 1109 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, the one or more processor cores 1107 can each process a different set of instructions 1109, which can include instructions that help to emulate other instruction sets. In at least one embodiment, the one or more processor cores 1107 can also include other processing devices, such as a digital signal processor (DSP).
[0106] In at least one embodiment, one or more processors 1102 include a cache memory 1104. In at least one embodiment, one or more processors 1102 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of one or more processors 1102. In at least one embodiment, one or more processors 1102 also use an external cache (e.g., a level three (L3) cache or a last level cache (LLC)) (not shown), and this external cache may be shared among one or more processor cores 1107 using known cache coherence techniques. In at least one embodiment, a register file 1106 is further included in one or more processors 1102, and the processor may include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 1106 may include general-purpose registers or other registers.
[0107] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transfer communication signals, such as address, data, or control signals, between one or more processors 1102 and other components in the system 1100. In at least one embodiment, one or more interface buses 1110 may be a processor bus, such as a version of the direct media interface (DMI) bus, in one embodiment. In at least one embodiment, one or more interface buses 1110 are not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between the memory device and other components of the processing system 1100, while the platform controller hub (PCH) 1130 provides a connection to I / O devices via a local I / O bus.
[0108] In at least one embodiment, the memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have suitable performance to be used as processor memory. In at least one embodiment, the memory device 1120 may be used as the system memory of the processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, the memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 among one or more processors 1102 to perform graphics and media operations. In at least one embodiment, the display device 1111 may be connected to the processor 1102. In at least one embodiment, the display device 1111 may include one or more of internal display devices, such as in a mobile electronic device or a laptop device or an external display device connected through a display interface (such as DisplayPort, etc.). In at least one embodiment, the display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) applications or augmented reality (AR) applications.
[0109] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to be connected to the memory device 1120 and one or more processors 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, a data storage device 1124 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1128 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1134 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to one or more interface buses 1110. In at least one embodiment, the audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1100. In at least one embodiment, the platform controller hub 1130 can also be connected to one or more Universal Serial Bus (USB) controllers 1142, which connect input devices, such as a keyboard and mouse 1143 combination, a camera 1144, or other USB input devices.
[0110] In at least one embodiment, instances of the memory controller 1116 and the platform controller hub 1130 can be integrated into a discrete external graphics processor, such as the external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 can be external to one or more processors 1102. For example, in at least one embodiment, the system 1100 can include an external memory controller 1116 and a platform controller hub 1130, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 1102.
[0111] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Below, in conjunction with Figure 7A and / orFigure 7B Provide details regarding inference and / or training logic 715. In at least one embodiment, part or all of the inference and / or training logic 715 may be incorporated into the graphics processor 1100. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in the graphics processor. Additionally, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than Figure 7A and / or Figure 7B the logic shown. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the graphics processor to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0112] Such components may be used for object generation and modification.
[0113] Figure 12 is a block diagram of a processor 1200 having one or more processor cores 1202A - 1202N, an integrated memory controller 1214, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, the processor 1200 may include additional cores, up to and including the additional core 1202N represented by the dashed box. In at least one embodiment, each of the one or more processor cores 1202A - 1202N includes one or more internal cache units 1204A - 1204N. In at least one embodiment, each processor core may also access one or more shared cache units 1206.
[0114] In at least one embodiment, the one or more internal cache units 1204A - 1204N and the one or more shared cache units 1206 represent the cache memory hierarchy within the processor 1200. In at least one embodiment, the one or more cache memory units 1204A - 1204N may include at least a level of instruction and data cache within each processor core and one or more levels of cache in a shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1206 and 1204A - 1204N.
[0115] In at least one embodiment, the processor 1200 may further include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, the one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1210 provides management functions for various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown).
[0116] In at least one embodiment, the one or more processor cores 1202A - 1202N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 1210 includes components for coordinating the one or more processor cores 1202A - 1202N during multi-threaded processing. In at least one embodiment, the system agent core 1210 may additionally include a power control unit (PCU) that includes logic and components for regulating the one or more power states of the one or more processor cores 1202A - 1202N and the graphics processor 1208.
[0117] In at least one embodiment, the processor 1200 further includes a graphics processor 1208 for performing graphics processing operations. In at least one embodiment, the graphics processor 1208 is coupled to one or more shared cache units 1206 and the system agent core 1210 that includes one or more integrated memory controllers 1214. In at least one embodiment, the system agent core 1210 further includes a display controller 1211 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1211 may also be a separate module coupled to the graphics processor 1208 via at least one interconnect, or may be integrated within the graphics processor 1208.
[0118] In at least one embodiment, a ring-based interconnect unit 1212 is used to couple the internal components of the processor 1200. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1208 is coupled to the ring-based interconnect unit 1212 via an I / O link 1213.
[0119] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnections, including a package I / O interconnection that facilitates communication between various processor components and a high-performance embedded memory module 1218 (e.g., an eDRAM module). In at least one embodiment, each of one or more processor cores 1202A - 1202N and the graphics processor 1208 uses the embedded memory module 1218 as a shared last-level cache.
[0120] In at least one embodiment, one or more processor cores 1202A - 1202N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, one or more processor cores 1202A - 1202N are heterogeneous in terms of instruction set architecture (ISA), where one or more processor cores 1202A - 1202N execute a common instruction set, while one or more other cores among one or more processor cores 1202A - 1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, in terms of microarchitecture, one or more processor cores 1202A - 1202N are heterogeneous, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1200 can be implemented on one or more chips or be implemented as a SoC integrated circuit.
[0121] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 715 are provided below in conjunction with Figure 7A and / or Figure 7B In at least one embodiment, part or all of the inference and / or training logic 715 can be incorporated into the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs embodied in Figure 12 the graphics processor 1208, one or more processor cores 1202A - 1202N, or other components. Additionally, in at least one embodiment, the inference and / or training operations described herein can be completed using logic other than Figure 7A and / or Figure 7B the logic shown. In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALUs of the graphics processor 1200 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0122] Such components can be used for object generation and modification.
[0123] Virtualized computing platform
[0124] Figure 13 is an example data flow diagram of a process 1300 for generating and deploying an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 1300 can be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. The process 1300 can be executed within a training system 1304 and / or a deployment system 1306. In at least one embodiment, the training system 1304 can be used to perform the training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for the deployment system 1306. In at least one embodiment, the deployment system 1306 can be configured to offload processing and computing resources in a distributed computing environment to reduce the infrastructure requirements of the facilities 1302. In at least one embodiment, one or more applications in the pipeline can use or invoke the services (e.g., inference, visualization, computing, AI, etc.) of the deployment system 1306 during application execution.
[0125] In at least one embodiment, some applications used in the advanced processing and inference pipeline can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, data 1308 (e.g., imaging data) generated at the facility 1302 (and stored on one or more picture archiving and communication system (PACS) servers at the facility 1302) can be used to train machine learning models at the facility 1302, imaging or sequencing data 1308 from another or more facilities can be used to train machine learning models, or a combination thereof. In at least one embodiment, the training system 1304 can be used to provide applications, services, and / or other resources to generate a working, deployable machine learning model for the deployment system 1306.
[0126] In at least one embodiment, the model registry 1324 can be supported by an object store that can support version control and object metadata. In at least one embodiment, the object store can be accessed from within a cloud platform via, for example, a cloud storage-compatible application programming interface (API). In at least one embodiment, the machine learning models within the model registry 1324 can be uploaded, listed, modified, or deleted by developers or partners of systems that interact with the API. In at least one embodiment, the API can provide access to methods that allow users with appropriate credentials to associate a model with an application such that the model can be executed as part of the containerized instantiation of the application.
[0127] In at least one embodiment, the training system 1304 ( Figure 13) may include the following scenarios: where the facility 1302 is training their own machine learning model, or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device, a sequencing device, and / or other types of devices may be received. In at least one embodiment, once the imaging data 1308 is received, AI-assisted annotation 1310 may be used to help generate annotations corresponding to the imaging data 1308 for use as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 1310 may include one or more machine learning models (e.g., a convolutional neural network (CNN)), and the machine learning model may be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, the AI-assisted annotation 1310 may then be used directly, or may be adjusted or fine-tuned using an annotation tool to generate the ground truth data. In at least one embodiment, the AI-assisted annotation 1310, the annotated data 1312, or a combination thereof may be used as the ground truth data for training the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as the (one or more) output model 1316 and may be used by the deployment system 1306 as described herein.
[0128] In at least one embodiment, the training pipeline can include the following scenario: where facility 1302 requires a machine learning model to perform one or more processing tasks for deploying one or more applications in system 1306, but facility 1302 may not currently have such a machine learning model (or may not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from model registry 1324. In at least one embodiment, model registry 1324 can include machine learning models that are trained to perform various different inference tasks on imaging data. In at least one embodiment, the machine learning models in model registry 1324 can be trained on imaging data from different facilities (e.g., a facility located remotely) rather than facility 1302. In at least one embodiment, the machine learning model may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a specific location, the training can be performed at that location or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data off-site. In at least one embodiment, once the model or a portion of the model has been trained at one location, the machine learning model can be added to model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in model registry 1324. In at least one embodiment, a machine learning model can then be selected from model registry 1324 (and referred to as output model 1316), and can be used in deployment system 1306 to perform one or more processing tasks for one or more applications of the deployment system.
[0129] In at least one embodiment, a scenario can include a facility 1302 that requires a machine learning model to perform one or more processing tasks for deploying one or more applications in a deployment system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have an optimized, efficient, or effective model). In at least one embodiment, due to population differences, robustness of training data for training the machine learning model, diversity of training data anomalies, and / or other issues with the training data, the machine learning model selected from the model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at the facility 1302. In at least one embodiment, AI-assisted annotation 1310 can be used to help generate annotations corresponding to the imaging data 1308 to be used as ground truth data for training or updating the machine learning model. In at least one embodiment, labeled clinical data 1312 can be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model can be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI-assisted annotation 1310, labeled data 1312, or a combination thereof) can be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, the trained machine learning model can be referred to as an output model 1316 and can be used by the deployment system 1306 as described herein.
[0130] In at least one embodiment, the deployment system 1306 may include software 1318, services 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, the deployment system 1306 may include a software "stack" such that the software 1318 may be built on top of the services 1320 and the services 1320 may be used to perform some or all of the processing tasks, and the services 1320 and the software 1318 may be built on top of the hardware 1322 and the hardware 1322 may be used to perform the processing, storage, and / or other computing tasks of the deployment system. In at least one embodiment, the software 1318 may include any number of different containers, where each container may perform an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) in a high-level processing and inference pipeline. In at least one embodiment, in addition to the containers that receive and configure the imaging data for use by each container and / or by the facility 1302 after being processed through the pipeline, the high-level processing and inference pipeline may be defined based on the selection of different containers that are desired or required for processing the imaging data 1308 (e.g., to convert the output back to a usable data type). In at least one embodiment, a combination of containers within the software 1318 (e.g., which form a pipeline) may be referred to as a virtual instrument (as described in more detail herein), and the virtual instrument may utilize the services 1320 and the hardware 1322 to perform some or all of the processing tasks of the applications instantiated in the containers.
[0131] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of the deployment system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may be preprocessed as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference tasks or other processing tasks in the pipeline to prepare the output data for the next application and / or to prepare the output data for transmission and / or use by the user (e.g., as a response to the inference request). In at least one embodiment, the inference tasks may be performed by one or more machine learning models, such as a trained or deployed neural network, and the models may include the output model 1316 of the training system 1304.
[0132] In at least one embodiment, the tasks of a data processing pipeline can be encapsulated in containers, where each container represents a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, a container or application can be published to a private (e.g., limited access) region of a container registry (described in more detail herein), and a trained or deployed model can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in a pipeline, the image can be used to generate a container for instantiation of the application for use by the user's system.
[0133] In at least one embodiment, a developer (e.g., a software developer, a clinician, a doctor, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publishing, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, on data from the first facility), where the SDK, as part of the system (e.g., Figure 12 system 1200 herein), can support at least some services 1320. In at least one embodiment, since a DICOM object can contain from one to hundreds of images or other data types, and due to the variability of the data, the developer is responsible for managing (e.g., setting up constructs for building preprocessing into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by the system 1300 (e.g., for accuracy), the application becomes available in the container registry for a user to select and / or implement to perform one or more processing tasks on data at the user's facility (e.g., a second facility).
[0134] In at least one embodiment, the developer can then share the application or container over a network for use by a system (e.g., Figure 13User access and use of the system 1300). In at least one embodiment, a completed and verified application or container can be stored in a container registry, and a related machine learning model can be stored in the model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) can browse the container registry and / or the model registry 1324 to obtain applications, containers, data sets, machine learning models, etc., select a desired combination of elements to include in a data processing pipeline, and submit an image processing request. In at least one embodiment, the request can include the input data necessary to execute the request (and in some examples, patient-related data), and / or can include a selection of the application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request can then be passed to one or more components (e.g., the cloud) of the deployment system 1306 to perform the processing of the data processing pipeline. In at least one embodiment, the processing performed by the deployment system 1306 can include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or the model registry 1324. In at least one embodiment, once the results are generated through the pipeline, the results can be returned to the user for reference (e.g., for viewing in a viewing application suite executed locally, on a local workstation, or on a terminal).
[0135] In at least one embodiment, to assist in processing or executing an application or container in the pipeline, the service 1320 can be utilized. In at least one embodiment, the service 1320 can include computing services, artificial intelligence (AI) services, visualization services, and / or other service types. In at least one embodiment, the service 1320 can provide functions common to one or more applications in the software 1318, and thus the functions can be abstracted as services that can be called or utilized by the applications. In at least one embodiment, the functions provided by the service 1320 can run dynamically and more efficiently, while also allowing the applications to process data in parallel (e.g., using the parallel computing platform 1230( Figure 12)) to scale well. In at least one embodiment, rather than requiring each application that shares the same functionality provided by the shared service 1320 to have a corresponding instance of the service 1320, the service 1320 can be shared among and within various applications. In at least one embodiment, by way of non-limiting example, the service can include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included, which can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data augmentation service can be further included, which can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other augmentation. In at least one embodiment, a visualization service can be used, which can add image rendering effects (e.g., ray tracing, rasterization, denoising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service can be included, which provides beamforming, segmentation, inference, imaging, and / or support for other applications within the pipeline of virtual instruments.
[0136] In at least one embodiment, in the case where the service 1320 includes an AI service (e.g., an inference service), as part of the execution of an application, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., an inference server) to perform one or more machine learning models or their processing. In at least one embodiment, in the case where another application includes one or more machine learning models for a segmentation task, the application can call the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, the software 1318 that implements the advanced processing and inference pipeline, which includes a segmentation application and an anomaly detection application, can be pipelined because each application can call the same inference service to perform one or more inference tasks.
[0137] In at least one embodiment, the hardware 1322 may include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient, specially built support for the software 1318 and the services 1320 in the deployment system 1306. In at least one embodiment, GPU processing may be implemented to perform local processing (e.g., at the facility 1302) within an AI / deep learning system, in a cloud system, and / or in other processing components of the deployment system 1306 to improve the efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, as a non-limiting example, with respect to deep learning, machine learning, and / or high-performance computing, the software 1318 and / or the services 1320 may be optimized for GPU processing. In at least one embodiment, at least some of the computing environments of the deployment system 1306 and / or the training system 1304 may be executed in a data center with GPU-optimized software (e.g., the hardware and software combination of the NVIDIA DGX system), one or more supercomputers, or high-performance computer systems. In at least one embodiment, as described herein, the hardware 1322 may include any number of GPUs, which may be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on NVIDIA's DGX system) may be used as a hardware abstraction and scaling platform to execute a cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform may integrate an application container cluster system or a coordination system (e.g., KUBERNETES) on multiple GPUs to achieve seamless scaling and load balancing.
[0138] Figure 14 is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, the system 1400 may be used to implement Figure 13 the process 1300 and / or other processes, including advanced processing and inference pipelines. In at least one embodiment, the system 1400 may include a training system 1304 and a deployment system 1306. In at least one embodiment, the training system 1304 and the deployment system 1306 may be implemented using the software 1318, the services 1320, and / or the hardware 1322, as described herein.
[0139] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) may be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 may be implemented locally (with respect to a healthcare facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1426 may be restricted to authorized users by establishing security measures or protocols. In at least one embodiment, the security protocol may include a network token, which may be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and may carry appropriate authorization. In at least one embodiment, the APIs of virtual instruments (described herein) or other instances of system 1400 may be restricted to a set of public IPs that have been audited or authorized for interaction.
[0140] In at least one embodiment, the various components of system 1400 may communicate with each other using any of a variety of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between the facilities and components of system 1400 (e.g., for sending inference requests, for receiving results of inference requests, etc.) may be conveyed via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.
[0141] In at least one embodiment, similar to what is described herein with respect to Figure 13 the training system 1304 may execute one or more training pipelines 1404. In at least one embodiment, where the deployment system 1306 will use one or more machine learning models in one or more deployment pipelines 1410, one or more training pipelines 1404 may be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, as a result of one or more training pipelines 1404, one or more output models 1316 may be generated. In at least one embodiment, one or more training pipelines 1404 may include any number of processing steps, such as but not limited to the transformation or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 may be used for different machine learning models used by the deployment system 1306. In at least one embodiment, similar to what is described with respect to Figure 13 the first example described, one or more training pipelines 1404 may be used for a first machine learning model, similar to what is described with respect to Figure 13One or more training pipelines 1404 of the second example described can be used for the second machine learning model, similar to the one or more training pipelines 1404 of the third example described being used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used according to the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models may already be trained and ready for deployment, so the training system 1304 may not perform any processing on the machine learning models, and one or more machine learning models can be implemented by the deployment system 1306. Figure 13 One or more training pipelines 1404 of the third example described can be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used according to the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models may already be trained and ready for deployment, so the training system 1304 may not perform any processing on the machine learning models, and one or more machine learning models can be implemented by the deployment system 1306.
[0142] In at least one embodiment, depending on the implementation or embodiment, the output model 1316 and / or the pre-trained model 1406 can include any type of machine learning model. In at least one embodiment and without limitation, the machine learning models used by the system 1400 can include those using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.
[0143] In at least one embodiment, one or more training pipelines 1404 can include AI-assisted annotation, as described herein with respect to at least Figure 14More detailed description. In at least one embodiment, the labeled clinical data 1312 (e.g., traditional annotations) can be generated by any number of techniques. In at least one embodiment, in some examples, the labels or other annotations can be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of application suitable for generating ground truth annotations or labels, and / or can be hand-drawn. In at least one embodiment, the ground truth data can be synthetically generated (e.g., generated from a computer model or rendering), real-world generated (e.g., designed and generated from real-world data), machine-automatically generated (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a tagger or annotation expert to define the location of the labels), and / or a combination thereof. In at least one embodiment, for each instance of the imaging data 1308 (or other data types used by the machine learning model), there can be corresponding ground truth data generated by the training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of the deployment pipeline 1410; supplementing or replacing the AI-assisted annotation included in the training pipeline 1404. In at least one embodiment, the system 1400 can include a multi-layer platform, and the multi-layer platform can include a software layer (e.g., software 1318) of a diagnostic application (or other application type), which can perform one or more medical imaging and diagnostic functions. In at least one embodiment, the system 1400 can be communicatively coupled to (e.g., via an encrypted link) the PACS server network of one or more facilities. In at least one embodiment, the system 1400 can be configured to access and reference data from the PACS server to perform operations such as training a machine learning model, deploying a machine learning model, image processing, inference, and / or other operations.
[0144] In at least one embodiment, the software layer can be implemented as a secure, encrypted, and / or certified API, through which an application or container can be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, the application can then call or execute one or more services 1320 to perform computing, AI, or visualization tasks associated with the respective application, and the software 1318 and / or the service 1320 can utilize the hardware 1322 to perform processing tasks in an effective and efficient manner. In at least one embodiment, a pair of DICOM adapters 1402A, 1402B can be used to send communications to or receive communications from the training system 1304 and the deployment system 1306.
[0145] In at least one embodiment, the deployment system 1306 may execute one or more deployment pipelines 1410. In at least one embodiment, one or more deployment pipelines 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) - including AI-assisted annotation, where the imaging data is generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, as described herein, the deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., virtual ultrasound instrument, virtual CT scan instrument, virtual sequencing instrument, etc.). In at least one embodiment, for a single device, there may be more than one deployment pipeline 1410, depending on the information desired from the data generated by the device. In at least one embodiment, in the case where an abnormality is expected to be detected from an MRI machine, there may be one or more first deployment pipelines 1410, and in the case where image enhancement is desired from the output of an MRI machine, there may be one or more second deployment pipelines 1410.
[0146] In at least one embodiment, the image generation application may include processing tasks that include using a machine learning model. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in the application that performs the processing task. In at least one embodiment, the application may be selectable and customizable, and by defining the construction of the application, the deployment and implementation of the application for a specific user are presented as a more seamless user experience. In at least one embodiment, by leveraging other features of the system 1400 (such as services 1320 and hardware 1322), one or more deployment pipelines 1410 can be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.
[0147] In at least one embodiment, the deployment system 1306 can include a user interface (“UI”) 1414 (e.g., a graphical user interface, a web interface, etc.), which can be used to select applications to be included in the deployment pipeline 1410, arrange the applications, modify or change the applications or their parameters or configurations, use and interact with the deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with the deployment system 1306. In at least one embodiment, although not shown with respect to the training system 1304, the UI 1414 (or a different user interface) can be used to select models to be used in the deployment system 1306, to select models for training or retraining in the training system 1304, and / or to otherwise interact with the training system 1304.
[0148] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 can be used to manage the interaction between the applications or containers of the deployment pipeline 1410 and the services 1320 and / or the hardware 1322. In at least one embodiment, the pipeline manager 1412 can be configured to facilitate interactions from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although shown as being included in the software 1318, this is not intended to be limiting, and in some examples, the pipeline manager 1412 can be included in the service 1320. In at least one embodiment, the application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) can include a container coordination system, which can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications (e.g., rebuilt applications, split applications, etc.) from the deployment pipeline 1410 with individual containers, each application can execute in a self - contained environment (e.g., at the kernel level) to improve speed and efficiency.
[0149] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed separately (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer), which can allow for focusing on and attending to the tasks of a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can assist in the communication and collaboration between different containers or applications. In at least one embodiment, as long as the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container), the application coordination system 1428 and / or the pipeline manager 1412 can facilitate the communication and resource sharing between and among each application or container. In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, perform load balancing, and determine the sharing of services or resources between and among the various applications or containers. In at least one embodiment, a scheduler can be used to track the resource requirements of an application or container, the current or planned use of those resources, and the resource availability. Thus, in at least one embodiment, the scheduler can allocate resources to different applications, and distribute resources between and among applications, taking into account the requirements and availability of the system. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine the resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or deferred processing), etc.
[0150] In at least one embodiment, services 1320 utilized and shared by applications or containers in deployment system 1306 may include compute services 1416, AI services 1418, visualization services 1420, and / or other service types. In at least one embodiment, an application may call (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize compute services 1416 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1416 may be utilized to perform parallel processing (e.g., using parallel computing platform 1430) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA's CUDA) may implement general-purpose computing on a GPU (GPGPU) (e.g., GPU / graphics 1422). In at least one embodiment, the software layer of parallel computing platform 1430 may provide access to the virtual instruction set and parallel computing elements of the GPU to execute compute kernels. In at least one embodiment, parallel computing platform 1430 may include memory, and in some embodiments, the memory may be shared among and within multiple containers and / or between and within different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within a container to use the same data for a shared memory segment from parallel computing platform 1430 (e.g., where multiple different stages of one application or multiple applications are processing the same information). In at least one embodiment, rather than copying data and moving the data to different locations in memory (e.g., read / write operations), the same data in the same location in memory may be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, since data is used to generate new data as a result of processing, this information about the new location of the data may be stored and shared among various applications. In at least one embodiment, the location of the data and the location of the updated or modified data may be part of how the definition of the payload in a container is understood.
[0151] In at least one embodiment, one or more AI services 1418 may be utilized to perform an inference service for executing a machine learning model associated with an application (e.g., the task is to perform one or more processing tasks of the application). In at least one embodiment, one or more AI services 1418 may utilize an AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, an application of the deployment pipeline 1410 may use one or more output models 1316 from the self-training system 1304 and / or other models of the application to perform inference on imaging data. In at least one embodiment, two or more examples of performing inference using an application coordination system 1428 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path, which may implement a higher service level agreement, such as for performing inference on an emergency request in an emergency situation or for a radiologist during a diagnostic process. In at least one embodiment, a second category may include a standard-priority path, which may be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1428 may allocate resources (e.g., service 1320 and / or hardware 1322) based on the priority path for different inference tasks of the AI service 1418.
[0152] In at least one embodiment, a shared memory may be installed to one or more AI services 1418 in the system 1400. In at least one embodiment, the shared memory may operate as a cache (or other storage device type), and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of the deployment system 1306 may receive the request, and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be input into a database, and if not already in the cache, the machine learning model may be located from the model registry 1324, a validation step may ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of the pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if an inference server has not been started to execute the model, the inference server may be started. Any number of inference servers may be started per model. In at least one embodiment, in a pull model where inference servers are clustered, the model may be cached whenever load balancing is beneficial. In at least one embodiment, the inference servers may be statically loaded into the corresponding distributed servers.
[0153] In at least one embodiment, an inference server running in a container may be used to perform inference. In at least one embodiment, an instance of the inference server may be associated with a model (and optionally with multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on a model is received, a new instance may be loaded. In at least one embodiment, when starting the inference server, the model may be passed to the inference server such that the same container may be used to serve different models as long as the inference server runs as different instances.
[0154] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., an instance of a hosted inference server) can be loaded (if not already loaded), and a launcher can be invoked. In at least one embodiment, preprocessing logic in the container can load, decode, and / or perform any additional preprocessing on the incoming data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is ready for inference, the container can perform inference on the data as needed. In at least one embodiment, this can include a single inference call on an image (e.g., a hand X-ray), or may require inference on hundreds of images (e.g., chest CTs). In at least one embodiment, the application can summarize the results before completion, which can include but is not limited to a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities can be assigned to different models or applications. For example, some models may have real-time (TAT less than 1 minute) priority, while other models may have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, the model execution time can be measured from the requesting agency or entity, and can include the collaborative network traversal time as well as the execution time of the inference service.
[0155] In at least one embodiment, the transfer of requests between the service 1320 and the inference application can be hidden behind a software development kit (SDK), and a robust transport can be provided via a queue. In at least one embodiment, requests will be placed in the queue via an API for an individual application / tenant ID combination, and the SDK will pull requests from the queue and provide the requests to the application. In at least one embodiment, the name of the queue can be provided in the environment from which the SDK will pick it up. In at least one embodiment, asynchronous communication via the queue can be useful as it can allow any instance of the application to pick up work when it is available. Results can be transferred back via the queue to ensure no data is lost. In at least one embodiment, the queue can also provide the ability to split the work, as the highest priority work can go into the queue connected to most instances of the application, while the lowest priority work can go into the queue connected to a single instance that processes tasks in the order received. In at least one embodiment, the application can run on a GPU-accelerated instance that is generated in the cloud 1426, and the inference service can perform inference on the GPU.
[0156] In at least one embodiment, a visualization service 1420 can be utilized to generate visualizations for viewing the output of an application and / or a deployment pipeline 1410. In at least one embodiment, the visualization service 1420 can utilize GPU / graphics 1422 to generate visualizations. In at least one embodiment, the visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, a virtualized environment can be used to generate a virtual interactive display or environment (e.g., a virtual environment) for interaction by system users (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, the visualization service 1420 can include an internal visualizer, movie, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).
[0157] In at least one embodiment, the hardware 1322 can include GPU / graphics 1422, an AI system 1424, a cloud 1426, and / or any other hardware for executing the training system 1304 and / or the deployment system 1306. In at least one embodiment, the GPU / graphics 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) can include any number of GPUs that can be used to perform processing tasks for any features or functions of the computing service 1416, the AI service 1418, the visualization service 1420, other services, and / or software 1318. For example, for the AI service 1418, the GPU / graphics 1422 can be used to perform preprocessing on imaging data (or other data types used by machine learning models), perform postprocessing on the output of machine learning models, and / or perform inference (e.g., to execute a machine learning model). In at least one embodiment, the cloud 1426, the AI system 1424, and / or other components of the system 1400 can use the GPU / graphics 1422. In at least one embodiment, the cloud 1426 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1424 can use GPUs, and one or more AI systems 1424 can be used to perform the cloud 1426 (or at least part of the tasks that are deep learning or inference). Similarly, although the hardware 1322 is shown as discrete components, this is not intended to be limiting, and any component of the hardware 1322 can be combined with or utilized by any other component of the hardware 1322.
[0158] In at least one embodiment, the AI system 1424 can include a specially constructed computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to the CPU, RAM, memory, and / or other components, features, or functions, the AI system 1424 (e.g., NVIDIA's DGX) can also include software (e.g., a software stack) that can use multiple GPUs / graphics 1422 to perform GPU-optimized operations. In at least one embodiment, one or more AI systems 1424 can be implemented in the cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1400.
[0159] In at least one embodiment, the cloud 1426 can include GPU-accelerated infrastructure (e.g., NVIDIA's NGC), which can provide a GPU-optimized platform for performing the processing tasks of the system 1400. In at least one embodiment, the cloud 1426 can include an AI system 1424 for performing one or more AI-based tasks of the system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, the cloud 1426 can be integrated with the application coordination system 1428 that utilizes multiple GPUs to achieve seamless scaling and load balancing between and within the applications and services 1320. In at least one embodiment, as described herein, the cloud 1426 can be responsible for executing at least some of the services 1320 of the system 1400, including one or more computing services 1416, one or more AI services 1418, and / or one or more visualization services 1420. In at least one embodiment, the cloud 1426 can perform inference on large and small batches (e.g., execute NVIDIA's TENSORRT), provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA's CUDA), execute the application coordination system 1428 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher-quality movie effects), and / or can provide other functions for the system 1400.
[0160] Figure 15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, the following can be used as non-limiting examples Figure 14System 1400 to perform process 1500. In at least one embodiment, process 1500 may utilize services and / or hardware as described herein. In at least one embodiment, the refined model 1512 generated by process 1500 may be executed by the deployment system for one or more containerized applications in the deployment pipeline.
[0161] In at least one embodiment, model training 1514 may include retraining or updating the initial model 1504 (e.g., pre-trained model) using new training data (e.g., new input data such as customer dataset 1506, and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset, deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously fine-tuned parameters (e.g., weights and / or biases) retained from a previous training, so training or retraining 1514 may not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 1514, when generating predictions on the new customer dataset 1506 by resetting or replacing the output or loss layer of the initial model 1504, the parameters of the new dataset may be updated and readjusted based on the loss calculation associated with the accuracy of the output or loss layer.
[0162] In at least one embodiment, the pre-trained model 1506 can be stored in a data store or registry. In at least one embodiment, the pre-trained model 1506 may have been trained at least in part at one or more facilities other than the facility where the execution process 1500 takes place. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers of different facilities, the pre-trained model 1506 may have been trained locally using locally generated customer or patient data. In at least one embodiment, a cloud and / or other hardware can be used to train the pre-trained model 1306, but confidential, privacy-protected patient data may not be transferred to, used by, or accessed by any component of the cloud (or other non-local hardware). In at least one embodiment, if patient data from more than one facility is used to train the pre-trained model 1506, the pre-trained model 1506 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, for example, where privacy issues have been released for customer or patient data (e.g., by waiver, for experimental use, etc.), or where customer or patient data is included in a public dataset, customer or patient data from any number of facilities can be used to train the pre-trained model 1506 locally and / or externally, such as in a data center or other cloud computing infrastructure.
[0163] In at least one embodiment, when selecting an application to use in a deployment pipeline, the user can also select a machine learning model for a specific application. In at least one embodiment, the user may not have a model to use, so the user can select a pre-trained model to use with the application. In at least one embodiment, the pre-trained model may not be optimized to generate accurate results on the customer dataset 1506 of the user's facility (e.g., based on patient diversity, demographics, type of medical imaging device used, etc.). In at least one embodiment, before deploying the pre-trained model into a deployment pipeline for use with one or more applications, the pre-trained model can be updated, retrained, and / or fine-tuned for use at individual facilities.
[0164] In at least one embodiment, a user may select a pre-trained model to be updated, retrained, and / or fine-tuned, and the pre-trained model may be referred to as the initial model 1504 of the training system in process 1500. In at least one embodiment, a customer dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by devices at a facility) may be used to perform model training (which may include but is not limited to transfer learning) on the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to the customer dataset 1506 may be generated by the training system 1304. In at least one embodiment, the ground truth data may be generated at least in part by clinicians, scientists, physicians, practitioners at the facility.
[0165] In at least one embodiment, in some examples, AI-assisted annotation may be used to generate ground truth data. In at least one embodiment, AI-assisted annotation (e.g., implemented using an AI-assisted annotation SDK) may utilize a machine learning model (e.g., a neural network) to generate ground truth data for suggestions or predictions for the customer dataset. In at least one embodiment, a user may use an annotation tool within a user interface (graphical user interface (GUI)) on a computing device.
[0166] In at least one embodiment, a user 1510 may interact with the GUI via a computing device 1508 to edit or fine-tune the annotation or auto-annotation. In at least one embodiment, a polygon editing feature may be used to move the vertices of a polygon to a more precise or fine-tuned location.
[0167] In at least one embodiment, once the customer dataset 1506 has associated ground truth data, the ground truth data (e.g., from AI-assisted annotation, manual labeling, etc.) may be used during model training to generate a refined model 1512. In at least one embodiment, the customer dataset 1506 may be applied to the initial model 1504 any number of times, and the ground truth data may be used to update the parameters of the initial model 1504 until an acceptable accuracy level is reached for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, the refined model 1512 may be deployed within one or more deployment pipelines at the facility to perform one or more processing tasks regarding medical imaging data.
[0168] In at least one embodiment, the refined model 1512 may be uploaded to the pre-trained model in a model registry for selection by another facility. In at least one embodiment, this process may be done at any number of facilities such that the refined model 1512 may be further refined any number of times on new datasets to generate a more general model.
[0169] Figure 15B is an example illustration of a client - server architecture 1532 for enhancing an annotation tool using a pre - trained annotation model according to at least one embodiment. In at least one embodiment, an AI - assisted annotation tool 1536 can be instantiated based on the client - server architecture 1532. In at least one embodiment, the AI - assisted annotation tool 1536 in an imaging application can assist a radiologist, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application can include software tools, as non - limiting examples, that help a user 1510 identify several extreme points on a specific organ of interest in an original image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of the specific organ. In at least one embodiment, the results can be stored as training data 1538 in a data store and used as (e.g., but not limited to) ground truth data for training. In at least one embodiment, when a computing device 1508 sends extreme points for AI - assisted annotation, for example, a deep learning model can receive this data as input and return an inference result for segmenting the organ or abnormality. In at least one embodiment, a pre - instantiated annotation tool (e.g., Figure 15B the AI - assisted annotation tool 1536 in) can be enhanced by making an API call (e.g., API call 1544) to a server, such as an annotation assistant server 1540, which can include a set of pre - trained models 1542 stored in, for example, an annotation model registry. In at least one embodiment, the annotation model registry can store pre - trained models 1542 (e.g., machine learning models, such as deep learning models) that are pre - trained to perform AI - assisted annotation for a specific organ or abnormality. In at least one embodiment, these models can be further updated by using a training pipeline. In at least one embodiment, as new labeled data is added, the pre - installed annotation tools can be improved over time.
[0170] Various embodiments can be described by the following clauses:
[0171] 1. A computer - implemented method, comprising:
[0172] Determining a target level of detail for a three - dimensional (3D) volume;
[0173] Using an image generation network and based at least on a current view representing the 3D volume, generating an updated view representing the 3D volume at the target level of detail;
[0174] Providing the updated view in response to the target level of detail;
[0175] Add the updated view to the image set associated with the 3D volume; and
[0176] Update the 3D volume based at least on the updated view.
[0177] 2. The computer-implemented method according to clause 1, wherein the 3D volume is represented by a Neural Radiance Field (NeRF).
[0178] 3. The computer-implemented method according to clause 1, wherein the image generation network is a diffusion model conditioned on text and images.
[0179] 4. The computer-implemented method according to clause 1, further comprising:
[0180] Receiving a prompt for the current view; and
[0181] Providing the prompt to a language model associated with the image generation network.
[0182] 5. The computer-implemented method according to clause 4, wherein the language model is a large language model (LLM) configured to generate an information hierarchy based at least on the prompt.
[0183] 6. The computer-implemented method according to clause 1, wherein the image generation network is a super-resolution model conditioned on the image having a resolution less than a threshold.
[0184] 7. The computer-implemented method according to clause 1, further comprising:
[0185] After receiving the updated view, removing one or more previous images from the image set; and
[0186] Updating the network associated with the 3D volume.
[0187] 8. The computer-implemented method according to clause 1, wherein the target level of detail is associated with an input command from a user in an interactive environment.
[0188] 9. The computer-implemented method according to clause 1, wherein the 3D volume is represented by a Neural Radiance Field (NeRF), the method further comprising:
[0189] Converting the NeRF to a mesh-based representation.
[0190] 10. The computer-implemented method according to clause 1, further comprising:
[0191] Receiving, at an associated LLM, a prompt requesting an information hierarchy of an object associated with the 3D volume;
[0192] Determine a plurality of sub - levels of the 3D volume based at least on the information hierarchy; and
[0193] Establish an order for the plurality of sub - levels associated with respective levels of each of the plurality of sub - levels.
[0194] 11. The computer - implemented method according to clause 10, further comprising:
[0195] Store the plurality of sub - levels;
[0196] Provide the object in response to a first command; and
[0197] Provide a sub - level of the plurality of sub - levels in response to a second command.
[0198] 12. A processor, comprising:
[0199] One or more processing units for:
[0200] Receive a request to generate an image using a Neural Radiance Field (NeRF);
[0201] Determine a target level of detail of the image according to a prompt associated with the request;
[0202] Determine that an image generated using the NeRF will not meet the target level of detail;
[0203] Generate a new image at the target level of detail through one or more diffusion models;
[0204] Provide the new image in response to the request; and
[0205] Add the new image to an image set associated with the NeRF.
[0206] 13. The processor according to clause 12, wherein the prompt is a text prompt, and wherein the one or more processing units are further for:
[0207] Provide the text prompt to a Large Language Model (LLM);
[0208] Receive a command from the LLM based at least on the text prompt; and
[0209] Provide the command to the one or more diffusion models.
[0210] 14. The processor according to clause 12, wherein the one or more diffusion models are conditioned on text and images.
[0211] 15. The processor according to clause 12, wherein at least one of the one or more diffusion models includes a super-resolution model conditional on the image having a resolution less than a threshold.
[0212] 16. The processor according to clause 12, wherein the processor is included in at least one of the following:
[0213] A system for performing simulation operations;
[0214] A system for performing simulation operations to test or verify autonomous machine applications;
[0215] A system for performing digital twin operations;
[0216] A system for performing optical transmission simulation;
[0217] A system for rendering graphical output;
[0218] A system for performing deep learning operations;
[0219] A system implemented using edge devices;
[0220] A system for generating or presenting virtual reality (VR) content;
[0221] A system for generating or presenting augmented reality (AR) content;
[0222] A system for generating or presenting mixed reality (MR) content;
[0223] A system including one or more virtual machines (VMs);
[0224] A system for performing conversational AI application operations;
[0225] A system for performing generative AI application operations;
[0226] A system for performing operations using a language model;
[0227] A system for performing one or more generative content operations using a large language model (LLM);
[0228] A system implemented at least partially in a data center;
[0229] A system for performing hardware testing using simulation;
[0230] A system for performing one or more generative content operations using a language model;
[0231] A system for synthetic data generation;
[0232] A collaborative content creation platform for 3D assets; or
[0233] A system implemented at least in part using cloud computing resources.
[0234] 17. A system comprising:
[0235] One or more processors, the one or more processors including processing circuitry configured to generate an output image having a finer level of detail (LOD) than an input image generated using a Neural Radiance Field (NeRF), and to update the NeRF using an image set including the output image.
[0236] 18. The system of clause 17, wherein the output image is generated by one or more diffusion models in response to a request.
[0237] 19. The system of clause 17, wherein the output image is at least one of a hallucinated image or an image of a higher resolution than the input image.
[0238] 20. The system of clause 17, wherein the system includes at least one of the following:
[0239] A system for performing simulation operations;
[0240] A system for performing simulation operations to test or validate an autonomous machine application;
[0241] A system for performing digital twin operations;
[0242] A system for performing light transport simulation;
[0243] A system for rendering a graphical output;
[0244] A system for performing deep learning operations;
[0245] A system implemented using an edge device;
[0246] A system for generating or presenting virtual reality (VR) content;
[0247] A system for generating or presenting augmented reality (AR) content;
[0248] A system for generating or presenting mixed reality (MR) content;
[0249] A system including one or more virtual machines (VMs);
[0250] A system for performing operations of a conversational AI application;
[0251] A system for performing operations of a generative AI application;
[0252] Systems for performing operations using a language model;
[0253] Systems for performing one or more generative content operations using a large language model (LLM);
[0254] Systems implemented at least in part in a data center;
[0255] Systems for performing hardware testing using simulation;
[0256] Systems for performing one or more generative content operations using a language model;
[0257] Systems for synthetic data generation;
[0258] A collaborative content creation platform for 3D assets; or
[0259] Systems implemented at least in part using cloud computing resources.
[0260] Other variations are within the spirit of the present disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative constructions, certain of its embodiments shown in the drawings have been described in detail above. However, it should be understood that the intention is not to limit the disclosure to the one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0261] Unless otherwise specified or clearly inconsistent with the context, in the context of describing the disclosed embodiments (especially in the context of the appended claims), the use of the terms "a", "an", "the", and similar referents should be construed to cover both the singular and the plural, rather than as a definition of the terms. Unless otherwise specified, the terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to"). The term "connected" (when unmodified refers to a physical connection) should be construed to include in part or in whole, attached to, or joined together, even with some intervening elements. Unless otherwise indicated herein, references to numerical ranges in this document are only intended as a shorthand method for referring separately to each individual value falling within the range, and each individual value is incorporated into the specification as if it were recited individually herein. Unless otherwise indicated or inconsistent with the context, the use of the term "set" (e.g., "set of items") or "subset" should be construed to include a non-empty set of one or more members. Further, unless otherwise indicated or inconsistent with the context, a "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but rather the subset and the corresponding set may be equal.
[0262] Unless otherwise expressly indicated or clearly contradicted by the context, conjunctive language such as the phrase “at least one of A, B, and C” or “at least one of A, B or C” is understood in context to typically mean that the items, clauses, etc., can be A or B or C, or any non - empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the conjunctive phrases “at least one of A, B, and C” and “at least one of A, B or C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “plurality” denotes a plural state (e.g., “a plurality of items” means multiple items). The number of items in a plurality is at least two, but can be more if expressly indicated or indicated by the context. Further, unless otherwise stated or clear from the context, the phrase “based on” means “at least partially based on” rather than “based solely on”.
[0263] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored, for example, in the form of a computer program on a computer-readable storage medium that includes multiple instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., propagating transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, and when the executable instructions are executed by one or more processors of a computer system (i.e., as a result of being executed), cause the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the multiple non-transitory computer-readable storage media lack all of the code, but the multiple non-transitory computer-readable storage media together store all of the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, e.g., the non-transitory computer-readable storage medium stores the instructions, and a main central processing unit (“CPU”) executes some instructions while a graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of the instructions.
[0264] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or jointly perform the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software that enables the implementation of the operations. Additionally, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices operating in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.
[0265] The use of any and all examples or exemplary language (e.g., "such as") provided herein is for illustrative purposes only to better clarify the embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential for practicing the disclosed subject matter.
[0266] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference to the extent that each reference is specifically and individually indicated to be incorporated by reference and the entire content thereof is set forth herein.
[0267] In the specification and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not intended as synonyms for each other. Instead, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.
[0268] Unless otherwise explicitly stated, it is understood that throughout the specification, terms such as "processing", "computing", "calculating", "determining", etc., refer to actions and / or processes of a computer or computing system or similar electronic computing device that process and / or transform data represented as a physical quantity (e.g., electronic) in the registers and / or memory of the computing system into other data similarly represented as a physical quantity in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0269] In a similar manner, the term "processor" may refer to any device or portion of a memory that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that may be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes that execute instructions sequentially or in parallel continuously or intermittently. The terms "system" and "method" may be used interchangeably herein as long as a system can embody one or more methods and a method can be considered a system.
[0270] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. There are various ways to obtain, acquire, receive, or input analog and digital data, such as by receiving data as an argument to a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transferring, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transferring, sending, or presenting analog or digital data can be implemented by transmitting the data as an input or output argument to a function call, an application programming interface, or a parameter of an interprocess communication mechanism.
[0271] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. Additionally, although specific assignments of responsibilities were defined above for purposes of discussion, the various functions and responsibilities may be assigned and partitioned differently depending on the circumstances.
[0272] Moreover, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A computer-implemented method comprising: Determine a target level of detail for a three-dimensional 3D volume; generating, using an image generation network and based at least on a current view representing the 3D volume, an updated view representing the 3D volume at the target level of detail; providing the updated view in response to the target level of detail; adding the updated view to a set of images associated with the 3D volume; as well as The 3D volume is updated based at least on the updated view.
2. The computer-implemented method of claim 1, wherein the 3D volume is represented by a Neural Radiance Field (NeRF).
3. The computer-implemented method of claim 1, wherein the image generation network is a diffusion model conditioned on text and images.
4. The computer-implemented method of claim 1 , further comprising: receiving a prompt for the current view; as well as The hint is provided to a language model associated with the image generation network.
5. The computer-implemented method of claim 4, wherein the language model is a Large Language Model (LLM) configured to generate an information hierarchy based at least on the prompt.
6. The computer-implemented method of claim 1, wherein the image generation network is a super-resolution model conditioned on an image having a resolution less than a threshold.
7. The computer-implemented method of claim 1 , further comprising: upon receiving the updated view, removing one or more previous images from the set of images; as well as A network associated with the 3D volume is updated.
8. The computer-implemented method of claim 1, wherein the target level of detail is associated with an input command from a user of an interactive environment.
9. The computer-implemented method of claim 1 , wherein the 3D volume is represented by a Neural Radiance Field (NeRF), the method further comprising: Convert the NeRF to a grid-based representation.
10. The computer-implemented method of claim 1 , further comprising: receiving, at an associated LLM, a prompt requesting an information hierarchy of objects associated with the 3D volume; determining a plurality of sub-levels of the 3D volume based at least on the information hierarchy; as well as An ordering is established for the plurality of sub-levels associated with a respective level of each of the plurality of sub-levels.
11. The computer-implemented method of claim 10, further comprising: storing the plurality of sub-levels; providing the object in response to a first command; as well as A sub-level of the plurality of sub-levels is provided in response to a second command.
12. A processor, comprising: One or more processing units for: receiving a request to generate an image using a neural radiation field NeRF; determining a target level of detail for the image based on a hint associated with the request; determining that an image generated using the NeRF will not meet the target level of detail; generating a new image of the target level of detail by one or more diffusion models; providing the new image in response to the request; as well as The new image is added to the image set associated with the NeRF.
13. The processor of claim 12, wherein the prompt is a text prompt, and wherein the one or more processing units are further configured to: Providing the text prompt to a large language model (LLM); receiving a command from the LLM based at least on the text prompt; and The commands are provided to the one or more diffusion models.
14. The processor of claim 12, wherein the one or more diffusion models are conditioned on text and images.
15. The processor of claim 12, wherein at least one of the one or more diffusion models comprises a super-resolution model conditioned on an image having a resolution less than a threshold.
16. The processor of claim 12, wherein the processor is included in at least one of: A system for performing simulation operations; Systems for performing simulated operations to test or validate autonomous machine applications; Systems for performing digital twin operations; A system for performing light transport simulations; Systems for rendering graphics output; Systems for performing deep learning operations; Systems implemented using edge devices; Systems for generating or presenting virtual reality (VR) content; Systems for generating or presenting augmented reality (AR) content; Systems for generating or presenting mixed reality (MR) content; A system comprising one or more virtual machines VM; Systems for performing conversational AI application operations; Systems for performing operations of generative AI applications; Systems that use language models to perform actions; A system for performing one or more generative content operations using a large language model (LLM); A system implemented at least in part in a data center; Systems that perform hardware testing using simulation; A system for performing one or more generative content operations using a language model; Systems for synthetic data generation; A collaborative content creation platform for 3D assets; or A system implemented at least in part using cloud computing resources.
17. A system comprising: One or more processors, the one or more processors comprising processing circuitry for generating an output image having a finer level of detail (LOD) than an input image generated using a neural radiation field (NeRF), and updating the NeRF using an image set including the output image.
18. The system of claim 17, wherein the output image is generated by one or more diffusion models in response to a request.
19. The system of claim 17, wherein the output image is at least one of a ghosted image or a higher resolution image relative to the input image.
20. The system of claim 17, wherein the system comprises at least one of the following: A system for performing simulation operations; Systems for performing simulated operations to test or validate autonomous machine applications; Systems for performing digital twin operations; A system for performing light transport simulations; Systems for rendering graphics output; Systems for performing deep learning operations; Systems implemented using edge devices; Systems for generating or presenting virtual reality (VR) content; Systems for generating or presenting augmented reality (AR) content; Systems for generating or presenting mixed reality (MR) content; A system comprising one or more virtual machines VM; Systems for performing operations of conversational AI applications; Systems for performing operations of generative AI applications; A system for performing operations using a language model; A system for performing one or more generative content operations using a large language model (LLM); A system implemented at least in part in a data center; A system for performing hardware testing using simulation; A system for performing one or more generative content operations using a language model; Systems for synthetic data generation; A collaborative content creation platform for 3D assets; or A system implemented at least in part using cloud computing resources.