Three-dimensional tomographic reconstruction pipeline
By using a neural network model for 3D tomographic reconstruction pipelines and self-supervised training, the problem of noise in tomographic images was solved, enabling the generation of high-quality 3D density volumes and 2D density images while reducing radiation dose.
Patent Information
- Application Number
- CN202111529491.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-01
- Filing Date
- 2021-12-14
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing tomographic images are often noisy, affecting their diagnostic value, while increasing the X-ray radiation dose is harmful to subjects. Current technologies are not effective in reducing noise and radiation dose.
A 3D tomographic reconstruction pipeline, including a first neural network model, a fixed-function back projection unit, and a second neural network model, is used to reconstruct 3D density volumes from tomographic images through self-supervised training, thereby reducing noise and radiation dose.
It achieves the reduction of X-ray radiation dose while reducing noise, generates high-quality 3D density volumes, and can cut from any angle to obtain clear 2D density images.
Smart Images

Figure CN114638927B_ABST
Abstract
Description
[0001] CLAIM
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 126,025, filed December 16, 2020, entitled “Three-Dimensional Reconstruction for Computed Tomography,” the entire contents of which are incorporated herein by reference. BACKGROUND
[0003] Tomographic images generated from x-ray machines, such as cone-beam scanning machines, provide valuable diagnostic information. Typically, tomographic images are noisy, which interferes with or reduces the diagnostic value of the tomographic images. Although increasing x-ray radiation dosage reduces noise, increasing dosage is harmful to the subject being scanned. There is a need to address these problems and / or other problems associated with the prior art. SUMMARY
[0004] Embodiments of the present disclosure relate to a three-dimensional (3D) tomographic reconstruction pipeline. Systems and methods are disclosed for constructing a 3D density volume from tomographic images, such as x-ray images. A tomographic image is a captured projection image of all structures of a subject or object (e.g., a human body) between a beam source and an imaging sensor. The beam effectively integrates along a path through the object, producing a tomographic image at the imaging sensor, where each pixel represents an attenuation along the path. In one embodiment, the 3D reconstruction pipeline includes a first neural network model, a fixed-function back-projection unit, and a second neural network model. Given information of the capture environment, a tomographic image is processed by the 3D reconstruction pipeline to produce a reconstructed 3D density volume of the object. In contrast to a set of two-dimensional (2D) slices, the entire 3D density volume is reconstructed, so that 2D density images can be computed by slicing through any portion of the 3D density volume at any angle.
[0005] A method, computer-readable medium, and system for 3D tomographic reconstruction are disclosed. The method includes the steps of processing, by a first neural network, a tomographic image to produce at least one channel of 2D features for each tomographic image, and computing 3D features by back-projecting the at least one channel of 2D features of the tomographic image according to characteristics of a physical environment used to capture the tomographic image. The 3D features are processed by a second neural network to produce a 3D density volume corresponding to the tomographic image.
[0006] A method, computer readable medium, and system for training a 3D tomographic reconstruction neural network system are disclosed. The method includes the steps of processing, by the neural network system, a 2D tomographic image of an object according to parameters to produce a 3D density volume of the object, where the 2D tomographic image is generated by a physical capture environment. The 3D density volume is projected based on characteristics of the capture environment to produce a simulated tomographic image corresponding to the 2D tomographic image, and the parameters of the neural network system are adjusted to reduce a difference between the simulated tomographic image and the 2D tomographic image. BRIEF DESCRIPTION OF DRAWINGS
[0007] The present system and method for 3D tomographic reconstruction pipeline is described in detail below with reference to the accompanying drawings, wherein:
[0008] FIG. 1A An environment suitable for use in implementing some embodiments of the present disclosure for generating tomographic images is illustrated.
[0009] FIG. 1B A block diagram of an example 3D tomographic reconstruction system suitable for use in implementing some embodiments of the present disclosure is illustrated.
[0010] FIG. 1C A flowchart of a method for 3D tomographic reconstruction suitable for use in implementing some embodiments of the present disclosure is illustrated.
[0011] FIG. 1D A block diagram of an example 2D density image generation system suitable for use in implementing some embodiments of the present disclosure including a 3D tomographic reconstruction system is illustrated. FIG. 1B
[0012] FIG. 1E A 2D density image generated from a reconstructed 3D density volume according to one embodiment is illustrated.
[0013] FIG. 1F Another 2D density image generated from a reconstructed 3D density volume according to one embodiment is illustrated.
[0014] FIG. 1G Yet another 2D density image generated from a reconstructed 3D density volume according to one embodiment is illustrated.
[0015] FIG. 2A A conceptual diagram illustrating back-projections of 2D tomographic images contributing to a 3D density volume according to one embodiment is illustrated.
[0016] FIG. 2B A conceptual diagram illustrating back-projections of additional 2D tomographic images also contributing to a 3D density volume according to one embodiment is illustrated.
[0017] FIG. 2C FIG. illustrates a flow diagram of a method suitable for implementing steps of the flow diagram shown in FIG. FIG. 1B
[0018] FIG. 3A FIG. illustrates a block diagram of an example 3D tomographic reconstruction system training configuration suitable for implementing some embodiments of the present disclosure.
[0019] FIG. 3B FIG. illustrates a flow diagram of a method suitable for implementing some embodiments of the present disclosure for training a 3D tomographic reconstruction system.
[0020] FIG. 4 FIG. illustrates an example parallel processing unit suitable for implementing some embodiments of the present disclosure.
[0021] FIG. 5A FIG. illustrates a conceptual diagram of a processing system implemented using a PPU suitable for implementing some embodiments of the present disclosure. FIG. 4
[0022] FIG. 5B FIG. illustrates an example system in which various previously described embodiments can be implemented.
[0023] FIG. 5C FIG. illustrates components of an example system that can be used for training and utilizing machine learning in at least one embodiment.
[0024] FIG. 6A FIG. illustrates a conceptual diagram of a graphics processing pipeline implemented using a PPU suitable for implementing some embodiments of the present disclosure. FIG. 4
[0025] FIG. 6B FIG. illustrates an example streaming system suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] Systems and methods related to 3D tomographic reconstruction are disclosed. In one embodiment, a 3D density volume (model) is constructed from tomographic images (e.g., x-ray images). A tomographic image is a captured projection image of all structures of an object (e.g., a human body) between a beam source and an imaging sensor. In one embodiment, the imaging sensor includes multiple rows of detector elements. The beam is effectively integrated along a path through the object, producing a tomographic image at the imaging sensor where each pixel represents attenuation along that path, producing a 2D projection image such as an x-ray image. In one embodiment, a 3D reconstruction pipeline includes a first neural network model, a fixed-function back projection unit, and a second neural network model. Given information of a captured environment, a tomographic image is processed by the 3D reconstruction pipeline to produce a reconstructed 3D density volume of the object. In contrast to a set of 2D slices, the entire 3D density volume is reconstructed, so that 2D density images can be computed by slicing through any portion of the 3D density volume at any angle.
[0027] In the context of the following description, several terms are defined as follows.
[0028] Tomographic image: A 2D projection image of a 3D volume of an object, subject, or body generated by a tomographic machine.
[0029] 3D density volume: A reconstructed 3D model generated from tomographic images.
[0030] Simulated tomographic image: A simulated projection or 2D density image generated from a 3D density volume to simulate a real tomographic image.
[0031] Slice: A planar portion of a 3D density volume, such as a 2D density image corresponding to a 2D plane that slices through or transverses the 3D density volume.
[0032] A slice is a reconstructed image at a 2D plane that illustrates what is inside the body at that plane, with no “fogging” projection of material between the x-ray source and the contents at that plane. Conventional back projection techniques reconstruct individual 2D slices from tomographic images using pixels selected based on beam and slice plane. In contrast, the entire 3D density volume is reconstructed from the tomographic images, and individual slices can be generated from the 3D density volume.
[0033] In one embodiment, projection images (or 2D features generated from projection images) used to generate a 3D density volume during 3D tomographic reconstruction are pre-filtered. Pre-filtering the projection images provides a set of projection images at varying resolutions, such as MIP (multum in parvo) maps that include varying levels of detail. Generating a set of pre-filtered projection images is efficient, and techniques such as bilinear and / or trilinear filtering can be used to sample one or more of the different pre-filtered projection images in the set to reduce aliasing of back projection data.
[0034] Systems and methods related to end-to-end training for a 3D tomographic reconstruction pipeline are disclosed. In contrast to conventional systems that employ supervised training, self-supervised training can be used to train a 3D tomographic reconstruction pipeline. Conventional supervised training requires 2D tomographic images and corresponding ground truth density data as training data. However, perfectly noise-free ground truth density data is not available. Rather than attempting to train the reconstruction pipeline using an estimated 3D density volume obtained by some other technique, self-supervised training instead generates simulated tomographic images from the 3D density volume output by the reconstruction pipeline. The simulated 2D input tomographic images can be compared to the input tomographic images used to generate the 3D density volume. The parameters of the 3D tomographic reconstruction pipeline can be learned to reduce the difference between the simulated and input tomographic images.
[0035] FIG. 1A An environment 100 for generating tomographic images suitable for implementing some embodiments of the present disclosure is illustrated. Different types of machines capture tomographic images according to the physical environment using different mechanisms and beam scan paths. For example, a first machine (not shown) can project planar (e.g., parallel) beams. A second machine, such as machine 110, can include a beam source 105 that projects a conical or pyramidal beam 108 onto a circular imaging sensor 112 disposed within machine 110. In one embodiment, imaging sensor 112 is flat rather than curved. In one embodiment, imaging sensor 112 rotates in coordination with beam source 105. In one embodiment, imaging sensor 112 rotates at the same speed as beam source 105.
[0036] Machine 110 moves beam 108 in a circle around object 120. Object 120 moves continuously through circular imaging sensor 112, resulting in a helical beam scan path 115. Another machine (not shown) can rotate a beam in a circle around an object and move the object through the circle alternately, generating several disconnected circular tomographic images at different points along the length of the object.
[0037] Beam 108 moves along beam scan path 115 and forms a projection image at imaging sensor 112. Given information specific to capture environment 100, a tomographic image can be back-projected to produce a reconstructed 2D density image or 3D density volume of the object. Characteristics of capture environment 100 can include the specific geometry, position, and / or orientation of beam source 105, beam scan path 115, imaging sensor 112, beam 108, and so on.
[0038] FIG. 1BA block diagram of an example 3D tomographic reconstruction system 125 suitable for implementing some embodiments of the present disclosure is illustrated. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be omitted altogether. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. The various functions described herein as being performed by entities can be carried out by hardware, firmware, and / or software. For instance, various functions can be implemented by a processor executing instructions stored in memory. Moreover, those of ordinary skill in the art will recognize that any system embodying the operations of the 3D tomographic reconstruction system 125 is within the scope and spirit of embodiments of the present disclosure.
[0039] The 3D tomographic reconstruction system 125 is a 3D reconstruction pipeline that includes a 2D neural network 140, a fixed-function back-projection unit 145, and a 3D density volume construction neural network 150. The 2D neural network 140 receives 2D tomographic images and generates at least one channel (per pixel) of 2D features for each tomographic image. The tomographic images are typically quite noisy (presenting as grainy, coarsely sampled, and / or including visual artifacts). While noise can be reduced by increasing the x-ray radiation dose used to capture the tomographic images, increasing x-ray radiation can be harmful to the subject. When conventional back-projection techniques are used to generate 2D density images (slices), the noise present in the tomographic images is also back-projected, resulting in noisy 2D density images. The 2D neural network 140 reduces noise when processing the tomographic images, yielding a 3D density volume with reduced noise. The ability to reduce noise can beneficially allow for lower x-ray doses.
[0040] First, each of the tomographic images is individually processed by one or more 2D neural networks 140 independently to generate at least one channel of 2D features. In one embodiment, multiple channels of 2D features are generated to provide higher dimensional data. The 2D tomographic images inherently include 3D information because they are projections with per-pixel attenuation values. Thus, each channel of 2D features produced encodes 3D information represented in the tomographic image.
[0041] The back-projection unit 145 "smears" the 2D features of each tomographic image along the beam 108 used to capture the tomographic image to generate 3D features. In one embodiment, the 3D features include associated attributes and voxels of a 3D density volume. FIG. 2A and FIG. 2BThe application is conceptually illustrated in FIG. 1. The calculations performed by the back-projection unit 145 can vary based on the particular capture environment, such as the capture environment 100. Having 2D tomographic images captured from a wide variety of different angles allows for the recovery of 3D data. The 2D features of each tomographic image can be independently (and in parallel) back-projected to compute voxel attributes. Importantly, the 2D features of multiple images can contribute to a single voxel. In contrast to conventional techniques that generate a single slice, more pixels of the tomographic images are typically utilized during the back-projection computation to produce a 3D density volume. In one embodiment, back-projection includes a high-pass filtering and a 2D image lookup operation. In another embodiment, back-projection can also include a ray-casting operation. Gather-based back-projection methods loop over 3D voxels and perform lookups from 2D tomographic images. Scatter-based back-projection methods loop over 2D pixels and perform ray-casting to 3D voxels. The back-projection technique is described in detail in the doctoral thesis of H. Turbell, "Cone-beam reconstruction using filtered back-projection," Uppsala University, February 2001, which is incorporated herein by reference in its entirety. The back-projection technique is described in detail in the doctoral thesis of H. Turbell, "Cone-beam reconstruction using filtered back-projection," Uppsala University, February 2001, which is incorporated herein by reference in its entirety.
[0042] Each of the 2D neural network 140 and the 3D density volume construction neural network 150 is a learned filter implemented using a neural network model. In contrast, the back-projection unit 145 performs a fixed function operation and does not require training. However, in one embodiment, the 2D neural network 140 and / or the 3D density volume construction neural network 150 is replaced with a fixed function filter.
[0043] The 3D density volume construction neural network 150 processes 3D features (voxels and attributes) to generate a 3D reconstruction of the subject as a 3D density volume. FIG. 1B The 3D density volume in FIG. 1 is a conceptual representation of the torso that has portions of the outermost layer and the underlying layers removed to illustrate that the 3D density volume represents the internal structure of the subject in contrast to a 3D mesh that consists only of the outermost layer.
[0044] During processing, the 3D density volume construction neural network 150 corrects for reconstruction errors introduced in the back-projection computation. The 3D density volume construction neural network 150 can reduce the remaining noise present in the 3D features, resulting in a noise-reduced 3D density volume. One advantage of reconstructing the entire 3D density volume is that 2D density images can be generated by slicing through any portion of the 3D density volume at any angle.
[0045] Conventional techniques process 2D tomographic images to generate individual slices of a particular transverse plane (e.g., cross-section). For example, a particular transverse x,y plane can correspond to different z coordinate values on a z-axis along which a subject moves through a scanner. While multiple slices can be generated by conventional techniques, each slice is reconstructed individually.
[0046] More illustrative information will now be set forth regarding various optional architectures and features with which the foregoing framework can be implemented in accordance with the user's desires. It should be strongly noted that the following information is set forth in
[0047] FIG. 1C A flow diagram illustrating a method 160 for 3D tomographic reconstruction suitable for implementing some embodiments of the present disclosure is shown. Each block of the method 160 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For instance, various functions can be carried out by a processor executing instructions stored in memory. The method can also be implemented as computer-usable instructions stored on a computer storage medium. The method can be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. In addition, the method 160 is described with respect to the system of FIG. 1B However, additionally or in the alternative, the method can be performed by any one system, or any combination of systems, including but not limited to those described herein. Further, those of ordinary skill in the art will recognize that any system performing the method 160 is within the scope and spirit of the present embodiments.
[0048] At step 165, the tomographic images are processed by a first neural network to produce at least one channel of 2D features for each tomographic image. In one embodiment, the first neural network is the 2D neural network 140. The tomographic images are comprised of attenuation values for each pixel, where the 2D features can include one or more feature values (channels) associated with each pixel.
[0049] At step 170, 3D features are computed by back-projecting at least one channel of 2D features of the tomographic image. The at least one channel of 2D features is back-projected according to characteristics of a physical environment used to capture the tomographic image. In one embodiment, back-projection unit 145 computes the 3D features. In one embodiment, the 3D features are voxels and associated attributes. In one embodiment, the physical environment used to capture the tomographic image includes a cone-beam computed tomography machine. In one embodiment, back-projection includes computing a footprint of a projection of a voxel or pixel, and accessing one or more pre-filtered versions of the 2D features or tomographic image according to at least one dimension of the footprint of the projection. In one embodiment, the footprint of the projection is computed for a pixel, such as when performing conventional slice reconstruction. Conceptually, a 3D voxel of the 3D features corresponds to a 2D pixel on a slice.
[0050] At step 175, the 3D features are processed by a second neural network to produce a 3D density volume corresponding to the tomographic image. In one embodiment, the second neural network is neural network 150 that constructs the 3D density volume. In one embodiment, noise present in the tomographic image is reduced in the 3D density volume. In one embodiment, the 3D density volume corresponds to a portion of a human body.
[0051] FIG. 1D FIGURE 1 illustrates a block diagram of an example 2D density image generation system including a 3D tomographic reconstruction system 125 suitable for use in implementing some embodiments of the present disclosure. FIG. 1B FIGURE 1 illustrates a block diagram of an example 2D density image generation system including a 3D tomographic reconstruction system 125 suitable for use in implementing some embodiments of the present disclosure.
[0052] FIG. 1E FIGURE 2 illustrates a 2D density image 152 generated from a reconstructed 3D density volume, according to one embodiment. A plane 154 is defined for generating the 2D density image 152 from the reconstructed 3D density volume of a human torso. As shown in FIGURE 2, the plane 154 transects the reconstructed 3D density volume to produce the 2D density image 162 with a diagonal orientation from approximately the clavicle to the middle of the spine. FIG. 1E FIGURE 2 illustrates a 2D density image 152 generated from a reconstructed 3D density volume, according to one embodiment. A plane 154 is defined for generating the 2D density image 152 from the reconstructed 3D density volume of a human torso. As shown in FIGURE 2, the plane 154 transects the reconstructed 3D density volume to produce the 2D density image 162 with a diagonal orientation from approximately the clavicle to the middle of the spine.
[0053] FIG. 1F FIGURE 3 illustrates another 2D density image 162 generated from a reconstructed 3D density volume, according to one embodiment. A plane 164 is defined for generating the 2D density image 162 from the reconstructed 3D density volume of a human torso. As shown in FIGURE 3, the plane 164 transects the reconstructed 3D density volume to produce the 2D density image 162 with a vertical orientation approximately aligned with the spine. FIG. 1F FIGURE 3 illustrates another 2D density image 162 generated from a reconstructed 3D density volume, according to one embodiment. A plane 164 is defined for generating the 2D density image 162 from the reconstructed 3D density volume of a human torso. As shown in FIGURE 3, the plane 164 transects the reconstructed 3D density volume to produce the 2D density image 162 with a vertical orientation approximately aligned with the spine.
[0054] FIG. 1GA further 2D density image 172 generated from the reconstructed 3D density volume is illustrated according to one embodiment. A plane 174 is defined for generating the 2D density image 172 from the reconstructed 3D density volume of the human torso. As shown in FIG. 1G The plane 174 intersects the reconstructed 3D density volume to produce the 2D density image 172 with a vertical orientation passing through the chest and offset to the left from the spine, as shown in
[0055] FIG. 2A and FIG. 2B A conceptual diagram illustrating back-projection of 2D tomographic images contributing to the generation of a 3D density volume according to one embodiment. In one embodiment, 2D features are generated for the tomographic images by the 2D neural network 140, and the 2D features are effectively smeared back along the paths of the respective beams by the back-projection unit 145 to produce 3D features. In one embodiment, each back-projected 2D tomographic image affects 3D voxels in the region corresponding to the beams used to produce the projected 2D tomographic image. For visualization purposes, partial 2D density images and a full accumulated 2D density image are used to represent the 3D features corresponding to a slice of the 3D density volume.
[0056] FIG. 2C A flowchart of a method suitable for implementing the steps 170 of the method 160 shown in FIG. 1B In step 170, 3D features are computed by back-projecting at least one channel of 2D features of the tomographic images, as previously described.
[0057] In step 235, a 3D pre-filter Gaussian distribution in the 3D voxel space is determined. The 3D pre-filter Gaussian distribution can be determined based on properties of the input tomographic images and the voxel grid, such as resolution, pixel / voxel spacing, or characteristics of the physical environment, such as the capture environment 100, and / or characteristics of the detector, such as the imaging sensor 112. Importantly, the content of the images is not used to determine the 3D pre-filter Gaussian distribution. In one embodiment, the pre-filter Gaussian distribution is sized according to the spacing of the 3D voxels. In step 240, a projection of the 3D pre-filter Gaussian distribution onto the detector is computed to produce a projected 2D Gaussian distribution at the detector. In one embodiment, the projection is computed according to characteristics of the physical environment in which the tomographic images were captured, i.e., the capture environment.
[0058] At step 245, texture space occupancy is computed based on the projected 2D Gaussian distribution. In one embodiment, the texture space for texture data is equivalent to the 2D feature space for at least one channel of 2D features. At step 250, 2D features generated for the tomographic image are sampled based on the computed texture space occupancy to compute 3D features. In one embodiment, texture coordinates and associated filtering information are determined for sampling the 2D features. Pre-filtering the 2D features of the projection images provides a set of 2D features with varying resolutions. Generating a set of pre-filtered 2D features is efficient, and techniques such as bilinear and / or trilinear filtering can be used to sample one or more of the different pre-filtered 2D features in the set to reduce aliasing of backprojection data. Conventional backprojection techniques typically sample the highest resolution 2D tomographic image and compute backprojection data using filtering for any size of occupancy. For example, when using nearest neighbor sampling, aliasing can cause line artifacts in backprojection volumes generated using conventional techniques. While the pre-filtering techniques are described in the context of 3D tomographic reconstruction system 125, the pre-filtering techniques can be used to improve the quality of 2D density images produced by conventional backprojection-based systems.
[0059] In one embodiment, one or more of steps 235, 240, 245, and 250 are performed in parallel for at least a portion of voxels in the 3D features. In one embodiment, one or more of steps 235, 240, 245, and 250 are performed in parallel for 2D features of at least a portion of the tomographic images. In contrast to conventional techniques, the backprojection operations performed by step 170 directly generate 3D features of the 3D density volume by backprojecting beams corresponding to each pixel in the tomographic images based on characteristics of the captured environment. Multiple pixels from different 2D tomographic images contribute to the 3D features. The characteristics can be used to determine the origin and direction of each beam. Thus, approximations relied upon by conventional backprojection techniques such as rebinning helical trajectory cone measurements into a flat Z-plane can be avoided. Furthermore, reconstruction errors caused by the combination of helical trajectories and backprojection can be corrected by the 3D density volume constructing neural network 150.
[0060] In general, the generation of the 3D density volume allows for the implementation of 2D density images computed for any plane that is transverse to the 3D density volume. Since the 3D tomographic reconstruction system 125 does not necessarily propagate noise present in the tomographic images to the 3D density volume, the radiation dose used to capture the tomographic images can be reduced.
[0061] End-to-end training for three-dimensional tomographic reconstruction pipelines
[0062] Conventional supervised learning techniques require a reference 3D density volume as a guide or ground truth output for use during training. 3D tomographic reconstruction is somewhat unique in that a reference 3D density volume cannot be directly measured. For example, providing reference density data for a human body requires physical sampling of the human body, which is impractical, if not impossible. Reference 3D density volumes constructed from tomographic images using conventional systems are not eligible as true references due to the presence of noise and other artifacts in the flawed reference 3D density volume. A neural network-based system being trained will simply learn to reproduce the artifacts present in the flawed reference 3D density volume rather than generate a higher quality 3D density volume.
[0063] For best results, supervised training should use 2D tomographic images and corresponding ground truth 3D density data that represent the inputs seen in production use, i.e., when the system being trained is deployed in a clinical environment. Thus, the 2D tomographic images typically include noise. Increasing the radiation dose can reduce the noise in the 2D tomographic images, but, unfortunately, completely noise-free ground truth 3D density data is not available. If noise-free or low-noise 2D tomographic images are available, then different types and amounts of noise and / or other corruption can be introduced in these 2D tomographic images to train the system for deployment in a clinical environment.
[0064] 3D tomographic reconstruction system 125 can be trained using self-supervised training. In contrast, conventional systems that include neural networks employ supervised training. The self-supervised training methods described herein can also be applied to conventional 3D tomographic reconstruction systems.
[0065] For self-supervised training, simulated 2D tomographic images are generated from 3D density volumes output from a reconstruction pipeline such as 3D tomographic reconstruction system 125. The 3D density volumes are projected according to the capture environment used to generate the captured 2D tomographic images to produce simulated 2D tomographic images. Each of these simulated tomographic images is compared to a corresponding (machine-generated) captured 2D tomographic image. The captured 2D tomographic images input to the reconstruction pipeline effectively serve as reference (ground truth) 2D tomographic images.
[0066] A loss function can be computed based on the difference between the simulated 2D tomographic images and the captured 2D tomographic images. The difference determined by the loss function is backpropagated to update the neural network model parameters of the reconstruction pipeline. The reconstruction pipeline learns to generate noiseless reconstructions even when the captured 2D tomographic images are noisy and / or corrupted. The ability of the reconstruction pipeline to learn to generate noiseless 3D density volumes can appear surprising and is explained by the noise-to-noise principle described by Lehtinen et al. in "Noise2Noise: Learning Image Restoration without Clean Data," in the International Conference on Machine Learning (ICML), October 2018. The noise in the tomographic images is zero-mean photon noise, and the expected difference for the optimal 3D density volume is minimized. In summary, the reconstruction pipeline can learn to remove noise from the images even when trained using only noisy input images. Thus, the reconstruction pipeline can learn to generate 3D density volumes that are less noisy or noiseless even when trained using only noisy tomographic images.
[0067] FIG. 3A A block diagram illustrating an example 3D tomographic reconstruction system training configuration 300 suitable for implementing some embodiments of the present disclosure is shown. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be wholly omitted depending on the context. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented in hardware, firmware, and / or software. For instance, various functions can be implemented by a processor executing instructions stored in memory. Moreover, those of ordinary skill in the art will appreciate that any system that performs the operations of the 3D tomographic reconstruction system training configuration 300 is within the scope and spirit of embodiments of the present disclosure.
[0068] The 3D tomographic reconstruction system training configuration 300 includes the 3D tomographic reconstruction system 125, a tomographic image simulator 310, and a loss minimization unit 320. The tomographic image simulator 310 receives a 3D density volume and generates simulated tomographic images corresponding to captured tomographic images. In one embodiment, the tomographic image simulator 310 performs a ray marching operation to generate the simulated tomographic images. The tomographic image simulator 310 simulates what the 3D density volume being reconstructed would produce when imaged in a capture environment. If the 3D density volume is an exact representation of the physical volume being imaged, then the simulated tomographic images should be similar to the captured tomographic images. In one embodiment, the simulated tomographic images match the captured tomographic images without noise.
[0069] The loss minimization unit 320 identifies a difference between the simulated tomographic image and the (captured) tomographic image to generate a training signal for updating the parameters of the 2D neural network 140 and the 3D density volume construction neural network 150. In one embodiment, the loss minimization unit 320 generates weight updates that minimize the difference using an L2 norm (least squares error). In one embodiment, self-supervised training is performed using as many available tomographic images as possible. In one embodiment, the number of input tomographic images is limited, for example, to introduce randomness in the training, to improve robustness to different physical settings, or to enable evaluation of the system with different validation sets. The tomographic images need not be associated with the same subject. Thus, a training dataset of captured tomographic images is readily available. The training dataset can include tomographic images captured using low radiation dose, with high noise levels, and / or tomographic images captured using higher radiation dose, with lower noise levels. In one embodiment, a different subset of tomographic images is used in each training iteration.
[0070] FIG. 3B A flowchart of a method 330 for training a 3D tomographic reconstruction system suitable for implementing some embodiments of the present disclosure is illustrated. Each block of the method 330 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For instance, various functions can be implemented by a processor executing instructions stored in memory. The method can also be implemented as computer-usable instructions stored on a computer storage media. The method can be provided by a standalone application, a service, or a hosted service (either independently or in combination with another hosted service or another product's plug-in, to name a few. Moreover, the method 330 is described with respect to the system of FIG. 1B and FIG. 3A However, additionally or alternatively, the method can be performed by any of the systems, or any combination of systems, including but not limited to those described herein. Moreover, those of ordinary skill in the art will recognize that any system performing the method 330 is within the scope and spirit of the embodiments of the present disclosure.
[0071] At step 335, a 2D tomographic image of a subject is processed by the neural network system in accordance with the parameters to produce a 3D density volume of the subject. In one embodiment, the 2D tomographic image is generated by a physical capture environment. In one embodiment, the neural network system is the 3D tomographic reconstruction system 125. In one embodiment, the entire 3D density volume is reconstructed. In one embodiment, the neural network system produces a set of slices of the 3D volume of the subject, rather than the entire 3D density volume. In one embodiment, the 3D density volume corresponds to a portion of a human body. In one embodiment, the physical capture environment includes a cone-beam computed tomography machine.
[0072] In one embodiment, the neural network system computes 3D data by backprojecting the 2D tomographic images according to the characteristics of the physically captured environment, and processing the 3D data by the neural network model to produce the 3D density volume. In one embodiment, the backprojecting includes computing a footprint of a projection of a pixel, and accessing one or more pre-filtered versions of the 2D tomographic images according to at least one dimension of the footprint of the projection.
[0073] In one embodiment, the neural network system produces the 3D density volume by processing the 2D tomographic images by a first neural network model to produce at least one channel of 2D features, computing three-dimensional features by backprojecting the at least one channel of 2D features according to the characteristics, and processing the 3D features by a second neural network to produce the 3D density volume corresponding to the 2D tomographic images. In one embodiment, the backprojecting includes computing a footprint of a projection of a pixel, and accessing one or more pre-filtered versions of the at least one channel of 2D features according to at least one dimension of the footprint of the projection.
[0074] At step 340, the 3D density volume is projected based on the characteristics of the physically captured environment to produce simulated tomographic images corresponding to the 2D tomographic images. In one embodiment, noise present in the 2D tomographic images is reduced in the simulated tomographic images. The projection operation can implement ray marching to integrate the 3D density volume along rays corresponding to pixels of the tomographic images, i.e., a discretized version of the standard volume attenuation line integral that describes how the tomographic images relate to the underlying 3D density volume. Formulas for the projection operation are detailed as Equations 2.1 and 2.2 in H. Turbell, “Cone-beam reconstruction using filtered backprojection” Ph.D. thesis, University of Sweden, February 2001.
[0075] At step 345, parameters of the neural network system are adjusted to reduce the difference between the simulated tomographic images and the 2D tomographic images. In one embodiment, the parameters are weights of the 2D neural network 140 and / or the 3D density volume construction neural network 150. In one embodiment, steps 335, 340, and 345 are repeated for additional 2D tomographic images of additional subjects. In one embodiment, steps 335, 340, and 345 are repeated several times for one or more subjects. In one embodiment, only a subset of the available 2D tomographic images is used at a time.
[0076] Self-supervised training of a neural network-based tomographic reconstruction system can be used without ground truth reference data. Instead of requiring reference 3D density data as for supervised training, simulated tomographic images are generated from 3D density data (e.g., an entire 3D density volume or a set of slices). The simulated tomographic images are generated during self-supervised training using 3D density data output by the neural network-based tomographic reconstruction system. Advantageously, simulated tomographic images can be generated corresponding to all available 2D tomographic images of a particular subject.
[0077] Parallel processing architecture
[0078] FIG. 4 A parallel processing unit (PPU) 400 according to one embodiment is illustrated. The PPU 400 can be used to implement the 3D tomographic reconstruction system 125 and / or the 3D tomographic reconstruction system training configuration 300. The PPU 400 can be used to implement one or more of the 2D neural network 140, the back projection unit 145, the 3D density volume construction neural network 150, the slice generation unit 185, the tomographic image simulator 310, and the loss minimization unit 320. In one embodiment, a processor such as the PPU 400 can be configured to implement a neural network model. The neural network model can be implemented as software instructions executed by the processor, or in other embodiments, the processor can include a matrix of hardware elements configured to process a set of inputs (e.g., electrical signals representing values) to generate a set of outputs that can represent activations of the neural network model. In other embodiments, the neural network model can be implemented as a combination of processing by the matrix of hardware elements and software instructions. Implementing the neural network model can include determining a set of parameters for the neural network model through, for example, supervised or unsupervised training of the neural network model, and or alternatively, performing inference using the set of parameters to process a new set of inputs.
[0079] In one embodiment, the PPU 400 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 400 is a latency hiding architecture designed to process many threads in parallel. A thread (e.g., an execution thread) is an instantiation of a set of instructions configured to be executed by the PPU 400. In one embodiment, the PPU 400 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data in order to generate two-dimensional (2D) image data for display on a display device. In other embodiments, the PPU 400 can be used to perform general purpose computations. Although one exemplary parallel processor is provided herein for purposes of illustration, it is specifically intended that such processor be illustrative of any processor which can be substituted for the processor as is known in the art.
[0080] One or more PPU 400s can be configured to accelerate thousands of high-performance computing (HPC), data center, cloud computing, and machine learning applications. PPU 400s can be configured to accelerate numerous deep learning systems and applications used in autonomous vehicles, simulations, computational graphics such as ray or path tracing, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.
[0081] like FIG. 4 As shown, PPU 400 includes an input / output (I / O) unit 405, a front-end unit 415, a scheduler unit 420, a job allocation unit 425, a hub 430, a crossbar (Xbar) 470, one or more general purpose processing clusters (GPCs) 450, and one or more memory partitioning units 480. PPU 400 can be connected to a host processor or other PPU 400 via one or more high-speed NVLink 410 interconnects. PPU 400 can be connected to a host processor or other peripheral devices via interconnect 402. PPU 400 can also be connected to local memory 404, which includes multiple memory devices. In one embodiment, local memory may include multiple dynamic random access memory (DRAM) devices. The DRAM devices may be configured as a high-bandwidth memory (HBM) subsystem, wherein multiple DRAM dies are stacked within each device.
[0082] The NVLink 410 interconnect enables the system to expand and include one or more PPUs 400 in conjunction with one or more CPUs, supporting cache coherency between the PPUs 400 and the CPU, as well as the CPU controller. Data and / or commands can be sent from or from the NVLink 410 to other units of the PPU 400 via hub 430, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). FIG. 5B A more detailed description of the NVLink 410.
[0083] The I / O unit 405 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) over the interconnect 402. The I / O unit 405 can communicate directly with the host processor via the interconnect 402, or through one or more intermediate devices such as a memory bridge. In one embodiment, the I / O unit 405 can communicate with one or more other processors, such as one or more PPUs 400, via the interconnect 402. In one embodiment, the I / O unit 405 implements a Peripheral Component Interconnect Express (PCIe) interface for communications over a PCIe bus, and the interconnect 402 is a PCIe bus. In alternate embodiments, the I / O unit 405 can implement other types of known interfaces for communicating with external devices.
[0084] The I / O unit 405 decodes packets of data received via the interconnect 402. In one embodiment, the packets of data represent commands configured to cause the PPU 400 to perform various operations. The I / O unit 405 transmits the decoded commands to various other units of the PPU 400 that the commands can specify. For example, some commands can be transmitted to the front-end unit 415. Other commands can be transmitted to the hub 430 or other units of the PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the I / O unit 405 is configured to route communications between and among various logical units of the PPU 400.
[0085] In one embodiment, a program executed by the host processor encodes a stream of commands in a buffer that provides a workload to the PPU 400 for processing. The workload can include a number of instructions and data to be processed by those instructions. The buffer is a region of memory that is accessible (e.g., read / write) by both the host processor and the PPU 400. For example, the I / O unit 405 can be configured to access the buffer in system memory connected to the interconnect 402 via memory requests transmitted over the interconnect 402. In one embodiment, the host processor writes the stream of commands to the buffer and the PPU 400 transmits a pointer to the beginning of the stream of commands. The front-end unit 415 receives the pointer to the stream or streams of commands. The front-end unit 415 manages the stream or streams, reading commands from the stream or streams and forwarding the commands to various units of the PPU 400.
[0086] The front-end unit 415 is coupled to a scheduler unit 420, which allocates tasks to be performed by the GPCs 450 to various GPCs 450. The scheduler unit 420 can track status information related to various tasks managed by the scheduler unit 420. The status can indicate which GPC 450 a task is assigned to, whether task is active or inactive, what priority the task has, etc. The scheduler unit 420 manages execution of a plurality of tasks on the one or more GPCs 450.
[0087] The scheduler unit 420 is coupled to a work distribution unit 425, which is configured to distribute tasks to be performed by the GPCs 450. The work distribution unit 425 can track a number of scheduled tasks received from the scheduler unit 420. In one embodiment, the work distribution unit 425 manages a pending task pool and an active task pool for each GPC 450. When a GPC 450 completes execution of a task, that task is evicted from the active task pool for the GPC 450, and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 450. If an active task on a GPC 450 has idled, e.g., while waiting for a data dependency to be resolved, then the active task can be evicted from the GPC 450 and returned to the pending task pool while another task is selected from the pending task pool and scheduled for execution on the GPC 450.
[0088] In one embodiment, a host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 400. In one embodiment, multiple compute applications are executed simultaneously by the PPU 400, and the PPU 400 provides isolation, quality of service (QoS), and independent address spaces for the multiple compute applications. An application can generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 400. The driver kernel outputs the tasks to one or more streams that are being processed by the PPU 400. Each task can include one or more related groups of threads, referred to herein as warps. In one embodiment, a warp includes 32 related threads that can be executed in parallel. Cooperative threads can refer to a plurality of threads that execute instructions of a task and can exchange data through shared memory. These tasks can be allocated to one or more processing units within a GPC 450, and instructions are scheduled for execution by at least one thread warp.
[0089] The work distribution unit 425 communicates with the one or more GPCs 450 via the XBar 470. The XBar 470 is an interconnect network coupling many of the units of the PPU 400 to other units of the PPU 400. For example, the XBar 470 can be configured to couple the work distribution unit 425 to a particular GPC 450. Although not explicitly shown, one or more other units of the PPU 400 can also be connected to the XBar 470 via the hub 430.
[0090] Tasks are managed by the scheduler unit 420 and dispatched by the work distribution unit 425 to the GPCs 450. The GPCs 450 are configured to process the tasks and generate results. The results can be consumed by other tasks within the GPCs 450, routed to different GPCs 450 via the XBar 470, or stored in the memory 404. The results can be written to the memory 404 via the memory partition unit 480, which implements a memory interface for reading from and writing to the memory 404. The results can be transmitted to another PPU 400 or CPU via the NVLink 410. In one embodiment, the PPU 400 includes a number U of memory partition units 480 equal to the number of separate and distinct memory devices that are coupled to the PPU 400. Each GPC 450 can include a memory management unit to provide translations of virtual addresses into physical addresses, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit provides one or more translation lookaside buffers (TLBs) for performing translations of virtual addresses into physical addresses in memory 404.
[0091] In one embodiment, the memory partition unit 480 includes a raster operations (ROP) unit, a level 2 (L2) cache, and a memory interface to the memory 404. The memory interface can implement a 32-bit, 64-bit, 128-bit, 1024-bit data bus for high-speed data transfer. The PPU 400 can connect to up to Y memory devices, such as high bandwidth memory stacks or graphics double data rate, version 5, synchronous dynamic random access memory, or other types of persistent storage. In one embodiment, the memory interface implements an HBM2 memory interface and Y is equal to half of U. In one embodiment, the HBM2 memory stacks are located on the same physical package as the PPU 400, providing significant power and area savings compared to a conventional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies and Y is equal to 4, with each HBM2 stack including two 129-bit channels per die for a total of 8 channels and a data bus width of 1024 bits.
[0092] In one embodiment, memory 404 supports single error correction dual error detection (SECDED) error correcting code (ECC) to protect data. ECC provides higher reliability for compute applications that are sensitive to data corruption. Reliability is particularly important in large scale cluster computing environments in which PPU 400 processes very large datasets and / or runs applications over extended periods.
[0093] In one embodiment, PPU 400 implements a multi-level memory hierarchy. In one embodiment, memory partition unit 480 supports a unified memory to provide a single unified virtual address space for CPU and PPU 400 memory, allowing data sharing between virtual memory systems. In one embodiment, the frequency at which PPU 400 accesses memory located on other processors is tracked such that memory pages that are frequently accessed on other PPU 400 are moved to the physical memory of the PPU 400 that is accessing those pages most frequently. In one embodiment, NVLink 410 supports an address translation service that allows PPU 400 to directly access CPU page tables and provide full access to CPU memory by PPU 400.
[0094] In one embodiment, a copy engine transfers data between multiple PPU 400 or between a PPU 400 and a CPU. The copy engine can generate a page fault for an address that is not mapped into a page table. Memory partition unit 480 can then service the page fault, map the address into a page table, and the copy engine can perform the transfer. In a conventional system, memory is pinned (e.g., unpageable) for multiple copy engine operations between multiple processors, greatly reducing the amount of available memory. With hardware page faults, an address can be passed to the copy engine without worrying about whether a memory page is resident, and the copy process is transparent.
[0095] Data from memory 404 or other system memory can be fetched by memory partition unit 480 and stored in L2 cache 460, which is on-chip and shared between various GPCs 450. As shown, each memory partition unit 480 includes a portion of the L2 cache associated with the respective memory 404. Lower level caches can then be implemented within various units within GPC 450. For example, each of the processing units within GPC 450 can implement a level one (LI) cache. The LI cache is private per processing unit and stores data and / or instructions cache-lined to the processing units. L2 cache 460 is coupled to the memory interface 470 and XBar 470, and data from the L2 cache can be fetched and stored in each of the LI caches for processing by the processing units.
[0096] In one embodiment, the processing units within each GPC 450 implement a SIMD (Single Instruction Multiple Data) architecture, wherein each thread in a group of threads (e.g., a warp) is configured to process a different element of a
[0097] A cooperative group is a programming model for organizing groups of threads that allows developers to express the granularity at which threads are communicating, allowing for more expressive and efficient parallel decomposition. The Cooperative Launch API supports synchronization between thread blocks used to execute parallel algorithms. Conventional programming models provide a single simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often wish to define groups of threads smaller than the thread block granularity in the form of collective group- wide function interfaces and synchronize within the defined groups in order to allow for greater performance, design flexibility, and software reuse.
[0098] Cooperative groups enable programmers to explicitly define groups of threads at sub-block and multi-block granularity (as small as a single thread), and perform collective operations such as synchronization on the threads in a cooperative group. This programming model supports clean composition across software boundaries, so that libraries and utility functions can safely synchronize within their local context without having to make assumptions about the aggregation. Cooperative group primitives allow for new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire grid of thread blocks.
[0099] Each processing unit includes a large number (e.g., 128, etc.) of different processing cores (e.g., functional units) that can be fully pipelined, single-precision, double-precision, and / or mixed-precision, and include floating point arithmetic logic units and integer arithmetic logic units. In one embodiment, the floating point arithmetic logic units implement the IEEE 754-2008 standard for floating point arithmetic. In one embodiment, the cores include 64 single-precision (32-bit) floating point cores, 64 integer cores, 32 double-precision (64-bit) floating point cores, and 8 tensor cores.
[0100] The tensor cores are configured to perform matrix operations. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as GEMM (matrix-matrix multiplication), for convolution operations during neural network training and inference. In one embodiment, each tensor core operates on 4x4 matrices and performs matrix multiplication and accumulation operations, D = A x B + C, where A, B, C, and D are 4x4 matrices.
[0101] In one embodiment, the matrix multiplication inputs A and B can be integer, fixed point, or floating point matrices, while the accumulation matrices C and D can be integer, fixed point, or floating point matrices of equal or higher bit-width. In one embodiment, the tensor cores operate on 1-bit, 4-bit, or 8-bit integer input data with 32-bit integer accumulation. An 8-bit integer matrix multiplication requires 1024 operations and results in a full precision product that is then accumulated with other intermediate products using 32-bit integer addition for an 8x8x16 matrix multiplication. In one embodiment, the tensor cores operate on 16-bit floating point input data with 32-bit floating point accumulation. A 16-bit floating point multiplication requires 64 operations and results in a full precision product that is then accumulated with other intermediate products using 32-bit floating point addition for a 4x4x4 matrix multiplication. In practice, the tensor cores are used to perform much larger two-dimensional or higher dimensional matrix operations composed of these smaller elements. APIs such as CUDA 9 C++ API expose specialized matrix load, matrix multiply and accumulate, and matrix store operations to efficiently use the tensor cores for CUDA-C++ programs. At the CUDA level, the interface at the warp level takes a 16x16 size matrix across all 32 threads of a warp.
[0102] Each processing unit can also include M special function units (SFUs) that perform special functions such as attribute evaluations, reciprocal square root, etc. In one embodiment, the SFUs can include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFUs can include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load texture maps (e.g., 2D arrays of texels) from memory 404 and sample these texture maps to produce sampled texture values for use by the shader programs executed by the processing units. In one embodiment, the texture maps are stored in shared memory that can include or contain an LI cache. The texture unit uses mip maps (e.g., texture maps with varying levels of detail) to perform texture operations such as filtering operations. In one embodiment, each processing unit includes two texture units.
[0103] Each processing unit also includes N load store units (LSUs) that implement load and store operations between shared memory and the register file. Each processing unit includes an interconnect network that connects each of the cores to the register file and connects the LSUs to the register file, shared memory. In one embodiment, the interconnect network is a crossbar that can be configured to connect any of the cores to any of the registers in the register file and to connect the LSUs to registers in the register file and to memory locations in shared memory.
[0104] Shared memory is an on-chip memory array that allows data storage and communication between processing units and between threads within a processing unit. In one embodiment, shared memory includes 128 KB of storage capacity and is on a path from each of the processing units to the memory partition unit 480. Shared memory can be used for cache reads and writes. One or more of shared memory, LI cache, L2 cache, and memory 404 are backing stores.
[0105] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory accesses. The capacity can be used as a cache by programs that do not use shared memory. For example, if shared memory is configured to use half of the capacity, then the remaining capacity can be used for texture and load / store operations. Integration into shared memory enables shared memory to be used as a high throughput conduit for streaming data while providing high bandwidth and low latency access for frequently reused data.
[0106] When configured for general purpose parallel computing, a simpler configuration can be used compared to graphics computing. In particular, the fixed function graphics processing units are bypassed, creating a much simpler programming model. In this general purpose parallel computing configuration, the work distribution unit 425 dispatches and allocates thread blocks directly to the processing units within the GPCs 450. The threads execute the same program, using unique thread IDs, in the compute to ensure that each thread uses a processing unit on which to execute the program and perform the computation, shared memory to communicate between threads, and the LSU to read and write global memory through shared memory and the memory partitioning unit 480. When configured for general purpose parallel computing, the processing units can also write back commands that the scheduler unit 420 can use to initiate new work on the processing units.
[0107] Each of the PPU 400 can include one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel visual core (PVC), a ray tracing (RT) core, a visual processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multi-processor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI), or a peripheral component interconnect express (PCIe) element, and / or the like, configured to perform the functions thereof.
[0108] The PPU 400 can be included inside a desktop computer, laptop computer, tablet computer, server computer, supercomputer, smart- phone (e.g., a wireless, hand-held device), personal digital assistant (PDA), digital camera, vehicle, head-mounted display, hand-held electronic device, and / or the like. In one embodiment, the PPU 400 is implemented on a single semiconductor die. In another embodiment, the PPU 400 is included in a system-on-a-chip (SoC) along with, for example, one or more of the other devices, such as an additional PPU 400, a memory 404, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog front end (DAFE), and / or the like.
[0109] In one embodiment, PPU 400 can be included on a graphics card that includes one or more memory devices. The graphics card can be configured to interface with a PCIe slot on a motherboard of a desktop computer. In another embodiment, PPU 400 can be an integrated graphics processing unit (iGPU) or parallel processor included in a chipset of a motherboard. In another embodiment, PPU 400 can be implemented in reconfigurable logic. In yet another embodiment, portions of PPU 400 can be implemented in reconfigurable logic.
[0110] Exemplary computing system
[0111] As developers expose and utilize more parallelism in applications such as artificial intelligence computing, systems with multiple GPUs and CPUs are used in a variety of industries. High performance GPU accelerated systems with tens to thousands of compute nodes are deployed in data centers, research facilities, and supercomputers to solve increasingly larger problems. As the number of processing devices within a high performance system increases, communication and data transfer mechanisms need to scale to support the increased bandwidth.
[0112] FIG. 5A Conceptual diagram of a processing system 500 implemented using PPU 400 in accordance with one embodiment. Exemplary system 500 can be configured to implement method 160 shown in FIG. 16 and / or method 330 shown in FIG. 33. Processing system 500 includes CPU 530, switch 510, and multiple PPUs 400, and respective memories 404. FIG. 4 FIG. 1C NVLinks 410 provide high-speed communication links between each of the PPUs 400. Although a specific number of NVLinks 410 and interconnects 402 connections are illustrated in FIG. 4, the number of connections to each PPU 400 and CPU 530 can vary. Switch 510 forms an interface between interconnects 402 and CPU 530. PPUs 400, memories 404, and NVLinks 410 can be located on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols to form an interface between various different connections and / or links. FIG. 3B
[0113] NVLinks 410 provide high-speed communication links between each of the PPUs 400. Although a specific number of NVLinks 410 and interconnects 402 connections are illustrated in FIG. 4, the number of connections to each PPU 400 and CPU 530 can vary. Switch 510 forms an interface between interconnects 402 and CPU 530. PPUs 400, memories 404, and NVLinks 410 can be located on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols to form an interface between various different connections and / or links. FIG. 5B
[0114] In another embodiment (not shown), NVLinks 410 provide one or more high-speed communication links between each PPU 400 and CPU 530, and switch 510 forms an interface between interconnect 402 and each PPU 400. PPU 400, memory 404, and interconnect 402 can be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 402 provides one or more communication links between each PPU 400 and CPU 530, and switch 510 forms an interface between each PPU 400 using NVLinks 410 to provide one or more high-speed communication links between PPUs 400. In another embodiment (not shown), NVLinks 410 provide one or more high-speed communication links between PPUs 400 and CPU 530 through switch 510. In yet another embodiment (not shown), interconnect 402 provides one or more communication links between each PPU 400 directly. One or more of the NVLink 410 high-speed communication links can be implemented as physical NVLink interconnects or on-chip or on-die interconnects using the same protocol as NVLink 410.
[0115] In the context of this specification, a single semiconductor platform can refer to a singular integrated circuit die with all components integrated on the die, or a plurality of integrated circuit dies interconnected together with a substrate which can be a single printed circuit board or a plurality of printed circuit boards interconnected together. In the context of this specification, a single semiconductor platform can also refer to a single package which can be a single semiconductor die in a package, or a plurality of semiconductor dies packaged together as a single package. In the context of this specification, a single semiconductor platform can also refer to a single semiconductor complex which can include a single integrated circuit die with all components integrated on the die, or a plurality of integrated circuit dies interconnected together with a substrate which can be a single printed circuit board or a plurality of printed circuit boards interconnected together. In the context of this specification, a single semiconductor platform can also refer to a plurality of semiconductor complexes, with each complex having all or some of the components integrated on a single die, and the complexes interconnected together by a substrate which can be a single printed circuit board, or a plurality of printed circuit boards interconnected together. Of course, the various embodiments described herein can be implemented in a single semiconductor platform or a plurality of semiconductor platforms in various combinations.
[0116] In one embodiment, the signaling rate of each NVLink 410 is 20-25 gigabits / second, and each PPU 400 includes six NVLink 410 interfaces (as shown in FIG. 5A FIG. 5B). Each NVLink 410 provides a data transfer rate of 25 gigabytes / second in each direction, for a total of 400 gigabytes / second for six links. NVLinks 410 can be used exclusively for PPU-to-PPU communication, or some combination of PPU-to-PPU and PPU-to-CPU when CPU 530 also includes one or more NVLink 410 interfaces. FIG. 5A In one embodiment, the signaling rate of each NVLink 410 is 20-25 gigabits / second, and each PPU 400 includes six NVLink 410 interfaces (as shown in FIG. 5A FIG. 5B). Each NVLink 410 provides a data transfer rate of 25 gigabytes / second in each direction, for a total of 400 gigabytes / second for six links. NVLinks 410 can be used exclusively for PPU-to-PPU communication, or some combination of PPU-to-PPU and PPU-to-CPU when CPU 530 also includes one or more NVLink 410 interfaces.
[0117] In one embodiment, the NVLink 410 allows direct loads / stores / atomic accesses from the CPU 530 to the memory 404 of each PPU 400. In one embodiment, the NVLink 410 supports coherency operations allowing data read from the memory 404 to be stored in the cache hierarchy of the CPU 530, reducing cache access latency for the CPU 530. In one embodiment, the NVLink 410 includes support for an address translation service (ATS) allowing the PPU 400 to directly access page tables within the CPU 530. One or more of the NVLinks 410 can also be configured to operate in a low power mode.
[0118] FIG. 5B FIGURE 1 illustrates an exemplary system 565 in which various previous embodiments can be implemented. The exemplary system 565 can be configured to implement the method 160 illustrated in FIGURE 1 and / or the method 330 illustrated in FIGURE 3. FIG. 1C FIG. 3B
[0119] As shown, a system 565 is provided that includes at least one central processing unit 530 coupled to a communication bus 575. The communication bus 575 can directly or indirectly couple one or more of the following: a main memory 540, a network interface 535, the CPU 530, a display device 545, an input device 560, a switch 510, and a parallel processing system 525. The communication bus 575 can be implemented using any suitable protocol, and can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The communication bus 575 can include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standards board (VESA) bus, a peripheral component
[0120] Although the components of the system 565 are shown as discrete components, one of skill in the art will recognize that the components of the system 565 can be implemented as one or more sets of instructions executed by one or more processors (e.g., the CPU 530) of the system 565. FIG. 5B different blocks are shown as being connected via a communication bus 575, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presentation component such as display device 545 can be considered an I / O component, e.g., input device 560 if the display is a touch screen. As another example, CPU 530 and / or parallel processing system 525 can include memory (e.g., main memory 540 can represent a storage device in addition to parallel processing system 525, CPU 530, and / or other components). In other words, FIG. 5B The computing device of FIG. 5 is merely illustrative. Distinctions are not made between “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “hand-held device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are contemplated FIG. 5B within the scope of the computing device of FIG. 5.
[0121] System 565 also includes main memory 540. Control logic (software) and data are stored in main memory 540, which can take the form of various computer-readable media. Computer-readable media can be any available media that can be accessed by system 565. By way of example, and not limitation, computer-readable media can comprise computer storage media and communication media.
[0122] Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, main memory 540 can store computer readable instructions such as an operating system (e.g., which represents programs and / or program elements). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and which can be accessed by system 565. When used in the context
[0123] A computer storage medium can include a computer-readable instruction, a data structure, a program module, or other data type in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include a wired medium such as a wired network or direct-wired connection, and a wireless medium such as sound, RF, infrared, and other wireless media. Combinations of the any of the above should also be included within the scope of computer readable media.
[0124] When executed, the computer program enables system 565 to perform various functions. CPU(s) 530 can be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Each of CPU(s) 530 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing multiple software threads concurrently. CPU(s) 530 can include any type of processors, and can include different types of processors depending on the type of system 565 being implemented (e.g., a fewer number of cores for mobile devices, and a greater number of cores for servers). For example, depending on the type of system 565, the processors can be Advanced RISC Machines (ARM) processors implemented using reduced instruction set computing (RISC) or x86 processors implemented using complex instruction set computing (CISC). System 565 can include one or more CPU(s) 530 in addition to one or more microprocessors or supplemental co-processors such as math co-processors.
[0125] In addition to or alternatively from CPU(s) 530, parallel processing module 525 can be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Parallel processing module 525 can be used by system 565 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, parallel processing module 525 can be used for general-purpose computing on GPUs (GPGPU). In embodiments, CPU(s) 530 and / or parallel processing module 525 can perform any combination of the described methods, processes, and / or portions thereof, discretely or jointly.
[0126] System 565 also includes input device 560, parallel processing system 525, and display device 545. Display device 545 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Display device 545 can receive data from other components (e.g., parallel processing system 525, CPU 530, etc.) and output that data (e.g., images, video, sound, etc.).
[0127] Network interface 535 can enable system 565 to be logically coupled to other devices, including input device 560, display device 545, and / or other components, some of which can be embedded in (e.g., integrated with) system 565. Illustrative input devices 560 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. Input devices 560 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some instances, inputs can be transmitted to an appropriate network element for further processing. A NUI can implement voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition both on screen and adjacent to the screen, air gestures, head tracking, and touch recognition (as described in more detail below) associated with a display of system 565. System 565 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. In addition, system 565 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable detection of motion. In some examples, the output of the accelerometers or gyroscopes can be used by system 565 to render immersive augmented reality or virtual reality.
[0128] In addition, system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable network, etc.) for communication purposes through network interface 535. System 565 can be included within a distributed network and / or cloud computing environment.
[0129] The network interface 535 can include one or more receivers, transmitters and / or transceivers that enable the system 565 to communicate with other computing devices via electronic communication networks including wired and / or wireless communication. The network interface 535 can include components and functionality allowing for communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0130] The system 565 can also include secondary storage (not shown). The secondary storage includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a magnetic tape drive, a compact disk drive, digital versatile disk (DVD) drive, recording device, universal serial bus (USB) flash memory, etc. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner. The system 565 can also include a hard-wired or battery-backed-up
[0131] Each of the foregoing modules and / or devices can even be located on a single semiconductor platform. Alternatively, the various different modules can be located in separate devices or components and / or in various combinations thereof. Although the exemplary embodiments have been described above in terms of medical applications, those skilled in the art will recognize that the embodiments of the present disclosure also apply to other applications and that the systems and methods described herein have a wide range of applications.
[0132] Example Network Environment
[0133] A network environment suitable for implementing embodiments of the present disclosure can include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) can be implemented on a processing system 500 and / or an example system 565, such as the example system 565 of FIG. 5, for example. Each device can include similar components, features, and / or functionality of the processing system 500 and / or the example system 565, for example. FIG. 5A FIG. 5B The client devices, servers, and / or other device types (e.g., each device) can be implemented on one or more instances of the processing system 500 and / or the example system 565, such as the example system 565 of FIG. 5, for example. Each device can include similar components, features, and / or functionality of the processing system 500 and / or the example system 565, for example.
[0134] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or a combination thereof. The network can include multiple networks or networks of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet, and / or the public switched telephone network (PSTN), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (among other components) can provide wireless connectivity.
[0135] A compatible network environment can include one or more peer-to-peer network environments, in which case servers can not be included in the network environment, and one or more client-server network environments, in which case one or more servers can be included in the network environment. In a peer-to-peer network environment, functionality described herein with respect to servers can be implemented on any number of client devices.
[0136] In at least one embodiment, a network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports a software layer and / or one or more applications of an application layer. The software or applications can include web-based service software or applications, respectively. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, without limitation, a free and open-source software web application framework type that, for example, can use a distributed file system for large-scale data processing (e.g., “big data”).
[0137] A cloud-based network environment can provide cloud computing and / or cloud storage that implement the computing and / or data storage functionality (or one or more portions thereof) described herein. Any of these different functionalities can be distributed across multiple locations from a central or core server (e.g., a central or core server of one or more data centers, which can be distributed across states, regions, countries, globally, and the like). If a connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0138] A client device can include FIG. 5A at least some of the components, features, and functions of the example processing system 500 and / or FIG. 5B The client device can be implemented as, include, or otherwise host an example system 565, for example and without limitation, a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, a spaceship, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these delineated devices, or any other suitable device.
[0139] Machine learning
[0140] Deep neural networks (DNNs) developed on processors such as the PPU 400 have been used for a wide variety of use cases, from self-driving cars to faster drug development, from automatically captioning images in an online image database to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning process of the human brain, which learns continuously, gets smarter continuously, and delivers more accurate results more quickly over time. A child is initially taught by an adult to correctly identify and classify a variety of different shapes, and eventually is able to identify shapes without any guidance. Similarly, a deep learning or neural learning system needs to be trained in object recognition and classification so that it gets smarter and more efficient at identifying basic objects, occluded objects, and so on, while also imparting context to the objects.
[0141] At the simplest level, neurons in the human brain watch a variety of inputs that are received, a level of importance is imparted to each of these inputs, and an output is delivered to other neurons to react. An artificial neuron or perceptron is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs that represent various features of an object that the perceptron is being trained to recognize and classify, and each of these features is imparted a certain weight based on the importance of that feature in defining the shape of the object.
[0142] Deep neural network (DNN) models include multiple layers of many connected nodes (e.g., perceptron, Boltzmann machine, radial basis function, convolutional layer, etc.) that can be trained with vast amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into different segments and looks for basic patterns such as lines and corners. The second layer assembles these lines to look for higher level patterns such as wheels, windshield, and mirrors. The next layer identifies the type of vehicle, and the final few layers generate a label for the input image that identifies the make and model of the particular car.
[0143] Once a DNN is trained, it can be deployed and used to identify and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include identifying handwritten numbers on checks deposited into an ATM machine, identifying images of friends in a photograph, delivering movie recommendations to over 50 million users, identifying and classifying different types of cars, pedestrians, and road hazards in a self-driving car, or translating human languages in real time.
[0144] During training, data flows through the DNN in a forward propagation phase until a prediction is made that indicates a label corresponding to the input. If the neural network did not correctly label the input, the error between the correct label and the predicted label is analyzed and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels the input and other inputs in the training data set. Training complex neural networks requires a large amount of parallel computing performance, including floating point multiplication and addition supported by PPU 400. Inference, which is less computationally intensive than training, is a latency sensitive process in which a trained neural network is applied to new inputs that it has not seen before to classify images, detect emotions, identify recommendations, recognize and translate languages, and generally infer new information.
[0145] Neural networks rely heavily on matrix math operations, and for both efficiency and speed, complex multi-layer networks require vast amounts of floating point performance and bandwidth. With thousands of processing cores optimized for matrix math operations and providing tens to hundreds of TFLOPS of performance, PPU 400 is a computing platform capable of providing the performance needed for deep neural network based artificial intelligence and machine learning applications.
[0146] Further, images generated applying one or more of the techniques disclosed herein can be used to train, test, or certify DNNs for recognizing objects and environments in the real world. Such images can include scenes of roadways, factories, buildings, urban environments, rural environments, humans, animals, and any other physical objects or real-world environments. Such images can be used to train, test, or certify DNNs employed in machines or robots for manipulating, handling, or modifying physical objects in the real world. Further, such images can be used to train, test, or certify DNNs employed in autonomous vehicles for navigating and moving the vehicle in the real world. Further, images generated applying one or more of the techniques disclosed herein can be used to convey information to users of such machines, robots, and vehicles.
[0147] FIG. 5C Components of an example system 555 that can be used to train and utilize machine learning are shown in accordance with at least one embodiment. As will be discussed, various components can be provided by a single computing system or various combinations of computing devices and resources that can be under the control of a single entity or multiple entities. Further, various aspects can be triggered, initiated, or requested by different entities. In at least one embodiment, training of a neural network can be directed by a vendor associated with a vendor environment 506, while in at least one embodiment, training can be requested by a customer or other user that has access to the vendor environment through a client device 502 or other such resource. In at least one embodiment, training data (or data to be analyzed by a trained neural network) can be provided by a vendor, user, or third-party content provider 524. In at least one embodiment, a client device 502 can be, for example, a vehicle or object to be navigated on behalf of a user that can submit a request and / or receive instructions that facilitate navigation of the device.
[0148] In at least one embodiment, a request can be submitted through at least one network 504 for receipt by a vendor environment 506. In at least one embodiment, a client device can be any suitable electronic and / or computing device that enables a user to generate and send such a request, such as but not limited to a desktop computer, a notebook computer, a computer server, a smart phone, a tablet computer, a game console (portable or otherwise), a computer processor, computing logic, and a set-top box. One or more networks 504 can include any suitable network or networks for transmitting requests or other such data, for example, can include the Internet, an intranet, an Ethernet network, a cellular network, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), an ad hoc network of direct wireless connections between peers, and the like.
[0149] In at least one embodiment, a request can be received at interface layer 508, which in this example can forward data to training and inference manager 532. Training and inference manager 532 can be a system or service that includes hardware and software for managing services and requests corresponding to data or content. In at least one embodiment, training and inference manager 532 can receive a request to train a neural network, and can provide data for the request to training module 512. In at least one embodiment, training module 512 can select an appropriate model or neural network to use, if not specified by the request, and can train the model using relevant training data. In at least one embodiment, training data can be a batch of data stored in a training data repository, received from client device 502, or obtained from third party vendor 524. In at least one embodiment, training module 512 can be responsible for training data. A neural network can be any appropriate network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Once a neural network is trained and successfully evaluated, a trained neural network can be stored to, for example, model repository 516, which can store different models or networks for users, applications, or services, etc. In at least one embodiment, there can be multiple models for a single application or entity, which can be utilized based on a number of different factors.
[0150] In at least one embodiment, at a subsequent point in time, a request for content (e.g., path determination) or data determined or influenced at least in part by a trained neural network can be received from client device 502 (or another such device). This request can include, for example, input data that is to be processed using a neural network to obtain one or more inference or other output values, classifications, or predictions, or input data can be received by interface layer 508 and directed to inference module 518, although different systems or services can also be used. In at least one embodiment, if not already stored locally to inference module 518, inference module 518 can obtain an appropriate trained network, such as a trained deep neural network (DNN) as discussed herein, from model store 516. Inference module 518 can provide data as input to the trained network, which can then generate one or more inferences as output. For example, this can include a classification of an input data instance. In at least one embodiment, the inferences can then be transmitted to client device 502 for display to or other communication with a user. In at least one embodiment, a user’s contextual data can also be stored to user contextual data store 522, which can include data about a user that can be used as network input to generate inferences or determine data returned to a user after an instance is obtained. In at least one embodiment, related data that can include at least some of input or inference data can also be stored to local database 534 for processing of future requests. In at least one embodiment, a user can use account information or other information to access resources or functionality of a vendor environment. In at least one embodiment, if allowed and available, user data can also be collected and used to further train models in order to provide more accurate inferences for future requests. In at least one embodiment, requests to machine learning application 526 executing on client device 502 can be received through a user interface and results displayed through the same interface. A client device can include resources such as a processor 528 and memory 562 for generating requests and processing results or responses, as well as at least one data storage element 552 for storing data for machine learning application 526.
[0151] In at least one embodiment, processor 528 (or processor of training module 512 or inference module 518) will be a central processing unit (CPU). However, as noted above, resources in such environments can utilize GPUs to process data for at least certain types of requests. GPUs, such as PPU 300, have thousands of cores designed to handle large parallel workloads, and thus have become popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building allows for training of larger, more complex models faster, offline generation of predictions means that request-time input features cannot be used, or must be prearranged for all features to generate predictions and store them in a lookup table to service real-time requests. If a deep learning framework supports CPU mode, and the model is small and simple enough to perform feedforward on a CPU with reasonable latency, a service on a CPU instance can host the model. In this case, training can be done offline on a GPU, and inference in real-time on a CPU. If a CPU approach is not feasible, a service can run on a GPU instance. However, because GPUs have different performance and cost characteristics than CPUs, running a service that offloads runtime algorithms to a GPU can require designing it differently than a CPU-based service.
[0152] In at least one embodiment, video data can be provided from client device 502 for augmentation in vendor environment 506. In at least one embodiment, video data can be processed for augmentation on client device 502. In at least one embodiment, video data can be streamed from third party content vendor 524 and augmented by third party content vendor 524, vendor environment 506, or client device 502. In at least one embodiment, video data can be provided from client device 502 for use as training data in vendor environment 506.
[0153] In at least one embodiment, supervised and / or unsupervised training can be performed by client device 502 and / or vendor environment 506. In at least one embodiment, a set of training data 514 (e.g., classified or labeled data) is provided as input to be used as training data. In at least one embodiment, training data can include instances of at least one type of object for which a neural network is to be trained, along with information identifying that object type. In at least one embodiment, training data can include a set of images, each image including a representation of one type of object, where each image also includes or is associated with labeling, metadata, classification, or other information identifying the type of object represented in the corresponding image. Various other types of data can also be used as training data, which can include textual data, audio data, video data, and so on. In at least one embodiment, training data 514 is provided as training input to training module 512. In at least one embodiment, training module 512 can be a system or service including hardware and software, such as one or more computing devices executing a training application, for training a neural network (or other model or algorithm, and so on). In at least one embodiment, training module 512 receives instructions or requests indicating a type of model to be used for training, which in at least one embodiment can be any appropriate statistical model, network, or algorithm useful for such purposes, which can include artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, and so on. In at least one embodiment, training module 512 can select an initial model or other untrained model from an appropriate repository, and train the model with training data 514, generating a trained model (e.g., a trained deep neural network) that can be used to classify or generate other such inferences on similar types of data. In at least one embodiment in which no training data is used, an initial model can still be selected for training on input data to each training module 512.
[0154] In at least one embodiment, a model can be trained in several different ways, which can depend in part on the type of model selected. In at least one embodiment, a machine learning algorithm can be provided with a set of training data, where the model is a model artifact created by a training process. In at least one embodiment, each instance of training data contains a correct answer (e.g., a classification) that can be referred to as a target or target attribute. In at least one embodiment, a learning algorithm finds patterns in the training data that map input data attributes to the target - the answer to be predicted - and the machine learning model is an output that captures these patterns. In at least one embodiment, a machine learning model can then be used to obtain predictions on new data for which a target is not specified.
[0155] In at least one embodiment, training and inference manager 532 can select from a set of machine learning models including binary classification, multi-class classification, generative, and regression models. In at least one embodiment, a model type to be used can depend at least in part on a target type to be predicted.
[0156] Graphics processing pipeline
[0157] In one embodiment, PPU 400 includes a graphics processing unit (GPU). PPU 400 is configured to receive commands that specify processing of shader programs. The graphics data can be defined by a set of primitives that form one or more graphics objects. A primitive includes, for example, data specifying vertices (e.g., in a model-space coordinate system) and attributes associated with each vertex of the primitive. PPU 400 can be configured to process the primitives to generate a frame buffer (e.g., pixel data for each of pixels in a display).
[0158] An application writes model data (e.g., attributes and a set of vertices) for a scene to a memory such as system memory or memory 404. The model data defines each of the objects that can be visible on a display. The application then makes an API call to a driver kernel that requests that the model data be rendered and displayed. The driver kernel reads the model data and writes commands to the one or more streams to perform operations that process the model data. These commands can reference different shader programs to be implemented on processing units within PPU 400, including one or more of a vertex shader, a hull shader, a domain shader, a geometry shader, and a pixel shader. For example, one or more of the processing units can be configured to execute a vertex shader program that processes a number of vertices defined by the model data. In one embodiment, these different processing units can be configured to execute different shader programs concurrently. For example, a first subset of processing units can be configured to execute a vertex shader program while a second subset of processing units can be configured to execute a pixel shader program. The first subset of processing units processes the vertex data to produce processed vertex data and writes the processed vertex data to L2 cache 460 and / or memory 404. After the processed vertex data is rasterized (e.g., transformed from three-dimensional data to two-dimensional data in screen space) to produce fragment data, the second subset of processing units executes a pixel shader to produce processed fragment data that is then blended with other processed fragment data and written to a frame buffer in memory 404. The vertex shader program and the pixel shader program can be executed concurrently, processing different data from the same scene in a pipelined fashion until all model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then transmitted to a display controller for display on a display device.
[0159] FIG. 6A is implemented by a PPU 400 according to one embodiment FIG. 4 A conceptual diagram of a graphics processing pipeline 600 implemented by the PPU 400 of As is known, pipeline architectures can more efficiently perform long-latency operations by breaking the operations into multiple stages, where the output of each stage is coupled to the input of the next successive stage. Thus, the graphics processing pipeline 600 receives input data 601 that is passed from one stage of the graphics processing pipeline 600 to the next to generate output data 602. In one embodiment, the graphics processing pipeline 600 can represent a graphics processing pipeline defined by an API. As an option, the graphics processing pipeline 600 can be implemented in the context of the functionality and architecture of the previous figures and / or one or more any subsequent figures.
[0160] As shown in FIG. 6A The graphics processing pipeline 600 includes a pipeline architecture that includes multiple stages. These stages include, but are not limited to, a data assembly stage 610, a vertex shading stage 620, a primitive assembly stage 630, a geometry shading stage 640, a viewport scale, cull, and clip (VSCC) stage 650, a rasterization stage 660, a fragment shading stage 670, and a raster operations stage 680. In one embodiment, the input data 601 includes commands that configure the processing units to implement the stages of the graphics processing pipeline 600 and configure geometric primitives (e.g., points, lines, triangles, quads, triangle strips, or fans, etc.) for processing by the stages. The output data 602 can include pixel data (e.g., color data) that is copied to a frame buffer in memory or other type of surface data structure.
[0161] The data assembly stage 610 receives input data 601 that specifies vertex data for high-order surfaces, primitives, etc. The data assembly stage 610 collects vertex data in a temporary storage or queue, for example, by receiving a command from a host processor that includes a pointer to a buffer in memory and reading the vertex data from the buffer. The vertex data is then passed to the vertex shading stage 620 for processing.
[0162] The vertex shading stage 620 processes vertex data by performing a set of operations (e.g., a vertex shader or program) on each of the vertices. A vertex can be specified, for example, as a 4-coordinate vector (e.g., <x, y, z, w>) associated with one or more vertex attributes (e.g., color, texture coordinates, surface normal, etc.). The vertex shading stage 620 can manipulate individual vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading stage 620 performs operations on vertex coordinates or other vertex attributes associated with a vertex. Such operations typically include lighting operations (e.g., modifying a color attribute of a vertex) and transformation operations (e.g., modifying a coordinate space of a vertex). For example, a vertex can be specified using coordinates in an object coordinate space, which are transformed by multiplying the coordinates by a matrix that converts the coordinates from the object coordinate space to a world space or a normalized-device-coordinate (NDC) space. The vertex shading stage 620 generates transformed vertex data that is passed to the primitive assembly stage 630.
[0163] The primitive assembly stage 630 collects vertices output by the vertex shading stage 620 and groups the vertices into geometric primitives for processing by the geometry shading stage 640. For example, the primitive assembly stage 630 can be configured to group every three consecutive vertices into a geometric primitive (e.g., a triangle) for passing to the geometry shading stage 640. In some embodiments, particular vertices can be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip can share two vertices). The primitive assembly stage 630 passes the geometric primitives (e.g., a set of associated vertices) to the geometry shading stage 640.
[0164] The geometry shading stage 640 processes the geometric primitives by performing a set of operations (e.g., a geometry shader or program) on the geometric primitives. Tessellation operations can generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading stage 640 can tessellate each geometric primitive into a finer mesh of two or more geometric primitives for processing by the remainder of the graphics processing pipeline 600. The geometry shading stage 640 passes the geometric primitives to the viewport SCC stage 650.
[0165] In one embodiment, graphics processing pipeline 600 can operate within a streaming multiprocessor, and vertex shading stage 620, primitive assembly stage 630, geometry shading stage 640, fragment shading stage 670, and / or hardware / software associated therewith can sequentially perform processing operations. In one embodiment, once the sequential processing operations are complete, viewport SCC stage 650 can utilize the data. In one embodiment, primitive data processed by one or more of the stages in graphics processing pipeline 600 can be written into a cache (e.g., an LI cache, a vertex cache, etc.). In such a case, in one embodiment, viewport SCC stage 650 can access the data in the cache. In one embodiment, viewport SCC stage 650 and rasterization stage 660 are implemented as fixed function circuitry.
[0166] Viewport SCC stage 650 performs viewport scaling, culling, and clipping of the geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of a viewer watching the scene and defines a viewing frustum that encloses the objects of the scene. The viewing frustum can include a viewing plane, a back plane, and four clipping planes. Any geometric primitive that is completely outside the viewing frustum can be culled (e.g., discarded) because it will not contribute to the final rendered scene. Any geometric primitive that is partially inside the viewing frustum and partially outside the viewing frustum can be clipped (e.g., transformed into a new geometric primitive that is enclosed within the viewing frustum). In addition, each geometric primitive can be scaled based on the depth of the viewing frustum. Then, all potentially visible geometric primitives are passed to rasterization stage 660.
[0167] Rasterization stage 660 converts 3D geometric primitives into 2D fragments (e.g., that can be used for display, etc.). Rasterization stage 660 can be configured to set up a set of plane equations with the vertices of the geometric primitive from which various attributes can be interpolated. Rasterization stage 660 can also compute a coverage mask for a plurality of pixels that indicates whether one or more sample locations of the pixel intercept the geometric primitive. In one embodiment, a z-test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have already been rasterized. Rasterization stage 660 generates fragment data (e.g., interpolated vertex attributes associated with particular sample locations of each covered pixel) that is passed to fragment shading stage 670.
[0168] The fragment shading stage 670 processes the fragment data by performing a set of operations (e.g., a fragment shader or program) on each of the fragments. The fragment shading stage 670 can generate pixel data (e.g., color values) for a fragment, such as by performing lighting operations or sampling texture maps using the fragment's interpolated texture coordinates. The fragment shading stage 670 generates pixel data, which is passed to the raster operations stage 680.
[0169] The raster operations stage 680 can perform various operations on the pixel data, such as performing alpha tests, stencil tests, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the raster operations stage 680 has completed processing of the pixel data (e.g., output data 602), the pixel data can be written to a render target, such as a frame buffer, color buffer, etc.
[0170] It should be appreciated that one or more additional stages can be included in the graphics processing pipeline 600 in addition to or instead of one or more of the stages described above. Various implementations of an abstract graphics processing pipeline can implement different stages. Further, in some embodiments, one or more of the stages described above can be excluded from the graphics processing pipeline (such as the geometry shading stage 640). Other types of graphics processing pipelines are contemplated to be within the scope of the present disclosure. Further, any of the stages of the graphics processing pipeline 600 can be implemented by one or more specialized hardware units within a graphics processor, such as the PPU 400. Other stages of the graphics processing pipeline 600 can be implemented by programmable hardware units, such as processing units within the PPU 400.
[0171] The graphics processing pipeline 600 can be implemented via an application program executed by a host processor such as a CPU. In one embodiment, a device driver can implement an application programming interface (API) that defines various functions that can be utilized by an application program to generate graphics data for display. The device driver is a software program that includes a plurality of instructions that control the operation of the PPU 400. The API provides an abstraction for programmers that allows programmers to generate graphics data utilizing specialized graphics hardware such as the PPU 400 without requiring the programmer to utilize the specific instruction set of the PPU 400. An application program can include API calls that are routed to the device driver of the PPU 400. The device driver interprets the API calls and performs various operations in response to the API calls. In some cases, the device driver can perform operations by executing instructions on the CPU. In other cases, the device driver can perform operations at least in part by initiating operations on the PPU 400 utilizing an input / output interface between the CPU and the PPU 400. In one embodiment, the device driver is configured to utilize the hardware of the PPU 400 to implement the graphics processing pipeline 600.
[0172] Various programs can be executed within the PPU 400 in order to implement the various stages of the graphics processing pipeline 600. For example, a device driver can initiate a kernel on the PPU 400 to execute the vertex shading stage 620 on one processing unit (or multiple processing units). The device driver (or an initial kernel executed by the PPU 400) can also initiate other kernels on the PPU 400 to execute other stages of the graphics processing pipeline 600, such as the geometry shading stage 640 and the fragment shading stage 670. In addition, some of the stages of the graphics processing pipeline 600 can be implemented on fixed- function hardware such as a rasterizer or a data assembler implemented within the PPU 400. It will be appreciated that results from one kernel can be processed by one or more intermediate fixed-function hardware units before being processed by a subsequent kernel on a processing unit.
[0173] Images generated using one or more of the techniques disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device may be directly coupled to the system or processor that generates or renders the image. In other embodiments, the display device may be indirectly coupled to the system or processor, for example, via a network. Examples of such networks include the Internet, mobile telecommunications networks, Wi-Fi networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, images generated by the system or processor can be streamed to the display device over the network. Such streaming allows, for example, video games or other applications that render images to execute on servers, data centers, or cloud-based computing environments, and the rendered images are transmitted and displayed on one or more user devices (e.g., computers, video game consoles, smartphones, other mobile devices, etc.) physically separate from the server or data center. Therefore, the techniques disclosed herein can be applied to enhance streamed images and services that stream images, such as NVIDIA GeForce Now (GFN), Google Stadia, etc.
[0174] Example Streaming System
[0175] FIG. 6B This is a schematic diagram of an example system 605 of a streaming system according to some embodiments of the present disclosure.
[0176] FIG. 6B Includes server 603 (which may include with FIG. 5A Example processing system 500 and / or FIG. 5B (Similar components, features and / or functions to exemplary system 565), client 604 (which may include similar ... FIG. 5A Example processing system 500 and / or FIG. 5B FIG. 5B The exemplary system 565 has similar components, features, and / or functions to the network 606 (which may be similar to the network described herein). In some embodiments of this disclosure, system 605 may be implemented.
[0177] In one embodiment, the streaming system 605 is a game streaming system, and the server 603 is a game server. In the system 605, for a game session, the client device 604 can receive input data in response to input of the input devices 626, send the input data to the server 603, receive encoded display data from the server 603, and display the display data on the display 624. In this way, computationally intensive computations and processing are offloaded to the server 603 (e.g., rendering of graphical output of the game session, especially ray or path tracing, performed by the GPU 615 of the server 603). In other words, the game session is streamed from the server 603 to the client device 604, reducing the requirements of the client device 604 for graphics processing and rendering.
[0178] For example, with respect to instantiation of a game session, the client device 604 can be displaying a frame of the game session on the display 624 based on receiving display data from the server 603. The client device 604 can receive input of one of the input devices 626, and in response generate input data. The client device 604 can send the input data to the server 603 via the communication interface 621 and over the network 606 (e.g., the Internet), and the server 603 can receive the input data via the communication interface 618. The CPU 608 can receive the input data, process the input data, and send data to the GPU 615 that causes the GPU 615 to generate a rendering of the game session. For example, the input data can represent movement of a user character in the game, firing a weapon, reloading, passing a ball, turning a vehicle, etc. The rendering component 612 can render the game session (e.g., representing a result of the input data), and the rendering capture component 614 can capture the rendering of the game session as display data (e.g., as image data of a frame capturing the rendering of the game session). The rendering of the game session can include lighting and / or shadow effects of ray or path tracing computed using one or more parallel processing units of the server 603 (e.g., a GPU, which can further employ use of one or more specialized hardware accelerators or processing cores to perform ray or path tracing techniques). The encoder 616 can then encode the display data to generate encoded display data, and the encoded display data can be sent to the client device 604 via the communication interface 618 over the network 606. The client device 604 can receive the encoded display data via the communication interface 621, and the decoder 622 can decode the encoded display data to generate display data. The client device 604 can then display the display data via the display 624.
[0179] Embodiments of the invention can be in view of the following clauses:
[0180] 1. A computer-implemented method, comprising:
[0181] processing, by a neural network system, a two-dimensional (2D) tomographic image of an object according to parameters to produce a three-dimensional (3D) density volume of the object, wherein the 2D tomographic image is generated by a physical capture environment;
[0182] projecting the 3D density volume based on characteristics of the physical capture environment to produce a simulated tomographic image corresponding to the 2D tomographic image; and
[0183] adjusting the parameters of the neural network system to reduce a difference between the simulated tomographic image and the 2D tomographic image.
[0184] 2. The computer-implemented method of clause 1, wherein the 3D density volume is fully reconstructed.
[0185] 3. The computer-implemented method of clause 1, wherein the 3D density volume comprises at least two layers of 3D voxels.
[0186] 4. The computer-implemented method of clause 1, further comprising producing a 2D density image corresponding to a slice through the 3D density volume.
[0187] 5. The computer-implemented method of clause 1, wherein noise present in the 2D tomographic image is reduced in the simulated tomographic image.
[0188] 6. The method of clause 1, further comprising:
[0189] processing, by the neural network system, an additional 2D tomographic image of an additional object according to the parameters to produce an additional 3D density volume of the additional object;
[0190] projecting the additional 3D density volume to produce an additional simulated tomographic image corresponding to the additional 2D tomographic image; and
[0191] adjusting the parameters of the neural network system to reduce a difference between the additional simulated tomographic image and the additional 2D tomographic image.
[0192] 7. The computer-implemented method of clause 1, wherein the 3D density volume corresponds to a portion of a human body.
[0193] 8. The computer-implemented method of clause 1, wherein the physical capture environment comprises a cone-beam computed tomography machine.
[0194] 9. The computer-implemented method of clause 1, wherein the neural network system produces the 3D density volume by:
[0195] computing 3D data by back-projecting the 2D tomographic images according to characteristics of the physical capture environment; and
[0196] processing the 3D data by a neural network model to produce the 3D density volume.
[0197] 10. The computer-implemented method of clause 9, wherein the back-projecting comprises computing a footprint of a projection of a pixel, and accessing one or more pre-filtered versions of the 2D tomographic images according to at least one dimension of the footprint of the projection.
[0198] 11. The computer-implemented method of clause 1, wherein the neural network system produces the 3D density volume by:
[0199] processing the 2D tomographic images by a first neural network model to produce at least one channel of 2D features;
[0200] computing 3D features by back-projecting the at least one channel of 2D features according to the characteristics; and
[0201] processing the 3D features by a second neural network to produce the 3D density volume corresponding to the 2D tomographic images.
[0202] 12. The computer-implemented method of clause 11, wherein the back-projecting comprises computing a footprint of a projection of a pixel, and accessing one or more pre-filtered versions of the at least one channel of 2D features according to at least one dimension of the footprint of the projection.
[0203] 13. The computer-implemented method of clause 1, wherein at least one of the processing, projecting, and conditioning steps is performed on a server or within a data center prior to the neural network system being streamed to a user device.
[0204] 14. The computer-implemented method of clause 1, wherein at least one of the processing, projecting, and conditioning steps is performed within a cloud computing environment.
[0205] 15. The computer-implemented method of clause 1, wherein at least one of the processing, projecting, and conditioning steps is performed to train, test, or certify a neural network employed in a machine, robot, or autonomous vehicle.
[0206] 16. The computer-implemented method of clause 1, wherein at least one of the processing, projecting, and conditioning steps is performed on a virtual machine comprising a portion of a graphics processing unit.
[0207] 17. A system comprising:
[0208] a memory storing two-dimensional (2D) tomographic images of an object, wherein the 2D tomographic images are generated by a physical capture environment; and
[0209] a processor connected to the memory, wherein the processor is configured to train a neural network system by:
[0210] execute the neural network system to process the 2D tomographic images according to parameters to produce a three-dimensional (3D) density volume of the object;
[0211] project the 3D density volume based on characteristics of the physical capture environment to produce a simulated tomographic image corresponding to the 2D tomographic images; and
[0212] adjust parameters of the neural network system to reduce a difference between the simulated tomographic image and the 2D tomographic images.
[0213] 18. The system of clause 17, wherein noise present in the 2D tomographic images is reduced in the simulated tomographic image.
[0214] 19. The system of clause 17, wherein the 3D density volume corresponds to a portion of a human body.
[0215] 20. The system of clause 17, wherein the physical capture environment comprises a cone-beam computed tomography machine.
[0216] 21. A non-transitory computer-readable medium storing computer instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:
[0217] processing two-dimensional (2D) tomographic images of an object according to parameters by a neural network system to produce a three-dimensional (3D) density volume of the object, wherein the 2D tomographic images are generated by a physical capture environment;
[0218] projecting the 3D density volume based on characteristics of the physical capture environment to produce a simulated tomographic image corresponding to the 2D tomographic images; and
[0219] adjusting parameters of the neural network system to reduce a difference between the simulated tomographic image and the 2D tomographic images.
[0220] 22. The non-transitory computer-readable medium of clause 21, wherein noise present in the 2D tomographic images is reduced in the simulated tomographic image.
[0221] It should be noted that the techniques described herein can embody themselves in executable instructions stored in a computer readable medium for use by or in connection with a processor-based instruction execution machine, system, apparatus, or device. Those skilled in the art will recognize that a variety of computer readable media can be used to store data for some embodiments. As used herein, "computer readable medium" includes one or more of any suitable media for storing the executable instructions for a computer program so that an instruction execution machine, system, apparatus, or device can read (or fetch) the instructions from the computer readable medium and execute the instructions to implement the described embodiments. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer readable media includes: portable computer disks; random access memories (RAM); read only memories (ROM); erasable programmable read only memories (EPROM); flash memory devices; and optical storage devices, including portable compact discs (CD), portable digital video discs (DVD), and so forth.
[0222] It should be understood that the arrangement of components shown in the figures is for illustrative purposes, and other arrangements are possible. For example, one or more of the elements described herein can be implemented in whole or in part as electronic hardware components. Other elements can be implemented in software, hardware, or a combination of software and hardware. Also, some or all of these other elements can be combined, some can be omitted entirely, and additional components can be added, while still implementing the functionality described herein. Accordingly, the subject matter described herein can be implemented in many different variations and all such variations are contemplated to be within the scope of the claims.
[0223] To facilitate understanding of the subject matter described herein, many aspects are described in the context of action sequences. Those skilled in the art will recognize that the various actions can be performed by specialized circuits or circuitry, by program instructions executed by one or more processors, or by a combination of both. The description of any sequence of actions herein does not necessarily imply that the particular order described for that sequence is the only order in which the sequence can be performed. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context.
[0224] The use of the terms "one," "a," "the" and analogous expressions are used in the context of describing the subject matter herein (particularly in the context of the following claims) and are to be interpreted to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by the context. The use of the term "at least one" followed by a list of one or more items (for example, "at least one of A and B") is to be interpreted as meaning one item from the list A or B or any combination of two or more of the items in the list A and B, unless otherwise indicated herein or clearly contradicted by the context. Furthermore, the foregoing description is for the purpose of illustration only and not for the purpose of limitation, as the scope of the present application is defined by the claims as set hereinafter, together with their equivalents. The use of any and all examples, or exemplary language (e.g., "such as") provided herein, is intended merely to better illuminate the subject matter and does not pose a limitation on the scope of the subject matter unless otherwise claimed. The use of the "based on," along with other similar phrases (e.g., "based on the") indicating a condition for bringing about a result in both the claims and the written description is not intended to foreclose any other condition for bringing about that result. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
Claims
1. A computer-implemented method, comprising: The tomographic images are processed by a first neural network to generate at least one channel of two-dimensional features for each tomographic image; Three-dimensional features are calculated by back-projecting at least one channel of the two-dimensional features of the tomographic image based on the characteristics of the physical environment used to capture the tomographic image. The three-dimensional features are processed by a second neural network to generate a three-dimensional density volume corresponding to the tomographic image; The three-dimensional density volume data is projected according to the capture environment to generate a projected two-dimensional image corresponding to the tomographic image; and The parameters of at least one of the first neural network and the second neural network are adjusted to reduce the difference between the projected two-dimensional image and the tomographic image.
2. The computer-implemented method of claim 1, wherein noise present in the tomographic image is reduced in the three-dimensional density volume.
3. The computer-implemented method of claim 1, wherein the three-dimensional features are voxels and associated attributes.
4. The computer-implemented method of claim 1, wherein the three-dimensional density body corresponds to a portion of the human body.
5. The computer-implemented method of claim 1, wherein the physical environment for capturing the tomographic image comprises a conical spiral computerized tomography machine.
6. The computer-implemented method of claim 1, wherein the back projection comprises: Calculate the occupancy of the pixel projection, and access one or more pre-filtered versions of the tomographic image based on at least one dimension of the occupancy of the projection.
7. The computer-implemented method of claim 1, wherein at least one of the steps of processing the tomographic image, calculating and processing the three-dimensional features is performed on a server or in a data center, and the three-dimensional density volume is streamed to a user device.
8. The computer-implemented method of claim 1, wherein at least one of the steps of processing the tomographic image, calculating the three-dimensional features, and processing the three-dimensional features is performed within a cloud computing environment.
9. The computer-implemented method of claim 1, wherein at least one of the steps of processing the tomographic image, calculating the three-dimensional features, and processing the three-dimensional features is performed to train, test, or certify a neural network used in a machine, robot, or autonomous vehicle.
10. The computer-implemented method of claim 1, wherein at least one of the steps of processing the tomographic image, calculating the three-dimensional features, and processing the three-dimensional features is performed on a virtual machine including a portion of a graphics processing unit.
11. A system comprising: A memory that stores tomographic images; A processor connected to the memory, wherein the processor is configured to: Execute a first neural network to generate at least one channel of two-dimensional features for each tomographic image; Three-dimensional features are calculated by back-projecting at least one channel of the two-dimensional features of the tomographic image based on the characteristics of the physical environment used to capture the tomographic image. A second neural network is executed to process the three-dimensional features and generate a three-dimensional density volume corresponding to the tomographic image; The three-dimensional density volume data is projected according to the capture environment to generate a projected two-dimensional image corresponding to the tomographic image; and The parameters of at least one of the first neural network and the second neural network are adjusted to reduce the difference between the projected two-dimensional image and the tomographic image.
12. The system of claim 11, wherein noise present in the tomographic image is reduced in the three-dimensional density volume.
13. The system of claim 11, wherein the three-dimensional features are voxels and associated attributes.
14. The system of claim 11, wherein the three-dimensional density volume corresponds to a portion of the human body.
15. The system of claim 11, wherein the physical environment for capturing the tomographic image comprises a conical spiral computerized tomography machine.
16. The system of claim 11, wherein the reverse projection comprises: Calculate the occupancy of the pixel projection, and access one or more pre-filtered versions of the tomographic image based on at least one dimension of the occupancy of the projection.
17. A non-transitory computer-readable medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform the following steps: The tomographic images are processed by a first neural network to generate at least one channel of two-dimensional features for each tomographic image; Three-dimensional features are calculated by back-projecting at least one channel of the two-dimensional features of the tomographic image based on the characteristics of the physical environment used to capture the tomographic image. The three-dimensional features are processed by a second neural network to generate a three-dimensional density volume corresponding to the tomographic image; The three-dimensional density volume data is projected according to the capture environment to generate a projected two-dimensional image corresponding to the tomographic image; and The parameters of at least one of the first neural network and the second neural network are adjusted to reduce the difference between the projected two-dimensional image and the tomographic image.
18. The non-transitory computer-readable medium of claim 17, wherein the back projection comprises: Calculate the occupancy of the pixel projection, and access one or more pre-filtered versions of the tomographic image based on at least one dimension of the occupancy of the projection.
Citation Information
Patent Citations
Three-dimensional multi-energy-spectrum CT reconstruction method and device based on neural network and storage medium
CN110544282A
System and method for training pseudo image data enhancement of machine learning model
CN115443481A