Arbitrary resolution diffusion super resolution on ray traced images
The hybrid approach of combining ray tracing with denoise diffusion super-resolution addresses the computational intensity and noise issues in ray tracing, resulting in high-quality, photo-realistic images with reduced processing time and improved cache efficiency.
Patent Information
- Application Number
- PCT/US2023/082852
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-12
AI Technical Summary
Existing ray tracing techniques are computationally intensive, have low cache hit rates, and often produce noisy imagery, which can be unsuitable for high-quality applications like movie-making and CGI.
A hybrid approach that combines ray tracing with denoise diffusion super-resolution, where a low-resolution image generated by ray tracing is processed using a denoise diffusion SR model to produce a high-resolution image with reduced noise and computational complexity.
This approach significantly reduces computational complexity and noise in imagery, enabling the generation of high-quality, photo-realistic images with improved cache hitting rates and faster processing times compared to conventional ray tracing methods.
Smart Images

Figure US2023082852_12062025_PF_FP_ABST
Abstract
Description
Arbitrary Resolution Diffusion Super Resolution on Ray Traced ImagesBACKGROUND
[0001] Photo-realistic imagery can be created from three-dimensional (3D) data, such as 3D meshes, textures and the like. Techniques such as ray tracing or path tracing are ty pical approaches used to create such imagery, which can be used for purposes such as animation or computer-generated (CG) effects. One significant drawback to tracing techniques is that they can be very' computationally intensive and take a long time to render suitable imagery. Another drawback is that tracing techniques are subject to a low cache hit rate. Also, such approaches may generate noisy imagery that can be unsuitable for the intended purpose or otherwise degrade the quality of the end product. Denoising methods may be applied to address quality concerns; however, this may further delay the rendering and add to the computational cost of the overall process.
[0002] Other imaging techniques may alternatively be used to generate photo-realistic imagery. These can include super-resolution (SR) techniques, where low resolution (LR) images can be up- sampled in order to obtain a high-quality or high resolution (HR) image. In addition to the resolution, SR techniques may need address image-related factors including blurring, noise, and / or compression that can impact image quality. A variant SR technique is super-resolution diffusion. This can include applying a neural network model that starts with an input LR image and constructs a corresponding high-resolution image using noise (e.g., Gaussian noise). For instance, the model may be trained via image corruption, where Gaussian noise is progressively added to a high-resolution image while slowly eliminating details in the image data until only pure noise remains. The model then learns to reverse the process, training the neural network to reverse the noise corruption process, beginning from pure noise and progressively removing noise to reach a particular distribution through the guidance of the input image. Super-resolution approaches can be challenging because multiple generated images may be inconsistent with a single starting image, and a conditional distribution of output images given the starting may not conform well to simple parametric distributions.BRIEF SUMMARY
[0003] The technology utilizes different aspects of ray tracing and super resolution diffusion to generate high quality imagery via a trained model. This approach limits the number of ray tracing iterations to avoid issues with computational complexity. This includes applying a low-resolution image of a first size output by the ray tracing process to a denoise diffusion SR model in order to generate a resultant high-resolution image of a selected size, which is larger than that of the low- resolution image. The larger image may be of arbitrary size. Fig. 1A illustrates an example low- resolution image of a first size generated by a ray tracing process, and Fig. IB illustrates an example high-resolution image of a second (greater) size after application of the low-resolution image to adenoise diffusion SR model. This hybrid approach is discussed further below. The technology may be particularly beneficial in a wide variety of use cases, including movie-making such as for animated content or CGI, rearranging on game engines, CAD / CAM programs, etc.
[0004] According to one aspect of the technology, an image processing system comprises me ory configured to store imagery data and one or more processors, which are configured to: retrieve at least a portion of the imagery data stored in the memory; perform an iterative tracing operation in a plurality of iterations to generate an initial image having a first resolution; perform, using the first image and a target scaling factor, an iterative super-resolution operation with diffusion and denoising to generate a final image having a second resolution greater than the first resolution, the iterative super-resolution operation including a plurality of iterations that each generate a respective intermediate image having a corresponding noise value; and output the final image.
[0005] The iterative super-resolution operation may be further performed using a noise strength value. Here, the noise strength value can be adjusted for one or more of the plurality of iterations of the iterative super-resolution operation. The adjustment of the noise strength value can reduce the corresponding noise value for each of the one or more iterations. By way of example, the noise strength value of the last iteration can be reduced to 0.
[0006] The iterative tracing operation can be a ray tracing operation or a path tracing operation. Alternatively or additionally to any of the above, the plurality of iterations of the iterative tracing operation can be a fixed number of iterations, or can be selected so that the first resolution achieves a selected resolution level. Moreover, the image processing system may be configured to vary at least one of the plurality of iterations of the iterative tracing operation or the second resolution of the final image.
[0007] Alternatively or additionally to any of the above, the tracing operation can be performed by the one or more processors via a low-resolution tracing module. The low -resolution tracing module may generate the target scaling factor in addition to the initial image having the first resolution that is less than the second resolution of the final image. The iterative super-resolution operation can be performed by the one or more processors via an iterative denoise diffusion super-resolution module. In addition, each intermediate image can have a corresponding Gaussian noise strength. Alternatively or additionally to any of the above, the output of the final image may include either or both of saving the final in the memory or transmitting the final image to a computing device.
[0008] According to another aspect of the technology, a computer-implemented method is provided. The method comprises: retrieving, by one or more processors at least a portion of imagery data stored in memory; performing, by the one or more processors, an iterative tracing operation in a plurality of iterations to generate an initial image having a first resolution; performing, by the one or more processors using the first image and a target scaling factor, an iterative super-resolution operationwith diffusion and denoising to generate a final image having a second resolution greater than the first resolution, the iterative super-resolution operation including a plurality of iterations that each generate a respective intermediate image having a corresponding noise value; and outputting the final image.
[0009] The iterative super-resolution operation may be further performed using a noise strength value. The iterative tracing operation may be either a ray tracing operation or a path tracing operation. The plurality of iterations of the iterative tracing operation can be selected so that the first resolution achieves a selected resolution level. Alternatively or additionally to any of the above, the method may further comprise vary ing at least one of the plurality of iterations of the iterative tracing operation or the second resolution of the final image.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0011] Figs. 1A-B illustrate example low and high-resolution images at different stages of an image processing approach in accordance with aspects of the technology.
[0012] Fig. 2A illustrates an example architecture for super-resolution.
[0013] Fig. 2B illustrates an exemplary ray tracing process.
[0014] Figs. 3A-B illustrate a hybrid image processing approach in accordance with aspects of the technology.
[0015] Figs. 4A-B illustrate a system for use with aspects of the technology.
[0016] Fig. 5 illustrates a method in accordance with aspects of the technology.DETAILED DESCRIPTIONOverview
[0017] The innovation discussed herein is able to generate photo-realistic imagery using input 3D data, according to a hybrid, multi-stage process. One stage of the process employs ray tracing (or path tracing) to produce a low-resolution image of a first size. This may be done in a selected number of ray (or path) tracing iterations, which provides the technical benefits of significantly reducing the computations complexity over conventional tracing operations, and also addresses the cache hitting issue that can be prevalent in conventional ray (or path) tracing.
[0018] Ray tracing is an approach that is used to simulate light and shadows in an image scene. It involves tracing the route for a virtual beam of light particles in the image scene as it would take in the real (physical) world. For efficiency, the process frequently works backward from a virtual camera to objects in the scene to a virtual light source. This approach is used to simulate virtual beams that seem to reflect off of objects in the image scene, casting realistic-seeming shadows and reflections. How the virtual beams reflect and cast shadows are at least partly dependent on properties of the object(s) in theimage scene, such as shape, color, reflectiveness, etc. New rays can be generated for each new object interaction. This is a recursive process that involves calculating the color of a light beam at its terminal point, based on the material properties of the surface associated with that point.
[0019] Path tracing is generally similar to ray tracing. It involves Monte Carlo processing to generate 3D scenes with realistic global illumination. Path tracing integrates all of the illuminations that reach a particular point on the surface of an object in the image scene. This approach recursively generates multiple rays for each pixel. Path tracing may more accurately simulate the physics of light Ilian ray tracing, but can take a long time to converge, and frequently generates noisier imagery in real applications. In addition, path tracing can be more realistic at dealing with special light effects such as caustics.
[0020] Further details about ray tracing can be found in the article by Turner Whitted, entitled “An improved illumination model for shaded display” in Communications of the ACM, 23(6):343-349, 1980, the entire disclosure of which is incorporated herein by reference. See also U.S. Patent Publication No. 2005 / 017968, the entire disclosure of which is incorporated herein by reference. Additional information about ray tracing and path tracing may be found in the article by Majercik et al., entitled “Dynamic Diffuse Global Illumination Resampling”, published August 11, 2021, the entire disclosure of which is incorporated herein by reference.
[0021] For instance, as explained in the Whitted article, a reflected intensity (I) can be determined according to the following:where la is the reflection due to ambient light, kj is a diffuse reflection constant, A is a unit surface normal, L}is a vector in the direction of the fhlight source, ksis the specular reflection coefficient, S is the intensity of light incident from a reflective (R) direction (where the reflection angle equals the incidence angle), ktis the transmission coefficient, and T is the intensity of light from the transmitted (P) direction. The process may be repeated recursively until all branches in a tree of rays terminate, or are truncated. A shader can be run, traversing the tree of rays to calculate intensity at each node. Rays can then be traced from the “viewer” to objects in the image scene.
[0022] Alternatively, as described in the Majercik article, outgoing radiance L can be described according to the equation L = Le4- T L. with the outgoing radiance being the sum of the radiance Leemitted by light sources plus the conveyed radiance T L, where f corresponds to a bidirectional scattering distribution function that indicates how surfaces in the scene convey radiance (which can incorporate diffuse and glossy components). In this approach, the entire scene surface may be considered as a light source. For instance, the process may first generate, on all scene surfaces, thecandidate samples. This can include computing direct and indirect illumination. One candidate may be selected from, e.g., a weighted reservoir that is proportional to each sample’s contribution. This stage may include spatial resampling for neighboring pixels, and also temporal resampling of one or more prior frames. Then the process can trace a single shadow ray to a selected point. Then, if it is visible, shade from the selected point according to Le, and then adding any contribution from the dynamic diffuse global illumination volume. The process may add diffuse indirectly reflected light (without direct reflections) and / or glossy contributions. Specular recursion may be performed for shading purposes.
[0023] As noted above, both ray tracing and path tracing can be computationally intensive. There are many calculations to perform, such as to find all the light paths and calculate all the corresponding rays. In addition, the cache hit rate can be greatly affected by how ray tracing or path tracing is performed. The cache hit rate is the ratio of successful cache hits to cache lookups. By way of example, the 3D geometry (3D mesh) of the scene may be aligned within the memory of the processing device (e.g., in an LI or L2 cache). The system calculates inside of this 3D geometry, to determine whether a ray hits a particular vertex in the polygon of the geometry. It may not be possible to predict where a next ray will hit inside the polygon, which can result in cache misses. A low cache hitting rate can be a significant bottleneck in the ray tracing or path tracing process. And as more iterations of that process are performed to increase the resolution of the output imagery, the more computationally intensive it becomes.
[0024] Denoising involves removing ‘’noise”, which can be artifacts or defects, from an image. By way of example, there may be blurring, image distortion, image cropping, incorrect colorization, and / or a gap in the image. Denoising the image to address such degradations can result in a sharper, more accurate image. A model comprising trained neural network may be employed to determine whether the image had any degradations, and then to remove such degradations. For an example approach to denoising, see U.S. Patent Publication No. 2023 / 0103638, the entire disclosure of which is incorporated herein by reference.
[0025] For instance, a denoising neural network may be trained on training data that has a set of image pairs. Each pair has a noisy image and a denoised version of that image. The model may generate a forward diffusion process by introducing noise, in an iterative manner, to the denoised image. The model can then learn a reverse diffusion process, which, once trained, can be applied in a denoising process. In one architecture, the model may be configured as an encoder-decoder network that has at least one encoder, at least one corresponding decoder, and a set of connections between selected layers of the encoder(s) and decoder(s). By way of example, the encoder-decoder network may comprise a set of self-attention refinement neural network layers.
[0026] According to one approach, in the first iteration, a noisy estimate of the denoised version of the noisy image is initialized to generate an initial estimate for the noisy image. Here, the method includes iteratively generating the forward diffusion process by predicting, at each iteration in a sequence and based on the current noisy estimate of the denoised version of the noisy image, noise data in order to predict a next noisy estimate for the denoised version of the noisy image. This can include updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current version with predicted noise data. In the next iteration, the next noisy estimate can be reinitialized as the current noisy estimate, which is provided as input to the diffusion model. The diffusion model is then able to predict updated noise data. Moreover, the updated noise data can be combined with the current noisy estimate to generate a next noisy estimate. One or more further iterations may be performed until a desired next noisy estimate is obtained. After the diffusion model generates the forward diffusion process, it learns a reverse diffusion process by inverting the forward diffusion process. Thus, the trained diffusion model may be be configured to predict the denoised version of the noisy image.
[0027] Aspects of denoising diffusion with super-resolution can be found, by way of example, in U.S. Patent No. 11.769.228. as well as in the article by Saharia et al entitled “Image Super-Resolution via Iterative Refinement”, published June 30, 2021, which are incorporated herein by reference in their entireties.
[0028] By way of example, image super-resolution involves creating a high-resolution image that is consistent with a low-resolution image used as the starting point. This process may be challenging because two or more created super-resolution images can be consistent with one input image. In addition, the conditional distribution of created images given the input typically may not suitably conform to simple parametric distributions. Various techniques may be employed to learn complex empirical image distributions, including normalizing flows (NFs), variational autoencoders (VAEs), and generative adversarial networks (GANs). However, such approaches may have certain undesirable constraints, or are otherwise too computationally intensive for viable implementation.
[0029] An alternative approach that may be used in the technology herein is super-resolution via repeated refinement. This approach functions by learning to transform a standard normal distribution into an empirical data distribution via a set of refinement steps. This may be done via a U-net architecture that is trained according to a denoising objective in order to remove, iteratively, varying amounts of noise from the output of the process. This can include minimizing a defined loss function, and can use a constant amount of inference steps regardless of the resolution of the output image.
[0030] A conditional denoising diffusion model may utilize a stored dataset of input-output image pairs, which may have multiple target images that are consistent with one source image. The systemmay learn a parametric approximation for an unknown conditional distribution, through stochastic iterative refinement, which maps a source image to a target image.
[0031] Fig. 2A is an example U-net architecture 200 for a neural network, for use in a superresolution approach. A low-resolution input image x (202) is interpolated to a target high resolution, and then concatenated with a noisy high-resolution image yt(204). The activation dimensions for an example task of 16x16 128x128 super-resolution is shown. The neural netw ork may be a convolutional neural network (CNN) based on a denoising diffusion probabilistic (DDPM) model.
[0032] As illustrates, in one step of the iteration from the first noisy high-resolution image 204. to a second noisy high-resolution image yt-i (206), the input low resolution input image 202 can be downsampled from 128* 128 at 208 to 64x64 at 210. then down to 8x8 at block 212. Then an output from this down sampling can be up-sampled from 8x8 at 214 to 64x64 at 216, and then to 128x 128 at 218. Skip connections can be used, such as a first skip connection 220 connecting 208 to 218, and a second skip connection 222 connecting 210 to 216.
[0033] The following table illustrates example task-specific architecture hyperparameters for a Linet model according to the architecture 200 of Fig. 2A.
[0034] In particular, this table provides example task specific architecture hyperparameters for the U-Net model. The leftmost column displays a super-resolution task. The next column shows the channel dimension associated with the super-resolution task. The third column presents a set of depth multipliers associated with the super-resolution task. The term “Channel Dim’' refers to the dimension of the first U-Net layer, while the “Depth Multipliers” denote multipliers for the subsequent resolutions.
[0035] This architecture may be trained and operated as follow s. For training of a neural network model, the system receives a plurality of image pairs as training data. Each pair includes an input image and at least one corresponding target version of that image. The neural network is trained using the training data, predicting an enhanced version of the input image. The training can include, e.g., applying a forward Gaussian-type diffusion process, which introduces Gaussian noise to at least one of the corresponding target versions of each of the pluralities of image pairs. This enforces iterative denoising of the input image. Once the model is trained, it can be saved in memory and / or sent to another device for implementation.
[0036] Implementation involves the system receiving or otherwise obtaining an input image. The trained neural network is applied to the input image to predict an enhanced version thereof by iteratively denoising the input image. By way of example, the iterative denoising may be done according to a reverse Markov chain associated with a forward Gaussian diffusion process. From this, a predicted enhanced version of the input image is generated. This image may be saved in memory and / or output by die system, e.g., for presentation to a user.Overall Architecture
[0037] Fig. 2B illustrates an example 250 for a conventional ray tracing process. As shown, stored 3D data 252 is applied to a ray tracing module 254 to generate an image frame 256 at a selected resolution. By way of example, the process performed by the ray tracing module 254 (e.g., in accordance with above ray tracing discussion) may require at a minimum of hundreds or thousands of iterations to achieve the selected resolution. In this example, 1,000 iterations were needed to render the image frame 256. As noted above, the ray tracing process may result in a noisy image. Thus, optionally a denoiser module 258 may process the output of the ray tracing module 254 to substantially reduce the noise for the output image.
[0038] In contrast, the hybrid innovation utilizes ray (or path) tracing to generate an initial low- resolution image having first, relatively small, size. This can be done in a few. e.g., tens or at most hundreds of iterations. Then a denoising diffusion super-resolution process is performed. This SR process is much faster than ray (or path) tracing, and can take advantage of the starting point (a low- resolution image) generated by the tracing process.
[0039] Fig. 3A illustrates an example hybrid approach 300. Here, similar to the example 200. stored 3D imagery data 302 is applied as input. In this situation, the 3D imagery data 302 is applied to a low -resolution ray tracing (or path tracing) module 304. This module 304 is configured to generate a low-resolution image 306 of a first, reduced size. The module 304 may be configured to perform a fixed number of tracing iterations or to run until an image 306 having a specific resolution is achieved. Alternatively, the system may vary the number of tracing iterations and / or resolution to be output by this module.
[0040] The image 306 and a target scaling factor 308 (“s") are provided to an iterative denoise diffusion super-resolution module 310. The target scaling factor, s, may be a fixed scaling factor or a variable scaling factor. For instance, the scaling factor may be changed depending on the type of imagery, a desired output resolution, the size of the image generated by the module 304 and / or other factors. This module 310 also takes a noise strength value 312 (“f ’), which can be adjusted by iteration. The module 310 is configured to implement a super-resolution approach with diffusion and denoising, as discussed above. It performs upscaling of the traced image 306 to a desired output resolution.
[0041] The module 310 performs a set of iterations on the upscaled version of the image. At each iteration there is a refined image 314 and a corresponding Gaussian noise strength 316. Some or all of the images 314 may be stored in memory or discarded after each iteration. The noise strength 316 is adjusted for each subsequent iteration by reducing the value t, thereby gradually reducing the noise. As shown in Fig. 3B, when t reaches 0, the module 310 outputs a final image 318 at the selected resolution. This is a high quality, denoised, super-resolution image. This final image 318 may be stored in system memory and / or sent to a computing device, e.g., for presentation on the display of a client device.
[0042] While this hybrid approach may involve a substantial amount of iterative computing, a technical benefit is an improved cache hitting rate over a conventional ray (or) path tracing process, since it may involve only a small fraction of the number of tracing iterations relative to those conventional processes. For instance, the ray tracing may only include performing 50 iterations (or more or less, such as no more than 100 iterations) as opposed to 1,000 iterations in the example of Fig. 2. Here, the module 310 may also perform on the order of 50 iterations, where the output final image 318 may be of the same or higher quality than the image 206.
[0043] Moreover, this approach provides another technical benefit, which is iterative refinement of the diffusion-based module 310. Depending on the image quality desired, the system may trade off betw een the number of tracing iterations performed by the module 304 versus the number of denoise diffusion iterations performed by the module 310. By way of example, the system may balance the computation between the two modules 304, 310 depending on the generated image quality and computation costs associated with the overall process.
[0044] In one scenario, the output, final image 318 is a fixed size super-resolution image. How ever, in another scenario, the final image 318 may be scaled to any selected size and / or resolution. This can be done by increasing the number of iterations performed by the module 310, depending on the scale factor for the size. Thus, the system may iterate until it drops below a particular noise threshold, or just iterates a specific number of times. Alternatively or additionally with any of the above, the system may be configured to generate a set of output images each having a different size and / or resolution.
[0045] The ray (or path) tracing module 304 and the SR module 310 may be configured as separate modules. In one approach, there may be set of modules 304 that support different tracing implementations, and / or a set of modules 310 that support different SR implementations. The system may select one module from each set, for instance depending on the type of imaging application involved. Alternatively, the modules 304 and 310 may be part of a combined module 320 (Fig. 3A). The module 320 may incorporate a trained machine learning model. Regardless of the configuration, the module(s) may be implemented by one of more processors of a computing system executing a set of instructions as discussed below.
[0046] Note that ray tracing or path tracing have very poor cache hitting rates, since those techniques randomly generate rays and the system does not know how each triangle will be hit by each ray. On the other hand, a denoising diffusion neural network handles 2D images, which the image can be cached easily, and in many cases the system handles neighboring pixels, which has very good cache hit rate. Moreover, since ray tracing is computationally expensive, the present system only performs partial ray tracing, and leaves some space left. The remaining pixels can be filled with an inpainting (diffusion) model. Since the system already know s on which pixels ray tracing is performed and which pixels are not ray traced, there is a perfect mask image for the system to perform inpainting.Example computing architecture
[0047] The models described herein may be trained on one or more tensor processing units (TPUs), graphics processing units (GPUs), CPUs or other computing architectures in order to implement the hybrid image processing approach discussed above in order to generate high quality 3D imagery.
[0048] One example computing architecture is shown in Figs. 4A and 4B. In particular. Figs. 4A and 4B are pictorial and functional diagrams, respectively, of an example system 400 that includes a plurality of computing devices and databases connected via a network. For instance, computing device(s) 402 may be a cloud-based server system. Databases 404, 406 and 408 may store, e.g.. the 3D image data or other input imagery, generated intermediate or output imagery', and / or trained models respectively. The server system may access the databases via network 410. Client devices may include one or more of a desktop computer 412 and a laptop or tablet PC 414, for instance to provide the input imagery, manage model training, and / or to view' or otherwise employ the output imagery (e.g., view generated images or use of such imagery' in movie-making or a game engine, or use in a different app / program).
[0049] As shown in Fig. 4B, each of the computing devices 402 and 412-414 may include one or more processors, memory', data and instructions. The memory stores information accessible by the one or more processors, including instructions and data that may be executed or otherwise used by the processor(s). The memory may be of any ty pe capable of storing information accessible by the processor(s), including a computing device-readable medium. The memory' is a non-transitory medium such as a hard-drive, memory card, optical disk, solid-state, etc. Systems may include different combinations of the foregoing, whereby different portions of the instructions and data are stored on different types of media. The instructions may be any set of instructions to be executed directly (such as machine code) or indirectly (such as scripts) by the processor(s). For example, tire instructions may be stored as computing device code on the computing device-readable medium. In that regard, the terms “instructions”, “modules” and “programs” may be used interchangeably herein. The instructions may be stored in object code format for direct processing by the processor, or in any other computingdevice language including scripts or collections of independent source code modules that are interpreted on demand or compiled in advance.
[0050] The processors may be any conventional processors, such as commercially available CPUs, TPUs, GPUs, etc. Alternatively , each processor may be a dedicated device such as an ASIC or other hardware-based processor. Although Fig. 4B functionally illustrates the processors, memory, and other elements of a given computing device as being within the same block, such devices may actually include multiple processors, computing devices, or memories that may or may not be stored within the same physical housing. Similarly, the memory may be a hard drive or other storage media located in a housing different from that of the processor(s), for instance in a cloud computing system of server 402. Accordingly, references to a processor or computing device will be understood to include references to a collection of processors or computing devices or memories that may or may not operate in parallel.
[0051] Reference to “one or more processors” herein includes situations where a set of processors (e.g., two or more CPUs, TPUs. GPUs or any combination thereof) may be configured to perform one or more operations. Any combination of such a set of processors may perform individual operations or a group of operations. Therefore, reference to “one or more processors” does not require that all processors in the set must perform all of the operations. Rather, unless expressly stated, any one of the one or more processors may perform different operations when a set of operations is indicated. For instance, different processors may perform specific operations. For example, a first processor performs one or more iterations of the ray or path tracing process, while a second processors performs one or more iterations of the denoise diffusion process.
[0052] Input data, such as one or more image pairs, may be operated on by the modules, models and processes described herein. The client devices may utilize such information in various apps or other programs to perform CAD / CAM operations, video animation, game design, etc. The computing devices may include all of the components normally used in connection with a computing device such as the processor and memory described above as well as a user interface subsystem for receiving input from a user and presenting information to the user (e.g., text, imagery and / or other graphical elements). The user interface subsystem may include one or more user inputs (e.g., at least one front (user) facing camera, a mouse, key board, touch screen and / or microphone) and one or more display devices (e.g., a monitor having a screen or any other electrical device that is operable to display information (e.g., text, imagery and / or other graphical elements). Other output devices, such as speaker(s) may also provide information to users.
[0053] The user-focused types of computing devices (e.g., 412-414) may communicate with a back-end computing system (e.g., server 402) via one or more networks, such as network 410. The network 410. and intervening nodes, may include various configurations and protocols including short range communication protocols such as Bluetooth™. Bluetooth LE™, the Internet, World Wide Web,intranets, virtual private networks, wide area networks, local networks, private networks using communication protocols proprietary to one or more companies, Ethernet, WiFi and HTTP, and various combinations of the foregoing. Such communication may be facilitated by any device capable of transmitting data to and from other computing devices, such as modems and wireless interfaces.
[0054] In one example, computing device 402 may include one or more server computing devices having a plurality of computing devices, e g., a load balanced server farm or cloud computing system, that exchange information with different nodes of a netw ork for the purpose of receiving, processing and transmitting the data to and from other computing devices. For instance, computing device 402 may include one or more server computing devices that are capable of communicating with any of the computing devices 412-414 via the network 410.
[0055] Fig. 5 illustrates an exemplary computer-implemented method 500 in accordance with the above-described features. This method includes, at block 502, retrieving, by one or more processors at least a portion of imagery data stored in memory. Then at block 504, the method includes performing, by the one or more processors, an iterative tracing operation in a plurality of iterations to generate an initial image having a first resolution. At block 506 the method includes performing, by the one or more processors using the first image and a target scaling factor, an iterative super-resolution operation with diffusion and denoising to generate a final image having a second resolution greater than the first resolution. The iterative super-resolution operation includes a plurality of iterations that each generate a respective intermediate image having a corresponding noise value. And at block 508. the method includes outputting the final image. This can include one or both of saving the final in the memory or transmitting the final image to a computing device.
[0056] Input imagery, generated 3D images and / or trained neural network models may be shared by the server with one or more of the client computing devices. Alternatively or additionally, the client device(s) may maintain their own databases that store such information.
[0057] Although the technology herein has been described with reference to particular embodiments, it is to be understood that these embodiments are merely illustrative of the principles and applications of the present technology. It is therefore to be understood that numerous modifications may be made to the illustrative embodiments and that other arrangements may be devised without departing from the spirit and scope of the present technology as defined by the appended claims.
Claims
CLAIMS1. An image processing system, comprising: memory configured to store imagery data; and one or more processors configured to: retrieve at least a portion of the imagery data stored in the memory; perform an iterative tracing operation in a plurality of iterations to generate an initial image having a first resolution; perform, using the first image and a target scaling factor, an iterative super-resolution operation with diffusion and denoising to generate a final image having a second resolution greater than the first resolution, the iterative super-resolution operation including a plurality of iterations that each generate a respective intermediate image having a corresponding noise value; and output the final image.
2. The image processing system of claim 1, wherein the iterative super-resolution operation is further performed using a noise strength value.
3. The image processing system of claim 2, wherein the noise strength value is adjusted for one or more of the plurality of iterations of the iterative super-resolution operation.
4. The image processing system of claim 3. wherein adjustment of the noise strength value reduces the corresponding noise value for each of the one or more iterations.
5. The image processing system of claim 3. wherein the noise strength value of the last iteration is 0.
6. The image processing system of claim 1, wherein the iterative tracing operation is a ray tracing operation.
7. The image processing system of claim 1, wherein the iterative tracing operation is a path tracing operation.8., The image processing system of claim 1, wherein the plurality of iterations of the iterative tracing operation is a fixed number of iterations.
9. The image processing system of claim 1, wherein the plurality of iterations of the iterative tracing operation is selected so that the first resolution achieves a selected resolution level.
10. The image processing system of claim 1, wherein the image processing system is configured to vary at least one of the plurality of iterations of the iterative tracing operation or the second resolution of the final image.
11. The image processing system of claim 1, wherein the tracing operation is performed by the one or more processors via a low -resolution tracing module.
12. The image processing system of claim 11, wherein the low-resolution tracing module generates the target scaling factor in addition to the initial image having the first resolution that is less than the second resolution of the final image.
13. The image processing system of claim 11, wherein the iterative super-resolution operation is performed by the one or more processors via an iterative denoise diffusion super-resolution module.
14. The image processing system of claim 13, wherein each intermediate image has a corresponding Gaussian noise strength.
15. The image processing system of claim 1, wherein output of the final image includes either save the final in the memory or transmit the final image to a computing device.
16. A computer-implemented method, comprising: retrieving, by one or more processors at least a portion of imagery data stored in memory; performing, by the one or more processors, an iterative tracing operation in a plurality of iterations to generate an initial image having a first resolution; performing, by the one or more processors using the first image and a target scaling factor, an iterative super-resolution operation with diffusion and denoising to generate a final image having a second resolution greater than the first resolution, the iterative super-resolution operation including a plurality of iterations that each generate a respective intermediate image having a corresponding noise value; and outputting the final image.
17. The computer-implemented method of claim 16, wherein the iterative super-resolution operation is further performed using a noise strength value.
18. The computer-implemented method of claim 16, wherein the iterative tracing operation is a ray tracing operation or a path tracing operation.
19. The computer-implemented method of claim 16, wherein the plurality of iterations of the iterative tracing operation is selected so that the first resolution achieves a selected resolution level.
20. The computer-implemented method of claim 16, further comprising varying at least one of the plurality of iterations of the iterative tracing operation or the second resolution of the final image.
Citation Information
Patent Citations
Image enhancement via iterative refinement based on machine learning models
US11769228B2
Differential stream of point samples for real-time 3D video
US20050017968A1
Image-to-Image Mapping by Iterative De-Noising
US20230103638A1
Information processing device, information processing method, information processing program, and information processing system
WO2022196200A1