A super-resolution rendering method and device integrating non-aligned spatiotemporal information
By fusing non-aligned spatiotemporal information and utilizing high-resolution geometric buffered data, combined with the estimation model of convolutional neural networks, the existing super-resolution rendering methods in complex texture detail recovery and temporal stability are solved, and the rendering effect with higher quality and higher performance is achieved.
Patent Information
- Application Number
- CN202411465556.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-10-21
AI Technical Summary
Existing super-resolution rendering methods do not perform well in restoring complex texture details, especially in high-resolution and high-frequency details, and are prone to flicker artifacts caused by time instability.
By fusing non-aligned spatiotemporal information, using high-resolution geometric buffer data and pre-integrated bidirectional reflection distribution function values, pseudo-illumination values and pseudo-visibility values of anti-aliased bidirectional reflection distribution function values, pseudo-illumination values and pseudo-visibility values in real time, and reducing artifacts through multi-frame information fusion to improve the temporal stability of the image.
Improves the recovery ability of high-frequency details, reduces flicker artifacts, improves the time stability of images, and significantly improves the quality and performance of super-resolution rendering.
Smart Images

Figure CN118967901B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the combination of computer graphics and image processing, and specifically relates to a super-resolution rendering method and device integrating non-aligned spatiotemporal information. Background Art
[0002] Over the past few years, the computational workload of real-time rendering has increased significantly with the popularity of high-resolution and high-refresh rate displays, realistic lighting effects, and advances in real-time ray tracing technology. The advent of real-time ray tracing further increases the computational overhead of rendering high-quality output. Even high-end consumer GPUs have difficulty rendering high-quality images at 144 FPS and 4K resolution. As a result, users have to make a trade-off between rendering quality, resolution, and refresh rate.
[0003] To address this challenge, researchers have proposed a variety of techniques to ease the computational burden. Real-time denoising techniques produce reasonable results by rendering images at a low tracking budget and trying to reduce the noise level of these images. Frame interpolation methods focus on reconstructing accurate shading results from historical frames with motion vectors to speed up rendering. Foveated rendering methods propose to reduce the resolution at the edge of the user's field of view without sacrificing perceived visual quality, thereby improving efficiency, but are only applicable to virtual reality headsets.
[0004] The most widely adopted and successful methods are super-resolution (SR) methods. Users can reduce the resolution of rendered images to reduce rendering time and upsample low-resolution (LR) rendered images to obtain the final high-resolution (HR) images. These methods utilize deep learning and artificial intelligence techniques to train models to predict and reconstruct high-resolution images. However, these methods mainly consider upsampling factors less than 2×2, which limits higher performance gains.
[0005] Existing RRSR (real-time super-resolution) methods mainly include traditional methods and deep learning-based methods. Temporal anti-aliasing upsampling (TAAU) simply uses temporal accumulation to perform SR. Edelsten et al. proposed deep learning super sampling (DLSS), which takes into account temporal and spatial information. However, their method depends on NVIDIA's hardware platform and there is no public technical information. Unlike DLSS, super-resolution sharp picture technology (FSR) does not rely on a specific hardware platform. It usually achieves SR through upsampling and edge sharpening, with limited quality improvement, which has been improved after integrating TAAU, namely FSR 2.0. Intel's XeSS is also a deep learning-based SR method for rendered images, mainly for production applications, and currently only demonstration samples are available. NSRR proposed by Xiao et al. uses UNet for SR reconstruction and processes after zero-sampling the features of multiple frames. However, NSRR will produce ghosting artifacts when processing scenes with fast-moving objects or cameras. Gu et al. proposed to extract edge features for SR interpolation, but did not consider temporal stability. Guo et al. trained two networks separately to classify pixels into different categories and then predict the weights of frame mixing. Yang et al. used an alternating sub-pixel sampling pattern during rasterization to create a small SR model that can run on mobile devices. Mercier et al. proposed a lightweight recurrent network and a game dataset for RRSR. Although these methods can achieve real-time performance, they cannot accurately recover complex texture details.
[0006] A recent advance, FuseSR proposed by Zhong et al., implements a real-time super-resolution technique that can provide high-fidelity 4×4 or even 8×8 upsampled reconstructions, with significantly improved quality and performance over existing works. The key improvement lies in leveraging HR spatial information and HR G-buffer to provide per-pixel cues for HR targets. While this intuitive modification breaks through the previous upsampling scale limitation and shows a promising path for practical large-scale SR, it also reveals another challenge: when the rendered frame is full of high-frequency details, temporal instability becomes more visually significant. This instability usually manifests as flickering artifacts, which seriously affects the user experience. Summary of the invention
[0007] In view of the above, the purpose of the present invention is to provide a super-resolution rendering method and device that integrates non-aligned spatiotemporal information, which provides higher high-frequency detail recovery capability through non-aligned spatiotemporal information, and at the same time effectively reduces these artifacts and improves the temporal stability of the image by introducing multi-frame information fusion in the time domain.
[0008] To achieve the above-mentioned purpose of the invention, an embodiment provides a super-resolution rendering method integrating non-aligned spatiotemporal information, comprising the following steps:
[0009] Get the low-resolution color image and high-resolution geometry buffer data of the current frame in real time;
[0010] Based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image, the anti-aliased bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value of the current frame are estimated in real time by non-aligned spatiotemporal feature fusion;
[0011] The high-resolution color image of the current frame is obtained by multiplying the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
[0012] Preferably, the pre-integrated bidirectional reflection distribution function value is obtained in real time and then participates in the estimation calculation, including: pre-integrating and calculating the bidirectional reflection distribution function value at different roughness values for the viewing angle direction determined by each combination of normal and illumination direction and storing it; when applied in real time, the pre-integrated bidirectional reflection distribution function value of the current frame is queried from the storage according to the roughness and viewing angle direction of each pixel provided by the high-resolution geometric buffer data of the current frame.
[0013] Preferably, an estimation model constructed according to a convolutional neural network is used to estimate the anti-aliased bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame in real time based on high-resolution geometric buffer data, pre-integrated bidirectional reflectance distribution function values, and low-resolution color images in a non-aligned spatiotemporal information manner. First, the high-resolution geometric buffer data and the pre-integrated bidirectional reflectance distribution function values are reduced to low resolution and then feature fused with the low-resolution color image in the low-resolution space. Then, information enhancement is extracted for the fused features. During the information enhancement extraction, multi-layer feature accumulation is performed and different mixing coefficients are used to mix the corresponding features of the same layer of the previous frame. Finally, the anti-aliased bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame are predicted and estimated based on the features after the information enhancement extraction process are upgraded to high resolution.
[0014] Preferably, the estimation model includes a spatiotemporal encoding module, an information enhancement module, and a decoding module.
[0015] The spatiotemporal coding module is used to mix the high-resolution auxiliary features composed of the high-resolution geometric buffer data and the pre-integrated bidirectional reflectance distribution function value with the previous frame of high-resolution color image according to a certain mixing coefficient and reduce the dimension to a low-resolution space, and then perform spatiotemporal coding with the extracted features of the low-resolution color image to obtain the coding features;
[0016] The information enhancement module is used to extract multi-layer features based on the coding features, and use different mixing coefficients to mix the corresponding features of the same layer of the previous frame when extracting certain layer features;
[0017] The decoding module is used to fuse the features output by the information enhancement module with the extracted features of the low-resolution color image and upgrade the dimension to the high-resolution space, and then estimate the anti-aliasing bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame through decoding.
[0018] Preferably, the dimensionality reduction process uses a pixel anti-shuffle operation, and the dimensionality increase process uses a pixel shuffle operation.
[0019] Preferably, when using different mixing coefficients to mix corresponding features of the same layer of the previous frame, a small mixing coefficient is used for a fast-moving object, and a large mixing coefficient is used for a slow-moving object.
[0020] Preferably, the convolutional neural network undergoes parameter optimization before performing estimation calculations, and the loss functions used in parameter optimization include first paradigm loss, perceptual loss, structural similarity loss based on the generated high-resolution color image and the high-resolution real image, and also include temporal stability loss based on the optical flow between two adjacent frames of high-resolution color images.
[0021] To achieve the above-mentioned purpose of the invention, an embodiment further provides a super-resolution rendering device integrating non-aligned spatiotemporal information, comprising a data acquisition unit, a real-time estimation unit, and a real-time synthesis unit;
[0022] The data acquisition unit is used to acquire the low-resolution color image and high-resolution geometric buffer data of the current frame in real time;
[0023] The real-time estimation unit is used to estimate the anti-aliasing bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame in real time based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image by non-aligned spatiotemporal feature fusion method;
[0024] The real-time synthesis unit is used to obtain a high-resolution color image of the current frame based on multiplication of the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
[0025] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned super-resolution rendering method that integrates non-aligned spatiotemporal information.
[0026] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the above-mentioned super-resolution rendering method integrating non-aligned spatiotemporal information.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] Based on the low-resolution color image of the current frame and the high-resolution geometric buffer data, the anti-aliasing bidirectional reflectance distribution function value, pseudo-illumination value, and pseudo-visibility value of the current frame are estimated in real time; then the high-resolution color image is fitted by the anti-aliasing bidirectional reflectance distribution function value, pseudo-illumination value, and pseudo-visibility value, so that the ability to restore high-frequency details can be improved through non-aligned low-resolution and high-resolution spatiotemporal information. On this basis, multi-frame information fusion in the time domain is introduced when estimating the anti-aliasing bidirectional reflectance distribution function value, pseudo-illumination value, and pseudo-visibility value of the current frame, which can effectively reduce these artifacts and improve the temporal stability of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 is a flow chart of a super-resolution rendering method integrating non-aligned spatiotemporal information provided by an embodiment;
[0031] Figure 2 is a schematic diagram of the structure of an estimation model constructed according to a convolutional neural network provided in an embodiment;
[0032] Figure 3 is a comparison diagram of super-resolution imaging results of a 4x4 task provided in an embodiment;
[0033] Figure 4 is a comparison diagram of super-resolution imaging results of the 8x8 task provided in the embodiment;
[0034] Figure 5 4 is a schematic diagram of the structure of a super-resolution rendering device that integrates non-aligned spatiotemporal information provided by an embodiment. DETAILED DESCRIPTION
[0035] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0036] The inventive concept of the present invention is that the existing methods perform poorly in restoring complex texture details, especially in scenes with high resolution and high frequency details. Traditional methods such as FSR only achieve SR through upsampling and edge sharpening, and the quality improvement is limited; although the deep learning-based methods of DLSS and NSRR take into account spatiotemporal information, they do not utilize high-resolution information and have limited high-frequency detail recovery. Although FuseSR utilizes HR spatial information and HR G-buffer, it is not good at shadow performance. To this end, an embodiment of the present invention provides a super-resolution rendering scheme that fuses non-aligned spatiotemporal information, which provides higher detail recovery capabilities by fusing non-aligned spatiotemporal information, and effectively solves the problem of insufficient high-frequency detail recovery. At the same time, the existing super-resolution methods are prone to flickering artifacts and temporal instability when processing dynamic scenes. For example, although FuseSR utilizes HR spatial information and HR G-buffer, it is prone to temporal instability in high-frequency detail scenes. The super-resolution rendering scheme of the present invention also effectively reduces these artifacts and improves the temporal stability of the image by introducing multi-frame information fusion in the time domain.
[0037] like Figure 1 As shown, the embodiment provides a super-resolution rendering method integrating non-aligned spatiotemporal information, comprising the following steps:
[0038] S1, obtain the low-resolution color image and high-resolution geometry buffer data of the current frame in real time.
[0039] In the embodiment, a low-resolution color image of the current frame can be obtained in real time, and high-resolution geometry buffer (G-buffer) data can also be obtained at the same time, wherein the high-resolution G-buffer, as a common rendering byproduct, contains rich scene information, such as depth, normals, and texture details, and its acquisition cost is significantly lower than the heavy shading and potential post-processing tasks (each 1080p frame takes only a few milliseconds). The embodiment uses the high-resolution G-buffer as a super-resolution clue to make the problem easier to handle.
[0040] In order to make better use of the high-resolution G-buffer, a pre-integrated BRDF demodulation method is used to explicitly filter out super-resolution details, and the segmentation and summation approximation is introduced into the super-resolution task, so that the color frame super-resolution problem is transformed into a demodulated irradiance super-resolution problem, namely:
[0041] ;
[0042] ;
[0043] ;
[0044] in, and denote the incident and outgoing directions respectively, Represents the bidirectional reflectance distribution function (BRDF) at the shading point, the lighting function describes the incident radiance at the shading point, is the angle of incidence The cosine term of is the outgoing radiance, is the anti-aliased BRDF value, which decomposes the target radiance into the product of a BRDF-like and irradiance-like term. This strategy successfully filters out high-frequency material details, thereby better preserving overall quality and details. Therefore, it is possible to predict the full outgoing radiance. Better prediction of pseudo irradiance .
[0045] Further simplifying the problem, the shadow is transformed from pseudo irradiance In order to process high-frequency material details and shadow information respectively, an extended split and sum approximation is proposed to reformulate the rendering equation as follows:
[0046] ;
[0047] ;
[0048] ;
[0049] in, is the radiance generated by the engine without calculating shadows (i.e. ignoring all occlusion). On the one hand, the pseudo lighting value Represents the integrated lighting information that has been decoupled from shadows. It contains screen-space geometry details and can be efficiently predicted using high-resolution G-buffers and motion vectors. On the other hand, the pseudo visibility value represents the integral visibility information, which is usually independent of the screen space information and can be implicitly processed by the neural shading flow. Once decoupled from each other, both can achieve better prediction results. Therefore, the super-resolution task becomes estimating the corresponding high-resolution and high resolution , and the BRDF value for anti-aliasing .
[0050] S2, based on high-resolution geometry buffer data, pre-integrated bidirectional reflectance distribution function values, and low-resolution color images, estimates the anti-aliased bidirectional reflectance distribution function values, pseudo illumination values, and pseudo visibility values of the current frame in real time through non-aligned spatiotemporal feature fusion.
[0051] In an embodiment, the pre-integrated BRDF value can be obtained at a very low cost by pre-calculation. Specifically, the BRDF value at different roughness values for the viewing angle direction determined by each normal and illumination direction combination is pre-integrated and stored, for example, in a two-dimensional lookup texture (LUT). In real-time application, the pre-integrated current frame BRDF value is queried from the storage according to the roughness and viewing angle direction of each pixel provided by the high-resolution geometry buffer data of the current frame.
[0052] In an embodiment, the anti-aliasing bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame can be estimated in real time based on the high-resolution geometric buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image through a non-aligned spatiotemporal feature fusion method. Specifically, depth data, normal data, etc. in the high-resolution geometric buffer data can be selected when making the estimation. An estimation model constructed by a convolutional neural network can also be used to align the size of the high-resolution features with the features extracted from the low-resolution color image by adjusting the size in the network before performing feature fusion.
[0053] Specifically, since convolutional neural networks take multi-resolution features (low-resolution images and high-resolution G-buffer) as input, in order to make full use of the high-resolution G-buffer input. It is necessary to align pixels that share the same screen space position to maintain the correct spatial association between multi-resolution input features. In order to align pixels that share the same screen space position, upsampling and pooling are two common strategies, the former aligning at the high-resolution level and the latter aligning at the low-resolution level. Considering that the performance of convolutional networks drops sharply with the increase of input resolution, upsampling alignment conflicts with the demand for real-time performance, and pooling is a more feasible option. Pooling operations are common in neural networks, including maximum pooling and average pooling. However, these pooling operations inevitably damage spatial details, such as high-resolution edges and textures, which are key information for improving super-resolution quality. Therefore, when adjusting the size (i.e., resolution), a pixel deshuffle operation is used to align high-resolution features to low resolution instead of a lossy pooling operation. It is found that pixel deshuffle can losslessly reduce high-resolution feature maps to low-resolution space, converting pixel-level spatial information into channel-level deep information. Specifically, this operation divides the high-resolution image into blocks of size r×r, where r is the downsampling factor, and connects the features of all pixels in each block to form a low-resolution version of the pixels, that is, converting an image of shape [C, H×r, W×r] into a low-resolution image of shape [C×r×r, H, W] without losing information.
[0054] The high-resolution G-buffer, pre-integrated bidirectional reflectance distribution function values, and high-resolution historical information frames are converted to low resolution using pixel de-shuffling operations to align multi-resolution features. The aligned high-resolution G-buffer features are fused with the features extracted from the low-resolution color image, and then dimensionalized to high-resolution space through pixel shuffling operations to predict anti-aliasing BRDF values, pseudo illumination values, and pseudo visibility values in high-resolution space.
[0055] Due to the rapid changes in shadows and visibility issues, artifacts such as ghosting and flickering are common challenges in temporal sampling. When the historical frame comes from an object different from the current foreground, ghosting occurs, resulting in overlapping traces of previous objects. On the other hand, when the weight of the historical frame is too low, flickering occurs, resulting in a lack of stability in the static area. When dealing with complex motion or illumination changes, the historical frame information will be unreliable, reducing the overall image quality. In order to solve this problem, the embodiment also performs information enhancement extraction on the fusion features in the low-resolution space, accumulates multiple layers of features and uses different mixing coefficients to mix the corresponding features of the previous frame to form a time accumulation of neural network features, avoiding the problems caused by traditional time anti-aliasing. After the features after information enhancement extraction are upgraded to high resolution, the bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the anti-aliasing of the current frame are predicted and estimated. Among them, when using different mixing coefficients to mix the corresponding features of the previous frame, a small mixing coefficient is used for objects with fast moving speeds, and a large mixing coefficient is used for objects with slow moving speeds.
[0056] like Figure 2 As shown, the estimation model constructed according to the convolutional neural network includes a spatiotemporal encoding module, an information enhancement module, and a decoding module, wherein the spatiotemporal encoding module is used to mix the high-resolution auxiliary features composed of the high-resolution geometric buffer data and the pre-integrated bidirectional reflectance distribution function values with the previous frame of high-resolution color image and reduce the dimension to a low-resolution space, and then perform spatiotemporal encoding with the extracted features of the low-resolution color image to obtain the encoded features; wherein, the encoder module (Encoder Block) is used to extract the features of the low-resolution space of the low-resolution color image, and the dimensionality reduction processing adopts the pixel unshuffling operation (Pixel Unshuffling).
[0057] The information enhancement module is used to extract multi-layer features based on the encoded features, and when extracting certain layer features, the features extracted from the same layer of the previous frame are mixed according to different mixing coefficients. The residual network block (ResidualBlock) is used for feature extraction of each layer, and the features extracted from the same layer of the previous frame are mixed with the extracted features of the same layer of the current frame according to different mixing coefficients for the following feature extraction. When mixing, a small mixing coefficient is used for fast-moving objects, and a large mixing coefficient is used for slow-moving objects.
[0058] The decoding module is used to fuse the features output by the information enhancement module with the extracted features of the low-resolution color image and upgrade the dimensions to the high-resolution space, and then estimate the anti-aliasing BRDF value, pseudo illumination value, and pseudo visibility value of the current frame through decoding. The dimension upgrade process uses pixel shuffling, and the decoding process is implemented using the output module (Output Block). During the decoding process, the Relu activation function can be used for the BRDF value and pseudo illumination value. For the pseudo visibility value, since it needs to output a value between 0 and 1, the Sigmod activation function is used.
[0059] like Figure 2 The estimation model based on the convolutional neural network shown in the figure needs to go through an end-to-end training process before it can be used. During training, the image dataset is augmented, converted and cropped to generate more sliced image training sets, and trained under the supervision of high-resolution real images. The loss function used during training includes the first paradigm loss L between the generated high-resolution color image and the high-resolution real image. c , the perceptual loss L s and the structural similarity loss L ssim , the temporal stability loss L between high-resolution color images predicted by two adjacent frames t .
[0060] Among them, the first paradigm loss L c Used to measure the difference between two images, the perceptual loss L s The pre-trained VGG-16 network can be used to extract the perceptual features of the generated high-resolution color image and the high-resolution real image, and then the perceptual loss and the temporal stability loss L can be calculated based on the perceptual features. t It is calculated based on the difference between the same features of the high-resolution color images predicted by two adjacent frames, ensuring that the features of the same object in the two frames are consistent as much as possible, and the same object between the two frames can be found through the motion vector. Structural similarity loss L ssim It is used to measure the color consistency of two images. The total loss L = L c + 0.5*L s + 0.5*L t + 0.05* L ssim .
[0061] S3, obtaining a high-resolution color image of the current frame based on multiplying the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
[0062] In the embodiment, based on the analysis of step S1, the present invention is based on super-resolution imaging of color frames Expanded split The approximate summation problem is Therefore, after estimating and predicting the anti-aliasing BRDF value , Pseudo Light Value , and pseudo-visibility values Finally, the three are multiplied together to obtain the high-resolution color image of the current frame.
[0063] In the embodiment, a large-scale dataset is constructed using four virtual scenes selected from the Unreal Engine market, including Kite and Showdown of Unreal Engine 4 and Slay and City of UnrealEngine 5. These scenes contain complex geometry and lighting conditions as well as dynamic objects. Each scene contains 1080 consecutive frames, of which 960 frames are used for training and 120 frames are used for testing. In order to demonstrate the high-quality rendering capability of the method of the present invention, two scenes from UE5 are rendered using real-time ray tracing to achieve photorealism. The resolution of the target high-resolution (HR) frame is set to 4K (i.e., 3840 × 2160), and the low-resolution (LR) frame is downscaled according to the scaling factor. All frames are generated using Unreal Engine 4 [2020b] and Unreal Engine 5
[2021] , and the required G-Buffer and pre-integrated BRDF are calculated by custom shaders.
[0064] When collecting data, it is difficult to cover all scenes, which leads to a limited number of training sets in practice. If you can generate various training data based on the existing data, you can achieve better super-resolution results. This is the purpose of data enhancement. Commonly used data enhancement techniques are: (1) Flipping: Flipping includes horizontal flipping and vertical flipping. (2) Rotation: Rotation is clockwise or counterclockwise rotation. Note that when rotating, it is best to rotate 90°-180°, otherwise there will be scale problems. (3) Cropping: Each large frame is cropped into multiple small blocks. These small blocks are used in the model training process to improve computational efficiency and the generalization ability of the model.
[0065] In the specific experiments, the network was implemented and trained using PyTorch [Paszke et al., 2019]. The training data was organized into blocks containing 8 consecutive frames, and the initial learning rate was set to 10 -4 , and gradually decays to 1×10 according to the cosine decay plan -5 All training and testing are performed on a single NVIDIA RTX 3090 GPU. Using the trained estimation model, super-resolution results are run in real time, such as Figure 3 The super-resolution imaging results of the 4x4 task shown in Figure 4 Super-resolution imaging results for the 8x8 task are shown.
[0066] The embodiment also performs performance evaluation and quality evaluation of the above method, specifically reporting the network's running time in milliseconds (ms) as a direct indicator of performance; in terms of quality, two widely used image metrics are used: peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). For PSNR and SSIM, the higher the value, the better the quality. The effect provided in an embodiment of the present invention is verified by providing these detailed metrics and data set descriptions, and the results are shown in the result performance evaluation table 1 and the running time table 2.
[0067] Table 1 Results Performance Evaluation Table
[0068]
[0069] Table 2 Operation schedule
[0070]
[0071] Ours represents the method of the present invention, and the method of the present invention is compared with several state-of-the-art super-resolution methods in academia and industry, including the SOTA single image super-resolution method LIIF [Chen et al. 2021], the real-time rendering super-resolution method FusesSR [Zhong et al. 2023], NSRR [Xiao et al. 2020], MNSS [Yang et al. 2023], and methods widely used in the gaming industry, including AMD's FidelityFX™ Super Resolution (FSR) [AMD 2021] and Intel Xe Super Sampling (XeSS) [Intel 2022]. 8x represents the 8x8 super-resolution task, and Ours, FusesSR, NSRR, and MNSS without a suffix represent the 4x4 super-resolution task.
[0072] Table 1 shows the method of the present invention ( Figure 2 The evaluation results of PSNR and SSIM of 4x4 and 8x8 super-resolution imaging tasks in different test scenarios are shown in the model (shown), and compared with other methods. The comparison results show that the method of the present invention is significantly better than other existing methods.
[0073] Table 2 compares the total runtime of the super-resolution task and the high-resolution G-Buffer generation time (in milliseconds) for different target resolutions. The results are tested on an NVIDIA RTX 3090 GPU and show that the runtime consumes less time than other methods.
[0074] like Figure 5 As shown, based on the same inventive concept, an embodiment also provides a super-resolution rendering device 50 that integrates non-aligned spatiotemporal information, including a data acquisition unit 51, a real-time estimation unit 52, and a real-time synthesis unit 53; wherein the data acquisition unit 51 is used to acquire the low-resolution color image and high-resolution geometric buffer data of the current frame in real time; the real-time estimation unit 52 is used to estimate the anti-aliasing bidirectional reflection distribution function value, pseudo-illumination value, and pseudo-visibility value of the current frame in real time based on the high-resolution geometric buffer data, the pre-integrated bidirectional reflection distribution function value, and the low-resolution color image through a non-aligned spatiotemporal feature fusion method; the real-time synthesis unit 53 is used to obtain the high-resolution color image of the current frame based on the multiplication of the anti-aliasing bidirectional reflection distribution function value, the pseudo-illumination value, and the pseudo-visibility value.
[0075] It should be noted that the super-resolution rendering device for integrating non-aligned spatiotemporal information provided in the above-mentioned embodiment should be illustrated by the division of the above-mentioned functional units when performing the super-resolution rendering method. The above-mentioned functions can be assigned to different functional units as needed, that is, the internal structure of the terminal or server is divided into different functional units to complete all or part of the functions described above. In addition, the super-resolution rendering device for integrating non-aligned spatiotemporal information provided in the above-mentioned embodiment and the embodiment of the super-resolution rendering method for integrating non-aligned spatiotemporal information belong to the same concept. The specific implementation process is detailed in the embodiment of the super-resolution rendering method for integrating non-aligned spatiotemporal information, which will not be repeated here.
[0076] Based on the same inventive concept, an embodiment further provides a computing device, including a memory and one or more processors, wherein an executable code is stored in the memory, and when the one or more processors execute the executable code, the method for super-resolution rendering that integrates non-aligned spatiotemporal information is implemented, specifically including the following steps:
[0077] S1, real-time acquisition of low-resolution color image and high-resolution geometry buffer data of the current frame;
[0078] S2, based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image, the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value of the current frame are estimated in real time by non-aligned spatiotemporal feature fusion;
[0079] S3, obtaining a high-resolution color image of the current frame based on multiplying the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
[0080] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the super-resolution rendering method of integrating non-aligned spatiotemporal information described in S1-S3 above. Of course, in addition to the software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0081] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the super-resolution rendering method for integrating non-aligned spatiotemporal information is implemented, which specifically includes the following steps:
[0082] S1, real-time acquisition of low-resolution color image and high-resolution geometry buffer data of the current frame;
[0083] S2, based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image, the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value of the current frame are estimated in real time by non-aligned spatiotemporal feature fusion;
[0084] S3, obtaining a high-resolution color image of the current frame based on multiplying the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
[0085] In the embodiment, computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.
[0086] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A super-resolution rendering method integrating non-aligned spatiotemporal information, characterized in that: The following steps are involved: Get the low-resolution color image and high-resolution geometry buffer data of the current frame in real time; Based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image, the anti-aliased bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value of the current frame are estimated in real time by non-aligned spatiotemporal feature fusion. The pseudo illumination value represents the integrated illumination information that has been decoupled from the shadow. It means that the calculation method is: ; in, Represents the radiance generated by the engine without calculating shadows. Represents the bidirectional reflectance distribution function value for anti-aliasing; The pseudo-visibility value represents the integral visibility information, using It means that the calculation method is: ; in, represents the outgoing radiance; The bidirectional reflectance distribution function value for anti-aliasing is It means that the calculation method is: ; in, and denote the incident and outgoing directions respectively, represents the bidirectional reflectance distribution function at the shading point, is the angle of incidence The cosine term of ; The high-resolution color image of the current frame is obtained by multiplying the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
2. The super-resolution rendering method for integrating non-aligned spatiotemporal information according to claim 1, characterized in that: The pre-integrated bidirectional reflectance distribution function value is obtained in real time and then used in the estimation calculation, including: The bidirectional reflectance distribution function values at different roughness values determined by the viewing angle direction determined by each combination of normal and illumination direction are calculated in advance and stored. When applied in real time, the pre-integrated bidirectional reflectance distribution function value of the current frame is queried from the storage according to the roughness and viewing angle direction of each pixel provided by the high-resolution geometry buffer data of the current frame.
3. The super-resolution rendering method for integrating non-aligned spatiotemporal information according to claim 1, characterized in that: An estimation model constructed according to a convolutional neural network is used based on high-resolution geometric buffer data, pre-integrated bidirectional reflectance distribution function values, and low-resolution color images, and the anti-aliasing bidirectional reflectance distribution function values, pseudo-illumination values, and pseudo-visibility values of the current frame are estimated in real time through non-aligned spatiotemporal information. First, based on the high-resolution geometric buffer data and the pre-integrated bidirectional reflectance distribution function values, the dimensions are reduced to low resolution and then feature fused with the low-resolution color image in the low-resolution space. Then, the fused features are subjected to information enhancement extraction. During the information enhancement extraction, multi-layer feature accumulation is performed and different mixing coefficients are used to mix the corresponding features of the same layer of the previous frame. Finally, based on the features after the information enhancement extraction, the dimensions are upgraded to high resolution and then the anti-aliasing bidirectional reflectance distribution function values, pseudo-illumination values, and pseudo-visibility values of the current frame are predicted and estimated.
4. The super-resolution rendering method for fusing non-aligned spatiotemporal information according to claim 3, characterized in that: The estimation model includes a spatiotemporal encoding module, an information enhancement module, and a decoding module. The spatiotemporal coding module is used to mix the high-resolution auxiliary features composed of the high-resolution geometric buffer data and the pre-integrated bidirectional reflectance distribution function value with the previous frame of high-resolution color image according to a certain mixing coefficient and reduce the dimension to a low-resolution space, and then perform spatiotemporal coding with the extracted features of the low-resolution color image to obtain the coding features; The information enhancement module is used to extract multi-layer features based on the coding features, and use different mixing coefficients to mix the corresponding features of the same layer of the previous frame when extracting certain layer features; The decoding module is used to fuse the features output by the information enhancement module with the extracted features of the low-resolution color image and upgrade the dimension to the high-resolution space, and then estimate the anti-aliasing bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame through decoding.
5. The super-resolution rendering method for fusing non-aligned spatiotemporal information according to claim 3 or 4, characterized in that: The dimensionality reduction process uses a pixel anti-shuffle operation, and the dimensionality increase process uses a pixel shuffle operation.
6. The super-resolution rendering method for fusing non-aligned spatiotemporal information according to claim 3 or 4, characterized in that: When using different mixing coefficients to mix the corresponding features of the same layer of the previous frame, a small mixing coefficient is used for fast-moving objects, and a large mixing coefficient is used for slow-moving objects.
7. The super-resolution rendering method integrating non-aligned spatiotemporal information according to claim 3, characterized in that: The convolutional neural network undergoes parameter optimization before performing estimation calculations. The loss functions used in parameter optimization include first paradigm loss, perceptual loss, and structural similarity loss based on the generated high-resolution color image and the high-resolution real image, as well as temporal stability loss based on the optical flow between two adjacent frames of high-resolution color images.
8. A super-resolution rendering device integrating non-aligned spatiotemporal information, characterized in that: It includes a data acquisition unit, a real-time estimation unit, and a real-time synthesis unit; The data acquisition unit is used to acquire the low-resolution color image and high-resolution geometric buffer data of the current frame in real time; The real-time estimation unit is used to estimate the anti-aliasing bidirectional reflectance distribution function value, pseudo illumination value, and pseudo visibility value of the current frame in real time based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image by non-aligned spatiotemporal feature fusion method; wherein the pseudo illumination value represents the integrated illumination information that has been decoupled from the shadow, and the pseudo visibility value is used to estimate the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value of the current frame in real time based on the high-resolution geometry buffer data, the pre-integrated bidirectional reflectance distribution function value, and the low-resolution color image by non-aligned spatiotemporal feature fusion method; wherein the pseudo illumination value represents the integrated illumination information that has been decoupled from the shadow, and the pseudo visibility value is used to estimate the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value represents the integrated illumination information that has been decoupled from the shadow, and the It means that the calculation method is: ; in, Represents the radiance generated by the engine without calculating shadows. Represents the bidirectional reflectance distribution function value for anti-aliasing; The pseudo-visibility value represents the integral visibility information, using It means that the calculation method is: ; in, represents the outgoing radiance; The bidirectional reflectance distribution function value for anti-aliasing is It means that the calculation method is: ; in, and denote the incident and outgoing directions respectively, represents the bidirectional reflectance distribution function at the shading point, is the angle of incidence The cosine term of ; The real-time synthesis unit is used to obtain a high-resolution color image of the current frame based on multiplication of the anti-aliasing bidirectional reflectance distribution function value, the pseudo illumination value, and the pseudo visibility value.
9. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the super-resolution rendering method that integrates non-aligned spatiotemporal information according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the super-resolution rendering method for fusing non-aligned spatiotemporal information according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Dental 3D scanner with angular-based shade matching
CN112912933A
Convolutional neural network-based real-time super-resolution method using extra rendering information
CN114820327A