Real-time super-resolution rendering method and device based on frame loop network
By integrating a frame-recurrent network with super-sampling algorithms and utilizing scene geometry data, the method addresses the challenge of applying video super-resolution in real-time rendering, achieving improved temporal stability and rendering quality.
Patent Information
- Application Number
- CN202510394469.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video super-scoring algorithms and real-time rendering technologies are difficult to directly apply to the real-time rendering pipeline, making it difficult to meet the needs of real-time rendering graphics programs in high resolution and high refresh rates.
The real-time rendering pipeline data is integrated with the supersampling algorithm, and through the frame loop network, the sliding window and attention weighting mechanism are combined with the characteristics of the reconstruction result of the previous frame to improve the image reconstruction of the current frame, improve the perception of three-dimensional space information and reduce artifacts.
It realizes the generation of high-resolution images in real-time rendering, improving time stability and image quality, especially in complex scenes.
Smart Images

Figure CN120318079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of graphics rendering and image super-resolution, and particularly to a real-time rendering super-resolution method and apparatus based on a frame recurrent network. Background Art
[0002] Real-time rendering graphics programs (such as games, virtual reality, etc.) have increasingly high requirements for high resolution and high refresh rate, and real-time super-resolution technology for rendering images has become very necessary in real-time rendering. However, existing video super-resolution algorithms and real-time rendering are in different data processing pipelines, which makes it difficult to directly apply them to the real-time rendering pipeline. Summary of the Invention
[0003] In view of the above technical problems, the present invention provides a real-time rendering super-resolution method and apparatus based on a frame recurrent network, which integrates real-time rendering pipeline data with a super-sampling algorithm to make the algorithm reconstruction result more excellent.
[0004] In a first aspect of the present invention, there is provided a real-time rendering super-resolution method based on a frame recurrent network, which is applicable to a real-time rendering graphics program and includes the following steps.
[0005] Taking a super-sampling network model as a basic architecture, obtaining input data of the network from a real-time rendering pipeline, and inputting the input data into the super-sampling network model for image reconstruction;
[0006] Introducing the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map, and introducing a fixed number of historical frames into the image reconstruction process of the current frame based on the sliding window to obtain a second feature map;
[0007] Merging and inputting the first feature map and the second feature map into a U-NET network to generate a target graphic.
[0008] In an optional embodiment, the introducing the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map includes:
[0009] Performing a feature extraction operation on the reconstruction result of the previous frame to obtain a third feature map, performing a reprojection operation on the third feature map according to the motion vector map of the current frame, extracting effective features for the reconstruction result from the third feature map through an attention weighting mechanism, and introducing the effective features into the image reconstruction process of the current frame to obtain a first feature map.
[0010] In an optional embodiment, the extracting effective features for the reconstruction result from the third feature map through an attention weighting mechanism includes:
[0011] Generate a weighted graph using the previous frame, the motion vector graph, the scene normal vector graph, and the scene depth graph before the current frame, and multiply the weighted graph by the feature graph extracted from the reconstruction result of the previous frame to extract the effective features.
[0012] In an alternative embodiment, the first feature map is obtained based on the following formula:
[0013]
[0014] is the reconstructed image of the previous frame, represents the feature information extracted from the reconstructed image, B represents the reprojection of the feature map to the current frame, and SRAttention represents the operation of performing attention weighting on the feature map.
[0015] In an alternative embodiment, introducing a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map, including:
[0016] Extract the fourth feature map of a fixed number of historical frame graphics, perform a reprojection operation on the fourth feature map, and the fourth feature map will be evaluated by a reweighting network to mask invalid features caused by dynamic occlusion, and the second feature map is obtained after evaluation.
[0017] In an alternative embodiment, the fourth feature map is evaluated by the reweighting network to mask invalid features caused by dynamic occlusion, including:
[0018] Generate a corresponding weighting matrix for each input fourth feature, and the weighting matrix represents the dynamic occlusion relationship between the fourth feature and the object in the current frame;
[0019] Multiply the fourth feature by the corresponding weighting matrix to obtain the second feature map.
[0020] In an alternative embodiment, the second feature map is obtained based on the following formula:
[0021]
[0022]
[0023] represents the historical frame image, and LRReweight represents the operation of performing reweighting on the feature map of the historical frame image; represents the feature information extracted from the reconstructed image, and B represents the reprojection of the feature map to the current frame.
[0024] In an alternative embodiment, the loss function of the super-sampling network model is:
[0025]
[0026] wherein I i SR represents the reconstruction result of the i-th frame, and I i HR represents the true high-resolution image of the i-th frame. The value of w is set to 0.1, and SSIM represents the structural similarity index.
[0027] In a second aspect of the present invention, a real-time rendering super-resolution device based on a frame recurrent network is provided, including:
[0028] An input module, which uses a super-sampling network model as a basic architecture, obtains input data of the network from a real-time rendering pipeline, and inputs the input data into the super-sampling network model for image reconstruction;
[0029] A reconstruction module, which introduces the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map, and introduces a fixed number of historical frames into the image reconstruction process of the current frame to obtain a second feature map;
[0030] An output module, which combines and inputs the first feature map and the second feature map into a U-NET network to generate a target graphic.
[0031] In a third aspect of the present invention, an electronic device is provided, including:
[0032] At least one processor and / or graphics processor; and at least one memory communicatively connected to the processor and / or the graphics processor, wherein: the memory stores program instructions executable by the processor and / or the graphics processor, and the processor or the graphics processor can execute the real-time rendering super-resolution method based on a frame recurrent network according to any one of claims 1 to 8 by invoking the program instructions.
[0033] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is run by a computer, it executes the real-time rendering super-resolution method based on a frame recurrent network according to the first aspect of the embodiments of the present invention.
[0034] The present invention utilizes the low-resolution scene geometry data generated in the real-time rendering pipeline to enhance the network's perception of three-dimensional space information; combines the frame recurrent network into the super-sampling algorithm, and improves the result of the current frame by introducing the features of the previous frame's reconstruction result, thereby achieving stability on the time scale.
[0035] The present invention also places an attention network in the feature extraction module of the previous reconstruction result to improve the effectiveness of the extracted features. Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a schematic flow chart of a real-time rendering super-resolution method based on a frame loop network in an embodiment of the present invention.
[0038] Figure 2 It is a network diagram of a real-time rendering super-resolution method based on a frame loop network in an embodiment of the present invention.
[0039] Figure 3 It is the first scenario test chart for comparison with other super-resolution methods in an embodiment of the present invention.
[0040] Figure 4 It is the second scenario test chart for comparison with other super-resolution methods in an embodiment of the present invention.
[0041] Figure 5 It is a box plot of the SSIM index for the scenario test of comparison with other super-resolution methods in an embodiment of the present invention.
[0042] Figure 6 It is a schematic diagram of a real-time rendering super-resolution device based on a frame loop network in an embodiment of the present invention. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0044] It should be understood that the terms "first", "second", and "third", etc. in the claims, the description, and the drawings of the present invention are used to distinguish different objects, rather than to describe a specific order. The terms "including" and "comprising" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0045] It should also be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification and claims of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should be further understood that the term " / and" used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0046] In the prior art, in the case of limited device conditions such as mobile terminals, the 3D scene intelligently renders low-resolution images, which no longer meets the market demand; inserting the present invention into the rendering process can obtain high-resolution images.
[0047] Please refer to Figure 1 , the present invention provides a real-time rendering super-resolution method based on a frame recurrent network, which is applicable to a real-time rendering graphics program and includes the following steps.
[0048] Step 100: Using a super-sampling network model as the basic architecture, obtain the input data of the network from the real-time rendering pipeline, and input the input data into the super-sampling network model for image reconstruction.
[0049] For an image video, the information recorded in each image frame is collected within a certain time period, while real-time rendering is an instantaneous image rendered by point sampling based on the pixels in the screen space at a certain time point. In real-time rendering, since it is impossible to directly perform an integration operation on the physically based rendering equation, the graphics program often discretely samples the ambient light in the virtual space to fit the effect of the object under the real environmental light. Therefore, the super-resolution task related to real-time rendering is often more challenging than video.
[0050] Therefore, the present invention proposes a rendering framework for real-time neural super-sampling, which can enable the super-sampling algorithm to make full use of the low-resolution scene geometry data generated in the real-time rendering pipeline to enhance the network's perception of the three-dimensional space scene information. In the present invention, the deferred rendering technology is adopted to improve the rendering rate of complex scenes. Different from the forward rendering technology (Forward), this deferred rendering technology splits the shading calculation of the object into two rendering processes.
[0051] One is the GbufferPass process, which generates a set of textures that record scene geometry information equivalent to the screen size. The other is the LightPass process, which traverses each pixel of the texture and uses the geometric data recorded by these pixels to perform physically based shading calculations. Inputting these data into each module of the neural super-sampling network can improve the adaptability of the entire network to real-time rendering super-sampling and the superiority of the reconstruction results. Therefore, integrating the data in the deferred rendering pipeline into the neural super-sampling network can enable the network to perceive richer three-dimensional scene information during training and evaluation.
[0052] Step 200: Introduce the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map. Introduce a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map.
[0053] Step 300: Merge and input the first feature map and the second feature map into a U-NET network to generate a target graph.
[0054] The present invention proposes a network framework based on unidirectional frame cycling. For real-time rendering super-sampling tasks, features of the reconstruction results of previous frames are introduced during the reconstruction process of the current frame to improve temporal stability.
[0055] The network framework based on unidirectional frame cycling extracts effective features from the reconstruction results of previous frames through an attention weighting mechanism and adds them to the reconstruction process of the current frame to obtain a first feature map; then uses a feature re-weighting module to filter out invalid features brought by dynamic occlusion in the historical frame images to obtain a second feature map.
[0056] Finally, use a U-NET network to generate the final reconstruction result based on the first feature map and the second feature map.
[0057] The following will be combined with Figure 2 Specifically illustrate the steps of the present invention.
[0058] The super-resolution algorithm based on a sliding window regards the target task as a super-resolution problem between multiple independent frames. These methods often process multiple historical frame images with a sliding window of a fixed length to generate the final result. However, the computational cost of this method is very high because each historical frame will be processed multiple times under the sliding window framework. In addition, regarding each reconstruction result as an independently output frame will greatly reduce the temporal stability between the result frames, resulting in unpredictable artifacts and flickers.
[0059] In order to improve the temporal stability of the network output result while reducing the artifact effect caused by the accumulation of frame cycle errors, we introduce and redesigned the frame cycle network as follows.
[0060] In the above step 200, the process of introducing the previous frame into the current frame for image reconstruction based on a sliding window to obtain a first feature map includes the following steps.
[0061] Perform a feature extraction operation on the reconstruction result of the previous frame to obtain a third feature map, perform a reprojection operation on the third feature map according to the motion vector map of the current frame, extract effective features for the reconstruction result from the third feature map through an attention weighting mechanism, and introduce the effective features into the image reconstruction process of the current frame to obtain a first feature map.
[0062] In the above step, the frame loop network performs a feature extraction operation on the reconstruction result of the previous frame, then performs a reprojection operation on the above features according to the motion vector map of the current frame, and then extracts the effective part for the reconstruction result from the feature map through an attention weighting module.
[0063] Specifically, for the feature information extracted and reprojected on the reconstruction result of the previous frame, it itself has an error estimated by the network. The present invention uses the motion vector map, scene normal vector map, and scene depth map before the previous frame and the current frame to generate a weighting map, and multiplies the weighting map by the feature map extracted from the reconstruction result of the previous frame to extract the effective features.
[0064] An attention weighting module is set in the frame loop network. The attention weighting module first uses the motion vector map, scene normal vector map, and scene depth map two frames before to generate a weighting map, and then multiplies the weighting map by the feature map extracted from the reconstruction result of the previous frame to extract the effective features.
[0065] In some embodiments, the first feature map is obtained based on the following formula:
[0066]
[0067] is the reconstructed image of the previous frame, represents the feature information extracted from the reconstructed image, B represents reprojection of the feature map to the current frame, and SRAttention represents the operation of performing attention weighting on the feature map.
[0068] Then, a fixed number of historical frames are introduced into the reconstruction result of the current frame based on a sliding window. In the above step 200, the process of introducing a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map includes the following steps.
[0069] Extract the fourth feature map of a fixed number of historical frame graphics, perform a reprojection operation on the fourth feature map, and the reweighting network will evaluate the fourth feature map to mask the invalid features in the motion vector map. After evaluation, the second feature map is obtained.
[0070] In the above steps, in each cycle of the sliding window, a fixed number of historical frame graphics are input into the frame loop part network and the feature maps of these historical frame images are extracted. Then, a reprojection operation is performed on the feature maps. After that, the reweighting network will evaluate the feature maps again to weaken the negative impact of the invalid feature information caused by dynamic occlusion on the reconstruction result.
[0071] Specifically, for each input fourth feature, a corresponding weighting matrix is generated, and the weighting matrix represents the dynamic occlusion relationship between the fourth feature and the object in the current frame; then, by multiplying the fourth feature by the corresponding weighting matrix, the second feature map is obtained.
[0072] The present invention sets up a reweighting network, and introduces the low-resolution scene normal information generated in the rendering pipeline into the reweighting network. Compared with only using the shading map and depth map, the scene normal information can not only show different detail information on the surface of the same object, but also reflect the different orientations between different objects at the same depth. The reweighting network generates a corresponding weighting matrix for each input historical frame feature, and the weighting matrix reflects the dynamic occlusion relationship between the target historical frame and the object in the current frame. After multiplying the feature information of the historical frame by the weighting matrix, the possible artifacts can be effectively masked. The reweighting network is processed by a three-layer convolutional neural network, and all the historical frame feature information after the reprojection operation is used as the input. The reweighting network predicts the weight for each historical frame, and then multiplies it by the feature map of each corresponding historical frame, so as to obtain the output result of the reweighting network and obtain the second feature map.
[0073] In some embodiments, the second feature map is obtained based on the following formula:
[0074]
[0075] represents the historical frame image, and LRReweight represents performing a reweighting operation on the feature map of the historical frame image; represents the feature information extracted from the reconstructed image, and B represents reprojection of the feature map to the current frame.
[0076] Furthermore, the loss function of the supersampling network model is:
[0077]
[0078] Where Ii SR Represents the reconstruction result of the i-th frame, I i HR Represents the true high-resolution image of the i-th frame. The value of w is set to 0.1, and SSIM represents the structural similarity index.
[0079] SSIM can measure the similarity between two images, thereby evaluating the image reconstruction. Compared with the traditional image quality measurement metrics, structural similarity belongs to a perceptual model, and its metrics in image quality measurement are more in line with the human eye's judgment of image quality.
[0080] Is the perceptual loss calculated from the pre-trained VGG-16 network. Using the pre-trained VGG-16 network can compare the high-frequency features obtained by convolving the reconstructed image with those obtained by convolving the real image, making their high-frequency information similar.
[0081] As can be seen from the above, the present invention utilizes the low-resolution scene geometry data generated in the real-time rendering pipeline to enhance the network's perception of three-dimensional space information; combines the frame recurrent network into the super-sampling algorithm, and improves the result of the current frame by introducing the features of the previous frame's reconstruction result, thereby achieving stability on the time scale. The present invention also places the attention network in the feature extraction module of the previous reconstruction result to improve the effectiveness of the extracted features.
[0082] The present invention can further optimize the areas in the image that are prone to artifacts from the generation principle of the rendered image. For example, for dynamic shadows and specular reflections, ordinary motion vector maps cannot capture their trends that change with the movement of objects and light sources. It may be necessary to customize the rendered pipeline to output a vector map that can describe this change to solve this problem.
[0083] By comparing with other models, the performance of the present invention is better. See the following table for details:
[0084] Table 1 Evaluation results of super-resolution of different methods in different 3D scenes
[0085]
[0086] Table 2 Model parameter sizes and running times of different methods (image size 320×180, upsampling factor 4×4, model accuracy is all half)
[0087]
[0088]
[0089] Table 3 Running Time of Each Module (Image Size 320×180, Upsampling Factor 4×4, Model Precision is Half)
[0090]
[0091] In Table 1, we compared the average values of three evaluation metrics (PSNR, SSIM, LPIPS) with other super-resolution methods (VESPCN, EDSR, TecoGAN, RBPN, NSRR). For each of these scenarios, the method was evaluated using 25 consecutively rendered frame images, and the average value of the obtained reconstruction results was obtained. Table 2 shows the comparison of the network model parameter sizes and running times between the present invention and other methods. Table 3 shows the running times of each module within the present invention.
[0092] Table 1 shows that the present invention is superior to other super-resolution methods in most cases. Especially when compared with NSRR, the present invention shows more excellent results in all three metrics. Among them, the PSNR is increased by an average of 0.45dB, the SSIM is increased by an average of 0.015, and the LPIPS is decreased by an average of 0.025. It can be seen from Tables 2 and 3 that compared with other methods, the super-sampling network model we designed not only achieves significant results but also effectively reduces the evaluation time, so it has more real-time application value.
[0093] Figure 3 and Figure 4 are test images of different scenarios. The reconstruction results of one frame extracted from the continuously rendered frame image sequence are directly used for visual effect comparison, and the evaluation metric values obtained by each super-resolution method are marked below the reconstruction results. In Figure 3 , the rendered images captured by the camera in the indoor three-dimensional scene are selected for testing. In Figure 4 , the rendered images captured by the camera in the outdoor three-dimensional scene are tested, and two different time periods are selected.
[0094] From Figure 3 it can be seen that when the objects in the three-dimensional scene are stacked simply or the colors are relatively uniform, the evaluation metric values of the reconstruction results obtained by using the present invention do not have a particularly large advantage compared with other video super-resolution methods. However, as the complexity of the scene in Figure 4 increases, the super-sampling network of the present invention can make full use of the low-resolution scene depth map and scene normal vector map provided by the real-time rendering pipeline to perceive the spatial information of the three-dimensional scene, thereby reconstructing better results. In addition, it can also be directly seen from the comparison with the reconstruction results of other methods that the present invention has a significant improvement in the anti-aliasing performance at the edges and the overall clarity of the image.
[0095] Figure 5It shows the box plot of the scenario test under the SSIM evaluation metric. Each element in each box plot represents the set of evaluation results of the super-sampling of a group of 25 consecutive rendered images; it can be seen from the size and position of the box that compared with other super-resolution methods, the present invention obtains better evaluation metrics as a whole, and its data is relatively more concentrated, especially in comparison with other real-time super-resolution algorithms. This further verifies that the frame loop network framework introduced in the present invention has significantly improved the temporal stability between the super-sampling results of consecutive frames.
[0096] Please refer to Figure 6 , the present invention also provides a real-time rendering super-resolution device based on a frame loop network, including: an input module 51, a reconstruction module 52, and an output module 53.
[0097] The input module 51 is used to use the super-sampling network model as the basic architecture, obtain the input data of the network from the real-time rendering pipeline, and input the input data into the super-sampling network model for image reconstruction. The input module 51 is a GPU in this embodiment, and the input is the rendering data of the GPU.
[0098] In some embodiments, a third feature map is obtained by performing a feature extraction operation on the reconstruction result of the previous frame, the third feature map is reprojected according to the motion vector map of the current frame, effective features for the reconstruction result are extracted from the third feature map through an attention weighting mechanism, and the effective features are introduced into the image reconstruction process of the current frame to obtain a first feature map. For the effective features, a weighted map is generated by using the previous frame, the motion vector map before the current frame, the scene normal vector map, and the scene depth map, and the weighted map is multiplied by the feature map extracted from the reconstruction result of the previous frame to extract the effective features.
[0099] The reconstruction module 52 is used to introduce the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map, and introduce a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map.
[0100] In some embodiments, a fourth feature map of a fixed number of historical frame graphics is extracted, the fourth feature map is reprojected, and the reweighting network will evaluate the fourth feature map to mask the invalid features in the motion vector map, and the second feature map is obtained after evaluation. A corresponding weighted matrix is generated for each input fourth feature, and the weighted matrix represents the dynamic occlusion relationship between the fourth feature and the object in the current frame; the second feature map is obtained by multiplying the fourth feature by the corresponding weighted matrix.
[0101] The output module 53 is used to merge the first feature map and the second feature map and input them into the U-NET network to generate a target graphic.
[0102] For the description of the real-time rendering super-resolution device based on the frame recurrent network, reference may be made to the above real-time rendering super-resolution method based on the frame recurrent network.
[0103] The present invention also provides an electronic device, comprising:
[0104] at least one processor and / or graphics processor; and at least one memory communicatively connected to the processor and / or the graphics processor, wherein: the memory stores program instructions executable by the processor and / or the graphics processor, and the processor or the graphics processor can execute the real-time rendering super-resolution method according to any one of claims 1 to 8 by invoking the program instructions.
[0105] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the real-time rendering super-resolution method based on the frame recurrent network described above is implemented.
[0106] It can be understood that the computer-readable storage medium may include: any entity or device capable of carrying a computer program, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), and a software distribution medium, etc. The computer program includes computer program code. The computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), and a software distribution medium, etc.
[0107] In certain embodiments of the present invention, the electronic device may include a controller or a processor. The controller is a single-chip microcomputer chip that integrates a processor, a memory, a communication module, etc. The processor may refer to the processor included in the controller. The processor may be a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0108] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0109] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time rendering super-resolution method based on a frame recurrent network, applicable to real-time rendering graphics programs, characterized in that, Including: Taking the super-sampling network model as the basic architecture, obtaining the input data of the network from the real-time rendering pipeline, and inputting the input data into the super-sampling network model for image reconstruction; Introducing the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map, and introducing a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map; Merging the first feature map and the second feature map and inputting them into a U-NET network to generate a target graph.
2. The real-time rendering super-resolution method based on a frame cyclic network according to claim 1, wherein The step of introducing the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map includes: Performing a feature extraction operation on the reconstruction result of the previous frame to obtain a third feature map, performing a reprojection operation on the third feature map according to the motion vector map of the current frame, extracting effective features for the reconstruction result from the third feature map through an attention weighting mechanism, and introducing the effective features into the image reconstruction process of the current frame to obtain a first feature map.
3. The real-time rendering super-resolution method based on a frame recurrent network according to claim 2, wherein The step of extracting effective features for the reconstruction result from the third feature map through an attention weighting mechanism includes: Using the previous frame, the motion vector map before the current frame, the scene normal vector map, and the scene depth map to generate a weighting map, and multiplying the weighting map by the feature map extracted from the reconstruction result of the previous frame to extract the effective features.
4. The real-time rendering super-resolution method based on a frame recurrent network according to claim 3, wherein Obtaining the first feature map based on the following formula: is the reconstructed image of the previous frame, represents the feature information extracted from the reconstructed image, B represents the reprojection of the feature map to the current frame, and SRAttention represents the operation of performing attention weighting on the feature map.
5. The real-time rendering super-resolution method based on a frame cyclic network according to claim 1 or 2, characterized in that The step of introducing a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map includes: Extracting a fourth feature map of a fixed number of historical frame graphs, performing a reprojection operation on the fourth feature map, and evaluating the fourth feature map through a reweighting network to mask invalid features caused by dynamic occlusion, and obtaining the second feature map after evaluation.
6. The real-time rendering super-resolution method based on a frame recurrent network according to claim 5, wherein, The step of evaluating the fourth feature map through a reweighting network to mask invalid features caused by dynamic occlusion includes: Generating a corresponding weighting matrix for each input fourth feature, and the weighting matrix represents the dynamic occlusion relationship between the fourth feature and the object in the current frame; Obtaining the second feature map by multiplying the fourth feature by the corresponding weighting matrix.
7. The real-time rendering super-resolution method based on a frame recurrent network according to claim 6, characterized in that, Obtaining the second feature map based on the following formula: represents the historical frame image, and LRReweight represents performing a reweighting operation on the feature map of the historical frame image; represents the feature information extracted from the reconstructed image, and B represents reprojection of the feature map onto the current frame.
8. The real-time rendering super-resolution method based on a frame cyclic network according to claim 1, wherein The loss function of the super-sampling network model is: Among them, I i SR represents the reconstruction result of the i-th frame, I i HR represents the true high-resolution image of the i-th frame. The value of w is set to 0.1, and SSIM represents the structural similarity index.
9. A real-time rendering super-resolution device based on a frame recurrent network, which is used to execute a real-time rendering graphics program, and is characterized in that, Including: An input module for taking the super-sampling network model as the basic architecture, obtaining the input data of the network from the real-time rendering pipeline, and inputting the input data into the super-sampling network model for image reconstruction; A reconstruction module for introducing the previous frame into the image reconstruction process of the current frame based on a sliding window to obtain a first feature map, and introducing a fixed number of historical frames into the image reconstruction process of the current frame based on a sliding window to obtain a second feature map; An output module for merging the first feature map and the second feature map and inputting them into a U-NET network to generate a target graph.
10. An electronic device, characterized in that, Including: At least one processor and / or graphics processor; and at least one memory communicatively connected to the processor and / or the graphics processor, wherein: the memory stores program instructions executable by the processor and / or the graphics processor, and the processor or the graphics processor can execute the real-time rendering super-resolution method based on a frame loop network according to any one of claims 1 to 8 by invoking the program instructions.
11. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is run by a computer, it executes the real-time rendering super-resolution method based on a frame loop network according to any one of claims 1 to 8.