A rendering optimization method and program product based on interpolation and super-resolution, and a storage medium
By using reprojection and correction techniques, features are extracted and super-resolution rendering frames are reconstructed, solving the problems of high computational load and long running time in the rendering process in existing technologies, and achieving efficient and high-quality rendering effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-09
AI Technical Summary
Existing technologies involve large amounts of computation and long execution time during the rendering process, and may result in distorted image content or errors in object generation, making it impossible to efficiently generate high-quality, high-resolution rendering frames.
By acquiring optimized super-resolution rendering frames and low-resolution rendering frames, reprojection and correction are performed, features are extracted and combined with high-resolution intermediate frame data to reconstruct super-resolution intermediate rendering frames, and intermediate frames are inserted into the frame display sequence to update low-resolution frames into super-resolution frames.
It achieves fast and high-quality frame generation and super-resolution calculation, reduces the amount of computation, generates super-resolution intermediate frames and super-resolution rendering frames with correct geometric structure and color changes, and shortens the running time.
Smart Images

Figure CN122175785A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer vision technology, and more particularly to a rendering optimization method, program product, and storage medium based on frame interpolation and super-resolution. Background Technology
[0002] The development of electronic technology has led to the widespread use of devices with graphics display capabilities, and has also given rise to many works and demands centered around computer-generated visual effects, such as in the video game industry and the film and television special effects industry. As creators and users continue to demand high-quality visuals and effects, the performance of related devices is gradually becoming insufficient to smoothly render and display the images. This problem arises partly because the content to be rendered in each frame is more complex, and partly because the increased resolution of each frame leads to an increase in the number of pixels required for color calculation, thus increasing the basic computational load.
[0003] Among related technologies, frame generation and super-resolution techniques can ensure the smoothness and quality of the image. Frame generation technology refers to predicting intermediate frames from existing rendered frames and inserting them into the frame display sequence, shortening the display time of each frame and improving the smoothness of the image. Super-resolution technology refers to generating corresponding high-resolution predicted frames from existing low-resolution rendered frames, replacing the original low-resolution rendered frames for display, and improving the clarity of the image.
[0004] However, existing methods and devices have shortcomings in terms of computational load, running time, and effect. For example, they have excessive computational load, excessive running time, distortion of image content, and incorrect generation of occluded objects. Summary of the Invention
[0005] To address the problems existing in the prior art, the purpose of this invention is to provide a rendering optimization method, program product, and storage medium based on frame interpolation and super-resolution that has a short running time and good rendering effect.
[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0007] A rendering optimization method based on frame interpolation and super-resolution includes:
[0008] Get the optimized super-resolution rendering frame at time t as the first rendering frame, and get the low-resolution rendering frame at time t+1 in the frame display sequence as the second rendering frame.
[0009] The first rendered frame and the upsampled second rendered frame are reprojected onto the intermediate frame and corrected together to obtain the corrected data of the high-definition intermediate frame.
[0010] Extract the corrected data of the first rendered frame, the high-definition intermediate frame, and the features of the second rendered frame, respectively.
[0011] The features of the first and second rendering frames are reprojected onto the intermediate frame time, and combined with the features of the high-definition intermediate frame correction data to reconstruct the super-resolution intermediate rendering frame.
[0012] The moment when the first rendering frame is reprojected onto the second rendering frame is corrected together with the upsampled second rendering frame to obtain the corrected data of the second rendering frame;
[0013] Extract features from the correction data of the second rendered frame;
[0014] Based on the characteristics of the corrected data of the second rendering frame and the characteristics of the second rendering frame, the super-resolution rendering frame at time t+1 is reconstructed.
[0015] The intermediate rendered frame is inserted after the super-resolution rendered frame at time t in the frame display sequence. The low-resolution rendered frame at time t+1 in the frame display sequence is updated to the super-resolution rendered frame at time t+1 and then displayed.
[0016] Furthermore, if the current time is the first time, then the first rendered frame is invalid data or the first low-resolution rendered frame in the frame display sequence.
[0017] Furthermore, the step of reprojecting the first rendered frame and the upsampled second rendered frame onto the intermediate frame time and correcting them together to obtain the corrected data of the high-definition intermediate frame specifically includes:
[0018] The second rendered frame is upsampled to obtain a high-definition second rendered frame;
[0019] The first rendered frame and the high-definition second rendered frame are reprojected onto the intermediate frame time respectively to obtain the first intermediate frame and the second intermediate frame;
[0020] Pixels at the same position in the first intermediate frame and the second intermediate frame are compared, selected, and / or merged to obtain the corrected data of the final high-definition intermediate frame.
[0021] Furthermore, the extraction of features from the first rendered frame, the high-definition intermediate frame correction data, and the second rendered frame specifically includes:
[0022] For the first rendered frame, the high-definition intermediate frame corrected data, and the second rendered frame, a convolutional encoder comprising n output resolutions decreasing progressively is used to extract a feature set ft={ft} of features with progressively decreasing resolutions and progressively increasing channel numbers. j | j=1,…,n}:
[0023] ft j =E j (ft j-1 ),j=1,…,n,
[0024] Among them, ftj、 ft j- E represents the features at levels j and j-1. j () represents the j-th convolutional encoder, where ft0 takes the value of the first rendered frame, or the high-definition intermediate frame correction data, or the second rendered frame.
[0025] Furthermore, the step of reprojecting the features of the first and second rendering frames onto the intermediate frame time and combining them with the features of the high-definition intermediate frame correction data to reconstruct the super-resolution intermediate rendering frame specifically includes:
[0026] The feature sets of the first rendering frame and the second rendering frame are respectively transformed to the intermediate frame time by reprojection to obtain the feature set of the first rendering frame and the feature set of the second rendering frame.
[0027] For the feature sets of the first reprojected rendering frame, the second reprojected rendering frame, and the intermediate frame correction data, starting from the nth-level feature with the lowest resolution in the feature set, a pre-trained decoder with progressively increasing resolution is used for decoding and reconstruction. During decoding, the output of the next higher-level decoder is used as input until the decoded feature output of the first-level decoder with the lowest resolution is obtained. The decoder is used to decode and reconstruct the decoded features based on the features of the first reprojected rendering frame, the second reprojected rendering frame, the high-definition intermediate frame correction data, and the next higher-level decoded features.
[0028] The pre-trained level 0 decoder decodes the decoding features output by the level 1 decoder to obtain the blending mask and color residual of the first and second reprojected rendering frames. The level 0 decoder is used to obtain the blending mask and color residual of the first and second reprojected rendering frames based on the decoding features output by the level 1 decoder.
[0029] Based on the blending mask and color residual, the super-resolution intermediate rendering frame is reconstructed.
[0030] Furthermore, the specific operations of decoding and reconstruction are as follows:
[0031]
[0032] In the formula, and for and Downsampling to a mapping relationship with the same resolution as the nth level feature. , These represent the reprojection mapping relationships at the times of the first rendering frame, the second rendering frame, and the intermediate frames, respectively. This represents the nth-level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. This represents the k-th level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. , For the nth and kth level decoding features output by the decoder, , , For the mapping residual of the decoder output, , , , For upsampling operation, These represent the functions of the nth and kth level decoders, respectively, obtained through training.
[0033] Furthermore, the calculation expressions for the blending mask and color residual are as follows:
[0034] (
[0035] In the formula, These represent the blending mask and color residual of the first and second reprojected render frames, respectively. The function representing the level-zero decoder is obtained through training. This represents the decoding characteristics output by the first-level decoder. , This represents a mapping relationship.
[0036] Furthermore, the calculation expression for the super-resolution intermediate rendering frame is as follows:
[0037]
[0038] In the formula, Indicates an intermediate rendering frame. This represents the first rendered frame of the reprojection. This indicates the second rendered frame of the reprojection.
[0039] A computer program product includes a computer program that, when executed by a processor, implements the above-described method.
[0040] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the above-described method.
[0041] Compared with the prior art, the beneficial effects of this invention are as follows: The rendering frame optimization method provided by this invention can generate super-resolution intermediate generated frames with correct geometric structure and color changes, as well as super-resolution rendering frames with correct color details, resulting in better performance. At the same time, by reusing the extracted rendering frame features between multiple executions of the method, the amount of computation is reduced, the running time is shorter, and fast, high-quality frame generation and super-resolution calculation are achieved. Attached Figure Description
[0042] Figure 1 This is a flowchart illustrating the rendering optimization method based on frame interpolation and super-resolution provided in an embodiment of the present invention.
[0043] Figure 2 This is a schematic diagram of the reuse feature extraction method and features in rendering optimization provided in this embodiment of the invention;
[0044] Figure 3 This is a schematic diagram of the reuse feature extraction method and features in the generation of intermediate rendering frames provided in this embodiment of the invention. Detailed Implementation
[0045] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0046] Example 1
[0047] This invention provides a rendering optimization method based on frame interpolation and super-resolution, such as... Figure 1 As shown, the method includes the following steps:
[0048] Step S101: Obtain the optimized super-resolution rendering frame at time t as the first rendering frame, and obtain the low-resolution rendering frame at time t+1 in the frame display sequence as the second rendering frame.
[0049] The aforementioned rendering frames typically refer to the data generated after the rendering engine software performs two complete rendering processes consecutively. The content of the frames includes data such as color and scene depth.
[0050] The aforementioned super-resolution first rendering frame refers to the super-resolution rendering frame result generated during the previous execution of the method in this embodiment. During the change of the rendered content over time, if the super-resolution first rendering frame is at a certain moment, the low-resolution second rendering frame should be at the next consecutive moment. Specifically, for consecutive moments t, t+1, t+2, and t+3, if the first rendering frame time of a certain execution of this method is t+1, then the second rendering frame time is t+2. It can also be determined that the rendering frame times of the previous execution are t and t+1, and the rendering frame times of the next execution are t+2 and t+3, respectively. For the first execution of this method, the super-resolution first rendering frame can be set to the actual high-resolution first frame, invalid data, or an upsampled low-resolution first rendering frame.
[0051] The upsampling operation described above aims to increase the resolution of the rendered frame from low-resolution to high-resolution. Common upsampling schemes such as nearest neighbor sampling and bilinear sampling can be used, or other data-driven filters, convolutional neural networks, or Transformer networks used for super-resolution. Unless otherwise specified, the upsampling operations in subsequent steps are performed in the same way.
[0052] Step S102: Reproject the super-resolution first rendering frame and the upsampled low-resolution second rendering frame to the intermediate frame time and correct them together to obtain the corrected data of the high-resolution intermediate frame.
[0053] Specifically, the second rendering frame is first upsampled to obtain a high-definition second rendering frame. Then, the first rendering frame and the high-definition second rendering frame are reprojected onto the intermediate frame time respectively to obtain the first intermediate frame and the second intermediate frame. Finally, the pixels at the same position in the first intermediate frame and the second intermediate frame are compared, selected and / or merged to obtain the correction data of the final high-definition intermediate frame.
[0054] The aforementioned intermediate frame time refers to the time between two adjacent moments in which the rendered content changes over time, with the same time interval as these two moments. For consecutive moments t and t+1, the intermediate frame time refers to moment t+0.5.
[0055] The reprojection operation described above refers to moving the frame content from one time step to another based on the relationship between the screen spatial positions of pixels at two adjacent moments. For time t, at position... A pixel, if it is at position t+1 Mapping relationship Generally expressed as The mapping relationship can be obtained from the rendering engine or through methods such as optical flow estimation. For frame content with only one channel, the reprojection operation is performed on both the width and height dimensions of the frame content; in this case, the pixel position is composed of the coordinates in the width and height directions. For frame content with multiple channels, the reprojection operation is performed independently for each channel. The frame content after the reprojection operation... ,in, For reprojection to time The frame content, For a moment The frame content, For reprojection operations, specifically referring to the time interval... Each screen space location Through the above mapping relationship, using Find its time Location and will In position The frame content as In position The frame content. Unless otherwise specified, the reprojection operation in subsequent steps is similar.
[0056] By making assumptions about the motion of pixels, the pixel position at the intermediate frame time t+0.5 can be estimated through the mapping relationship of rendering frame time. For example, assuming that the pixel moves at a constant speed, the estimated position is... and The relationship is At the same time, it can be estimated that it is related to The relationship is Based on these two relationships, two mapping relationships to intermediate frames can be generated respectively. and For ease of writing, they will be referred to as follows: and .use and The contents of the first rendered frame and the high-definition second rendered frame are reprojected onto the intermediate frame to obtain the first intermediate frame. Second intermediate frame .
[0057] When estimating the position of a pixel in an intermediate frame, multiple pixels may be in the same position. In this case, you can take the pixel closest to the rendering camera by using pixel depth, or you can select one pixel or mix multiple pixels as the result by comparing the color and depth of the pixels.
[0058] The above correction operation is used to correct the first intermediate frame. Second intermediate frame The selection can be made by comparing the depth and color of the two reprojection results, or by using a neural network to select or blend the reprojection results of the two frames. The result of the correction operation can be color information, a mask indicating the effective data in the two reprojection results, or other information guiding the sampling from the reprojection results.
[0059] Step S103: Extract the corrected data of the first rendering frame, the high-definition intermediate frame, and the features of the second rendering frame.
[0060] The aforementioned feature extraction operations can be performed using a series of encoders with progressively decreasing resolution, rule-based downsampling, or other feature extraction and matching methods.
[0061] For example, a four-level convolutional encoder with progressively decreasing resolution. Extract features and input the original resolution frame data ft0 (first rendered frame, high-definition intermediate frame corrected data, or second rendered frame) into the first-level encoder. This yields first-level features with reduced resolution and increased channel number. Then the first-level features are input into the second-level encoder. This yields second-level features with lower resolution and more channels. This process is repeated to obtain a four-level feature set where the resolution decreases progressively and the number of channels increases progressively. ,but
[0062]
[0063] To achieve a uniform resolution, the first rendered frame, the high-definition intermediate frame correction data, and the second rendered frame can also be downgraded to low-definition resolution using an additional encoder or pixel shuffle, or by using an additional encoder sequence, such as a five-level encoder combination. The encoder is trained using a conventional training method, which will not be elaborated further. The encoder's network structure can be a conventional encoder network structure.
[0064] Step S104: Reproject the features of the first and second rendering frames to the intermediate frame time, and combine them with the features of the high-definition intermediate frame correction data to reconstruct the super-resolution intermediate rendering frame.
[0065] Specifically, the feature sets of the first and second rendering frames are transformed to the intermediate frame time using reprojection, resulting in the reprojected feature sets of the first and second rendering frames. For the feature sets of the reprojected first and second rendering frames, and the feature sets of the intermediate frame correction data, starting from the nth-level feature with the lowest resolution, a pre-trained decoder with progressively increasing resolution is used for decoding and reconstruction. During decoding, the output of the next higher-level decoder is used as input, until the decoded feature output of the first-level decoder with the lowest resolution is obtained. In this process, the decoder is used to decode based on the features of the reprojected first rendering frame, the reprojected second rendering frame, the high-definition intermediate frame correction data, and the decoding features of the next higher level, and reconstruct the decoding features; a pre-trained level zero decoder is used to decode the decoding features output by the level 1 decoder to obtain the blended mask and color residual of the reprojected first rendering frame and the reprojected second rendering frame; the level zero decoder is used to obtain the blended mask and color residual of the reprojected first rendering frame and the reprojected second rendering frame based on the decoding features output by the level 1 decoder; and the super-resolution intermediate rendering frame is reconstructed based on the blended mask and color residual.
[0066] For example, in the case of n=4, the decoder is a four-level decoder. Starting from the fourth-level features, it splices the features of the super-resolution first rendering frame, the high-definition intermediate frame correction data, and the low-definition second rendering frame, and uses a four-level decoder with progressively increasing resolution. Starting with the fourth-level decoder, which has the lowest resolution, decoding and reconstruction proceed step by step upwards. The inputs and outputs of each decoder level are as follows:
[0067] (
[0068] in, and for and Downsampling to a mapping relationship with the same resolution as the nth level feature. , These represent the reprojection mapping relationships at the times of the first rendering frame, the second rendering frame, and the intermediate frames, respectively. This represents the nth-level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. This represents the k-th level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. , For the nth and kth level decoding features output by the decoder, , , For the mapping residual of the decoder output, , , , For upsampling operation, Let represent the functions of the nth and kth level decoders, respectively. These functions are obtained through training using a standard network training method, which will not be elaborated further. The decoder structure is a standard decoder network structure.
[0069] Then, the blending mask and color residuals are obtained based on the level 0 decoder:
[0070] (
[0071] In the formula, These represent the blending mask and color residual of the first and second rendering frames, respectively. The blending mask is used to blend the super-resolution first rendering frame of the reprojection with the upsampled low-resolution second rendering frame of the reprojection. This is for color residuals, used to supplement details after mixing. The function representing the level-zero decoder is obtained through training. This represents the decoding characteristics output by the first-level decoder. , This represents a mapping relationship.
[0072] The calculation expression for the intermediate rendering frame of the super-resolution is:
[0073]
[0074] In the formula, Indicates an intermediate rendering frame. This represents the first rendered frame of the reprojection. This indicates the second rendered frame of the reprojection.
[0075] The reprojection of features can be analogous to the reprojection of rendered content, treating the width and height of features as the width and height of the rendered content image, and the channels of features as the channels of the rendered content image.
[0076] Step S105: The moment when the first rendering frame is reprojected onto the second rendering frame is corrected together with the upsampled second rendering frame to obtain the corrected data of the second rendering frame.
[0077] The reprojection and correction operations described above are similar to those in step S102. In practice, the same correction method can be reused, or different correction methods can be used.
[0078] Step S106: Extract the features of the second rendering frame correction data.
[0079] The above feature extraction operation is similar to step S103. In specific implementation, the same extraction method, some or all of the encoder weights can be reused, or different extraction methods can be used.
[0080] Step S107: Based on the characteristics of the second rendering frame correction data and the characteristics of the second rendering frame, the super-resolution rendering frame at time t+1 is reconstructed.
[0081] The process of feature reconstruction described above is similar to step S104. In practice, the same reconstruction method, some or all of the decoder weights can be reused, or different extraction methods can be used. Depending on the number of dependent features, the decoder input can be adjusted, or empty input or repeated input can be used.
[0082] Step S108: Insert the intermediate rendering frame after the super-resolution rendering frame at time t in the frame display sequence, update the low-resolution rendering frame at time t+1 in the frame display sequence to the super-resolution rendering frame at time t+1, and display it.
[0083] In this invention, rendering frames are reused, such as... Figure 2 As shown, the feature extraction algorithms for high-definition or super-resolution frames can be generalized, saving on the difficulty of method execution and the space occupied by computer code. This also facilitates parallelization and other computer programming techniques to extract features from these frames simultaneously, achieving feature extraction method reuse. The extracted features from the low-definition second rendering frame are reprojected to the intermediate time step and participate in the decoding and reconstruction of the super-resolution intermediate generated frame. Simultaneously, they can also directly participate in the decoding and reconstruction of the super-resolution second rendering frame, achieving feature reuse.
[0084] If only the intermediate frame generation portion of this method is used, and the input frames are standardized to a fixed resolution, excluding the process of super-resolution variation, a greater degree of reuse can be achieved. For example... Figure 3 As shown, each execution of the intermediate frame generation method only inputs one additional frame. The second rendering frame and its features during each execution can be retained as the first rendering frame and its features during the second execution of the method by means of caching or saving in the computer storage device. Theoretically, this can reduce the execution time by nearly one-third.
[0085] Example 2
[0086] This invention also provides a computer program product, such as an app on a mobile phone or tablet, or an installer on a computer. This product includes a computer program / instructions that, when executed by a processor, implement the method described in Embodiment 1. The code for the computer-executable program used to perform the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0087] Example 3
[0088] This invention provides a storage medium containing a computer-executable program, which, when executed by a computer processor, is used to perform the method of Embodiment 1.
[0089] The storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0090] Of course, the computer-executable program in the storage medium provided in the embodiments of the present invention is not limited to the above-described method operations, but can also perform related operations in the methods provided in any embodiment of the present invention.
[0091] It should be understood that the embodiments and descriptions above are only the principles, main features and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope of the invention, and all such changes and modifications fall within the protection scope of the present invention.
Claims
1. A rendering optimization method based on frame interpolation and super-resolution, characterized in that, include: Get the optimized super-resolution rendering frame at time t as the first rendering frame, and get the low-resolution rendering frame at time t+1 in the frame display sequence as the second rendering frame. The first rendered frame and the upsampled second rendered frame are reprojected onto the intermediate frame and corrected together to obtain the corrected data of the high-definition intermediate frame. Extract the corrected data of the first rendered frame, the high-definition intermediate frame, and the features of the second rendered frame, respectively. The features of the first and second rendering frames are reprojected onto the intermediate frame time, and combined with the features of the high-definition intermediate frame correction data to reconstruct the super-resolution intermediate rendering frame. The moment when the first rendering frame is reprojected onto the second rendering frame is corrected together with the upsampled second rendering frame to obtain the corrected data of the second rendering frame; Extract features from the correction data of the second rendered frame; Based on the characteristics of the corrected data of the second rendering frame and the characteristics of the second rendering frame, the super-resolution rendering frame at time t+1 is reconstructed. The intermediate rendered frame is inserted after the super-resolution rendered frame at time t in the frame display sequence. The low-resolution rendered frame at time t+1 in the frame display sequence is updated to the super-resolution rendered frame at time t+1 and then displayed.
2. The rendering optimization method based on frame interpolation and super-resolution according to claim 1, characterized in that, If the current time is the first time, then the first rendered frame is invalid data or the first low-resolution rendered frame in the frame display sequence.
3. The rendering optimization method based on frame interpolation and super-resolution according to claim 1, characterized in that, The step of reprojecting the first rendered frame and the upsampled second rendered frame onto the intermediate frame time and correcting them together to obtain the corrected data of the high-definition intermediate frame specifically includes: The second rendered frame is upsampled to obtain a high-definition second rendered frame; The first rendered frame and the high-definition second rendered frame are reprojected onto the intermediate frame time respectively to obtain the first intermediate frame and the second intermediate frame; Pixels at the same position in the first intermediate frame and the second intermediate frame are compared, selected, and / or merged to obtain the corrected data of the final high-definition intermediate frame.
4. The rendering optimization method based on frame interpolation and super-resolution according to claim 1, characterized in that, The extraction of features from the first rendered frame, the high-definition intermediate frame correction data, and the second rendered frame specifically includes: For the first rendered frame, the high-definition intermediate frame corrected data, and the second rendered frame, a convolutional encoder with n output resolutions decreasing progressively is used to extract a feature set ft={ft} of features with n progressively decreasing resolutions and progressively increasing channel numbers. j | j=1,…,n}: ft j =E j (ft j-1 ),j=1,…,n, Among them, ft j、 ft j- E represents the features at levels j and j-1. j () represents the j-th convolutional encoder, where ft0 takes the value of the first rendered frame, or the high-definition intermediate frame correction data, or the second rendered frame.
5. The rendering optimization method based on frame interpolation and super-resolution according to claim 1, characterized in that, The process of reprojecting the features of the first and second rendering frames onto the intermediate frame time and combining them with the features of the high-definition intermediate frame correction data to reconstruct the super-resolution intermediate rendering frame specifically includes: The feature sets of the first rendering frame and the second rendering frame are respectively transformed to the intermediate frame time by reprojection to obtain the feature set of the first rendering frame and the feature set of the second rendering frame. For the feature sets of the first reprojected rendering frame, the second reprojected rendering frame, and the intermediate frame correction data, starting from the nth-level feature with the lowest resolution in the feature set, a pre-trained decoder with progressively increasing resolution is used for decoding and reconstruction. During decoding, the output of the next higher-level decoder is used as input until the decoded feature output of the first-level decoder with the lowest resolution is obtained. The decoder is used to decode and reconstruct the decoded features based on the features of the first reprojected rendering frame, the second reprojected rendering frame, the high-definition intermediate frame correction data, and the next higher-level decoded features. The pre-trained level 0 decoder decodes the decoding features output by the level 1 decoder to obtain the blending mask and color residual of the first and second reprojected rendering frames. The level 0 decoder is used to obtain the blending mask and color residual of the first and second reprojected rendering frames based on the decoding features output by the level 1 decoder. Based on the blending mask and color residual, the super-resolution intermediate rendering frame is reconstructed.
6. The rendering optimization method based on frame interpolation and super-resolution according to claim 5, characterized in that, The specific operations for decoding and reconstruction are as follows: , In the formula, and for and Downsampling to a mapping relationship with the same resolution as the nth level feature. , These represent the reprojection mapping relationships at the times of the first rendering frame, the second rendering frame, and the intermediate frames, respectively. This represents the nth-level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. This represents the k-th level feature set of the reprojected first rendered frame, the reprojected second rendered frame, and the high-definition intermediate frame corrected data. , For the nth and kth level decoding features output by the decoder, , , For the mapping residual of the decoder output, , This is a mapping relationship. This is a mapping relationship. For upsampling operation, These represent the functions of the nth and kth level decoders, respectively, obtained through training.
7. The rendering optimization method based on frame interpolation and super-resolution according to claim 6, characterized in that, The calculation expressions for the blending mask and color residual are as follows: ( , In the formula, These represent the blending mask and color residual of the first and second reprojected render frames, respectively. The function representing the level-zero decoder is obtained through training. This represents the decoding characteristics output by the first-level decoder. For mapping relationship, This represents a mapping relationship.
8. The rendering optimization method based on frame interpolation and super-resolution according to claim 7, characterized in that, The calculation expression for the intermediate rendering frame of the super-resolution is: , In the formula, Indicates an intermediate rendering frame. This represents the first rendered frame of the reprojection. This indicates the second rendered frame of the reprojection.
9. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements the method of any one of claims 1-8.
10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that: The computer program / instructions, when executed by a processor, implement the method of any one of claims 1-8.