A Real-Time Super-Resolution Reconstruction Method and System Based on Historical Feature Fusion

By extracting and fusion of current frames and historical frames, and combining motion compensation technology, real-time super-resolution reconstruction with low computing power and high frame rate is achieved, solving the problems of poor real-time performance and detailed timing breakage in the existing technology, and improving picture quality and versatility.

CN114155152BActive Publication Date: 2025-06-10INNOSILICON MICROELECTRONICS (ZHUHAI) CO LTD +1

Patent Information

Application Number
CN202111513198.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-12
Publication Date
2025-06-10
Estimated Expiration
2041-12-12

AI Technical Summary

Technical Problem

The existing super-resolution technology has shortcomings in real-time, detail timing breakage, false colors and versatility, especially in video streaming scenarios, which are difficult to achieve efficient real-time super-resolution reconstruction.

Method used

A real-time super-resolution reconstruction method based on historical feature fusion is adopted to generate high-resolution reconstruction images by extracting and fusion linear and nonlinear features of the current frame and historical frames, combined with motion compensation technology.

Benefits of technology

Real-time super-resolution reconstruction with low computing power requirements and high frame rate is realized, which avoids delay and artifact problems in traditional methods, and improves picture quality and motion compensation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155152B_ABST
    Figure CN114155152B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time super-resolution reconstruction method and system based on historical feature fusion. The method includes constructing a network model; constructing the network model includes: extracting linear features of the current frame to obtain a linear feature map of the current frame; extracting non-linear features and performing feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion; wherein, the historical frames are the first N frames before the current frame, and N is a natural number; the extraction of non-linear features of the historical frames is based on the images after motion compensation of the historical frames; based on the non-linear feature map after feature fusion and the linear feature map of the current frame, a super-resolution reconstructed image of the current frame is obtained. The present invention solves a series of problems existing in current super-resolution technologies, such as poor real-time performance, detail temporal breakage, false colors, poor generality, etc., and takes into account both real-time performance and image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a real-time super-resolution reconstruction method and system based on historical feature fusion. Background Art

[0002] Super-resolution technology is a technical means of transforming a low-resolution image into a high-resolution image, and is currently gradually applied to scenarios such as videos and games. Traditional super-resolution technologies include bilinear and bicubic interpolation methods. Although these super-resolution methods represented by interpolation methods have low requirements for computing power, their essence is still to fill in colors locally, and the effect is not very satisfactory.

[0003] Since the SRCNN network was proposed, neural networks have been introduced to solve the super-resolution problem. Subsequently, super-resolution neural networks for images and super-resolution neural networks for videos have been continuously proposed, and records have been continuously refreshed in some metrics, improving the maturity and usability of super-resolution technology. However, currently, real-time super-resolution networks have not been deeply explored. There are some applications that enhance the TV picture quality or perform super-resolution on game scenes to obtain performance improvements. However, these applications are either super-resolution networks for images, which can only perform frame-by-frame super-resolution in video stream scenarios, easily causing problems such as temporal discontinuity of picture details and false colors; or they use a large number of motion vectors during GPU rendering as inputs, have a strong coupling with the GPU, and are not traditional frame-based super-resolution applications. Their structures are complex and their integration and specialization are high. Summary of the Invention

[0004] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a real-time super-resolution reconstruction method and system based on historical feature fusion. Based on the current frame features and accurately combining the motion conditions of historical frames for feature compensation, a series of problems existing in current super-resolution technologies such as poor real-time performance, temporal discontinuity of details, false colors, and poor generality are solved, taking into account both real-time performance and picture quality.

[0005] To achieve the above object, according to one aspect of the present invention, a super-resolution reconstruction method is provided, including constructing a network model; constructing the network model includes: performing linear feature extraction on the current frame to obtain a linear feature map of the current frame; performing non-linear feature extraction and feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion; where the historical frames are the first N frames before the current frame, and N is a natural number; the non-linear feature extraction of the historical frames is performed based on the images after motion compensation of the historical frames; based on the non-linear feature map after feature fusion and the linear feature map of the current frame, a super-resolution reconstructed image of the current frame is obtained.

[0006] In some embodiments, the non-linear feature extraction of historical frames includes: for each frame in the historical frames, after splicing with the current frame, an offset field compensation is performed using a convolution kernel without an activation function to obtain a frame motion compensation matrix of the historical frames, and then the frame motion compensation matrix of the historical frames is superimposed on the historical frames to obtain the historical frames after motion compensation; non-linear feature extraction is performed on the historical frames after motion compensation to obtain N non-linear feature maps of the historical frames after motion compensation.

[0007] In some embodiments, the feature fusion includes: after splicing N non-linear feature maps of the historical frames after motion compensation and the non-linear feature maps of the current frame, non-linear feature extraction is performed to obtain non-linear feature maps after feature fusion.

[0008] In some embodiments, obtaining a super-resolution reconstructed image of the current frame based on the non-linear feature maps after feature fusion and the linear feature maps of the current frame includes: superimposing the non-linear feature maps after feature fusion and the linear feature maps of the current frame and performing pixel rearrangement, and then using a convolution kernel without an activation function to correct the pixels of the image after pixel rearrangement to obtain a super-resolution reconstructed image of the current frame.

[0009] In some embodiments, the method further includes constructing a data set; constructing the data set includes: according to real-time motion data, rendering output videos at an initial resolution and a target resolution, constructing a data set at the initial resolution and a data set at the target resolution, and dividing the data set at the initial resolution and the data set at the target resolution into a training set and a test set according to a predetermined ratio.

[0010] In some embodiments, the method further includes training a network model; training the network model includes: using the image data at the initial resolution and the target resolution in the training set, calculating a loss value according to a preset loss function, and completing the training of the network model when the loss value and / or the accuracy of the training set and / or the accuracy of the test set reach a preset standard.

[0011] According to another aspect of the present invention, a super-resolution reconstruction method is provided. After quantifying the network model trained by the above method, it is deployed to a hardware platform to achieve the target function.

[0012] According to another aspect of the present invention, a super-resolution reconstruction system is provided, including a network model construction module; the network model construction module includes: a linear feature extraction module, configured to perform linear feature extraction on the current frame to obtain a linear feature map of the current frame; a motion compensation module, configured to perform motion compensation on the historical frame to obtain the historical frame after motion compensation; a non-linear feature extraction module, configured to perform non-linear feature extraction on the current frame to obtain a non-linear feature map of the previous frame, and further configured to perform non-linear feature extraction on the historical frame after motion compensation to obtain a non-linear feature map of the historical frame after motion compensation; a non-linear feature fusion module, configured to splice the non-linear feature map after motion compensation and the non-linear feature map of the current frame, and then perform non-linear feature fusion to obtain a non-linear feature map after feature fusion; an image processing module, configured to obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.

[0013] According to yet another aspect of the present invention, an electronic device is provided, including the above-mentioned super-resolution reconstruction system.

[0014] According to yet another aspect of the present invention, an electronic device is provided, including: a processor; a memory communicatively connected to the processor; the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is enabled to execute the above method.

[0015] According to yet another aspect of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.

[0016] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects are achieved:

[0017] 1. Low computing power requirement and high frame rate for super-resolution reconstruction under the same hardware conditions. For example, the computing power required to super-resolve an original image of (540*960*3) pixels to an image of (1080*1920*3) pixels is only 29.55 GFLOPs, and theoretically, each T of computing power can render 33.85 frames. Compared with the DUF video super-resolution network in 2018 CVPR, the real-time super-resolution network RTSR2 (Real-time Super Resolution II) proposed by the present invention only requires 1 / 17 of the computing power of the DUF network.

[0018] 2. Overcoming the traditional design, super-resolution reconstruction can be performed in real time. The traditional VSR (Video SuperResolution) method needs to call N historical frames and N future frames to super-resolution the current frame, which will inevitably cause a delay of at least N frames for the super-resolution reconstruction of the current frame. The RTSR2 network proposed in the present invention only uses the current frame and N historical frames to render the picture for prediction. The design avoids the call to future frames and thus reduces the delay. In addition, since the computing power required for super-resolution reconstruction of each frame is low, the delay in super-resolution reconstruction of the current frame after rendering the initial resolution is small, which fully guarantees the real-time requirements of the algorithm.

[0019] 3. High picture quality and excellent detail performance. Using the high-frequency information of historical frames, the RTSR2 network is rich in detail filling and effectively suppresses artifacts.

[0020] 4. Accurate motion compensation and coherent motion images. In the model, each historical frame is spliced ​​with the current frame and the frame motion compensation matrix is ​​calculated through convolution. Each historical frame is compared with the current frame, so the obtained motion compensation matrix is ​​flexible and accurate, reducing the tearing of the picture and the jumping of details between frames.

[0021] 5. Accurate data set. Different from the traditional method of constructing data sets by downsampling high-definition recorded videos, the present invention constructs data sets by synchronous rendering. Based on the same game information, frame-by-frame rendering is performed to obtain data sets of different resolutions that are synchronized in time sequence. This is more in line with actual application conditions, improves the accuracy of model training, and thus effectively improves the accuracy of super-resolution reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of a real-time super-resolution reconstruction method based on multi-frame fusion according to an embodiment of the present invention;

[0023] Figure 2 is a schematic diagram of a process of building a network model according to an embodiment of the present invention;

[0024] Figure 3 is a schematic diagram of a process of constructing a data set according to an embodiment of the present invention;

[0025] Figure 4 yes Figure 2 The specific implementation process diagram of the process of building a network model is shown;

[0026] Figure 5 yes Figure 4 The specific implementation process diagram of the process of building a network model is shown;

[0027] Figure 6 is a structural schematic diagram of a real-time super-resolution reconstruction system based on multi-frame fusion according to an embodiment of the present invention;

[0028] Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.

[0030] As Figure 1 shown, the real-time super-resolution reconstruction method based on historical feature fusion according to an embodiment of the present invention includes the following steps:

[0031] Step 1: Construct a data set.

[0032] Specifically, according to real-time motion data (including control instructions) and model files, render output videos at different resolutions, construct data sets at different resolutions, and divide the data sets at different resolutions into a training set and a test set according to a predetermined ratio.

[0033] In some embodiments, the different resolutions include a first resolution, a second resolution, and a third resolution. Among them, the different resolutions are all specified resolutions generated by motion data and model files. The second resolution is greater than the first resolution, and the third resolution is greater than the second resolution. In some embodiments, the second resolution is an integer multiple of the first resolution, and the third resolution is an integer multiple of the first resolution and the second resolution. In some embodiments, the second resolution is 2 times the first resolution, and the third resolution is 4 times the first resolution and 2 times the second resolution. For example, the first resolution is 540P, the second resolution is 1080P, and the third resolution is 2160P.

[0034] In some embodiments, a first dataset at a first resolution, a second dataset at a second resolution, and a third dataset at a third resolution are constructed. In some embodiments, when performing super-resolution reconstruction with an input (initial resolution) of the first resolution and an output (target resolution) of the second resolution, the first dataset and the second dataset are divided into a training set and a test set according to a predetermined ratio. It can be understood that in the embodiments of the present invention, the target resolution is greater than the initial resolution. For example, when performing 2x super-resolution reconstruction with an input of 540P and an output of 1080P, the 540P dataset and the 1080P dataset are divided into a training set and a test set according to a predetermined ratio. In some embodiments, when performing super-resolution reconstruction with an input of the second resolution and an output of the third resolution, the second dataset and the third dataset are divided into a training set and a test set according to a predetermined ratio. For example, when performing 2x super-resolution reconstruction with an input of 1080P and an output of 2160P, the 1080P dataset and the 2160P dataset are divided into a training set and a test set according to a predetermined ratio. In some embodiments, when performing super-resolution reconstruction with an input of the first resolution and an output of the third resolution, the first dataset and the third dataset are divided into a training set and a test set according to a predetermined ratio. For example, when performing 4x super-resolution reconstruction with an input of 540P and an output of 2160P, the 540P dataset and the 2160P dataset are divided into a training set and a test set according to a predetermined ratio.

[0035] In some embodiments, the data proportion of the training set is greater than that of the test set. For example, the initial resolution dataset is divided into a training set and a test set at a ratio of 7:3. At the same time, the target resolution data corresponding to the initial resolution data in the training set is also included in the training set, and the target resolution data corresponding to the initial resolution data in the test set is also included in the test set.

[0036] Step 2: Construct a network model.

[0037] As Figure 2 shown, non-linear feature extraction and feature fusion are performed on the current frame and historical frames to obtain a non-linear feature map after feature fusion. Based on the non-linear feature map after feature fusion and the linear feature map of the current frame, a super-resolution reconstructed image of the current frame is obtained.

[0038] Among them, the historical frames are the first N frames before the current frame, N is a natural number, and the non-linear feature extraction of the historical frames is based on the graphics after motion compensation of the historical frames. Define the current frame as F(i), and the historical frames as F(i - t), where t = 1,..., N.

[0039] In some embodiments, for each frame F(i - t) in the historical frames, where t = 1,..., N, after concatenating it with the current frame F(i), an offset field compensation is performed using a convolution kernel without an activation function to obtain the motion warping matrix MWM(i - t) of the historical frame F(i - t). Then, the motion warping matrix MWM(i - t) of the historical frame F(i - t) is superimposed on the historical frame F(i - t) to obtain the motion-compensated historical frame F’(i - t).

[0040] In some embodiments, a non-linear feature extraction is performed on the motion-compensated historical frame F’(i - t) using a convolution kernel activated by a non-linear function to obtain the non-linear feature map after motion compensation of the historical frame F(i - t). A non-linear feature extraction is performed on the current frame F(i) using a convolution kernel activated by a non-linear function to obtain the non-linear feature map of the current frame F(i). After concatenating the non-linear feature maps after motion compensation of the historical frames F(i - 1) to F(i - N) and the non-linear feature map of the current frame F(i), a non-linear feature extraction is performed using a convolution kernel activated by a non-linear function to obtain the non-linear feature map after feature fusion. The non-linear feature map after feature fusion is the unified expression of the non-linear features of the current frame and the historical frames, removing some invalid feature information.

[0041] In some embodiments, a linear feature extraction is performed on the current frame F(i) using a convolution kernel without an activation function to obtain the linear feature map of the current frame F(i). The concatenated non-linear feature map and the linear feature map of the current frame F(i) are superimposed and pixel shuffle is performed to obtain the super-resolution reconstructed image of the current frame.

[0042] In some embodiments, constructing the network model further includes performing pixel correction on the image after pixel shuffle. For example, pixel correction is performed on the image after pixel shuffle using a convolution kernel without an activation function to obtain the super-resolution reconstructed image of the current frame, completing the super-resolution reconstruction of the current frame.

[0043] In some embodiments, the non-linear function can be, for example, tanh, sigmoid, or relu.

[0044] In this step, the linear feature map of the current frame corresponds to low-frequency information. Specifically, the low-frequency information refers to the information of the areas in the image where the color / gray level changes slowly, and such information widely exists in the color block filling areas inside the edges; the non-linear feature maps that are temporally continuous correspond to high-frequency information. Specifically, the high-frequency information refers to the information of the areas in the image where the color / gray level changes rapidly, and is generally closely related to the edge / texture information in the image; through linear and non-linear feature extraction, the color block distribution and detail distribution of the image are obtained simultaneously, thereby effectively improving the super-resolution reconstruction quality of the model.

[0045] The following describes in detail the real-time super-resolution reconstruction method based on multi-frame fusion of the present invention with specific examples. Set the resolution magnification factor as u, the number of basic convolution kernels as n, the size of the basic convolution kernel as k, the convolution step size as a, and the activation function as AFunc.

[0046] As Figure 4 shown, a multi-branch neural network is adopted. For the current frame F(i), linear convolution is performed using a 3au kernel with a size of k*k to extract features for backup. The activation function AFunc is used to perform an n-kernel convolution with a kernel size of k*k and a 3au 2 , 2 , 2 , 2 kernel convolution on the current frame to extract non-linear features. For each frame F(i-t) in the historical N frames F(i-1) to F(i-N), after concatenating it with the current frame F(i), linear convolution is performed, including a 6-kernel convolution with a kernel size of k*k and a 3-kernel convolution with a kernel size of k*k, to obtain the frame motion compensation matrix MWM(i-t). Then, the offset field compensation matrix MWM(i-t) of the historical frame F(i-t) is added to F(i-t) to obtain the historical frame F’(i-t) after inter-frame motion compensation. Next, the activation function AFunc is used to perform an n-kernel convolution with a kernel size of k*k and a 3au 2 kernel convolution on F’(i-t) to extract the non-linear feature map of the historical frame F(i-t) after motion compensation. 2 The non-linear feature maps of the historical N frames F(i-1) to F(i-N) after motion compensation and the non-linear feature map of the current frame F(i) are concatenated and then the activation function AFunc is used again to perform an n-kernel convolution with a kernel size of k*k and a 3au

[0047] kernel convolution to obtain all non-linear features. The non-linear feature map is added to the linear feature of the current frame and then the feature map is pixel-rearranged to the required enlarged size. Finally, through a feature correction module composed of a 3-kernel linear convolution with a kernel size of k*k, the real-time reconstruction result of the super-resolution frame is obtained. 2 The non-linear feature maps of the historical N frames F(i-1) to F(i-N) after motion compensation and the non-linear feature map of the current frame F(i) are concatenated and then the activation function AFunc is used again to perform an n-kernel convolution with a kernel size of k*k and a 3au kernel convolution to obtain all non-linear features. The non-linear feature map is added to the linear feature of the current frame and then the feature map is pixel-rearranged to the required enlarged size. Finally, through a feature correction module composed of a 3-kernel linear convolution with a kernel size of k*k, the real-time reconstruction result of the super-resolution frame is obtained.

[0048] So far, the RTSR2 network construction is completed. The RTSR2 network uses the current frame and the historical N frames for feature extraction (N only needs to satisfy ≥1). During the process of linear and non-linear convolutions, the number n of basic convolution kernels has no fixed value (generally, n = 2a, a≥4, or n = 3a*u 2 , where a is a positive integer), and several convolution blocks are omitted at the ellipsis in Figure 4 , which can be added or deleted during actual design. The RTSR2 network has the characteristic of being easy to modify. It can adopt a streamlined version for mobile or embedded devices, or an enlarged version for scenarios with demanding quality requirements such as workstations and cinemas. The RTSR2 network collects the linear and non-linear features of the current frame, uses the non-linear features of the historical frames with precise motion compensation as an aid, and incorporates the features refined from the historical frames based on the features of the current frame, improving the detail performance of the picture.

[0049] Step 3: Train the network model.

[0050] Using the image data of the initial resolution and the target resolution in the training set, calculate the loss value according to the preset loss function. When the loss value and the accuracies of the training set and the test set reach the preset standards, retain the model to complete the training of the model.

[0051] In some embodiments, the Huber loss function is used to calculate the loss value to enhance the robustness of the squared error loss function to outliers. The expression of the loss value is as follows:

[0052]

[0053] Among them, y is the image data of the target resolution in the training set, that is, the target output. x is the image data of the initial resolution in the training set, the f() function is the model calculation process, and f(x) is the super-resolution reconstruction image result calculated by x through the neural network model. δ is a custom global parameter. When the deviation value |y - f(x)| is less than δ, the squared error is used. When the deviation value |y - f(x)| is greater than δ, the linear error is used.

[0054] In some embodiments, the training set constructed in Step 1 is used to train the network model constructed in Step 2 using the backpropagation algorithm. That is, the initial resolution (low resolution) data of the training set is input into the model, the model calculates and outputs the predicted high-resolution image, and then the predicted high-resolution image and the target resolution image (such as Ground Truth) of the training set are substituted into the loss function to calculate the loss value (loss). Then, the loss value is backpropagated to adjust the parameters of the model, and the above calculation process is repeated and iterated continuously until the loss value and the accuracies of the training set and the test set all reach the preset standards to complete the training of the model.

[0055] In some embodiments, the accuracy rate is the ratio of the number of pixels for which super-resolution reconstruction from the initial resolution to the target resolution is achieved using the neural network model to the total number of pixels. Specifically, the accuracy rate of the training set is obtained by inputting the initial resolution data in the training set into the model, comparing the resolution data output by the model with the corresponding target resolution data in the training set, and dividing the number of pixels with the same resolution by the total number of pixels in the training set, which is the accuracy rate of the training set. Similarly, the accuracy rate of the test set is obtained by inputting the initial resolution data in the test set into the model, comparing the resolution data output by the model with the corresponding target resolution data in the test set, and dividing the number of pixels with the same resolution by the total number of pixels in the test set, which is the accuracy rate of the test set.

[0056] The accuracy rate of the test set is approximate to that in the real scenario, and the performance of the model in the real scenario can be predicted. In addition, the accuracy rate of the test set can also reflect the robustness of the model. When the accuracy rate of the test set is approximate to that of the training set, the model has strong robustness; when the accuracy rate of the test set differs greatly from that of the training set, the model has poor robustness.

[0057] In some embodiments, the preset criteria include: loss value ≤ 10 -5 , accuracy rate of the training set ≥ 99.9%, and accuracy rate of the test set ≥ 99.9%.

[0058] Step 4: Deployment of the network model.

[0059] After quantifying the network model trained in Step 3 and deploying it to the hardware platform, the target function can be achieved.

[0060] Taking the example of super-resolution from 540P resolution (original resolution) to 1080P resolution (target resolution), using 2 historical frames (N = 2), and the game Honor of Kings, the real-time super-resolution reconstruction method based on multi-frame fusion of the present invention will be described in detail below.

[0061] The resolution magnification factor u = 2, the number of basic convolution kernels n = 32, the size of the basic convolution kernel k = 3, the convolution stride a = 1, and the activation function AFunc is tanh.

[0062] Combined with Figure 1 , the implementation process of the embodiment of the present invention is as follows:

[0063] Step 1: Construction of the data set.

[0064] As Figure 3As shown in the figure, an innovative method for constructing a dataset is proposed for game scenarios when building a dataset. Different from the common method of downsampling from high-definition videos to obtain different resolutions, the embodiments of the present invention render the output videos of the Honor of Kings game at different resolutions (540P, 1080P, 2160P) according to the in-game motion data (such as control instructions) and model files, construct a dataset between different multiple relationships (2 two-fold datasets, 1 four-fold dataset), and divide it into a training set and a test set according to a ratio of 7:3. The dataset combined with 540P and 1080P is selected for this training.

[0065] Step 2: Construct a network model.

[0066] As Figure 2 shown, linear and non-linear features of the current frame are extracted. For historical frames, after being concatenated with the current frame, the compensation amount of the motion offset field is calculated through linear convolution, and the offset field compensation and the historical frame are added to obtain a corrected frame with the same pose as the current frame image content. After non-linear feature extraction is performed on all corrected frames, they are concatenated with the non-linear features of the current frame, and non-linear features are extracted again. The non-linear features and the linear features of the aforementioned current frame are added to obtain a feature map rich in high and low frequency information. Next, pixel rearrangement is performed to restore the feature map to the resolution size of the image to be reconstructed. Finally, further correction of the features is performed on the image after pixel rearrangement to obtain the super-resolution reconstruction result of the current frame.

[0067] Specifically, in combination with Figure 5 , the RTSR2 super-resolution neural network uses a 12-core convolution with a convolution kernel size of 3*3 in 1 layer to extract linear features of the current frame F(i), and uses a 32-core convolution with a convolution kernel size of 3*3 in 2 layers activated by the tanh function and a 12-core convolution with a convolution kernel size of 3*3 in 1 layer to extract non-linear features of the current frame.

[0068] For any one of the historical two frames F(i-1) and F(i-2), t = 1, 2, the following method is used for processing. First, it is concatenated with the current frame F(i), and then linear convolution is performed through a 6-core convolution with a convolution kernel size of 3*3 in 1 layer and a 3-core convolution with a convolution kernel size of 3*3 in 1 layer to obtain a frame motion compensation matrix MWM(i-t). Then, the offset field compensation matrix MWM(i-t) is added to F(i-t) to obtain the historical frame F’(i-t) after inter-frame motion compensation. Next, a non-linear convolution group of F’(i-t) is extracted by using a 32-core convolution with a convolution kernel size of 3*3 in 2 layers activated by the tanh function and a 12-core convolution with a convolution kernel size of 3*3 in 1 layer activated by the tanh function to obtain a non-linear feature map after motion compensation.

[0069] After motion compensation of the historical frames F(i - 1) and F(i - 2), the non - linear feature maps are concatenated with the non - linear feature map of the current frame F(i), and then feature extraction is performed again on the concatenated result by a non - linear convolution group to obtain the total non - linear feature map. This non - linear feature map is added to the linear feature of the current frame, and then the pixel - rearranged feature map is resized to the required enlarged size. Finally, through a feature correction module containing a 3 - kernel linear convolution with a convolution kernel size of 3 * 3, the real - time reconstruction result of the super - resolution frame is obtained.

[0070] Step 3: Train the network model.

[0071] The loss function uses the Huber loss function to enhance the robustness of the squared error loss function to outliers. The expression of the Huber loss function is as follows:

[0072]

[0073] Use the dataset constructed in Step 1 to train the aforementioned model using the backpropagation algorithm. Use the dataset constructed in Step 1 to train the aforementioned model using the backpropagation algorithm, and iterate until the loss value L ≤ 10 -5 and the accuracies of both the training set and the test set are ≥ 99.9%, and then retain this model.

[0074] Step 4: Deploy the network model.

[0075] After quantizing the model trained in Step 3, deploy it to the hardware acceleration platform, so that the mobile processor can obtain smooth, stable, high - resolution and high - quality real - time results while rendering the King of Glory game and maintain the low - load level of the native rendering at 540P. This model can be applied to scenarios such as game video recording, game live streaming, and real - time resolution improvement during games involved in the super - resolution dataset.

[0076] As Figure 6 shown, the embodiment of the present application also provides a real - time super - resolution reconstruction system based on multi - frame fusion, including:

[0077] A dataset construction module, which is used to render output videos at different resolutions according to real - time motion data, construct datasets at different resolutions, and divide the datasets at different resolutions into a training set and a test set according to a predetermined ratio.

[0078] In some embodiments, the different resolutions include a first resolution, a second resolution, and a third resolution. Among them, the different resolutions are all specified resolutions generated from motion data and model files. The second resolution is greater than the first resolution, and the third resolution is greater than the second resolution. In some embodiments, the second resolution is an integer multiple of the first resolution, and the third resolution is an integer multiple of both the first resolution and the second resolution. In some embodiments, the second resolution is 2 times the first resolution, and the third resolution is 4 times the first resolution and 2 times the second resolution. For example, the first resolution is 540P, the second resolution is 1080P, and the third resolution is 2160P.

[0079] In some embodiments, a first dataset at the first resolution, a second dataset at the second resolution, and a third dataset at the third resolution are constructed. In some embodiments, for super-resolution reconstruction with an input (initial resolution) of the first resolution and an output (target resolution) of the second resolution, the first dataset and the second dataset are divided into a training set and a test set according to a predetermined ratio. For example, for 2-fold super-resolution reconstruction with an input of 540P and an output of 1080P, the 540P dataset and the 1080P dataset are divided into a training set and a test set according to a predetermined ratio. In some embodiments, for super-resolution reconstruction with an input of the second resolution and an output of the third resolution, the second dataset and the third dataset are divided into a training set and a test set according to a predetermined ratio. For example, for 2-fold super-resolution reconstruction with an input of 1080P and an output of 2160P, the 1080P dataset and the 2160P dataset are divided into a training set and a test set according to a predetermined ratio. In some embodiments, for super-resolution reconstruction with an input of the first resolution and an output of the third resolution, the first dataset and the third dataset are divided into a training set and a test set according to a predetermined ratio. For example, for 4-fold super-resolution reconstruction with an input of 540P and an output of 2160P, the 540P dataset and the 2160P dataset are divided into a training set and a test set according to a predetermined ratio.

[0080] In some embodiments, the proportion of data in the training set is greater than that in the test set. For example, the initial resolution dataset is divided into a training set and a test set at a ratio of 7:3. At the same time, the target resolution data corresponding to the initial resolution data in the training set is also included in the training set, and the target resolution data corresponding to the initial resolution data in the test set is also included in the test set.

[0081] A network model construction module, which is used to extract linear features from the current frame to obtain a linear feature map of the current frame, perform non-linear feature extraction and feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion, and obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.

[0082] In some embodiments, the network model construction module includes a linear feature extraction module, which is used to extract linear features from the current frame F(i) using a convolutional kernel without an activation function.

[0083] In some embodiments, the network model construction module includes a motion compensation module, which is used to splice the historical frame F(i - t) and the current frame F(i), perform offset field compensation using a convolutional kernel without an activation function to obtain a frame motion compensation matrix MWM(i - t) of the historical frame F(i - t), and then superimpose the frame motion compensation matrix MWM(i - t) of the historical frame F(i - t) on the historical frame F(i - t) to obtain the historical frame F'(i - t) after motion compensation.

[0084] In some embodiments, the network model construction module further includes a non-linear feature extraction module, which is used to extract non-linear features from the current frame F(i) using a convolutional kernel activated by a non-linear function to obtain a non-linear feature map of the current frame. The non-linear feature extraction module is also used to extract non-linear features from the historical frame F'(i - t) after motion compensation using a convolutional kernel activated by a non-linear function to obtain a non-linear feature map of the historical frame F(i - t) after motion compensation.

[0085] The network model construction module further includes a non-linear feature fusion module, which is used to splice N non-linear feature maps after motion compensation and the non-linear feature map of the current frame, and perform non-linear feature fusion using a convolutional kernel activated by a non-linear function to obtain a non-linear feature map after feature fusion. The non-linear feature map after feature fusion is the unified expression of the non-linear features of the current frame and historical frames, removing some invalid feature information.

[0086] In some embodiments, the network model construction module further includes an image processing module, which is used to obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.

[0087] In some embodiments, the image processing module further includes a pixel rearrangement module, which is used to superimpose the non-linear feature map after feature fusion and the linear feature map of the current frame and perform pixel rearrangement (Pixel Shuffle) to obtain a super-resolution reconstructed image of the current frame.

[0088] In some embodiments, the image processing module further includes a pixel correction module for correcting the pixels of the image after pixel rearrangement. For example, a convolutional kernel without an activation function is used to correct the pixels of the image after pixel rearrangement to obtain a super-resolution reconstructed image of the current frame, thus completing the super-resolution reconstruction of the current frame.

[0089] In some embodiments, the non-linear function can be, for example, tanh, sigmoid or relu.

[0090] The network model training module is used to calculate the loss value according to a preset loss function by using the image data of the initial resolution and the target resolution in the training set. When the loss value, and the accuracies of the training set and the test set reach the preset standards, the model is retained to complete the training of the model.

[0091] In some embodiments, the Huber loss function is adopted to calculate the loss value to enhance the robustness of the squared error loss function to outliers. The expression of the loss value is as follows:

[0092]

[0093] Where y is the image data of the target resolution in the training set, that is, the target output. x is the image data of the initial resolution in the training set, the f() function is the model calculation process, and f(x) is the result of the super-resolution reconstructed image calculated by x through the neural network model. δ is a custom global parameter. When the deviation value |y - f(x)| is less than δ, the squared error is adopted; when the deviation value |y - f(x)| is greater than δ, the linear error is adopted.

[0094] In some embodiments, the network model constructed by the network model construction module is trained by using the backpropagation algorithm with the dataset constructed by the dataset construction module. That is, the initial resolution (low resolution) data of the training set is input into the model, the model calculates and outputs the predicted high-resolution image, and then the predicted high-resolution image and the target resolution image (Ground Truth) of the training set are substituted into the loss function to calculate the loss value (loss). Then the loss value is backpropagated to adjust the parameters of the model, and the above calculation process is repeated, and the iteration is continuously performed until the loss value and accuracy of the training set and the loss value and accuracy of the test set reach the preset standards to complete the training of the model.

[0095] In some embodiments, the preset standards include: the loss value of the training set ≤ 10 -5 , the accuracy of the training set ≥ 99.9%, the loss value of the test set ≤ 10 -5 , the accuracy of the test set ≥ 99.9%.

[0096] The network model deployment module is used to quantize the network model trained by the network model training module and then deploy it to the hardware platform, so as to achieve the target function.

[0097] Figure 7 FIG. is a block diagram of an electronic device according to an embodiment of the present application. An embodiment of the present application also provides an electronic device, as Figure 7 shown. The electronic device includes: at least one processor 701, and a memory 703 communicatively connected to the at least one processor 701. Instructions executable by the at least one processor 701 are stored in the memory 703. The instructions are executed by the at least one processor 701. When the processor 701 executes the instructions, the driving scenario reconstruction method in the above embodiment is implemented. The number of the memory 703 and the processor 701 can be one or more. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0098] The electronic device may further include a communication interface 705 for communicating with external devices and performing data interaction and transmission. Each device is interconnected using different buses and can be mounted on a common motherboard or otherwise as required. The processor 701 can process instructions executed within the electronic device, including instructions for storing graphical information in the memory or on the memory to display a graphical user interface (GUI) on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as, as a server array, a set of blade servers, or a multi-processor system). The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.

[0099] Optionally, in a specific implementation, if the memory 703, the processor 701, and the communication interface 705 are integrated on a single chip, the memory 703, the processor 701, and the communication interface 705 can communicate with each other through an internal interface.

[0100] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0101] The embodiment of the present application provides a computer-readable storage medium (such as the above-mentioned memory 703), which stores computer instructions. When the program is executed by the processor, the method provided in the embodiment of the present application is implemented.

[0102] Optionally, the memory 703 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device of the driving scene reconstruction method, etc. In addition, the memory 703 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 703 may optionally include a memory remotely set relative to the processor 701, and these remote memories may be connected to the electronic device of the driving scene reconstruction method through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0103] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0104] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.

[0105] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more (two or more) executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.

[0106] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device).

[0107] It should be understood that various parts of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiments.

[0108] In addition, in each embodiment of this application, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0109] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A super-resolution reconstruction method, characterized in that, it includes constructing a network model; The constructing of the network model includes: Performing linear feature extraction on the current frame to obtain a linear feature map of the current frame; For each frame in the historical frames, after splicing with the current frame, using a convolution kernel without an activation function for offset field compensation to obtain a frame motion compensation matrix of the historical frame, and then superimposing the frame motion compensation matrix of the historical frame with the historical frame to obtain the historical frame after motion compensation; Performing non-linear feature extraction on the historical frame after motion compensation to obtain N non-linear feature maps of the historical frame after motion compensation; After splicing N non-linear feature maps of the historical frame after motion compensation and the non-linear feature map of the current frame, performing non-linear feature extraction to obtain a non-linear feature map after feature fusion; wherein, the historical frame is the first N frames before the current frame, and N is a natural number; Based on the non-linear feature map after feature fusion and the linear feature map of the current frame, obtaining a super-resolution reconstructed image of the current frame.

2. The super-resolution reconstruction method according to claim 1, characterized in that, The obtaining of the super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame includes: superimposing the non-linear feature map after feature fusion and the linear feature map of the current frame and performing pixel rearrangement, and then using a convolution kernel without an activation function to correct the pixels of the image after pixel rearrangement to obtain a super-resolution reconstructed image of the current frame.

3. The super-resolution reconstruction method according to claim 1 or 2, characterized in that, it further includes constructing a data set; The constructing of the data set includes: According to real-time motion data, rendering output videos at the initial resolution and the target resolution, constructing a data set at the initial resolution and a data set at the target resolution, and dividing the data set at the initial resolution and the data set at the target resolution into a training set and a test set according to a predetermined ratio.

4. The super-resolution reconstruction method according to claim 3, characterized in that, it further includes training the network model; The training of the network model includes: Using the image data at the initial resolution and the target resolution in the training set, calculating a loss value according to a preset loss function, and completing the training of the network model when the loss value and / or the accuracy rate of the training set and / or the accuracy rate of the test set reach a preset standard.

5. A super-resolution reconstruction method, characterized in that, After quantifying the network model trained by the method described in claim 4, deploying it to a hardware platform to achieve the target function.

6. A super-resolution reconstruction system, characterized in that, it includes a network model construction module; The network model construction module includes: A linear feature extraction module for performing linear feature extraction on the current frame to obtain a linear feature map of the current frame; A motion compensation module, for each frame in the historical frames, after splicing with the current frame, using a convolution kernel without an activation function for offset field compensation to obtain a frame motion compensation matrix of the historical frame, and then superimposing the frame motion compensation matrix of the historical frame with the historical frame to obtain the historical frame after motion compensation; A non-linear feature extraction module, which is used to perform non-linear feature extraction on the current frame to obtain the non-linear feature map of the previous frame, and is also used to perform non-linear feature extraction on the historical frame after motion compensation to obtain the non-linear feature map of the historical frame after motion compensation; A non-linear feature fusion module, which is used to splice the non-linear feature map after motion compensation and the non-linear feature map of the current frame, and then perform non-linear feature fusion to obtain the non-linear feature map after feature fusion; An image processing module, which is used to obtain the super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.

7. An electronic device, characterized in that, it includes the super-resolution reconstruction system described in claim 6; alternatively, the electronic device includes: a processor; a memory communicatively connected to the processor; the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is enabled to execute the method described in any one of claims 1 to 5.

8. A computer-readable storage medium storing computer instructions, characterized in that, when the computer instructions are executed by a processor, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Multi-frame image super-resolution reconstruction method based on optical flow motion estimation algorithm

    CN111696035A

  • Methods and Systems for Radial Basis Function Neural Network With Hammerstein Structure Based Non-Linear Interference Management in Multi-Technology Communications Devices

    US20160071007A1

Cited By

  • Convolutional neural network-based real-time super-resolution method using extra rendering information

    CN114820327A