A Real-time Super-resolution Reconstruction Method and System Based on Multi-frame Fusion
By extracting and fusion of current frames and historical frames, and combining inter-frame motion compensation, high-quality super-resolution image reconstruction in real-time scenes is achieved, solving the problems of real-time and picture continuity in the prior art.
Patent Information
- Application Number
- CN202111194772.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-13
AI Technical Summary
It is difficult for existing super-resolution technology to achieve high-quality image reconstruction in real-time scenarios, especially in games and monitoring scenarios. Traditional methods will lead to discontinuity and jitter in the picture, and the calculation volume is large, making it difficult to achieve real-time output.
A real-time super-resolution reconstruction method based on multi-frame fusion is adopted to generate high-resolution image reconstruction by extracting and fusion linear and nonlinear features of current frames and historical frames, combined with inter-frame motion compensation.
On the premise of ensuring real-time performance, the quality and visual effect of image reconstruction are improved, computing power needs are reduced, and the jump between frames is avoided, making the appearance more natural.
Smart Images

Figure CN113947528B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a real-time super-resolution reconstruction method and system based on multi-frame fusion. Background Art
[0002] Super-resolution technology is a technical means of transforming a low-resolution image into a high-resolution image, and is currently gradually applied to scenarios such as videos and games. Traditional super-resolution technologies include bilinear and bicubic interpolation methods. Although these super-resolution methods represented by interpolation methods have low requirements for computing power, their essence is still to fill in colors locally, and the effect is not very satisfactory.
[0003] Whether it is super-resolution technology for images or videos, it cannot effectively solve real-time super-resolution tasks (scenarios such as games and monitoring). For super-resolution technology for images, because it can only perform super-resolution frame by frame, applying it to real-time scenarios will cause discontinuity and jitter of the image. For super-resolution technology optimized for videos, because it usually needs to use the information of the previous and next frames of the video stream, it will cause at least N-frame delay; in addition, because the super-resolution network optimized for videos usually has a complex structure and a large amount of computation, the large-scale super-resolution process has high requirements for hardware and it is difficult to achieve real-time output. Summary of the Invention
[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a real-time super-resolution reconstruction method and system based on multi-frame fusion, which effectively solves the defects of the existing super-resolution technology in terms of real-time performance, and improves the quality and visual effect of image reconstruction on the premise of ensuring real-time performance.
[0005] To achieve the above object, according to one aspect of the present invention, a super-resolution reconstruction method is provided, including constructing a network model; constructing the network model includes: performing linear feature extraction on the current frame to obtain a linear feature map of the current frame; performing non-linear feature extraction and feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion; the historical frames are the first N frames of the current frame, where N is a natural number; based on the non-linear feature map after feature fusion and the linear feature map of the current frame, a super-resolution reconstructed image of the current frame is obtained.
[0006] In some embodiments, performing linear feature extraction on the current frame includes: performing linear feature extraction on the current frame using a convolution kernel without an activation function; performing non-linear feature extraction on the current frame and historical frames includes: performing non-linear feature extraction on the current frame using a convolution kernel activated by a non-linear function; performing non-linear feature extraction on the historical frames using a convolution kernel activated by a non-linear function.
[0007] In some embodiments, feature fusion includes: compensating for the inter-frame motion between each of the historical frames and the current frame using a convolutional kernel activated by a non-linear function to determine a compensated non-linear feature map; after concatenating the compensated non-linear feature map and the non-linear feature map of the current frame, performing non-linear feature fusion using a convolutional kernel activated by a non-linear function.
[0008] In some embodiments, obtaining a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame includes: superimposing the non-linear feature map after feature fusion and the linear feature map of the current frame and performing pixel rearrangement, and then using a convolutional kernel without an activation function to correct the pixels of the image after pixel rearrangement to obtain the super-resolution reconstructed image of the current frame.
[0009] In some embodiments, the method further includes constructing a data set; constructing the data set includes: rendering output videos at an initial resolution and a target resolution according to real-time motion data, constructing a data set at the initial resolution and a data set at the target resolution, and dividing the data set at the initial resolution and the data set at the target resolution into a training set and a test set according to a predetermined ratio.
[0010] In some embodiments, the method further includes training a network model; training the network model includes: using the image data at the initial resolution and the target resolution in the training set, calculating a loss value according to a preset loss function, and completing the training of the network model when the loss value and / or the accuracy rate of the training set and / or the accuracy rate of the test set reach a preset standard; wherein, the accuracy rate of the training set is the ratio of the number of pixel points that are reconstructed from the initial resolution to the target resolution using the network model in the training set to the total number of pixel points in the training set, and the accuracy rate of the test set is the ratio of the number of pixel points that are reconstructed from the initial resolution to the target resolution using the network model in the test set to the total number of pixel points in the test set.
[0011] In some embodiments, calculating the loss value using the Huber loss function is as follows:
[0012]
[0013] wherein, y is the image data at the target resolution in the training set or the test set, x is the image data at the initial resolution in the training set or the test set, f(x) is the result of the super-resolution reconstructed image calculated by x through the network model, and δ is a custom global parameter.
[0014] According to another aspect of the present invention, there is provided a super-resolution reconstruction method. After quantifying the network model trained by the above method, it is deployed to a hardware platform to achieve the target function.
[0015] According to another aspect of the present invention, there is provided a super-resolution reconstruction system, including a network model construction module; the network model construction module includes: a linear feature extraction module, configured to perform linear feature extraction on the current frame to obtain a linear feature map of the current frame; a non-linear feature extraction module, configured to perform non-linear feature extraction and feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion; the historical frames are the first N frames before the current frame, where N is a natural number; an image processing module, configured to obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.
[0016] In some embodiments, the linear feature extraction module is configured to perform linear feature extraction on the current frame using a convolutional kernel without an activation function; the non-linear feature extraction module is configured to perform non-linear feature extraction on the current frame using a convolutional kernel activated by a non-linear function, and is further configured to perform non-linear feature extraction on each of the first N frames before the current frame using a convolutional kernel activated by a non-linear function.
[0017] In some embodiments, the image processing module includes a motion compensation module, and the motion compensation module is configured to compensate for the inter-frame motion between each of the first N frames before the current frame and the current frame using a convolutional kernel activated by a non-linear function to obtain N compensated non-linear feature maps.
[0018] In some embodiments, the image processing module further includes a non-linear feature fusion module, and the non-linear feature fusion module is configured to splice the N compensated non-linear feature maps and the non-linear feature map of the current frame, and then perform non-linear feature fusion using a convolutional kernel activated by a non-linear function.
[0019] In some embodiments, the image processing module further includes a pixel rearrangement module, and the pixel rearrangement module is configured to superimpose the non-linear feature map after feature fusion and the linear feature map of the current frame and perform pixel rearrangement.
[0020] In some embodiments, the image processing module further includes a pixel correction module, and the pixel correction module is configured to perform pixel correction on the image after pixel rearrangement using a convolutional kernel without an activation function to obtain a super-resolution reconstructed image of the current frame.
[0021] In some embodiments, the system further includes a dataset construction module, configured to render output videos at an initial resolution and a target resolution according to real-time motion data, construct a dataset at the initial resolution and a dataset at the target resolution, and divide the dataset at the initial resolution and the dataset at the target resolution into a training set and a test set according to a predetermined ratio.
[0022] In some embodiments, the system further includes a network model training module, which is configured to use the image data of the initial resolution and the target resolution in the training set to calculate a loss value according to a preset loss function, and complete the training of the network model when the loss value and / or the accuracy of the training set and / or the accuracy of the test set reach a preset standard.
[0023] In some embodiments, the Huber loss function is used to calculate the loss value as follows:
[0024]
[0025] where y is the image data of the target resolution in the training set or the test set, x is the image data of the initial resolution in the training set or the test set, f(x) is the super-resolution reconstruction image result calculated by x through the network model, and δ is a custom global parameter.
[0026] According to another aspect of the present invention, there is provided an electronic device including the above super-resolution reconstruction system.
[0027] According to another aspect of the present invention, there is provided an electronic device, including: a processor; a memory communicatively connected to the processor; the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is enabled to execute the above method.
[0028] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.
[0029] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects are obtained:
[0030] 1. Real-time performance. Traditional VSR (Vedio Super Resolution) methods need to extract the features of future frames, while the technical solution of the present invention integrates historical frame data, avoiding the lag caused by extracting the features of future frames.
[0031] 2. Low computing power required for super-resolution of each frame. For example, the computing power required to super-resolve an original picture of (540*960*3) pixels into a picture of (1080*1920*3) pixels is 12.24 GFLOPs, and theoretically 81.70 frames can be rendered per T of computing power. The computing power required for the technical solution of the present invention to perform the same task is only about 1 / 57 of the computing power required by the DUF (2018 CVPR) network.
[0032] 3. The details between frames are rich and coherent. Due to the acquisition of the high-frequency texture information of historical frames, the super-resolution result of the technical solution of the present invention has richer detail performance, more accurate color restoration, and effectively suppresses the jumping phenomenon of details between frames, making the visual experience more natural.
[0033] 4. The dataset is more accurate. Different from the traditional method of constructing a dataset by downsampling a high-definition recorded video, the present invention uses a synchronous rendering method to construct a dataset. According to the same target information (such as game information), different-resolution datasets with temporal synchronization are rendered frame by frame, which is more suitable for actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flowchart of a real-time super-resolution reconstruction method based on multi-frame fusion according to an embodiment of the present invention;
[0035] Figure 2 is a schematic flowchart of constructing a network model according to an embodiment of the present invention;
[0036] Figure 3 is a schematic flowchart of constructing a dataset according to an embodiment of the present invention;
[0037] Figure 4 is Figure 2 a specific implementation flowchart of the process of constructing a network model shown;
[0038] Figure 5 is Figure 4 a specific implementation flowchart of the process of constructing a network model shown;
[0039] Figure 6 is a schematic structural diagram of a real-time super-resolution reconstruction system based on multi-frame fusion according to an embodiment of the present invention;
[0040] Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.
[0042] As Figure 1 shown, the real-time super-resolution reconstruction method based on multi-frame fusion according to an embodiment of the present invention includes the following steps:
[0043] Step 1: Construct a dataset.
[0044] Specifically, according to real-time motion data, render output videos at different resolutions, construct datasets at different resolutions, and divide the datasets at different resolutions into a training set and a test set according to a predetermined ratio.
[0045] In some embodiments, the different resolutions include a first resolution, a second resolution, and a third resolution. Among them, the different resolutions are all specified resolutions generated by motion data and a model file. The second resolution is greater than the first resolution, and the third resolution is greater than the second resolution. In some embodiments, the second resolution is an integer multiple of the first resolution, and the third resolution is an integer multiple of the first resolution and the second resolution. In some embodiments, the second resolution is 2 times the first resolution, and the third resolution is 4 times the first resolution and 2 times the second resolution. For example, the first resolution is 540P, the second resolution is 1080P, and the third resolution is 2160P.
[0046] In some embodiments, construct a first dataset at the first resolution, a second dataset at the second resolution, and a third dataset at the third resolution. In some embodiments, when performing super-resolution reconstruction with an input (initial resolution) of the first resolution and an output (target resolution) of the second resolution, divide the first dataset and the second dataset into a training set and a test set according to a predetermined ratio. It can be understood that in the embodiments of the present invention, the target resolution is greater than the initial resolution. For example, when performing 2-fold super-resolution reconstruction with an input of 540P and an output of 1080P, divide the dataset of 540P and the dataset of 1080P into a training set and a test set according to a predetermined ratio. In some embodiments, when performing super-resolution reconstruction with an input of the second resolution and an output of the third resolution, divide the second dataset and the third dataset into a training set and a test set according to a predetermined ratio. For example, when performing 2-fold super-resolution reconstruction with an input of 1080P and an output of 2160P, divide the dataset of 1080P and the dataset of 2160P into a training set and a test set according to a predetermined ratio. In some embodiments, when performing super-resolution reconstruction with an input of the first resolution and an output of the third resolution, divide the first dataset and the third dataset into a training set and a test set according to a predetermined ratio. For example, when performing 4-fold super-resolution reconstruction with an input of 540P and an output of 2160P, divide the dataset of 540P and the dataset of 2160P into a training set and a test set according to a predetermined ratio.
[0047] In some embodiments, the data proportion of the training set is greater than that of the test set. For example, the initial resolution dataset is divided into a training set and a test set in a ratio of 7:3. At the same time, the target resolution data corresponding to the initial resolution data in the training set is also included in the training set, and the target resolution data corresponding to the initial resolution data in the test set is also included in the test set.
[0048] Step 2: Construct a network model.
[0049] As Figure 2 shown, linear feature extraction is performed on the current frame to obtain the linear feature map of the current frame. Nonlinear feature extraction and feature fusion are performed on the current frame and historical frames to obtain the non - linear feature map after feature fusion. Based on the non - linear feature map after feature fusion and the linear feature map of the current frame, the super - resolution reconstructed image of the current frame is obtained.
[0050] In some embodiments, a convolution kernel without an activation function is used to perform linear feature extraction on the current frame F(i). In some embodiments, a convolution kernel activated by a non - linear function is used to perform non - linear feature extraction on the current frame F(i), and a convolution kernel activated by a non - linear function is used to perform non - linear feature extraction on the previous N (N satisfies N≥1, for example, Figure 2 in which N = 2) frames F(i - N)~F(i - 1) of the current frame respectively to obtain N + 1 temporally continuous non - linear feature maps.
[0051] In some embodiments, a convolution kernel activated by a non - linear function is used to compensate the inter - frame motion between each of the previous N frames of the current frame and the current frame. After splicing the N compensated non - linear feature maps with the non - linear feature map of the current frame, a convolution kernel activated by a non - linear function is used for non - linear feature fusion. The non - linear feature map after feature fusion is the unified expression of the non - linear features of the current frame and historical frames, removing some invalid feature information.
[0052] In some embodiments, the non - linear feature map after feature fusion and the linear feature map of the current frame are superimposed and pixel rearrangement (Pixel Shuffle) is performed to obtain the super - resolution reconstructed image of the current frame.
[0053] In some embodiments, constructing the network model further includes performing pixel correction on the image after pixel rearrangement. For example, a convolution kernel without an activation function is used to perform pixel correction on the image after pixel rearrangement to obtain the super - resolution reconstructed image of the current frame, completing the super - resolution reconstruction of the current frame.
[0054] In some embodiments, the non - linear function can be, for example, tanh, sigmoid or relu.
[0055] In this step, the linear feature map of the current frame corresponds to low-frequency information. Specifically, low-frequency information refers to the information of regions in the image where the color / gray level changes slowly, and such information widely exists in the color block filling regions inside the edges. The non-linear feature maps that are temporally continuous correspond to high-frequency information. Specifically, high-frequency information refers to the information of regions in the image where the color / gray level changes rapidly, and is generally closely related to the edge / texture information in the image. Through linear and non-linear feature extraction, the color block distribution and detail distribution of the image are obtained simultaneously, thereby effectively improving the super-resolution reconstruction quality of the model.
[0056] Step 3: Train the network model.
[0057] Using the image data of the initial resolution and the target resolution in the training set, calculate the loss value according to the preset loss function. When the loss value and the accuracies of the training set and the test set reach the preset standards, retain the model to complete the training of the model.
[0058] In some embodiments, the Huber loss function is used to calculate the loss value to enhance the robustness of the squared error loss function to outliers. The expression of the loss value is as follows:
[0059]
[0060] Where y is the image data of the target resolution in the training set, that is, the target output. x is the image data of the initial resolution in the training set, the f() function is the model calculation process, and f(x) is the super-resolution reconstruction image result calculated by x through the neural network model. δ is a custom global parameter. When the deviation value |y - f(x)| is less than δ, the squared error is used. When the deviation value |y - f(x)| is greater than δ, the linear error is used.
[0061] In some embodiments, the training set constructed in Step 1 is used to train the network model constructed in Step 2 using the backpropagation algorithm. That is, the initial resolution (low-resolution) data of the training set is input into the model, the model calculates and outputs the predicted high-resolution image, and then the predicted high-resolution image and the target resolution image (such as Ground Truth) of the training set are substituted into the loss function to calculate the loss value (loss). Then, the loss value is backpropagated to adjust the parameters of the model, and the above calculation process is repeated, and the iteration is continuously performed until the loss value and the accuracies of the training set and the test set both reach the preset standards to complete the training of the model.
[0062] In some embodiments, the accuracy rate is the ratio of the number of pixels for which super-resolution reconstruction from the initial resolution to the target resolution is achieved using a neural network model to the total number of pixels. Specifically, the accuracy rate of the training set is obtained by inputting the initial resolution data in the training set into the model, comparing the resolution data output by the model with the corresponding target resolution data in the training set, and dividing the number of pixels with the same resolution by the total number of pixels in the training set, which is the accuracy rate of the training set. Similarly, the accuracy rate of the test set is obtained by inputting the initial resolution data in the test set into the model, comparing the resolution data output by the model with the corresponding target resolution data in the test set, and dividing the number of pixels with the same resolution by the total number of pixels in the test set, which is the accuracy rate of the test set.
[0063] The accuracy rate of the test set is approximately the same as that in the real scenario, and the performance of the model in the real scenario can be predicted. In addition, the accuracy rate of the test set can also reflect the robustness of the model. When the accuracy rate of the test set is approximately the same as that of the training set, the model has strong robustness; when the accuracy rate of the test set and the training set differ significantly, the robustness of the model is poor.
[0064] In some embodiments, the preset criteria include: loss value ≤ 10 -5 , accuracy rate of the training set ≥ 99.9%, accuracy rate of the test set ≥ 99.9%.
[0065] Step 4: Deployment of the network model.
[0066] After quantizing the network model trained in Step 3 and deploying it to a hardware platform, the target function can be achieved.
[0067] The real-time super-resolution reconstruction method based on multi-frame fusion of the present invention will be described in detail below with specific examples. Set the resolution magnification factor as u, the number of basic convolution kernels as n, the size of the basic convolution kernel as k, the convolution stride as 1, and the activation function as AFunc. Taking the example of super-resolution from 540P resolution (original resolution) to 1080P resolution (target resolution), using 2 historical frames, and the King of Glory game, the resolution magnification factor u = 2, the number of basic convolution kernels n = 16 or 3, the size of the basic convolution kernel k = 3 or 5, and the activation function AFunc is tanh.
[0068] Combined with Figure 1 , the implementation process of the embodiment of the present invention is as follows:
[0069] Step 1: Construction of the dataset.
[0070] As Figure 3As shown in the figure, an innovative dataset construction method is proposed for the game scenario when constructing the dataset. Different from the commonly used method of downsampling from high-definition videos to obtain different resolutions, in the embodiments of the present invention, the output videos of the Honor of Kings game at different resolutions (540P, 1080P, 2160P) are rendered according to the in-game motion data (such as control instructions) and model files, and a dataset between different multiple relationships (2 two-fold datasets, 1 four-fold dataset) is constructed, and it is divided into a training set and a test set according to a ratio of 7:3.
[0071] Step 2: Construct a network model.
[0072] As Figure 4 shown in the figure, a multi-branch convolutional neural network is adopted. A convolutional kernel activated by the tanh function is used to extract non-linear features of the current frame F(i), a convolutional kernel without an activation function is used to extract linear features of the current frame F(i), and at the same time, convolutional kernels activated by the tanh function are used to extract non-linear features of the previous two frames F(i - 2) and F(i - 1) respectively, and convolutional kernels activated by the tanh function are used to compensate for the inter-frame motion between each of the two frames before the current frame and the current frame. After feature extraction and motion compensation, the three non-linear feature maps are spliced and then a convolutional kernel activated by the tanh function is used for non-linear feature fusion. The non-linear feature map after feature fusion and the linear feature map of the current frame are superimposed and pixel rearrangement (Pixel Shuffle) is performed, and finally a convolutional kernel without an activation function is used to correct the pixels of the image after pixel rearrangement, so as to complete the super-resolution reconstruction of the current frame.
[0073] In this step, the current frame and the historical N frames are used for feature extraction (N satisfies N≥1). During the process of linear and non-linear convolution, the number n of basic convolutional kernels is not limited (generally n = 2 a , a≥4 or n = 3a*u 2 , where a is a positive integer). It can be understood that several convolutional blocks are omitted at the ellipsis in Figure 4 . In actual design, they can be added according to requirements. This network collects the linear and non-linear features of the current frame, collects the non-linear features of the historical frames, and improves the detail performance of the picture based on the current frame. In addition, the computing power required by this network is low.
[0074] Specifically, as Figure 5As shown in the figure, 16 3x3 convolutional kernels without activation functions are used to extract linear features from the current frame. A convolutional layer consisting of 16 3x3 convolutional kernels activated by tanh is constructed, followed by a convolutional layer consisting of 16 3x3 convolutional kernels activated by tanh, which are used to extract non-linear features from the current frame F(i), historical frames F(i - 1) and F(i - 2) respectively. After the feature extraction of historical frames F(i - 1) and F(i - 2) is completed, the motion compensation function is completed through 2 identical branches (2 convolutional layers consisting of 16 3x3 convolutional kernels activated by tanh).
[0075] Then, the non-linear features of the current frame F(i) and the motion-compensated non-linear features are concatenated, and non-linear feature fusion is performed through a convolutional layer consisting of 16 3x3 convolutional kernels activated by tanh. After being superimposed with the linear features of the current frame, pixel rearrangement is performed to preliminarily reconstruct the super-resolution result of the current frame. Finally, the final super-resolution reconstruction result is obtained after feature correction through a convolutional layer consisting of 3 3x3 convolutional kernels without activation functions.
[0076] Step 3: Train the network model.
[0077] Using the image data of the initial resolution and the target resolution in the training set, the loss value is calculated. Here, the Huber loss function is used to enhance the robustness of the squared error loss function to outliers.
[0078] The model constructed in Step 1 is trained using the backpropagation algorithm with the dataset mentioned above, and the iteration continues until the loss value ≤ 10 -5 , the accuracy rate of the training set ≥ 99.9%, the accuracy rate of the test set ≥ 99.9%, and this model is retained.
[0079] Step 4: Deploy the network model.
[0080] After quantifying the model trained in Step 3, it is deployed to the hardware acceleration platform, enabling the mobile processor to obtain smooth and stable high-resolution and high-quality real-time results while maintaining the low load level of the native rendering of 540P when rendering the King of Glory game. This model can be applied to the games involved in the dataset for scenarios such as real-time super-resolution and game video super-resolution.
[0081] As Figure 6 shown, the embodiment of the present application also provides a real-time super-resolution reconstruction system based on multi-frame fusion, including:
[0082] A dataset construction module, which is used to render the output videos at different resolutions according to real-time motion data, construct datasets at different resolutions, and divide the datasets at different resolutions into a training set and a test set according to a predetermined ratio.
[0083] In some embodiments, the different resolutions include a first resolution, a second resolution, and a third resolution. Among them, the different resolutions are all specified resolutions generated from motion data and model files. The second resolution is greater than the first resolution, and the third resolution is greater than the second resolution. In some embodiments, the second resolution is an integer multiple of the first resolution, and the third resolution is an integer multiple of both the first resolution and the second resolution. In some embodiments, the second resolution is 2 times the first resolution, and the third resolution is 4 times the first resolution and 2 times the second resolution. For example, the first resolution is 540P, the second resolution is 1080P, and the third resolution is 2160P.
[0084] In some embodiments, a first data set at the first resolution, a second data set at the second resolution, and a third data set at the third resolution are constructed. In some embodiments, for super-resolution reconstruction where the input (initial resolution) is the first resolution and the output (target resolution) is the second resolution, the first data set and the second data set are divided into a training set and a test set according to a predetermined ratio. For example, for 2-fold super-resolution reconstruction where the input is 540P and the output is 1080P, the 540P data set and the 1080P data set are divided into a training set and a test set according to a predetermined ratio. In some embodiments, for super-resolution reconstruction where the input is the second resolution and the output is the third resolution, the second data set and the third data set are divided into a training set and a test set according to a predetermined ratio. For example, for 2-fold super-resolution reconstruction where the input is 1080P and the output is 2160P, the 1080P data set and the 2160P data set are divided into a training set and a test set according to a predetermined ratio. In some embodiments, for super-resolution reconstruction where the input is the first resolution and the output is the third resolution, the first data set and the third data set are divided into a training set and a test set according to a predetermined ratio. For example, for 4-fold super-resolution reconstruction where the input is 540P and the output is 2160P, the 540P data set and the 2160P data set are divided into a training set and a test set according to a predetermined ratio.
[0085] In some embodiments, the proportion of data in the training set is greater than that in the test set. For example, the initial resolution data set is divided into a training set and a test set at a ratio of 7:3. At the same time, the target resolution data corresponding to the initial resolution data in the training set is also included in the training set, and the target resolution data corresponding to the initial resolution data in the test set is also included in the test set.
[0086] A network model construction module, which is used to extract linear features from the current frame to obtain a linear feature map of the current frame, perform non-linear feature extraction and feature fusion on the current frame and historical frames to obtain a non-linear feature map after feature fusion, and obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.
[0087] In some embodiments, the network model construction module includes a linear feature extraction module, which is used to extract linear features from the current frame F(i) using a convolutional kernel without an activation function. In some embodiments, the network model construction module further includes a non-linear feature extraction module, which is used to extract non-linear features from the current frame F(i) using a convolutional kernel activated by a non-linear function, and use a convolutional kernel activated by a non-linear function to extract non-linear features from each of the previous N (N satisfies ≥1, for example, Figure 2 in which N = 2) frames F(i-N) to F(i-1) of the current frame respectively to obtain N+1 temporally continuous non-linear feature maps.
[0088] In some embodiments, the network model construction module includes a motion compensation module, which is used to compensate the inter-frame motion between each of the previous N frames of the current frame and the current frame using a convolutional kernel activated by a non-linear function to obtain N compensated non-linear feature maps. The network model construction module further includes a non-linear feature fusion module, which is used to splice the N compensated non-linear feature maps with the non-linear feature map of the current frame and then perform non-linear feature fusion using a convolutional kernel activated by a non-linear function. The non-linear feature map after feature fusion is the unified expression of the non-linear features of the current frame and historical frames, removing some invalid feature information.
[0089] In some embodiments, the network model construction module further includes a pixel rearrangement module, which is used to superimpose the non-linear feature map after feature fusion and the linear feature map of the current frame and perform pixel rearrangement (Pixel Shuffle) to obtain a super-resolution reconstructed image of the current frame.
[0090] In some embodiments, the network model construction module further includes a pixel correction module, which is used to correct the pixels of the image after pixel rearrangement. For example, use a convolutional kernel without an activation function to correct the pixels of the image after pixel rearrangement to obtain a super-resolution reconstructed image of the current frame, completing the super-resolution reconstruction of the current frame.
[0091] In some embodiments, the non-linear function can be, for example, tanh, sigmoid or relu.
[0092] A network model training module is used to calculate a loss value by using the image data of the initial resolution and the target resolution in the training set according to a preset loss function. When the loss value, as well as the accuracies of the training set and the test set, reach the preset standards, the model is retained to complete the training of the model.
[0093] In some embodiments, the Huber loss function is used to calculate the loss value to enhance the robustness of the squared error loss function to outliers. The expression of the loss value is as follows:
[0094]
[0095] Where y is the image data of the target resolution in the training set, that is, the target output. x is the image data of the initial resolution in the training set, the f() function is the model calculation process, and f(x) is the super-resolution reconstruction image result calculated by x through the neural network model. δ is a custom global parameter. When the deviation value |y - f(x)| is less than δ, the squared error is used. When the deviation value |y - f(x)| is greater than δ, the linear error is used.
[0096] In some embodiments, the network model constructed by the network model construction module is trained using the backpropagation algorithm with the dataset constructed by the dataset construction module. That is, the initial resolution (low resolution) data of the training set is input into the model, the model calculates and outputs the predicted high-resolution image, and then the predicted high-resolution image and the target resolution image (Ground Truth) of the training set are substituted into the loss function to calculate the loss value (loss). Then, the loss value is backpropagated to adjust the parameters of the model, and the above calculation process is repeated and iterated continuously until the loss value and accuracy of the training set, as well as the loss value and accuracy of the test set, reach the preset standards to complete the training of the model.
[0097] In some embodiments, the preset standards include: the loss value of the training set ≤ 10 -5 , the accuracy of the training set ≥ 99.9%, the loss value of the test set ≤ 10 -5 , the accuracy of the test set ≥ 99.9%.
[0098] A network model deployment module is used to quantize the network model trained by the network model training module and then deploy it to the hardware platform to achieve the target function.
[0099] Figure 7 It is a structural block diagram of an electronic device according to an embodiment of the present application. An embodiment of the present application also provides an electronic device, such as Figure 7As shown, the electronic device includes: at least one processor 701, and a memory 703 communicatively connected to the at least one processor 701. Instructions executable by the at least one processor 701 are stored in the memory 703. The instructions are executed by the at least one processor 701. When the processor 701 executes the instructions, the driving scenario reconstruction method in the above embodiments is implemented. The number of the memory 703 and the processor 701 may be one or more. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0100] The electronic device may further include a communication interface 705 for communicating with external devices and performing data interaction and transmission. Each device is interconnected using different buses and may be mounted on a common motherboard or otherwise installed as required. The processor 701 may process instructions executed within the electronic device, including instructions for storing graphical information in the memory or on the memory for displaying a graphical user interface (GUI) on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses may be used together with multiple memories and multiple memories. Similarly, multiple electronic devices may be connected, with each device providing some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown herein, but it does not mean that there is only one bus or one type of bus.
[0101] Optionally, in a specific implementation, if the memory 703, the processor 701, and the communication interface 705 are integrated on a single chip, the memory 703, the processor 701, and the communication interface 705 may communicate with each other through an internal interface.
[0102] It should be understood that the above-mentioned processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0103] An embodiment of the present application provides a computer-readable storage medium (such as the above-mentioned memory 703), which stores computer instructions. When the program is executed by the processor, the method provided in the embodiment of the present application is implemented.
[0104] Optionally, the memory 703 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device of the driving scenario reconstruction method, etc. In addition, the memory 703 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 703 may optionally include a memory remotely set relative to the processor 701, and these remote memories may be connected to the electronic device of the driving scenario reconstruction method through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0105] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0106] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0107] Any process or method description represented in a flowchart or described otherwise herein can be understood to represent code of an executable instruction including one or more (two or more) modules, segments, or portions for implementing specific logical functions or processes. And the scope of the preferred embodiments of this application includes additional implementations, where functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.
[0108] The logic and / or steps represented in a flowchart or described otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices.
[0109] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiments.
[0110] In addition, each functional unit in various embodiments of this application can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0111] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A super-resolution reconstruction method, characterized in that, it includes constructing a network model; The constructing of the network model includes: Performing linear feature extraction on the current frame to obtain a linear feature map of the current frame; Performing non-linear feature extraction on the current frame to obtain a non-linear feature map of the current frame; Performing non-linear feature extraction on historical frames, compensating the inter-frame motion between each historical frame and the current frame using a convolution kernel activated by a non-linear function, and determining a compensated non-linear feature map; After splicing the compensated non-linear feature map and the non-linear feature map of the current frame, performing non-linear feature fusion using a convolution kernel activated by a non-linear function to obtain a non-linear feature map after feature fusion; The historical frames are the first N frames of the current frame, where N is a natural number; Based on the non-linear feature map after feature fusion and the linear feature map of the current frame, obtaining a super-resolution reconstructed image of the current frame.
2. The super-resolution reconstruction method according to claim 1, characterized in that, The performing of linear feature extraction on the current frame includes: performing linear feature extraction on the current frame using a convolution kernel without an activation function; The performing of non-linear feature extraction on the current frame includes: performing non-linear feature extraction on the current frame using a convolution kernel activated by a non-linear function; The performing of non-linear feature extraction on historical frames includes: performing non-linear feature extraction on the historical frames using a convolution kernel activated by a non-linear function.
3. The super-resolution reconstruction method according to claim 1, characterized in that, The obtaining of the super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame includes: superimposing the non-linear feature map after feature fusion and the linear feature map of the current frame and performing pixel rearrangement, and then performing pixel correction on the image after pixel rearrangement using a convolution kernel without an activation function to obtain a super-resolution reconstructed image of the current frame.
4. The super-resolution reconstruction method according to any one of claims 1 to 3, characterized in that, it further includes constructing a data set; The constructing of the data set includes: According to real-time motion data, rendering output videos at an initial resolution and a target resolution, constructing a data set at the initial resolution and a data set at the target resolution, and dividing the data set at the initial resolution and the data set at the target resolution into a training set and a test set according to a predetermined ratio.
5. The super-resolution reconstruction method according to claim 4, characterized in that, it further includes training the network model; The training of the network model includes: Using the image data at the initial resolution and the target resolution in the training set, calculating a loss value according to a preset loss function, and completing the training of the network model when the loss value and / or the accuracy rate of the training set and / or the accuracy rate of the test set reach a preset standard.
6. A super-resolution reconstruction method, characterized in that, After quantizing the network model trained by the method described in claim 5, deploying it to a hardware platform to achieve the target function.
7. A super-resolution reconstruction system, characterized in that, it includes a network model construction module; The network model construction module includes: A linear feature extraction module, configured to perform linear feature extraction on a current frame to obtain a linear feature map of the current frame; A non-linear feature extraction module, configured to perform non-linear feature extraction on the current frame to obtain a non-linear feature map of the current frame; further configured to perform non-linear feature extraction on historical frames, and use a convolutional kernel activated by a non-linear function to compensate for the inter-frame motion between each of the historical frames and the current frame to determine a compensated non-linear feature map; further configured to, after splicing the compensated non-linear feature map and the non-linear feature map of the current frame, use a convolutional kernel activated by a non-linear function to perform non-linear feature fusion to obtain a non-linear feature map after feature fusion; the historical frames are the first N frames before the current frame, where N is a natural number; An image processing module, configured to obtain a super-resolution reconstructed image of the current frame based on the non-linear feature map after feature fusion and the linear feature map of the current frame.
8. An electronic device, characterized in that it includes the super-resolution reconstruction system according to claim 7; alternatively, the electronic device includes: a processor; a memory communicatively connected to the processor; the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the processor is enabled to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on deep residual network and storage medium
CN110689483A
Image super-resolution reconstruction method and device
CN113191955A