Video Super-Resolution Reconstruction Method and Device Based on Deep Learning

By querying the preset feature table on the end-side device and performing mean processing, video super-resolution reconstruction based on deep learning is realized, solving the problem of poor video reconstruction results caused by limited hardware configuration of the device, and achieving fast and efficient video reconstruction effect.

CN115205108BActive Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210544518.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-06-20
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

In the prior art, when the video super-resolution reconstruction method based on deep learning is run on the end-side device, it is difficult to achieve good high-resolution video reconstruction effect, mainly due to the limited hardware configuration, it is difficult to support the operation of the model.

Method used

A video super-resolution reconstruction method based on deep learning is proposed. By obtaining low-resolution video frames, querying preset feature tables, obtaining feature maps, and performing mean processing, the super-resolution reconstruction of video frames is realized. This method reduces the computational load of the end-side device by constructing a pre-trained super-score model.

Benefits of technology

Fast and efficient super-resolution video frame reconstruction is achieved on the end-side device, and the reconstruction effect is good, solving the problem of poor video reconstruction results caused by limited hardware configuration of the equipment in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205108B_ABST
    Figure CN115205108B_ABST
Patent Text Reader

Abstract

One or more embodiments of this specification provide a video super-resolution reconstruction method and apparatus based on deep learning, including: obtaining a low-resolution video frame; querying a preset first feature table according to the low-resolution video frame to obtain a plurality of first feature maps; querying a preset second feature table according to the plurality of first feature maps to obtain a plurality of second feature maps; performing a mean processing on the plurality of second feature maps to obtain a super-resolution video frame. The method of this embodiment can achieve fast and efficient super-resolution video frame reconstruction on a terminal device, and the reconstruction effect is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of video processing technologies, and in particular, to a video super-resolution reconstruction method and apparatus based on deep learning. Background Art

[0002] Video super-resolution (Video Super-Resolution) reconstruction is a process of using software or hardware to process a low-resolution video to obtain a high-resolution video. The reconstructed high-resolution video can greatly improve the user's viewing experience.

[0003] Existing video super-resolution reconstruction methods can be divided into two types. One is a traditional video processing method, and the other is a video super-resolution reconstruction method based on deep learning. The high-resolution video reconstruction effect of the former is worse than that of the latter. The model constructed by the latter based on deep learning depends on the powerful floating-point computing power and storage capacity provided by a graphics processing unit. Since most video-playing end-side devices (such as surveillance cameras, televisions, mobile devices, etc.) have limited hardware configurations, it is difficult to support the operation of the model and difficult to achieve a good high-resolution video reconstruction effect. Summary of the Invention

[0004] In view of this, the purpose of one or more embodiments of this specification is to propose a video super-resolution reconstruction method and apparatus based on deep learning to solve the problem of poor high-resolution video reconstruction effect on end-side devices.

[0005] Based on the above purpose, one or more embodiments of this specification provide a video super-resolution reconstruction method based on deep learning, including:

[0006] Obtain a low-resolution video frame;

[0007] According to the low-resolution video frame, query a preset first feature table to obtain a plurality of first feature maps;

[0008] According to the plurality of first feature maps, query a preset second feature table to obtain a plurality of second feature maps; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model, the super-resolution model includes a first convolutional layer, a second convolutional layer, and an upsampling layer, the first stage table is constructed according to all the first input data and the corresponding all the first output data of the first convolutional layer, and the second stage table is constructed according to all the second input data and the corresponding all the second output data of the second convolutional layer and the upsampling layer;

[0009] Perform a mean process on the plurality of second feature maps to obtain a super-resolution video frame.

[0010] Optionally, the receptive field of the first convolutional layer is the first receptive field; according to the low-resolution video frame, query a preset first feature table to obtain a plurality of first feature maps, including:

[0011] For each first receptive field of the low-resolution video frame, query the first feature table to obtain the output data corresponding to each first receptive field;

[0012] According to the output data, obtain a plurality of first feature maps.

[0013] Optionally, the receptive field of the second convolutional layer is the second receptive field; according to a plurality of first feature maps, query a preset second feature table to obtain a plurality of second feature maps, including:

[0014] According to each second receptive field of each first feature map, query the second feature table to obtain the output data corresponding to each second receptive field;

[0015] According to the output data of each first feature map, obtain the corresponding second feature map.

[0016] Optionally, before obtaining the low-resolution video frame, further include:

[0017] Obtain a low-resolution video frame sample and a high-resolution video frame sample of the same video frame;

[0018] Use the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model.

[0019] Optionally, the initial super-resolution model includes an initial first convolutional layer, an initial second convolutional layer, and an initial upsampling layer;

[0020] Using the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model, including:

[0021] Input the low-resolution video frame sample into the initial first convolutional layer, and the initial first convolutional layer outputs a plurality of first feature map samples;

[0022] Input the plurality of first feature map samples into the initial second convolutional layer, the output result of the initial second convolutional layer is input into the initial upsampling layer, and the initial upsampling layer outputs a plurality of second feature map samples;

[0023] Perform mean processing on the plurality of second feature map samples to obtain a high-resolution video frame prediction result;

[0024] Determine the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample, and adjust the parameters of the initial super-resolution model according to the difference degree;

[0025] Retrain the initial super-resolution model according to the adjusted parameters until the training end condition is met. Take the initial super-resolution model that meets the training end condition as the super-resolution model, and take the initial first convolutional layer, the initial second convolutional layer, and the initial upsampling layer corresponding to the initial super-resolution model that meets the training end condition as the first convolutional layer, the second convolutional layer, and the upsampling layer respectively.

[0026] Optionally, inputting the low-resolution video frame sample into the initial first convolutional layer includes:

[0027] Perform padding processing on the low-resolution video frame sample to obtain a padded low-resolution video frame sample;

[0028] Input the padded low-resolution video frame sample into the initial first convolutional layer; wherein, the size of the first feature map sample is the same as that of the low-resolution video frame sample;

[0029] Inputting multiple first feature map samples into the initial second convolutional layer includes:

[0030] Perform padding processing on the first feature map sample to obtain a padded first feature map sample;

[0031] Input the padded first feature map sample into the initial second convolutional layer; wherein, the output result of the initial second convolutional layer is the same as the size of the low-resolution video frame sample.

[0032] Optionally, the receptive field of the first convolutional layer is the first receptive field; after obtaining the super-resolution model, it further includes:

[0033] Construct all the input data of the first receptive field, input all the input data of the first receptive field into the first convolutional layer, obtain all the output data corresponding to the first receptive field, and construct the first feature table according to all the input data and the corresponding all the output data of the first receptive field.

[0034] Optionally, the receptive field of the second convolutional layer is the second receptive field; after obtaining the super-resolution model, it further includes:

[0035] Construct all the input data of the second receptive field, input all the input data of the second receptive field into the second convolutional layer, input the output result of the second convolutional layer into the upsampling layer, obtain all the output data corresponding to the second receptive field, and construct the second feature table according to all the input data and the corresponding all the output data of the second receptive field.

[0036] Optionally, constructing the first feature table according to all the input data and corresponding all the output data of the first receptive field, including:

[0037] Performing integerization processing on all the output data of the first receptive field to obtain integerized all the output data;

[0038] Constructing the first feature table according to all the input data of the first receptive field and the corresponding integerized all the output data;

[0039] Constructing the second feature table according to all the input data and corresponding all the output data of the second receptive field, including:

[0040] Performing integerization processing on all the output data of the second receptive field to obtain integerized all the output data;

[0041] Constructing the second feature table according to all the input data of the second receptive field and the corresponding integerized all the output data.

[0042] The embodiment of the present specification further provides a video super-resolution reconstruction device based on deep learning, including:

[0043] An acquisition module, configured to acquire a low-resolution video frame;

[0044] A first look-up table module, configured to query a preset first feature table according to the low-resolution video frame to obtain a plurality of first feature maps;

[0045] A second look-up table module, configured to query a preset second feature table according to a plurality of first feature maps to obtain a plurality of second feature maps; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model, the super-resolution model includes a first convolutional layer, a second convolutional layer and an upsampling layer, the first stage table is constructed according to the first all input data and corresponding first all output data of the first convolutional layer, and the second stage table is constructed according to the second all input data and corresponding second all output data of the second convolutional layer and the upsampling layer;

[0046] A mean processing module, configured to perform mean processing on a plurality of second feature maps to obtain a super-resolution video frame.

[0047] As can be seen from the above, the method and apparatus for video super-resolution reconstruction based on deep learning provided by one or more embodiments of this specification obtain a low-resolution video frame, query a preset first feature table according to the low-resolution video frame to obtain a plurality of first feature maps, query a preset second feature table according to the plurality of first feature maps to obtain a plurality of second feature maps, and perform mean processing on the plurality of second feature maps to obtain a super-resolution video frame. The method of this embodiment can achieve fast and efficient super-resolution video frame reconstruction on a terminal device, and the reconstruction effect is good. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only one or more embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 is a schematic flowchart of the method for one or more embodiments of this specification;

[0050] Figure 2 is a flowchart of the method for another embodiment of this specification;

[0051] Figure 3 is a flowchart of the method for training a model for one or more embodiments of this specification;

[0052] Figure 4A 、 4B are respectively flowcharts of the methods for generating the first and second feature tables for one or more embodiments of this specification;

[0053] Figure 5 is a block diagram of the device structure for one or more embodiments of this specification;

[0054] Figure 6 is a block diagram of the electronic device structure for one or more embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following will further elaborate on this disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0056] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of this specification should have the ordinary meaning understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in one or more embodiments of this specification do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0057] As Figure 1 , 2 shown, one or more embodiments of this specification provide a video super-resolution reconstruction method based on deep learning, including:

[0058] S101: Obtain a low-resolution video frame;

[0059] In this embodiment, first, a low-resolution video frame to be processed is obtained. The low-resolution video frame can be a frame in a low-resolution video stored in an electronic device, or a frame in a low-resolution video captured by an image acquisition unit configured in the electronic device, or a frame in a low-resolution video received by the electronic device from other devices. The acquisition method of the low-resolution video frame is not limited.

[0060] In some ways, the electronic device can be a fixed terminal such as a television or a monitoring device, or a mobile terminal such as a tablet computer or a smart phone. The specific form of the electronic device is not limited. The hardware configuration and performance of such electronic devices are limited, and in some application scenarios, there is a functional requirement for video super-resolution processing.

[0061] S102: Query a preset first feature table according to the low-resolution video frame to obtain a plurality of first feature maps;

[0062] S103: Query a preset second feature table according to the plurality of first feature maps to obtain a plurality of second feature maps; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model. The super-resolution model includes a first convolutional layer, a second convolutional layer, and an upsampling layer. The first-stage table is constructed according to the first all input data and the corresponding first all output data of the first convolutional layer. The second-stage table is constructed according to the second all input data and the corresponding second all output data of the second convolutional layer and the upsampling layer;

[0063] In this embodiment, a first feature table and a second feature table are stored in the electronic device. When converting a low-resolution video frame into a high-resolution video frame, first query the first feature table according to the low-resolution video frame to obtain multiple first feature maps obtained by the query; then, query the second feature table according to each first feature map to obtain multiple second feature maps. Both the first feature table and the second feature table are generated according to a pre-constructed super-resolution model. Compared with the method of inputting a low-resolution video frame into the model and outputting a super-resolution video frame by the model, the table lookup method can greatly improve the video frame processing efficiency.

[0064] S104: Perform mean processing on multiple second feature maps to obtain a super-resolution video frame.

[0065] In this embodiment, after obtaining multiple second feature maps by querying the second feature table, perform mean processing based on the multiple second feature maps to obtain the super-resolution video frame corresponding to the conversion of the low-resolution video frame. Thus, the method of this embodiment only needs to sequentially query the first feature table and the second feature table according to the low-resolution video frame, and then perform mean processing based on the query results. Compared with inputting each low-resolution video frame into the model and having the model process the low-resolution video frame and output the corresponding high-resolution video frame, it can greatly improve the video frame processing efficiency, and the data processing performance of table lookup and mean processing has low requirements for the configuration and performance of the electronic device, and can be suitable for end-side devices to achieve fast and efficient video frame processing.

[0066] In some embodiments, to ensure the reconstruction effect from a low-resolution video frame to a high-resolution video frame and the running efficiency on an end-side device, the super-resolution model includes a first convolutional layer and a second convolutional layer with a relatively small receptive field. The size of the receptive field determines the length of the first feature table and the second feature table, and also determines the query efficiency of the first feature table and the second feature table. For example, if the pixel value range of the video frame is 0 - 255 and the receptive field size is k, there may be 256 k kinds of input data. Optionally, to ensure the super-resolution reconstruction effect, the first receptive field of the first convolutional layer is 1×2, the receptive field of the second convolutional layer is 2×1, and the receptive field of the super-resolution model is approximately 2×2.

[0067] In some embodiments, the receptive field of the first convolutional layer is the first receptive field; according to the low-resolution video frame, query the preset first feature table to obtain multiple first feature maps, including:

[0068] For each first receptive field of the low-resolution video frame, query the first feature table to obtain the output data corresponding to each first receptive field;

[0069] According to the output data, obtain multiple first feature maps.

[0070] In this embodiment, the constructed first feature table is indexed by the first receptive field of the low-resolution video frame. According to a predetermined receptive field step size, each first receptive field of the low-resolution video frame is determined. Based on each first receptive field, the first feature table is queried to obtain the output data corresponding to each first receptive field. For example, if the size of the first receptive field is 1×2, and the pixel values of the pixel at the first row and first column and the pixel at the first row and second column of the low-resolution video frame are 0 and 0 respectively, then the first feature table is queried with (0, 0) as the index to obtain the corresponding output data.

[0071] The output data obtained by querying the first feature table according to the first receptive field corresponds to a pixel value of each first feature map. For example, when the first feature table is queried with (0, 0) as the index, the obtained output data is (0, 1, 0, 1). Each value in the output data corresponds to a pixel value in one of the four first feature maps. 0 corresponds to the pixel value at the first row and first column of the first feature map, 1 corresponds to the pixel value at the first row and first column of the second feature map, 0 corresponds to the pixel value at the first row and first column of the third feature map, and 1 corresponds to the pixel value at the first row and first column of the fourth feature map. Thus, the output data obtained by querying the first feature table according to the low-resolution video frame can correspond to multiple first feature maps.

[0072] In some embodiments, the receptive field of the second convolutional layer is the second receptive field; according to multiple first feature maps, a preset second feature table is queried to obtain multiple second feature maps, including:

[0073] According to each second receptive field of each first feature map, the second feature table is queried to obtain the output data corresponding to each second receptive field;

[0074] Based on the output data of each first feature map, the corresponding second feature map is obtained.

[0075] In this embodiment, after querying the first feature table according to the low-resolution video frame, multiple first feature maps are obtained. Then, further according to each first feature map, the second feature table is queried to obtain multiple second feature maps. For each first feature map, according to a predetermined receptive field step size, each second receptive field of the first feature map is determined. Based on each second receptive field, the second feature table is queried to obtain the output data corresponding to each second receptive field. After obtaining all the output data of the first feature map, the second feature map corresponding to the first feature map can be determined according to the output data.

[0076] After obtaining multiple second feature maps, the pixel values of the corresponding pixels of the multiple second feature maps are averaged. The super-resolution video frame is composed of the pixel values after the averaging process, thus completing the reconstruction of the super-resolution video frame from the low-resolution video frame through table lookup and averaging process. The reconstruction processing speed is fast, the efficiency is high, and the reconstruction effect is good.

[0077] In some embodiments, the method for training a super-resolution model includes:

[0078] Obtain a low-resolution video frame sample and a high-resolution video frame sample of the same video frame;

[0079] Use the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model.

[0080] In this embodiment, the first feature table and the second feature table are stored in the electronic device instead of the super-resolution model, and the reconstructed super-resolution video frame is directly obtained by looking up the table for the low-resolution video frame. To obtain the first feature table and the second feature table, a super-resolution model capable of reconstructing a low-resolution video frame into a super-resolution video frame is pre-trained, and the first feature table and the second feature table are constructed and generated according to the super-resolution model. To train the super-resolution model, first construct an initial super-resolution model, and obtain the low-resolution video frame sample and the high-resolution video frame sample for training the model, and use the low-resolution video frame sample and the high-resolution video frame sample to train the initial super-resolution model. After training, the super-resolution model is obtained.

[0081] It can be understood that the electronic device for training the super-resolution model can be a device with relatively high resource configurations such as computing and storage, that is, the super-resolution model is trained on a high-configuration device (for example, a server, a computing center, etc.), the first feature table and the second feature table are constructed based on the super-resolution model, and the first feature table and the second feature table are directly used on the edge device to reconstruct the super-resolution video frame by the table lookup method. In this way, the high-configuration device only needs to train the super-resolution model once, and the super-resolution video frame reconstruction can be realized on multiple distributed edge devices. There is no need for the high-configuration device to execute the video frame reconstruction task, and the edge device can quickly and efficiently complete the low-resolution video frame reconstruction task.

[0082] In some embodiments, the initial super-resolution model includes an initial first convolutional layer, an initial second convolutional layer, and an initial upsampling layer;

[0083] Using the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model includes:

[0084] Input the low-resolution video frame sample into the initial first convolutional layer, and the initial first convolutional layer outputs multiple first feature map samples;

[0085] Input the multiple first feature map samples into the initial second convolutional layer, and the output result of the initial second convolutional layer is input into the initial upsampling layer, and the initial upsampling layer outputs multiple second feature map samples;

[0086] Perform mean processing on the multiple second feature map samples to obtain the high-resolution video frame prediction result;

[0087] Determine the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample, and adjust the parameters of the initial super-resolution model according to the difference degree;

[0088] According to the adjusted parameters, retrain the initial super-resolution model until the training end condition is satisfied. Take the initial super-resolution model that meets the training end condition as the super-resolution model, and take the initial first convolutional layer, the initial second convolutional layer, and the initial upsampling layer of the initial super-resolution model that meets the training end condition as the first convolutional layer, the second convolutional layer, and the upsampling layer respectively.

[0089] As Figure 3 shown, in this embodiment, the constructed initial super-resolution model includes an initial first convolutional layer, an initial second convolutional layer, and an initial upsampling layer. The method for training the initial super-resolution model is to input the low-resolution video frame sample into the initial first convolutional layer. After being processed by the initial first convolutional layer, multiple first feature map samples are obtained. Then, the multiple first feature map samples are input into the initial second convolutional layer. After being processed by the initial second convolutional layer, multiple intermediate second feature map samples are obtained. The multiple intermediate second feature map samples are input into the initial upsampling layer. After being processed by the initial upsampling layer, multiple second feature map samples are output. Based on the multiple second feature map samples, an averaging process is performed to obtain the high-resolution video frame prediction result. Compare the obtained high-resolution video frame prediction result with the high-resolution video frame sample, calculate the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample, and then adjust the parameters of the initial super-resolution model according to the difference degree. For example, adjust the parameters of the initial first convolutional layer, the initial second convolutional layer, and the initial upsampling layer. After the parameter adjustment, input the low-resolution video frame sample into the model with adjusted parameters again according to the above process, and obtain the corresponding high-resolution video frame prediction result, and calculate the difference degree; repeat the above process. When the number of iterative training reaches a preset value, or the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample obtained in a certain iteration is less than the preset difference degree threshold, the model training ends. The model after training is the super-resolution model. The initial first convolutional layer corresponding to this super-resolution model is the first convolutional layer, the initial second convolutional layer is the second convolutional layer, and the initial upsampling layer is the upsampling layer.

[0090] Optionally, the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample is determined according to the mean square error value of the two.

[0091] In some embodiments, inputting the low-resolution video frame sample into the initial first convolutional layer includes:

[0092] Perform padding processing on the low-resolution video frame sample to obtain a padded low-resolution video frame sample;

[0093] Input the low-resolution video frame samples after padding into the initial first convolutional layer; among them, the size of the first feature map samples is the same as that of the low-resolution video frame samples;

[0094] Input multiple first feature map samples into the initial second convolutional layer, including:

[0095] Perform padding processing on the first feature map samples to obtain the padded first feature map samples;

[0096] Input the padded first feature map samples into the initial second convolutional layer; among them, the output result of the initial second convolutional layer is the same size as the low-resolution video frame samples.

[0097] In this embodiment, the super-resolution video frame is magnified by a predetermined multiple relative to the low-resolution video frame. To magnify the low-resolution video frame by a predetermined multiple, the first feature map output by the first convolutional layer is the same size as the low-resolution video frame, the intermediate second feature map output by the second convolutional layer is the same size as the low-resolution video frame, and then the intermediate second feature map is magnified by a predetermined multiple through the upsampling layer. Since the feature map obtained after the convolutional layer extracts features from the image is smaller than the original image, therefore, to ensure that the feature maps output by the two convolutional layers are the same size as the low-resolution video frame, it is necessary to perform padding processing on the low-resolution video frame and the first feature map to ensure that the feature maps output by the convolutional layer are the same size as the low-resolution video frame.

[0098] In some embodiments, the first convolutional layer includes four convolutional kernels, and the number of output channels is four. After the low-resolution video frame is processed by the first convolutional layer, four first feature maps are output. To ensure that the size of the first feature map is the same as that of the low-resolution video frame, it is necessary to preprocess the low-resolution video frame and pad the low-resolution video frame so that the four first feature maps after feature extraction are the same size as the low-resolution video frame. Optionally, mirror padding can be performed on the low-resolution video frame, and the specific methods of image padding and filling are not limited.

[0099] In some embodiments, the second convolutional layer includes sixteen convolutional kernels, and the number of output channels is sixteen. After the first feature map is processed by the second convolutional layer, sixteen intermediate second feature maps are output. To ensure that the size of the intermediate second feature map is the same as that of the low-resolution video frame, it is necessary to preprocess the intermediate second feature map and pad the intermediate second feature map so that the sixteen intermediate second feature maps after feature extraction are the same size as the low-resolution video frame.

[0100] In some ways, the upsampling layer adopts the Pixel Shuffle function, which can remap the feature values on the sixteen channels of the intermediate second feature map output by the second convolutional layer to the space of the feature map, obtaining a second feature map with an area enlarged by 16 times, that is, a feature map with a side length enlarged by 4 times. Thus, the number of output channels of the second convolutional layer can be determined according to the multiple of super-resolution, and the number of output channels of the second convolutional layer is equal to the square of the super-resolution magnification factor.

[0101] In some ways, before looking up the table for the low-resolution video frame, it is necessary to pad the low-resolution video frame, and look up the table according to the padded low-resolution video frame to obtain the reconstructed super-resolution video frame.

[0102] In some embodiments, after obtaining the super-resolution model, it further includes:

[0103] Construct all the input data of the first receptive field, input all the input data of the first receptive field into the first convolutional layer, obtain all the output data corresponding to the first receptive field, and construct the first feature table according to all the input data of the first receptive field and the corresponding all output data;

[0104] Construct all the input data of the second receptive field, input all the input data of the second receptive field into the second convolutional layer, input the output result of the second convolutional layer into the upsampling layer, obtain all the output data corresponding to the second receptive field, and construct the second feature table according to all the input data of the second receptive field and the corresponding all output data.

[0105] Combine Figure 4A 、 4B As shown, in this embodiment, after training the super-resolution model, the first feature table and the second feature table can be constructed based on the super-resolution model. Specifically, according to the size of the first receptive field of the first convolutional layer, construct all possible input data of the first receptive field, input all possible input data into the first convolutional layer, obtain the corresponding output data, and establish the first feature table according to the correspondence between the input data and the output data. Similarly, according to the size of the second receptive field of the second convolutional layer, construct all possible input data of the second receptive field, input all possible input data into the second convolutional layer, obtain the corresponding output data, and establish the second feature table according to the correspondence between the input data and the output data.

[0106] Correspondingly, when looking up the table, according to a predetermined step size, query the first feature table in turn according to each first receptive field of the low-resolution video frame to obtain the corresponding output result; then, according to the predetermined step size, query the second feature table in turn according to the second receptive field of each first feature map to obtain the corresponding output result.

[0107] In some embodiments, a first feature table is constructed based on all the input data and the corresponding all output data of the first receptive field, including:

[0108] Perform integerization processing on all the output data of the first receptive field to obtain the integerized all output data;

[0109] Construct a first feature table according to all the input data of the first receptive field and the corresponding integerized all output data;

[0110] According to all the input data of the second receptive field and the corresponding all output data, construct a second feature table, including:

[0111] Perform integerization processing on all the output data of the second receptive field to obtain the integerized all output data;

[0112] Construct a second feature table according to all the input data of the second receptive field and the corresponding integerized all output data.

[0113] In this embodiment, considering that the data processed by the convolutional layer may not be an integer, in order to facilitate the construction of the feature table, it is necessary to process the output data of the convolutional layer into an integer. For example, multiply the output data of the convolutional layer by 255 and then take the integer to obtain the output data after integerization processing. Then, based on all possible input data and the corresponding integerized output data, construct the first feature table and the second feature table.

[0114] The video super-resolution reconstruction method provided in this embodiment is applied to an edge device. The edge device stores a first feature table and a second feature table. For each frame of the low-resolution video to be processed, query the first feature table and the second feature table in sequence, and then perform averaging processing on the obtained multiple second feature maps to obtain the super-resolution video frame reconstructed for the current frame; after processing each low-resolution video frame, a reconstructed super-resolution video is obtained. This embodiment can achieve fast and efficient super-resolution video reconstruction on edge devices with low configuration, and the reconstruction effect is good.

[0115] It should be noted that the method of one or more embodiments of this specification can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and multiple devices cooperate with each other to complete it. In this case of a distributed scenario, one of these multiple devices can only execute one or more steps of the method of one or more embodiments of this specification, and these multiple devices will interact with each other to complete the described method.

[0116] It should be noted that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0117] As Figure 5 shown, an embodiment of this specification also provides a video super-resolution reconstruction device based on deep learning, including:

[0118] An acquisition module, configured to acquire low-resolution video frames;

[0119] A first look-up table module, configured to query a preset first feature table according to the low-resolution video frames to obtain a plurality of first feature maps;

[0120] A second look-up table module, configured to query a preset second feature table according to the plurality of first feature maps to obtain a plurality of second feature maps; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model, the super-resolution model includes a first convolutional layer, a second convolutional layer, and an upsampling layer, the first-stage table is constructed according to the first all input data and the corresponding first all output data of the first convolutional layer, and the second-stage table is constructed according to the second all input data and the corresponding second all output data of the second convolutional layer and the upsampling layer;

[0121] An average processing module, configured to perform average processing on the plurality of second feature maps to obtain a super-resolution video frame.

[0122] For the convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing one or more embodiments of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0123] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0124] Figure 6 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0125] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0126] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called by the processor 1010 for execution.

[0127] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0128] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0129] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0130] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0131] The electronic device in the above embodiment is used to implement the corresponding method in the foregoing embodiment and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0132] The computer-readable media of this embodiment include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device.

[0133] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; within the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present specification as described above, and for the sake of brevity, they are not provided in detail.

[0134] In addition, for simplicity of illustration and discussion, and so as not to make one or more embodiments of this specification difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making one or more embodiments of this specification difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of this specification will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that one or more embodiments of this specification can be implemented without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.

[0135] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.

[0136] One or more embodiments of this specification are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of this disclosure.

Claims

1. A method for video super-resolution reconstruction based on deep learning, characterized in that, Including: Obtain a low-resolution video frame; According to the low-resolution video frame, query a preset first feature table to obtain a plurality of first feature maps, including: for each first receptive field of the low-resolution video frame, query the first feature table to obtain output data corresponding to each first receptive field; according to the output data, obtain a plurality of first feature maps; According to the plurality of first feature maps, query a preset second feature table to obtain a plurality of second feature maps, including: according to each second receptive field of each first feature map, query the second feature table to obtain output data corresponding to each second receptive field; according to the output data of each first feature map, obtain the corresponding second feature map; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model, the super-resolution model includes a first convolutional layer, a second convolutional layer and an upsampling layer, the first feature table is constructed according to the first all input data and the corresponding first all output data of the first convolutional layer, and the second feature table is constructed according to the second all input data and the corresponding second all output data of the second convolutional layer and the upsampling layer; the receptive field of the first convolutional layer is the first receptive field, and the receptive field of the second convolutional layer is the second receptive field; Perform mean processing on the plurality of second feature maps to obtain a super-resolution video frame.

2. The method according to claim 1, characterized in that, Before obtaining the low-resolution video frame, it further includes: Obtain a low-resolution video frame sample and a high-resolution video frame sample of the same video frame; Use the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model.

3. The method according to claim 2, characterized in that, The initial super-resolution model includes an initial first convolutional layer, an initial second convolutional layer and an initial upsampling layer; Using the low-resolution video frame sample and the high-resolution video frame sample to train the constructed initial super-resolution model to obtain the super-resolution model, including: Input the low-resolution video frame sample into the initial first convolutional layer, and the initial first convolutional layer outputs a plurality of first feature map samples; Input the plurality of first feature map samples into the initial second convolutional layer, and the output result of the initial second convolutional layer is input into the initial upsampling layer, and the initial upsampling layer outputs a plurality of second feature map samples; Perform mean processing on the plurality of second feature map samples to obtain a high-resolution video frame prediction result; Determine the difference degree between the high-resolution video frame prediction result and the high-resolution video frame sample, and adjust the parameters of the initial super-resolution model according to the difference degree; According to the adjusted parameters, retrain the initial super-resolution model until the training end condition is met, and use the initial super-resolution model that meets the training end condition as the super-resolution model, and the initial first convolutional layer, the initial second convolutional layer and the initial upsampling layer corresponding to the initial super-resolution model that meets the training end condition are respectively used as the first convolutional layer, the second convolutional layer and the upsampling layer.

4. The method according to claim 3, characterized in that, Inputting the low-resolution video frame sample into the initial first convolutional layer includes: Perform padding processing on the low-resolution video frame sample to obtain a padded low-resolution video frame sample; Input the low-resolution video frame sample after padding into the initial first convolutional layer; wherein, the size of the first feature map sample is the same as that of the low-resolution video frame sample; Inputting multiple first feature map samples into the initial second convolutional layer includes: Perform padding processing on the first feature map sample to obtain a padded first feature map sample; Input the padded first feature map sample into the initial second convolutional layer; wherein, the output result of the initial second convolutional layer is the same size as the low-resolution video frame sample.

5. The method according to claim 3 or 4, characterized in that, The receptive field of the first convolutional layer is the first receptive field; After obtaining the super-resolution model, it further includes: Construct all the input data of the first receptive field, input all the input data of the first receptive field into the first convolutional layer, obtain all the output data corresponding to the first receptive field, and construct the first feature table according to all the input data and the corresponding all output data of the first receptive field.

6. The method according to claim 5, characterized in that, The receptive field of the second convolutional layer is the second receptive field; After obtaining the super-resolution model, it further includes: Construct all the input data of the second receptive field, input all the input data of the second receptive field into the second convolutional layer, input the output result of the second convolutional layer into the upsampling layer, obtain all the output data corresponding to the second receptive field, and construct the second feature table according to all the input data and the corresponding all output data of the second receptive field.

7. The method according to claim 6, characterized in that, Constructing the first feature table according to all the input data and the corresponding all output data of the first receptive field includes: Perform integerization processing on all the output data of the first receptive field to obtain integerized all output data; Construct the first feature table according to all the input data of the first receptive field and the corresponding integerized all output data; Constructing the second feature table according to all the input data and the corresponding all output data of the second receptive field includes: Perform integerization processing on all the output data of the second receptive field to obtain integerized all output data; Construct the second feature table according to all the input data of the second receptive field and the corresponding integerized all output data.

8. A video super-resolution reconstruction device based on deep learning, characterized in that, It includes: An acquisition module for acquiring low-resolution video frames; A first look-up table module for querying a preset first feature table according to the low-resolution video frame to obtain multiple first feature maps, including: for each first receptive field of the low-resolution video frame, query the first feature table to obtain the output data corresponding to each first receptive field; and obtain multiple first feature maps according to the output data; A second look-up table module, configured to query a preset second feature table according to a plurality of first feature maps to obtain a plurality of second feature maps, including: querying the second feature table according to each second receptive field of each first feature map to obtain output data corresponding to each second receptive field; obtaining a corresponding second feature map according to the output data of each first feature map; wherein, the first feature table and the second feature table are generated according to a pre-constructed super-resolution model, the super-resolution model includes a first convolutional layer, a second convolutional layer and an upsampling layer, the first feature table is constructed according to the first all input data and the corresponding first all output data of the first convolutional layer, and the second feature table is constructed according to the second all input data and the corresponding second all output data of the second convolutional layer and the upsampling layer; the receptive field of the first convolutional layer is a first receptive field, and the receptive field of the second convolutional layer is a second receptive field; A mean processing module, configured to perform mean processing on a plurality of second feature maps to obtain a super-resolution video frame.

Citation Information

Patent Citations

  • Binocular picture super-resolution reconstruction method based on multi-dimensional parallax prior

    CN113393382A

  • Image reconstruction method, image reconstruction model training method, device and equipment

    CN114022357A