An apparatus and method for non-integer multiple super-resolution reconstruction of video images

The method addresses the challenges of non-integer scale super-resolution in video images by using motion and texture analysis to select and combine super-resolution networks, achieving efficient and accurate detail recovery with reduced computational cost.

CN114240760BActive Publication Date: 2025-07-15SHANGHAI FULLHAN MICROELECTRONICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111640472.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-15
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, in the non-integer multiple super-resolution reconstruction of video images, the detailed texture cannot be accurately mapped, which is too expensive to calculate, especially in the video super-resolution method.

Method used

Through the block processing of image frames and motion information, combined with motion probability estimation and texture complexity classification, a super-resolution network with appropriate calculation costs is selected for non-integer multiple super-resolution reconstruction of image blocks, and weighted fusion is performed to finally generate high-quality video images.

Benefits of technology

Improves the detailed resilience capability of video images for non-integer multiples super-resolution reconstruction, reduces the computational cost of deep neural networks, improves performance and reduces the computational burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240760B_ABST
    Figure CN114240760B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for non-integer multiple super-resolution reconstruction of video images. The method includes: S1, obtaining an image frame and a motion information map of the image frame; S2, respectively performing block processing at the same position; S3, performing motion probability estimation; S4, classifying the texture complexity of each image block; S5, selecting a super-resolution network based on the motion probability estimation and the texture complexity classification label; S6, designing multiple super-resolution networks and training them using the same data set; S7, using K super-resolution networks to perform forward inference on the selected super-resolution network and outputting non-integer multiple super-resolution image blocks of each image block; S8, performing motion-based weighted fusion on the super-resolution results of each image block and the super-resolution results of the corresponding block in the previous frame to obtain the super-resolution output of the current frame; S9, splicing and fusing the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image super-resolution reconstruction, and particularly to a video image non-integer multiple super-resolution reconstruction device and method. Background Art

[0002] Super-resolution methods are a class of computer vision tasks that use low-resolution images to reconstruct high-resolution images. They are widely applied in devices such as video security, smartphones, and mobile cameras to generate clear high-resolution images and improve the output quality of various devices.

[0003] Among non-integer multiple super-resolution methods, there are those based on traditional interpolation techniques. For example, the Chinese patent application with the publication number CN102881000A provides a super-resolution method for video images. It first performs integer multiple upsampling to obtain a high-resolution image, and then performs a non-integer multiple interpolation-based method on the intermediate high-resolution image. However, in this method, the loss of image details in the first step of upsampling will be transmitted to the second step, resulting in an increase in the error of the non-integer interpolation method, making it impossible to accurately map the detailed textures in the low-resolution image to the high-resolution image.

[0004] Currently, super-resolution methods based on deep convolutional neural networks (CNNs) have greatly improved the performance of super-resolution. For example, single-image super-resolution methods represented by SRCNN and DCNN. However, when solving the non-integer multiple super-resolution problem, it is necessary to first use a high-magnification super-resolution network to obtain an intermediate image with integer multiple high resolution, and then use a downsampling method to reduce the integer multiple high-resolution intermediate image by an integer multiple to indirectly achieve the purpose of non-integer multiple magnification. Since the texture detail features of the high-resolution network output image continue to have an increased error after downsampling, the detailed textures in the low-resolution image still cannot be accurately mapped to the high-resolution image.

[0005] In addition, since the super-resolution network model needs to use larger feature maps, it requires a greater computational cost compared to other computer vision tasks using CNNs. And video super-resolution methods mainly adopt a frame-by-frame extraction method, using one frame of picture as input, or multiple frames input simultaneously. For example, the Chinese patent application with the publication number CN109767383A provides a method for video super-resolution using a convolutional neural network. By using two-stage motion compensation to construct a video super-resolution system, since each input frame needs to perform super-resolution on the entire image and there are multiple frames of data, the computational amount of the network is even more huge. Summary of the Invention

[0006] To overcome the deficiencies of the above-mentioned existing technologies, the object of the present invention is to provide a video image non-integer multiple super-resolution reconstruction device and method, so as to improve the ability of video image non-integer multiple super-resolution to restore details, and at the same time reduce the computational cost of the deep neural network super-resolution method.

[0007] To achieve the above and other objects, the present invention proposes a video image non-integer multiple super-resolution reconstruction method, including the following steps:

[0008] Step S1, obtaining an image frame and an image frame motion information map;

[0009] Step S2, performing block processing on the current frame image and the motion information map at the same position;

[0010] Step S3, performing motion probability estimation according to the motion information map corresponding to each image block;

[0011] Step S4, classifying the texture complexity of each image block;

[0012] Step S5, selecting a super-resolution network based on the motion probability estimation and texture complexity classification label of the image block;

[0013] Step S6, designing multiple super-resolution networks with different costs and training them using the same data set;

[0014] Step S7, using the trained K super-resolution networks, according to the selection in step S5, selecting the corresponding super-resolution network for each image block for forward inference, and outputting the non-integer multiple super-resolution image block corresponding to each image block;

[0015] Step S8, performing motion-based weighted fusion on the super-resolution results of each image block and the super-resolution results of the corresponding block in the previous frame to obtain the super-resolution output of the current frame video image;

[0016] Step S9, splicing and fusing the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.

[0017] Preferably, in step S2, for the current frame image, a single-frame image is obtained from the video stream, the image is first subjected to a boundary expansion operation, and an overlapping segmentation method is used, and then the image is segmented into several identical sub-blocks. The motion information map is single-channel, and block operation is performed using the same segmentation parameters as the image block.

[0018] Preferably, step S3 further includes:

[0019] Step S300, performing morphological operations of erosion first and then dilation on the motion information of each image block to obtain a preprocessed motion information map;

[0020] Step S301: Perform histogram statistics on the preprocessed motion information graph to obtain the frequency distribution.

[0021] Step S302: Configure the weighting coefficients according to the motion information at each position.

[0022] Step S303: Calculate the motion probability of the image blocks based on the frequency distribution.

[0023] Preferably, in step S301, first normalize the motion information C at each position ij to the range of 0 - 1, and then set the number of groups of the histogram to k, so the group interval size is 1 / k. Traverse C at all positions ij , and count them into the corresponding groups respectively to obtain k frequencies.

[0024] Preferably, step S4 further includes:

[0025] Step S400: Perform binarization processing on each image block;

[0026] Step S401: Extract the image texture complexity features of each image block;

[0027] Step S402: Construct a texture complexity classification model based on the extracted image texture complexity features.

[0028] Preferably, in step S401, the 5 extracted image texture complexity features are the maximum value feature, the mean value feature, the sample average value, the sample standard deviation, and the entropy value respectively.

[0029] Preferably, the preprocessing process of the training data of the texture complexity classification model is as follows:

[0030] Randomly extract multiple groups of training images from the super - resolution training set. Each group of training images consists of a low - resolution image and the corresponding high - resolution image, and perform random block division on each group of images to obtain n groups of image block data ((LR1,HR1),…,(LR n ,HR n ));

[0031] Use step S400 and step S401 to preprocess the high - resolution image blocks in each group of training image blocks to obtain the training feature data (F 1 ,…,F n );

[0032] Obtain the texture complexity labels corresponding to the feature vectors. First, downsample the high - resolution image blocks of the training image blocks to the same size as the low - resolution images, and then calculate the PSNR value between the downsampled image blocks and the original low - resolution image blocks and denote it as P i , i = 1,…,n;

[0033] Based on the PSNR distribution of all training image patches, the texture complexity is labeled. Suppose there are L categories of image texture complexity. Arrange the PSNR of all training data from largest to smallest, and divide it into L groups according to the number of data. The labels of each group of image patches are set to 1, …, L in sequence, indicating that the texture ranges from simple to complex;

[0034] The finally obtained training data and labels (F i , y i ), where y i = 1, …, L, representing the label of the corresponding data.

[0035] Preferably, in step S6, the non-integer multiple super-resolution network includes:

[0036] A low-resolution image patch input unit, which is responsible for inputting the image patch into the super-resolution network;

[0037] A low-resolution feature reconstruction unit, which is used to extract features from the low-resolution image and perform feature extraction without changing the image size;

[0038] A non-integer multiple feature super-resolution unit, which is used to perform fast and accurate learning on the high-frequency features of non-integer multiple super-resolution through bidirectional feature reconstruction for the input low-resolution features, so as to better restore the detailed texture of the non-integer super-resolution image from the low-resolution image and improve the performance of non-integer multiple super-resolution;

[0039] A high-resolution feature reconstruction unit, which is used to extract the non-integer multiple super-resolution image from the non-integer multiple super-resolution features and perform feature channel compression without changing the feature size;

[0040] A high-resolution image patch output unit, which is used to output the non-integer multiple super-resolution image after the forward inference of the network.

[0041] Preferably, the non-integer multiple feature super-resolution unit includes:

[0042] An input unit, which is used to perform channel compression on the output low-resolution features of the low-resolution feature reconstruction unit to obtain low-resolution input features, and the degree of compression decreases as the network calculation amount increases;

[0043] An initial non-integer multiple upsampling unit and a non-integer upsampling unit, and a dense connection layer composed of a dense residual connection structure is provided between the two units to achieve fast and accurate learning of the high-frequency features of non-integer multiple super-resolution;

[0044] The non-integer multiple super-resolution unit, which is composed of an upsampling unit, is used to perform non-integer multiple super-resolution on the bidirectional sampling features, thereby avoiding the loss of details when directly performing non-integer multiple super-resolution on the input features.

[0045] To achieve the above object, the present invention also provides a video image non-integer multiple super-resolution reconstruction device, including:

[0046] An image frame and image motion information acquisition unit, which is used to acquire an image frame and an image frame motion information map;

[0047] A block operation unit, which is used to perform block processing at the same position on the current frame image and the motion information map respectively;

[0048] An image block motion estimation unit, which is used to perform motion probability estimation according to the motion information map corresponding to each image block;

[0049] An image block texture complexity classification unit, which is used to classify the texture complexity of each image block;

[0050] A super-resolution network selection unit, which is used to select a super-resolution network based on the motion probability estimation and texture complexity classification label of the image block;

[0051] A non-integer multiple super-resolution network construction and training unit, which is used to design multiple super-resolution networks with different costs and train them using the same data set;

[0052] An image block super-resolution calculation unit, which is used to use the trained K super-resolution networks, and according to the selection of the super-resolution network selection unit, select the corresponding super-resolution network for each image block to perform forward inference, and output the non-integer multiple super-resolution image block corresponding to each image block;

[0053] An image block time-domain weighting unit, which is used to perform motion-based weighted fusion on the super-resolution result of each image block and the super-resolution result of the corresponding block in the previous frame to obtain the super-resolution output of the current frame video image;

[0054] An image block super-resolution output fusion unit, which is used to splice and fuse the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.

[0055] Compared with the prior art, the video image non-integer super-resolution method and device of the present invention comprehensively select and use non-integer multiple super-resolution networks with different computational amounts according to the motion information and texture complexity of the image, reducing the cost of video super-resolution calculation. At the same time, the present invention proposes a non-integer multiple feature super-resolution unit for constructing a super-resolution network, so that the super-resolution network constructed by this unit can better recover the detailed texture of the non-integer super-resolution image from the low-resolution image, and improve the performance of non-integer multiple super-resolution. Description of the Drawings

[0056] Figure 1 It is a flowchart of the steps of a method for non-integer multiple super-resolution reconstruction of video images according to the present invention;

[0057] Figure 2 It is the basic structure diagram of the non-integer multiple super-resolution network in a specific embodiment of the present invention;

[0058] Figure 3 It is the structure diagram of the non-integer multiple feature super-resolution unit in a specific embodiment of the present invention;

[0059] Figure 4 It is the system architecture diagram of a non-integer multiple super-resolution reconstruction device for video images according to the present invention;

[0060] Figure 5 It is the structure diagram of the image block motion estimation unit in a specific embodiment of the present invention;

[0061] Figure 6 It is the structure diagram of the image block texture complexity classification unit in a specific embodiment of the present invention;

[0062] Figure 7 It is the flowchart of an embodiment of the present invention. Detailed Embodiments

[0063] The following describes the embodiments of the present invention through specific specific examples and in conjunction with the drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific examples, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0064] Figure 1 It is a flowchart of the steps of a method for non-integer multiple super-resolution reconstruction of video images according to the present invention. As Figure 1 shown, a method for non-integer multiple super-resolution reconstruction of video images according to the present invention includes the following steps:

[0065] Step S1, obtaining an image frame and an image frame motion information map.

[0066] In the present invention, the method can be applied after a video decoder or an ISP (Image Signal Processor). If it is applied after the video decoder, the obtained image frames are from the output after decoding by the video decoding module of the terminal device. Therefore, encoding-related information of the corresponding coding blocks can be obtained, including inter-frame motion information. Since the motion information in the encoder is in the size of image coding blocks, it is necessary to first merge the encoding information of all coding blocks in the current frame to obtain the initial image motion information map of the full frame, and then perform Gaussian filtering on this motion information map to reduce the motion estimation errors caused by noise in the coding blocks. If it is selected to be placed after the ISP (Image Signal Processor), single-frame image data can be obtained without decoding, and the motion information map of the current frame image can be obtained from the temporal noise reduction module.

[0067] Step S2: Perform block processing on the current frame image and the motion information map at the same positions.

[0068] Specifically, the block processing of the current frame image is as follows: Obtain a single-frame image from the video stream. First, perform a boundary extension operation on the image, using an overlapping segmentation method, and then divide the image into several identical sub-blocks. Specifically, the image segmentation needs to meet the following conditions:

[0069] W + 2α1 = m * w - (m - 1) * β1, H + 2α2 = n * h - (n - 1) * β2

[0070] Where the width and height of the image are W and H respectively, the compensation sizes on the upper and lower sides and the left and right sides of the image are α1 and α2 respectively, and mirror mapping is used for boundary compensation, that is, the pixel values within the image are symmetrically mapped to the compensation area along the central axis of the image edge. w and h respectively represent the width and height of the image sub-blocks, m and n respectively represent the number of image sub-blocks in the horizontal and vertical directions, and β1 and β2 represent the overlapping distance between image blocks.

[0071] Since a single-frame image has three channels of R, G, and B, and the RGB three-channel image has three maps of the full-frame size, three block operations need to be performed. The motion information map is a single channel, and block operations are performed using the same segmentation parameters as the image blocks.

[0072] Step S3: Perform motion probability estimation according to the motion information map corresponding to each image block.

[0073] In the present invention, the motion probability of an image block is estimated based on its motion information. The smaller the probability, the quieter the image block. Therefore, more super-resolution outputs of the corresponding image block in the previous frame can be used as a reference, so that the image block can select a super-resolution model with lower computational complexity without considering the texture complexity, and the super-resolution effect of the image block is not affected.

[0074] Specifically, step S3 further includes:

[0075] Step S300: Perform a morphological operation of erosion followed by dilation on the motion information of each image block to obtain a preprocessed motion information map.

[0076] Specifically, assume the size of the motion information map of the current image block is w×h. Then, perform a morphological operation of erosion followed by dilation on the motion information of this image block to obtain a preprocessed motion information map C, and the motion information at each position is C ij , where i = 1, …, h, j = 1, …, w.

[0077] Step S301: Perform histogram statistics on the preprocessed motion information map to obtain a frequency distribution.

[0078] Specifically, first normalize the motion information C ij at each position to the range of 0 - 1. Then, set the number of groups of the histogram to k, so the bin size is 1 / k. Traverse all positions of C ij , and count them into the corresponding groups respectively to obtain k frequencies.

[0079] Step S302: Configure the weighting coefficients according to the motion information at each position.

[0080] For different values of C ij indicating different degrees of motion, so for larger values of C ij the contribution to the motion is greater. If μ x represents the weighting coefficient, and x = 1, …, k, then it satisfies μ x+1 = μ x + 1.

[0081] Step S303: Calculate the motion probability of the image block based on the frequency distribution.

[0082] In step S303, calculate the motion probability of the image block based on the frequency distribution as follows:

[0083]

[0084] where c x represents the frequency of the x-th group. The smaller the value in this group, the smaller the weight. Refer to the configuration k = 5.

[0085] Step S4: Classify the texture complexity of each image block.

[0086] Generally speaking, for image blocks with high complexity, the difficulty of super-resolution restoration is greater, and a non-integer multiple super-resolution network with a greater computational cost should be adopted; for image blocks with low complexity, the difficulty of super-resolution restoration is smaller, and a network with a smaller computational cost can meet the requirements. Specifically, step S4 further includes:

[0087] Step S400: Binarize each image block.

[0088] Since the calculation of texture complexity features is mainly carried out in the luminance domain, in the present invention, the three-channel color image block is first converted into a luminance image block, and the formula is:

[0089] Y = 0.299R + 0.587G + 0.114B,

[0090] Assume that the luminance value range at each position is: 0,..., g, and V represents the maximum gray value of the image. Then g binary images are generated respectively as where δ is the binarization threshold taking values from 1 to g.

[0091] Step S401: Extract the image texture complexity features for each image block.

[0092] If the method of the present invention is applied after a video decoder, then the g binary images obtained in step S400 can be used to construct a complexity statistical value of the image based on the difference between the surrounding neighborhood points and their own values:

[0093]

[0094] where represents the value at the i-th row and j-th column in the binary image I δ and represents the value at the i-th row and (j - 1)-th column in the binary image I δ . w and h respectively represent the width and height of the image sub-block. The five output image texture complexity features are: the maximum value feature f1 = max(C(δ)), the mean value feature sample average sample standard deviation entropy value

[0095] If the method of the present invention is applied after an ISP, then the detail layer D can be obtained from modules such as sharpening. The detail layer represents the texture distribution in the image to a certain extent. First, the detail layer is divided into blocks in the same way as the image frame, and the size is still w × h. Then the extracted texture complexity features are: the maximum value f1 = max(D), the mean value absolute difference standard deviation entropy value D(i, j) represents the value at the i-th row and j-th column in the detail layer.

[0096] Step S402: Construct a texture complexity classification model based on the extracted image texture complexity features.

[0097] Specifically, a texture complexity classification model, i.e., a logistic multi-classification model, is constructed based on the extracted image texture complexity features. The logistic multi-classification model is as follows:

[0098] Let the probability belonging to each class be:

[0099]

[0100] where F represents the 6-dimensional feature vector (f1,..., f s , 1) of the input image patch, and the model parameters are unknown, and l takes values 1,..., L - 1. y represents the class label of the image patch, and P(y = l|F) represents the conditional probability that the texture complexity of the image patch is class l under the condition of the input feature vector F.

[0101] Using one of the feature extraction methods in step S401, represent the image texture complexity with 5 features and input it into the texture complexity classification model to output the classification label of the image patch texture complexity. The smaller the label, the simpler the texture, and the larger the label, the more complex the texture.

[0102] Specifically, the preprocessing process of the training data of the texture complexity classification model is as follows:

[0103] 1. Randomly extract multiple groups of training images from the super-resolution training set. Each group of training images consists of a low-resolution image and the corresponding high-resolution image, and each group of images is randomly blocked to obtain n groups of image patch data ((LR1, HR1),..., (LR n , HR n ));

[0104] 2. Use steps S400 and S401 to preprocess the high-resolution image patches in each group of training image patches to obtain training feature data (F 1 ,..., F n ). The present invention supports training the model using different feature data according to the location, where F 1 , F n respectively represent the feature vectors (f1 n ,..., f5 1 ) and (f1 1 ,..., f5 n ) composed of 5 complexity texture features extracted from the HR1 and HR n image patches;

[0105] 3. Obtain the texture complexity label corresponding to the feature vector. First, downsample the high-resolution image patches of the training image patches to the same size as the low-resolution, and then calculate the PSNR value between the downsampled image patches and the original low-resolution image patches and record it as Pi , where \(i = 1,\cdots,n\);

[0106] 4. Based on the PSNR distribution of all training image patches, perform label partitioning on the texture complexity. Assume that there are \(L\) categories of image texture complexity. Arrange the PSNR of all training data from largest to smallest, and divide it into \(L\) groups according to the number of data. The labels of each group of image patches are sequentially set as \(1,\cdots,L\), indicating the texture from simple to complex.

[0107] 5. The finally obtained training data and labels \((F i , y i ), where \(y i = 1,\cdots,L\), representing the labels corresponding to the data.

[0108] For the obtained training data set \(T=\{(F 1 , y 1 ), (F 2 , y 2 ), \cdots, (F n , y n )\}, the parameters \(w_1,\cdots,w L-1 \) of the model are obtained by applying maximum likelihood estimation. Since maximum likelihood estimation is a general technique, it will not be elaborated here.

[0109] In the present invention, the preprocessing of the training data of the model and the parameter estimation process can be completed offline. During formal operation, only the image patch needs to be input. After extracting the features through step S400 and step S401, the probabilities that the texture complexity categories of the image patch belong to \(1,\cdots,L\) can be calculated:

[0110]

[0111] Select the label corresponding to the maximum probability as the model output, then the texture complexity category of the image patch can be obtained, denoted as \(l c .

[0112] Step S5, select the super-resolution network based on the motion probability estimation and texture complexity classification label of the image patch.

[0113] In the present invention, based on the comprehensive judgment of the motion probability estimation and texture complexity classification label of the image patch, the network with what kind of computational cost is used. There are a total of \(K\) preset non-integer multiple super-resolution networks, and the computational costs increase sequentially from small to large according to the numbers. The image patch outputs the motion probability \(p m , where \(p m \in(0,1)\); step S4 outputs the complexity category \(l c , \(l c∈(1, L), according to the motion probability and texture complexity, select the preference in the network size, set the texture complexity weight as μ, then the selected super-resolution network number:

[0114]

[0115] where round represents the floor function,

[0116] Step S6, design multiple super-resolution networks with different costs and train them using the same dataset.

[0117] In the present invention, design K non-integer multiple super-resolution networks with different computational costs. Each non-integer multiple super-resolution network has the same basic structure as Figure 2 shown. According to the super-resolution network number output in step S5, select the corresponding network for forward inference to output the corresponding super-resolution image. The main difference between the networks with different computational costs lies in the number of feature channels in each unit. Therefore, the number of convolution kernels used is also different, which in turn leads to different network computational complexities. Generally speaking, the smaller the computational cost, the smaller the number of channels.

[0118] As Figure 2 shown, the non-integer multiple super-resolution network specifically includes:

[0119] The low-resolution image patch input unit 201 is responsible for inputting the image patch into the super-resolution network;

[0120] The low-resolution feature reconstruction unit 202 is used to extract features from the low-resolution image and perform feature extraction without changing the format size. In the specific embodiment of the present invention, the convolution layer and its related feature extraction combination in the existing traditional super-resolution network technology can be used.

[0121] The non-integer multiple feature super-resolution unit 203 is used to perform fast and accurate learning on the high-frequency features of non-integer multiple super-resolution through bidirectional feature reconstruction for the input low-resolution features, so as to better restore the detailed texture of the non-integer super-resolution image from the low-resolution image and improve the performance of non-integer multiple super-resolution.

[0122] Specifically, as Figure 3 shown, the non-integer multiple feature super-resolution unit 203 includes:

[0123] The input unit 203-1 is used to compress the channels of the low-resolution features output by the low-resolution feature reconstruction unit 202 to obtain low-resolution input features, and the degree of compression decreases as the network computational amount increases;

[0124] The initial non-integer multiple upsampling and downsampling unit 203-2 and the non-integer upsampling and downsampling unit 203-3 have a dense connection layer composed of a dense residual connection structure between the units to achieve fast and accurate learning of the high-frequency features of non-integer multiple super-resolution. In a specific embodiment of the present invention, the non-integer multiple feature super-resolution network unit is introduced by taking non-integer multiple upsampling as an example. It should be understood that in the non-integer multiple super-resolution unit structure of the present invention, the convolution layer and the transposed convolution layer can be swapped to achieve the downsampling function.

[0125] The initial non-integer multiple upsampling and downsampling unit 203-2 and the non-integer upsampling and downsampling unit 203-3 are mainly composed of a number of upsampling units and downsampling units. The upsampling unit is used to magnify the features, and its internal is composed of a traditional transposed convolution layer 1 and a convolution layer 1. The transposed convolution layer 1 magnifies the feature area by x times, where N is a positive integer. The convolution layer 1 reduces the features by y times, where The data feature area after passing through the upsampling unit will become times of the original; in addition, to avoid the problem of high-frequency feature loss caused by only using the upsampling unit, a low-magnification downsampling unit composed of a transposed convolution layer 2 (magnifying times) and a convolution layer 2 (reducing times) is connected in series after each upsampling unit, where n2 >= 2, n2 ∈ N, m2 >= 3, m2 ∈ N and n2 < m2. This structure has a two-way feature reconstruction function from low resolution to high resolution and from high resolution to low resolution. To ensure that the number of input features of each unit structure is consistent with the input, when there is a dense residual connection, a 1×1 convolution layer structure is required between the sampling units to compress the multiple features. At the same time, to ensure the applicability of the present invention, one or more upsampling unit-downsampling unit structures can be connected in series in the non-integer upsampling and downsampling unit 203-3 according to requirements. There is only one in the schematic diagram here;

[0126] The non-integer multiple super-resolution unit 203-4 is composed of upsampling units and is used for non-integer multiple super-resolution of the two-way sampling features, avoiding the loss of details when directly performing non-integer multiple super-resolution on the input features.

[0127] In the non-integer multiple feature super-resolution unit, the denseness can be divided into three categories: 1. The low-resolution input feature directly crosses the connection between the initial non-integer multiple upsampling unit 203-2 and the non-integer upsampling unit 203-3, as shown by the marked line "a" in the figure. Its function is to enable local residual learning in the 'amplify-reduce' cascaded unit structure; 2. The residual connection between the output of the upsampling unit and the output of the subsequent amplification unit, as shown by the marked line "b" in the figure. Its function is to fuse the output features of multiple amplification units and accelerate the feature learning of the downsampling unit; 3. The residual connection between the output of the downsampling unit and the output of the subsequent downsampling unit, as shown by the marked line "c" in the figure. Its function is to fuse the output features of multiple downsampling units and accelerate the low-resolution feature learning.

[0128] The high-resolution feature reconstruction unit 204 is used to extract the non-integer multiple super-resolution image from the non-integer multiple super-resolution feature and perform feature channel compression without changing the feature size. The combination of the output convolutional layer and its related convolutional layers in the existing traditional super-resolution network technology can be used.

[0129] The high-resolution image block output unit 205 is used to output the non-integer multiple super-resolution image after the forward inference of the network.

[0130] Step S7: Using the trained K super-resolution networks, according to the selection in step S5, select the corresponding super-resolution network for each image block for forward inference, and output the non-integer multiple super-resolution image block corresponding to each image block.

[0131] Specifically, for each image block, select the corresponding super-resolution network for forward inference according to the number calculated in step S5, and output the non-integer multiple super-resolution image block corresponding to each image block.

[0132] Step S8: Perform motion-based weighted fusion on the super-resolution results of each image block and the super-resolution results of the corresponding blocks in the previous frame to obtain the super-resolution output of the current frame video image.

[0133] Specifically, for a certain non-integer multiple super-resolution image block SR t , its motion information is obtained by the encoder as C t , and the maximum value that can be obtained is C max , and the non-integer multiple super-resolution image block of the corresponding image block output in the previous frame is SR t-1 . Then first perform non-integer multiple super-resolution of C t by traditional bicubic interpolation to obtain SC t , so that its size is consistent with the super-resolved image block. Then use the pixel value SR t at each position of SR tThe value of (i, j) and the pixel value SR at the corresponding position in the previous frame t-1 Perform weighting on (i, j) based on motion information to obtain the value SR fused based on the motion map t (i, j) = (SR t (i, j) * SC t (i * j) + SR t-1 (i, j) * (C max - SC t (i * j))) / C max .

[0134] Step S9: Stitch and fuse the super-resolution results of all image patches in the current frame to obtain the full-frame super-resolution result of the current frame

[0135] As can be seen from step S2, there are overlapping regions between each image patch. Therefore, during the fusion and stitching process, it is necessary to average the overlapping regions to eliminate the block effect

[0136] Figure 4 This is the system architecture diagram of a video image non-integer multiple super-resolution reconstruction device of the present invention. As Figure 4 shown, a video image non-integer multiple super-resolution reconstruction device of the present invention includes:

[0137] An image frame and image motion information acquisition unit 101, configured to acquire an image frame and an image frame motion information map

[0138] In the present invention, the device of the present invention can be used after a video decoder or an ISP (image signal processor). If it is used after a video decoder, the image frame acquired by the image frame and image motion information acquisition unit 101 comes from the output after decoding by the video decoding module of the terminal device. Therefore, encoding-related information corresponding to the encoding block can be obtained, including inter-frame motion information. Since the motion information in the encoder is in the size of the image coding block, it is necessary to first merge the encoding information of all coding blocks in the current frame to obtain the initial image motion information map of the full frame, and then perform Gaussian filtering processing on this motion information map to reduce the motion estimation error caused by noise in the coding block; if it is selected to be placed after the ISP (image signal processor), single-frame image data can be acquired without decoding, and the motion information map of the current frame image can be acquired from the time-domain noise reduction module

[0139] A block operation unit 102, configured to perform block processing on the current frame image and the motion information map at the same position

[0140] Specifically, the block operation unit 102 performs block processing on the current frame image as follows: Obtain a single-frame image from the video stream, first perform a boundary extension operation on the image, use an overlapping segmentation method, and then divide the image into several identical sub-blocks. Specifically, the image segmentation needs to meet the following conditions:

[0141] W + 2α1 = m * w - (m - 1) * β1, H + 2α2 = n * h - (n - 1) * β2

[0142] Where the width and height of the image are W and H respectively, the compensation sizes on the upper and lower, left and right sides of the image are α1 and α2 respectively, and boundary compensation uses mirror mapping, that is, the pixel values within the image are symmetrically mapped to the compensation area along the image edge as the central axis. w and h respectively represent the width and height of the image sub - block, m and n respectively represent the number of image sub - blocks in the horizontal and vertical directions, and β1 and β2 represent the overlapping distance between image blocks.

[0143] Since a single - frame image has three channels of R, G, and B, three block - dividing operations are required; the motion information map is single - channel, and block - dividing operations are performed using the same segmentation parameters as the image blocks.

[0144] The image - block motion estimation unit 103 is used to estimate the motion probability according to the motion information map corresponding to each image block.

[0145] In the present invention, the image - block motion estimation unit 103 estimates the motion probability of an image block according to the motion information of the image block. The smaller the probability, the quieter the image block is. Then, more super - resolution outputs of the corresponding image block in the previous frame can be used as a reference, so that the image block can select a super - resolution model with lower computational complexity without considering the texture complexity, and the super - resolution effect of the image block is not affected.

[0146] Specifically, as Figure 5 shown, the image - block motion estimation unit 103 further includes:

[0147] The image - block motion information pre - processing unit 103 - 1 is used to perform morphological operations of erosion first and then dilation on the motion information of each image block to obtain a pre - processed motion information map.

[0148] Specifically, assuming that the size of the motion information map of the current image block is w × h, morphological operations of erosion first and then dilation are performed on the motion information of the image block to obtain a pre - processed motion information map C, and the motion information at each position is C ij , i = 1,..., h, j = 1,..., w.

[0149] The histogram statistics unit 103 - 2 of motion information is used to perform histogram statistics on the pre - processed motion information map to obtain a frequency distribution.

[0150] Specifically, first normalize the motion information C ij at each position to the 0 - 1 operation, then set the number of groups of the histogram as k, then the group interval size is 1 / k, and traverse all positions of C ij, count them separately into the corresponding groups to obtain k frequencies.

[0151] The weighted coefficient configuration unit 103-3 configures the weighted coefficients according to the motion information at each position.

[0152] For different values of C ij It represents different degrees of motion. Therefore, for a larger value of C ij the greater the contribution to the motion. If μ x represents the weighted coefficient, x = 1,..., k, then it satisfies μ x+1 = μ x +1.

[0153] The image block motion probability calculation unit 103-4 calculates the image block motion probability based on the frequency distribution.

[0154] In the image block motion probability calculation unit 103-4, the image block motion probability is calculated based on the frequency distribution as follows:

[0155]

[0156] where c x represents the frequency of the x-th group. The smaller the value in this group, the smaller the weight. Refer to the configuration k = 5.

[0157] The image block texture complexity classification unit 104 is used to classify the texture complexity of each image block.

[0158] Generally speaking, the super-resolution restoration of image blocks with high complexity is more difficult, and a non-integer multiple super-resolution network with a greater computational cost should be adopted; the super-resolution restoration of image blocks with low complexity is less difficult, and a network with a smaller computational cost can meet the requirements. Specifically, as Figure 6 shown, the image block texture complexity classification unit 104 further includes:

[0159] The image block binarization processing unit 104-1 is used to perform binarization processing on each image block.

[0160] Since the texture complexity feature calculation is mainly carried out in the luminance domain, in the present invention, the three-channel color image block is first converted into a luminance image block, and the formula is:

[0161] Y = 0.299R + 0.587G + 0.114B,

[0162] Assume that the luminance value range at each position is: 0,..., g, and V represents the maximum gray value of the image. Then g binary images are generated respectively as where δ is the binarization threshold value taking 1,..., g.

[0163] The image block texture complexity feature extraction unit 104-2 is used to extract the image texture complexity features for each image block.

[0164] If the device of the present invention is used after the video decoder, the g binary images obtained by the image block binarization processing unit 104-1 can be used to construct the complexity statistical value of the image based on the difference between the surrounding domain points and their own values:

[0165]

[0166] where represents the value at the i-th row and j-th column of the binary image I δ in, represents the binary image I δ in the value at the i-th row and (j-1)-th column. w and h respectively represent the width and height of the image sub-block, and the five output image texture complexity features are: the maximum value feature f1 = max(C(δ)), the mean feature sample average sample standard deviation entropy value

[0167] If the device of the present invention is used after the ISP, the detail layer D can be obtained from modules such as sharpening. The detail layer represents the texture distribution in the image to a certain extent. First, the detail layer is subjected to the same block processing as the image frame, with the size still being w×h, and then the extracted texture complexity features are: the maximum value f1 = max(D), the mean absolute difference standard deviation entropy value D(i, j) represents the value at the i-th row and j-th column in the detail layer.

[0168] The texture complexity classification model generation unit 104-3 is used to construct a texture complexity classification model based on the extracted image texture complexity features.

[0169] Specifically, the texture complexity classification model generation unit 104-3 constructs a texture complexity classification model based on the extracted image texture complexity features, that is, a logistic multi-classification model. The logistic multi-classification model is as follows:

[0170] Let the probability belonging to each category be: where F represents the 6-dimensional feature vector (f1,..., f s , 1) of the input image block, and the model parameters Unknown, where l takes values from 1 to L - 1. y represents the class label of the image patch, and P(y = l|F) represents the conditional probability that the texture complexity of the image patch is of class l under the condition of the input feature vector F.

[0171] Using one of the feature extraction methods in the image patch texture complexity feature extraction unit 104 - 2, the image texture complexity is represented by 5 features and input into the texture complexity classification model to output the texture complexity classification label of the image patch. The smaller the label, the simpler the texture, and the larger the label, the more complex the texture.

[0172] Specifically, the pre - processing process of the training data of the texture complexity classification model is as follows:

[0173] 1. Randomly extract multiple groups of training images from the super - resolution training set. Each group of training images consists of a low - resolution image and the corresponding high - resolution image, and each group of images is randomly divided into blocks to obtain n groups of image patch data ((LR1, HR1),..., (LR n , HR n ));

[0174] 2. Use the image patch binarization processing unit 104 - 1 and the image patch texture complexity feature extraction unit 104 - 2 to pre - process the high - resolution image patches in each group of training image patches to obtain training feature data (F 1 ,..., F n ). The present invention supports training the model using different feature data according to the location, where F 1 , F n respectively represent the feature vectors (f1 n ,..., f5 1 ), (f1 1 ,..., f5 n ,..., f5 n ) composed of 5 complexity texture features extracted from the HR1 and HR

[0175] 3. Obtain the texture complexity labels corresponding to the feature vectors. First, down - sample the high - resolution image patches of the training image patches to the same size as the low - resolution ones, and then calculate the PSNR value between the down - sampled image patches and the original low - resolution image patches, denoted as P i , i = 1,..., n;

[0176] 4. Based on the PSNR distribution of all training image patches, divide the texture complexity labels. Assume that the image texture complexity has a total of L classes. Arrange the PSNR of all training data from large to small and divide it into L groups according to the number of data. The labels of each group of image patches are set to 1,..., L in sequence, indicating that the texture ranges from simple to complex.

[0177] 5. The finally obtained training data and labels (F i , y i ), where y i = 1,..., L, represents the label corresponding to the data.

[0178] For the obtained training data set T = {(F 1 , y 1 ), (F 2 , y 2 ),..., (F n , y n )}, the parameters w1,..., w L-1 of the model are obtained by applying maximum likelihood estimation. Since maximum likelihood estimation is a common technique, it will not be elaborated here.

[0179] In the present invention, the preprocessing of the training data and the parameter estimation process of the model can be completed offline. During formal operation, only the image patch needs to be input. After extracting features through the image patch binarization processing unit 104-1 and the image patch texture complexity feature extraction unit 104-2, the probabilities that the image patch texture complexity categories belong to 1,..., L can be calculated Select the label corresponding to the maximum probability as the model output, then the image patch texture complexity category of the image patch can be obtained, denoted as l c .

[0180] The super-resolution network selection unit 105 is used to select a super-resolution network based on the motion probability estimation and texture complexity classification label of the image patch.

[0181] In the present invention, based on the motion probability estimation and texture complexity classification label of the image patch, a network with what kind of computational cost is comprehensively judged. There are a total of K non-integer multiple super-resolution networks preset, and the computational costs increase in order from smallest to largest in terms of numbering. The image patch outputs the motion probability p m , where p m ∈ (0, 1); step S4 outputs the complexity category l c , l c ∈ (1, L). According to the motion probability and texture complexity, the preference in the network size is further selected, and the texture complexity weight is set as μ. Then the selected super-resolution network number:

[0182]

[0183] where round represents the floor function,

[0184] The non-integer multiple super-resolution network construction and training unit 106 is used to design multiple super-resolution networks with different costs and train them using the same data set.

[0185] In the present invention, K non-integer multiple super-resolution networks with different computational costs are designed. Each non-integer multiple super-resolution network has the same basic structure as Figure 2 shown. According to the super-resolution network number output by the super-resolution network selection unit 105, the corresponding network is selected for forward inference to output the corresponding super-resolution image. The main difference between the networks with different computational costs lies in the number of feature channels in each unit. Therefore, the number of convolution kernels used is also different, which in turn leads to different network computational complexities. Generally speaking, the smaller the computational cost, the smaller the number of channels.

[0186] Since the non-integer multiple super-resolution network has been described above, it will not be elaborated here.

[0187] The image patch super-resolution calculation unit 107 is used to utilize the trained K super-resolution networks. According to the selection of the super-resolution network selection unit 105, the corresponding super-resolution network is selected for each image patch for forward inference, and the non-integer multiple super-resolution image patch corresponding to each image patch is output.

[0188] Specifically, for each image patch, the number calculated by the super-resolution network selection unit 105 is used to select the corresponding super-resolution network for forward inference, and the non-integer multiple super-resolution image patch corresponding to each image patch is output.

[0189] The image patch temporal weighting unit 108 is used to perform motion-based weighted fusion on the super-resolution result of each image patch and the super-resolution result of the corresponding patch in the previous frame to obtain the super-resolution output of the current frame video image.

[0190] Specifically, for a certain non-integer multiple super-resolution image patch SR t output in a certain frame, its motion information is obtained by the encoder as C t , and the maximum value that can be obtained is C max . The non-integer multiple super-resolution image patch of the corresponding patch output in the previous frame is SR t-1 . Then, first perform non-integer multiple super-resolution of C t by traditional bicubic interpolation to obtain SC t , so that its size is consistent with the super-resolution image patch. Then use the pixel value SR t at each position of SR t (i, j) and the pixel value SR t-1 (i, j) at the corresponding position in the previous frame for weighting based on the motion information to obtain the value SR t (i, j) of the motion map fusion = (SR t (i, j) * SC t (i * j) + SR t-1 (i, j) * (C max-SC t (i * j))) / C max 。

[0191] The image block super-resolution output fusion unit 109 is used to splice and fuse the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.

[0192] It can be seen from the block operation unit 102 of the image and motion information map that there are overlapping regions between each image block. Therefore, during the fusion splicing process, it is necessary to average the overlapping regions to eliminate the block effect.

[0193] Embodiment

[0194] In this embodiment, taking a video image non-integer multiple super-resolution device placed after the decoder as an example, as Figure 7 shown, it includes:

[0195] Step 301: Transmit the video stream data to the video decoder using a transmission medium. The transmission medium includes but is not limited to coaxial cables, twisted pairs, and optical fibers. The video decoder is in the video receiving end device, including but not limited to devices with image display functions such as computers and mobile phones.

[0196] Step 302: Use the video decoder to obtain the YUV format image of the current frame, and use the relevant image format conversion device to obtain the R, G, and B channel data of the current frame.

[0197] Step 303: Use the decoding information output by the video decoder to extract the motion information of the encoding block corresponding to the current frame, splice all the encoding information of the current image, and use a 5×5 Gaussian rate filter to obtain the complete motion information map of the current frame.

[0198] Step 304: Perform a block processing on the image frame. After expanding the boundary, it is divided into overlapping image blocks. In this embodiment, taking the video stream image frame size of 1920×1080 as an example, a feasible image boundary expansion and image segmentation operation is given as follows: expand the image edges by 22 pixels respectively, and the pixel values of the compensation region are symmetric about the mid-axis of the edge within the image. The obtained image size is 1964×1124. The image is segmented into 32×32 sub-image blocks, and the overlapping region size between each image block is.

[0199] Step 305: Perform the same block operation on the motion image information map, and the specific operation is the same as Step 304.

[0200] Step 306: Perform motion estimation on each motion information tile. Taking a motion information tile of size 32×32 as an example, first perform a 3×3 morphological erosion operation on the image to remove noise points in the motion information; then use a 3×3 morphological dilation operation to compensate for some holes in the motion area. Then perform a histogram statistics of the motion information. Set the number of groups to 5, and the maximum value that can be obtained at each position in the motion information map is 1, so the size of each group is 0.2. Set the weighted coefficients of each group to 1 / 15,..., 5 / 15 respectively. The formula for calculating the motion probability of the image block is

[0201] Step 307: Perform texture classification on each image block. First, perform binarization processing on the image block. In this embodiment, the image is 8-bit, so the maximum gray value is 255, and 255 binary images can be obtained. Each image block has 1024 pixel points. Calculate the statistical value C(δ) of each binary image. After obtaining 255 statistical values, calculate four features respectively. Input (f1,..., f5, 1) into the texture classification model to obtain the classification label l of the image block c 。

[0202] Step 308: Pre-train the texture classification model before officially implementing the super-resolution device of the present invention to obtain model parameters. The method for obtaining training data for magnifying video images by 2.25 times in this embodiment is as follows: First, randomly select 512 groups of pictures from the super-resolution dataset. Each group of training images consists of a low-resolution image and a corresponding high-resolution image. Randomly divide the two pictures in each group into blocks and randomly cut them into 128 pairs of image blocks with sizes of 32×32 and 48×48 respectively. Perform bicubic downsampling on all 65536 high-resolution training image blocks to a size of 32×32, and calculate the PSNR value between the downsampled image block and the low-resolution image block in the same group, and arrange them in ascending order. Taking the setting of the image texture complexity as 3 categories as an example, divide the image blocks into 3 groups according to the PSNR values of the above image blocks, and the labels of each group are 1, 2, and 3 respectively

[0203] Step 308: Use the maximum likelihood estimation method to estimate 12 parameters, and the probability that the image block belongs to each category is Select the label corresponding to the maximum probability as the texture complexity category of the image block

[0204] Step 309: Select a super-resolution network with an appropriate calculation cost by integrating the motion probability of the image block and the texture complexity classification label. In this example, the texture complexity weight is 0.5. Given that the motion estimation probability of a certain image block is p m 、texture complexity classification label l c In this case, the selected super-resolution network number is

[0205] Step 310: Select the corresponding network for forward inference according to the super-resolution network number selected for the image patch. In this embodiment, three networks with different computational complexities are given:

[0206] 1. The low-resolution input and output units are the same;

[0207] 2. The input image sizes of the low-resolution feature reconstruction units are the same, all being 32×32×3. Among them, each of the multiple convolutional layers has convolution kernels of size 3×3, where represents the label of the network, n i represents the current number of convolution kernels of the network with small computational complexity. Multiple residual blocks are connected. Each residual block has two convolutional layers. When there are 2 residual blocks, there are a total of 4 convolutional layers to further extract the deep information in the image. Specifically, the structure of each residual block is as follows: a convolutional layer, a batch normalization layer, a non-linear activation function layer, a convolutional layer, and a batch normalization layer are connected in series in sequence. The input and output of the residual block are directly connected to extract local residual features. When the input data passes through the above-mentioned low-resolution feature reconstruction module, low-resolution deep features of size can be obtained.

[0208] 3. In this embodiment, a 2.25-fold super-resolution unit setting is given. Therefore, in the upsampling unit of each sub-module, the convolution kernel size k1 of the deconvolution layer 1 is 7, the stride s1 is 3, and the padding p1 is 2, that is, this layer uses convolution kernels of size in a three-dimensional convolution, where m i represents the number of convolutions of the network with the smallest computational complexity. The output image of the current layer is enlarged by 9 times (area). Then, the convolution kernel size k2 of the convolution layer 1 is set to 6, the stride s2 is 2, and the padding p2 is 2, that is, this layer uses convolution kernels of size in a three-dimensional convolution. After passing through this layer, the area of the data amplitude surface will become 0.25 times. Therefore, when the data passes through the downsampling unit, the amplitude surface becomes 2.25 times. Similarly, the parameters of the deconvolution layer 2 and the convolution layer 2 of the downsampling unit need to be set, and the above parameters can be brought in to obtain the required convolution kernel size. Obviously, the parameters of the deconvolution and convolution layers in the upsampling unit are the same as those of the convolution and deconvolution layers in the downsampling unit.

[0209] 4. The input feature size of the high-resolution feature reconstruction unit is where m last represents the number of convolutions in the last convolutional layer of the network with the smallest computational complexity in 3. In this embodiment, it is composed of two convolutional layers, and each layer is composed of It is composed of 3×3 small convolutional kernels. The size of the finally output super-resolution image is 48×48×3.

[0210] Step 311: Before officially implementing the super-resolution device of the present invention, train the non-integer multiple super-resolution network to obtain network parameters. Similar to the way of obtaining training image patches in 308, the difference is that there is no need to calculate the training data labels. The network inputs low-resolution image patches and outputs high-resolution image patches. Each group of image patches is a group of training data; use a loss function to measure the difference between the input and output training data. In the embodiments of the present invention, the balanced error MSE is used as the loss function. There are many relevant technical materials for this technology, which will not be described in detail here; use the backpropagation algorithm to update the parameters of the network with different calculation costs respectively. In the examples of the present invention, the stochastic gradient descent algorithm is used, and it can be not specifically limited according to the amount of training data and other situations.

[0211] Step 312: Use the corresponding super-resolution output image patches of the previous frame and the super-resolution image patches output by the network of the current frame, and perform weighting according to the motion weight. In this embodiment, the maximum value of the image motion information obtained by the encoder is 255. Then the output time-domain weighted super-resolution image patch is the product of the corresponding pixel of the current frame's image patch and the motion information value plus the product of the corresponding pixel of the previous frame's image patch and 255 minus the motion information value, and then the result is divided by 255.

[0212] Step 313: Stitch and fuse the image patches. In this embodiment, the size of the overlapping area of the image patches is 4. Then the image pixel values in the overlapping area are obtained by averaging the overlapping pixels of the two image patches to eliminate the block effect.

[0213] The above embodiments only illustrate the principle and efficacy of the present invention by way of example, rather than limiting the present invention. Any person skilled in the art can modify and change the above embodiments without departing from the spirit and scope of the present invention. Therefore, the scope of the right to protection of the present invention shall be as listed in the claims.

Claims

1. A method for non-integer multiple super-resolution reconstruction of video images, characterized in that, It includes the following steps: Step S1, obtaining an image frame and a motion information map of the image frame; Step S2, performing block processing on the current frame image and the motion information map at the same positions respectively; Step S3, estimating the motion probability according to the motion information map corresponding to each image block; Step S4, classifying the texture complexity of each image block; Step S5, select a super-resolution network based on the motion probability estimation and texture complexity classification label of the image patch; there are a total of K preset super-resolution networks, and the image patch outputs the motion probability p after step S3 m , where p m ∈(0, 1); step S4 outputs the complexity category l c , l c ∈(1, L), set the texture complexity weight to μ, then the selected super-resolution network number: where round represents the floor function, Step S6, designing multiple super-resolution networks with different costs and training them using the same dataset; the network computing complexities corresponding to the super-resolution networks with different costs are different, and the smaller the computing cost, the smaller the number of channels; Step S7, using the trained K super-resolution networks, according to the selection in Step S5, selecting the corresponding super-resolution network for each image block for forward inference, and outputting the non-integer multiple super-resolution image block corresponding to each image block; Step S8, performing motion-based weighted fusion on the super-resolution results of each image block and the super-resolution results of the corresponding blocks in the previous frame to obtain the super-resolution output of the current frame video image; Step S9, splicing and fusing the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.

2. The video image non-integer multiple super-resolution reconstruction method according to claim 1, characterized in that In Step S2, for the current frame image, a single-frame image is obtained from the video stream. First, a boundary expansion operation is performed on the image, and an overlapping segmentation method is used, and then the image is segmented into several identical sub-blocks. The motion information map is single-channel, and block operation is performed using the same segmentation parameters as the image blocks.

3. A method for non-integer multiple super-resolution reconstruction of video images according to claim 2, characterized in that Step S3 further includes: Step S300, performing morphological operations of erosion first and then dilation on the motion information of each image block to obtain a preprocessed motion information map; Step S301, performing histogram statistics on the preprocessed motion information map to obtain a frequency distribution; Step S302, configuring the weighting coefficients according to the motion information at each position; Step S303, calculating the motion probability of the image block based on the frequency distribution.

4. The method for non-integer multiple super-resolution reconstruction of video images according to claim 3, wherein In step S301, first normalize the motion information C at each position ij to the range of 0-1, and then set the number of groups of the histogram to k. The group interval size is 1 / k. Traverse C at all positions ij , and count them into the corresponding groups respectively to obtain k frequencies.

5. The method for non-integer multiple super-resolution reconstruction of video images according to claim 4, characterized in that, Step S4 further includes: Step S400, performing binarization processing on each image block; Step S401, extracting image texture complexity features for each image block; Step S402, constructing a texture complexity classification model based on the extracted image texture complexity features.

6. The method for non-integer multiple super-resolution reconstruction of video images according to claim 5, characterized in that In Step S401, the 5 extracted image texture complexity features are respectively the maximum value feature, the mean value feature, the sample average value, the sample standard deviation, and the entropy value.

7. The method for non-integer multiple super-resolution reconstruction of video images according to claim 6, characterized in that The training data preprocessing process of the texture complexity classification model is as follows: Randomly extract multiple groups of training images from the super-resolution training set. Each group of training images consists of a low-resolution image and the corresponding high-resolution image, and each group of images is randomly divided into blocks to obtain n groups of image block data ((LR1, HR1), …, (LR n , HR n )); Preprocess the high-resolution image patches in each group of training image patches using step S400 and step S401 to obtain training feature data (F 1 , …, F n ); To obtain the texture complexity label corresponding to the feature vector, first downsample the high-resolution image patch of the training image patch to the same size as the low-resolution image, and then calculate the PSNR value between the downsampled image patch and the original low-resolution image patch, denoted as P i , i = 1, …, n; Based on the PSNR distribution of all training image blocks, the texture complexity is labeled and divided. Suppose there are L classes of image texture complexity. The PSNRs of all training data are arranged from large to small and divided into L groups according to the number of data. The labels of each group of image blocks are sequentially set as 1,..., L, indicating that the texture ranges from simple to complex; The finally obtained training data and labels (F i , y i ), where y i = 1, …, L, represents the label of the corresponding data.

8. A method for non-integer multiple super-resolution reconstruction of video images according to claim 7, characterized in that, In Step S6, the non-integer multiple super-resolution network includes: A low-resolution image block input unit, which is responsible for inputting the image block into the super-resolution network; A low-resolution feature reconstruction unit, which is used to extract features from the low-resolution image and perform feature extraction without changing the format size; A non-integer multiple feature super-resolution unit is used to perform fast and accurate learning on the high-frequency features of non-integer multiple super-resolution for the input low-resolution features through bidirectional feature reconstruction, so as to better recover the detailed texture of the non-integer multiple super-resolution image from the low-resolution image and improve the performance of non-integer multiple super-resolution; A high-resolution feature reconstruction unit is used to extract the non-integer multiple super-resolution image from the non-integer multiple super-resolution features and perform feature channel compression without changing the feature size; A high-resolution image block output unit is used to output the non-integer multiple super-resolution image after the forward inference of the network.

9. A method for non-integer multiple super-resolution reconstruction of video images according to claim 8, characterized in that, The non-integer multiple feature super-resolution unit includes: An input unit is used to perform channel compression on the output low-resolution features of the low-resolution feature reconstruction unit to obtain low-resolution input features, and the degree of compression decreases with the increase of the network calculation amount; An initial non-integer multiple upsampling unit and a non-integer upsampling unit, and a dense connection layer composed of a dense residual connection structure is provided between the two units to achieve fast and accurate learning of the high-frequency features of non-integer multiple super-resolution; A non-integer multiple super-resolution unit, which is composed of an upsampling unit, is used to perform non-integer multiple super-resolution on the bidirectional sampling features, thereby avoiding detail loss when directly performing non-integer multiple super-resolution on the input features.

10. A video image non-integer multiple super-resolution reconstruction device, characterized in that, It includes: An image frame and image motion information acquisition unit is used to acquire the image frame and the image frame motion information map; A block operation unit is used to perform block processing at the same position on the current frame image and the motion information map respectively; An image block motion estimation unit is used to perform motion probability estimation according to the motion information map corresponding to each image block; An image block texture complexity classification unit is used to classify the texture complexity of each image block; The super-resolution network selection unit is used to select a super-resolution network based on the motion probability estimation and texture complexity classification label of the image patch; there are a total of K preset super-resolution networks, and the image patch outputs the motion probability p after step S3 m , where p m ∈(0,1); step S4 outputs the complexity category l c , l c ∈(1,L). Set the texture complexity weight to μ, then the selected super-resolution network number is: where round represents the floor function, A non-integer multiple super-resolution network construction and training unit is used to design multiple super-resolution networks with different costs and train them using the same dataset; the network calculation complexities corresponding to the super-resolution networks with different costs are different, and the smaller the calculation cost, the smaller the number of channels; An image block super-resolution calculation unit is used to use the trained K super-resolution networks, and according to the selection of the super-resolution network selection unit, select the corresponding super-resolution network for each image block to perform forward inference and output the non-integer multiple super-resolution image block corresponding to each image block; An image block time-domain weighting unit is used to perform motion-based weighted fusion on the super-resolution results of each image block and the super-resolution results of the corresponding block in the previous frame to obtain the super-resolution output of the current frame video image; An image block super-resolution output fusion unit is used to splice and fuse the super-resolution results of all image blocks in the current frame to obtain the full-frame super-resolution result of the current frame.

Citation Information

Patent Citations

  • Super-resolution method, device and equipment for video image

    CN102881000A

  • Method and apparatus for video super resolution using convolutional neural network

    CN109767383A

  • Video super-resolution method based on multi-frame attention mechanism progressive fusion

    CN112991183A

  • Video encoding and decoding method and device, computer device, and storage medium

    WO2019242491A1